Command-line usage

datalad-fuse adds three commands to DataLad:

datalad fusefs

Mount a dataset with FUSE, so that any program can read its files.

datalad fsspec-head

Print the first lines or bytes of a file.

datalad fsspec-cache-clear

Remove the on-disk cache of fetched data.

This page shows how to use them; Command line reference lists all of their options. Each command is also available from Python, see DataLad commands.

Mounting a dataset: datalad fusefs

$ mkdir mnt
$ datalad fusefs -d path/to/dataset --foreground mnt

mounts the dataset at path/to/dataset on the existing, empty directory mnt (the mount point). Without -d, the dataset containing the current directory is mounted. Create the mount point outside of the dataset: a mount point inside it would show up in the mount itself, and as an untracked directory in datalad status.

The --foreground (-f) option is currently required: the command keeps running for as long as the dataset is mounted. Run it in a terminal of its own, or in a tmux/screen session. When running it in the background of a shell (appending &), wait until the mount is ready before using it, e.g. with until mountpoint -q mnt; do sleep 0.1; done.

To unmount, press Ctrl-C in the terminal where datalad fusefs runs (for a background job, bring it to the foreground with fg first), or run:

$ fusermount -u mnt

What you see in the mount

  • The same directory tree as in the dataset’s working tree.

  • Annexed files appear as regular files, whether or not their content is present. For a file whose content is not present, the size comes from its annex key and the modification time is the date of the dataset’s last commit; reading it fetches the needed parts from a remote URL (see How it works). A file whose content is present shows the size, permissions and modification time of that content.

  • The .git directory is hidden (see --mode-transparent below).

  • Installed subdatasets are included; uninstalled ones are empty directories.

  • Files cannot be written, created or deleted (see Read-only access for a few exceptions).

Options

--caching ondisk

Keep fetched data in an on-disk cache for reuse, also by later mounts (see Caching). The default, none, only buffers data in memory while a file is open.

--backends <list>

Comma-separated, priority-ordered backends to read remote files with, e.g. --backends fsspec. The default is remfile,fsspec (see Backends).

--allow-other

Let other users access the mount; by default, only the user who mounted it can. This requires the line user_allow_other in /etc/fuse.conf.

--mode-transparent

Show the .git directories. Annexed files whose content is not present then appear as the symlinks they are in the dataset, and the targets of these symlinks under .git/annex/objects/ can be read, with their content fetched as needed. Files under .git can also be written to.

Clearing the cache on exit

The configuration option datalad.fusefs.cache-clear makes datalad fusefs remove on-disk caches when it exits:

visited

Clear the caches of the (sub)datasets that were accessed in the mount. This only has an effect if the mount used --caching ondisk.

recursive

Clear the caches of the mounted dataset and all its installed subdatasets.

Set it like any DataLad or git configuration option, permanently (e.g. git config --global datalad.fusefs.cache-clear visited) or for a single call:

$ datalad -c datalad.fusefs.cache-clear=visited fusefs -d ds --foreground --caching ondisk mnt

Peeking into a file: datalad fsspec-head

Prints the first lines (10 by default) or bytes of a file to standard output, fetching only what is needed. It is a quick way to check that the content of a file can be reached, or to look at the header of a file:

$ datalad fsspec-head -d path/to/dataset -n 5 data/participants.tsv
$ datalad fsspec-head -d path/to/dataset -c 8 data/recording.nwb | od -c

Note

Relative paths are interpreted relative to the top directory of the dataset, not to the current directory.

The output is the raw content of the file, without any result rendering, so it can be piped into other tools. --caching ondisk stores the fetched data in the dataset’s cache, and --backends chooses the backends to read with, as for datalad fusefs above.

Because it reports errors directly, datalad fsspec-head is also a handy way to check which backend handles a file, with debug logging:

$ datalad -l debug fsspec-head -d ds -c 8 sub-01/sub-01_ecephys.nwb 2>&1 | grep backend
[DEBUG] sub-01/sub-01_ecephys.nwb: opening via backend remfile

Clearing the cache: datalad fsspec-cache-clear

Removes the on-disk caches of a dataset (.git/datalad/cache/, one directory per backend):

$ datalad fsspec-cache-clear -d path/to/dataset

Add -r (--recursive) to also clear the caches of all installed subdatasets.

Choosing the backends

datalad fusefs and datalad fsspec-head both take --backends, a comma-separated, priority-ordered list of the backends to read remote files with (see Backends). Without it, the configuration option datalad.fusefs.backends is used, and failing that the default remfile,fsspec. Set the option like any DataLad or git configuration option, for a dataset, globally, or for a single call:

$ git config datalad.fusefs.backends fsspec            # in a dataset
$ git config --global datalad.fusefs.backends fsspec   # everywhere
$ datalad -c datalad.fusefs.backends=fsspec fsspec-head -d ds -c 8 file.nwb

datalad fsspec-cache-clear needs no such option: it clears the caches of all backends.