datalad-fuse
Open files of DataLad datasets and git-annex repositories without downloading them first.
When you clone a DataLad dataset or a
git-annex repository, you get all file
names and the full history, but the content of the large (“annexed”) files
stays where it was published: for example in the S3 storage of DANDI or OpenNeuro, or on a
web server. datalad get downloads such files in full before you can open
them.
datalad-fuse lets you open them right away instead. It looks up the URLs
that git-annex has recorded for a file and fetches only the parts of the file
that are actually read, with
fsspec or, for NWB/HDF5 files and
if installed, remfile (see
Backends).
Is it a good fit?
datalad-fuse pays off when you need a fraction of large files:
some arrays or tables out of NWB/HDF5 files, without the raw data;
the headers of many imaging files;
small metadata or sidecar files spread over a large dataset;
a quick look into a dataset that would not fit on your disk.
When you need whole files, for example the voxel data of compressed NIfTI
(.nii.gz) images, which must be decompressed from the start, datalad
get is usually faster, and keeps the files for later.
There are two ways to use it:
FUSE mount |
Python |
|
|---|---|---|
What |
|
|
Works with |
Any program: |
Python libraries that accept file objects: h5py, pynwb, pandas, json, … |
Needs |
Linux with FUSE installed |
Just Python (nothing to install beyond |
Use Python if the library you read files with accepts Python file objects (h5py, PyNWB, pandas, …); use the FUSE mount for tools that need a file name: MATLAB, FSL, command-line tools, or Python functions that only accept paths.
A quick look
Get a dataset, here a dandiset from the DANDI Archive, which is small to clone (about 2 MB) although its files add up to 1.86 GB:
$ datalad clone https://github.com/dandisets/000582
(OpenNeuro datasets are cloned the same way, e.g. datalad clone
https://github.com/OpenNeuroDatasets/ds000001.)
Mount it. datalad fusefs keeps running until the dataset is unmounted:
$ mkdir mnt
$ datalad fusefs -d 000582 --foreground mnt
Then, in another terminal, use any tool on its files, e.g. h5ls from the
HDF5 tools, and unmount when done:
$ h5ls mnt/sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb
acquisition Group
analysis Group
...
units Group
$ fusermount -u mnt
Or open files directly from Python:
from contextlib import closing
import h5py
import pynwb
from datalad_fuse.adapter import DatasetAdapter
# caching=False: keep fetched data in memory only (see "Caching")
with closing(DatasetAdapter("000582", caching=False)) as dsa:
with dsa.open("sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb") as f:
with h5py.File(f, "r") as h5, pynwb.NWBHDF5IO(file=h5) as io:
nwbfile = io.read()
print(nwbfile.units.to_dataframe())
The Tutorial: streaming NWB data from DANDI walks through both in more detail.
Getting started
Development