Python usage
From Python, files of a dataset can be opened without any FUSE mount, as file objects that many libraries accept in place of a file name. This page covers:
DatasetAdapter: open files of a single dataset;RemoteFilesystemAdapter: open files across a dataset and its subdatasets;the
dataladcommands ofdatalad-fuse, called from Python;mounting a dataset with FUSE from Python.
Note
Up to 0.6.0, the adapters lived in datalad_fuse.fsspec, and
RemoteFilesystemAdapter was called
FsspecAdapter. Both old names still work, with a
DeprecationWarning:
from datalad_fuse.fsspec import FsspecAdapter # deprecated
from datalad_fuse.adapter import RemoteFilesystemAdapter # use this
The examples use the dandiset cloned in the Tutorial: streaming NWB data from DANDI, and are run from the directory containing it:
nwb_path = "sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb"
Opening files of a dataset
DatasetAdapter takes the path to a dataset (or any
git-annex repository), and opens files by their path relative to the
dataset’s top directory:
from contextlib import closing
from datalad_fuse.adapter import DatasetAdapter
with closing(DatasetAdapter("000582", caching=False)) as dsa:
with dsa.open(nwb_path) as f: # binary mode, like open(..., "rb")
print(f.read(8)) # b'\x89HDF\r\n\x1a\n'
with dsa.open("dandiset.yaml", "rt") as f: # text mode
print(f.readline())
Important
Many libraries (h5py, PyNWB, zarr, …) read data lazily, only when you
access them. The file object must stay open until all data you need have
been read, so do the reading inside the with blocks, or keep the
objects open while working interactively (see the Tutorial: streaming NWB data from DANDI).
Reading after the file was closed fails, with h5py with an error
mentioning “identifier is not of specified type”.
open() returns a seekable, read-only file object:
for files read from disk (not annexed, or with content present), a regular Python file object;
for files read from a URL, an object from the backend that opened it, which fetches data as it is read: an fsspec file object from the
fsspecbackend, aRemfileWrapperfromremfile(see Backends).
Text mode ("r" or "rt") accepts encoding (default "utf-8") and
errors arguments, as the built-in open() does. Writing is not
supported.
The caching argument is required: True keeps fetched data in an
on-disk cache inside the dataset, False only buffers them in memory while
a file is open (see Caching). With caching=True, dsa.clear()
removes the dataset’s cache.
The optional backends argument chooses which backends to read remote
files with, as a comma-separated, priority-ordered string:
# read every file with fsspec, even if remfile is installed
DatasetAdapter("000582", caching=False, backends="fsspec")
Without it, the datalad.fusefs.backends configuration option is used, and
failing that the default "remfile,fsspec"
(DEFAULT_BACKENDS). See Backends
for what the backends do.
Note
The backends argument is newer than the 0.6.0 release.
The adapter starts git annex processes to answer its queries;
close(), called by contextlib.closing() above, stops them.
Passing file objects to other libraries
Anything that reads from a Python file object can read from these. For example:
import json
import h5py
import pandas as pd
with closing(DatasetAdapter("path/to/dataset", caching=True)) as dsa:
# HDF5, and formats based on it such as NWB
with dsa.open("data/recording.h5") as f, h5py.File(f, "r") as h5:
signal = h5["signal"][:1000]
# tabular text data
with dsa.open("participants.tsv", "rt") as f:
participants = pd.read_csv(f, sep="\t")
# JSON
with dsa.open("dataset_description.json", "rt") as f:
description = json.load(f)
Libraries that need a file name rather than a file object cannot use these objects; use a FUSE mount for them (see Mounting from Python).
Inspecting files
get_file_state tells whether a file is
annexed and whether its content is present, and returns its git-annex key as
an AnnexKey:
with closing(DatasetAdapter("000582", caching=False)) as dsa:
state, key = dsa.get_file_state(nwb_path)
print(state) # FileState.NO_CONTENT
print(key.size, key.backend) # 15657857 SHA256E
for url in dsa.get_urls(str(key)):
print(url)
which prints the URLs that will be tried, in order:
FileState.NO_CONTENT
15657857 SHA256E
https://api.dandiarchive.org/api/assets/2b9e441b-56bc-4be2-893e-0e02d22d239d/download/
https://dandiarchive.s3.amazonaws.com/blobs/26a/22c/26a22c31-09bc-43a4-9187-edc7394ed12c?versionId=__7hm7itizkF8RCsvO.Fidzi7Lqd1OMu
The possible states are described in How it works.
AnnexKey can also parse and format keys on its own:
from datalad_fuse.utils import AnnexKey
key = AnnexKey.parse("SHA256E-s15657857--43b3b435b953d22e276acc494af2926b63deaf15d2834531b2a87d08a8458a09.nwb")
print(key.backend, key.size, key.suffix) # SHA256E 15657857 .nwb
print(str(key)) # the key again
Datasets with subdatasets
RemoteFilesystemAdapter works on a dataset together
with its installed subdatasets: for each path, it finds the (sub)dataset that
contains it and uses a DatasetAdapter for that
dataset. It is a context manager:
from pathlib import Path
from datalad_fuse.adapter import RemoteFilesystemAdapter
root = Path("path/to/superdataset").resolve()
with RemoteFilesystemAdapter(root, caching=False) as fsa:
path = root / "subdataset" / "data" / "file.nwb"
print(fsa.get_file_state(path))
print(fsa.is_under_annex(path))
with fsa.open(path) as f:
header = f.read(1024)
Important
Use an absolute root, and absolute paths under it, as above. Paths
relative to root, or to the current directory, are not supported.
Besides open(), get_file_state() and is_under_annex(), it offers
get_commit_datetime() (the date of the last commit of the dataset
containing a path) and resolve_dataset() (the
DatasetAdapter and relative path used for a path).
DataLad commands
The commands of datalad-fuse are available as functions in
datalad.api and as methods of datalad.api.Dataset, with the same
options as on the command line (see Command-line usage):
from datalad.api import Dataset
ds = Dataset("000582")
res = ds.fsspec_head(nwb_path, bytes=8, result_renderer="disabled")
print(res[0]["data"]) # b'\x89HDF\r\n\x1a\n'
ds.fsspec_cache_clear(recursive=True)
Like all DataLad commands, they return result records (dictionaries);
fsspec_head puts the fetched bytes into the data field of its result.
Mounting from Python
datalad.api.fusefs (or Dataset.fusefs) mounts a dataset, and does not
return until it is unmounted. To work with the mount from the same Python
program, run it in a separate process:
from multiprocessing import Process
import os
import subprocess
import time
from datalad.api import fusefs
def main():
os.makedirs("mnt", exist_ok=True)
mount = Process(
target=fusefs,
args=("mnt",),
kwargs={"dataset": "000582", "foreground": True, "caching": "ondisk"},
)
mount.start()
while mount.is_alive() and not os.path.ismount("mnt"):
time.sleep(0.1)
try:
# any code or tool can now open files under mnt/
path = "mnt/sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb"
with open(path, "rb") as f:
print(f.read(8))
finally:
subprocess.run(["fusermount", "-u", "mnt"], check=True)
mount.join()
if __name__ == "__main__": # required by multiprocessing
main()
Put this code in a script: the if __name__ == "__main__" guard is required
where multiprocessing starts new processes by re-importing the main
module, which is the default on macOS and, from Python 3.14 on, on Linux.
To pass other FUSE mount options, mount the file system class
DataLadFUSE directly with fusepy; keyword arguments of
FUSE() become mount options. DataLadFUSE needs the absolute path of
the dataset, without symbolic links, as returned by os.path.realpath():
import os
from fuse import FUSE
from datalad_fuse.fuse_ import DataLadFUSE
FUSE(
DataLadFUSE(os.path.realpath("000582"), caching=True),
"mnt",
foreground=True,
ro=True,
)