tools: compare_subjects.py

compare_subjects.py compares two FastSurfer subject directories by content, for example two runs of the same input, or the results before and after an update of FastSurfer, FreeSurfer or the hardware. diff and checksums do not help here, because the files record timestamps, the command line and versions in their headers. The script compares what the files contain instead (voxels, surfaces, per-vertex data, transforms, statistics and annotations) and reports how large each difference is.

The script is in the tools folder of a FastSurfer checkout and of the Docker and Singularity images (/fastsurfer/tools); the macOS package does not include it. It needs FastSurfer’s Python environment.

Usage:

python3 <fastsurfer_home>/tools/compare_subjects.py <subject_dir_1> <subject_dir_2>

It prints one line per file that differs or exists on one side only, and a summary. The exit code is 0 if everything compared is identical, 1 if there are differences, and 2 if it could not compare at all.

Full commandline interface of tools/compare_subjects.py

Compare two FastSurfer subject directories by content.

diff and checksums are not usable for this: the file headers record timestamps, the command line, the user name and the FreeSurfer version, so two runs that produced exactly the same measurements still differ byte for byte. This compares what the files mean instead, and reports how large each difference is, which is what turns “the runs look slightly different” into a number.

What is compared, per output type:

  • volumes (mri/*.mgz, mri/*.nii.gz): the voxel array, and the geometry (the affine) separately, so a pure header change is not reported as a data change

  • surfaces (surf/?h.white, ?h.pial, …): the vertex coordinates, the face list and the volume geometry in the header, again separately. The vertex count is checked first, because a different count means an earlier step diverged and a per-vertex comparison would be meaningless

  • per-vertex data (surf/?h.curv, ?h.thickness, ?h.w-g.pct.mgh, …): the scalar array

  • transforms (mri/transforms/*.lta, *.xfm): the RMS distance between the two transforms in mm, and the source and destination volume geometry separately

  • statistics (stats/*.stats): the numeric columns, ignoring the header comment block

  • annotations (label/*.annot): the per-vertex label array

  • the rest of surf/ (callosum.surf, autodet.gw.stats.?h.dat, *.w): the surface, the text or the bytes, whichever the format allows

Transforms are read with neuroreg, which FastSurfer depends on anyway, so that the distance is measured on the RAS-to-RAS form and a transform stored as vox-to-vox in one run and as RAS-to-RAS in the other is not reported as a difference. Without it, the files fall back to a text comparison.

Typical uses are checking whether two runs of the same input agree, and checking whether a change to FastSurfer, to FreeSurfer or to the hardware moved any measurement.

Exit codes, so a caller can tell a finding from a failure: 0 when everything compared is identical, 1 when it ran and found differences, 2 when it could not run at all, such as a path that holds no subject or a pair with nothing comparable in it.

usage: tools/compare_subjects.py [-h] [--subject SUBJECT] [--all] a b

Positional Arguments

a

first subject directory, or a directory holding one

b

second subject directory, or a directory holding one

Named Arguments

--subject

subject id, if the two paths hold more than one subject

--all

also list the files that are identical, not only the differences

Default: False