Website mapping guide for the existing data standard
Data standardization for this project is already complete. The website JSON is only an index mapping for public search, sample cards, and detail views. It does not replace or redefine the original standard. Real raw data, processed volumes, and their authoritative metadata remain stored according to the project's existing specifications.
The sample index contains six project-supplied images, all in Validation. Names follow source filenames. Original image pixels are unchanged; the webpage adjusts their presentation rotation, framing and brightness, with an unfiltered full-original view. Volume metadata remains null and no volume attachments are available.
1. Integration order
- Establish the existing standard's name, version, field dictionary, and authoritative maintenance source. Have the project lead confirm which fields may be public.
- Map original fields → website fields, documenting conversion rules, units, and missing-value meanings. Extend the website mapping layer if necessary; do not change existing scientific meaning to fit the page.
- Copy
data/sample.template.jsonand enter metadata for one real sample approved for public release. The template contains no valid experimental values. Its presetis_demo: falseis intended for preparing a real record; it is not proof of authenticity and does not make the template publishable. - Check types, required fields, and enumerations against
data/metadata.schema.json, then validate the actual files and scientific meaning. Passing Schema validation does not establish correctness or permission to publish. - Put approved records in the top-level array of
data/samples.json, replacing or clearly separating DEMO records. Useis_demo: falseonly for genuine real samples. - Preview through an HTTP server and check filters, details, units, missing values, and resource states before mapping records in bulk.
The current Schema and template are authoritative for field types, nesting, and allowed values. When changing them, also update assets/js/app.js and this guide.
Current JSON structure
The root of samples.json is a sample array; the root of sample.template.json is one sample object. metadata.schema.json uses JSON Schema Draft 2020-12 and validates an array, so wrap a single template as [template] before validation. Do not submit the template object as the entire index or add an extra samples property around it.
Each sample has the following fields. Fields described as nullable must still retain their field names. Allowing null during preparation does not mean release review has passed.
| Field | Current type and meaning |
|---|---|
id |
Nonempty string; letters, digits, periods, underscores, and hyphens only, starting with a letter or digit. Check uniqueness separately before release. |
title / description |
Public title and description; accept strings or bilingual { "zh": "中文", "en": "English" } objects. Titles must not be empty. |
is_demo |
Boolean; true for synthetic interface demonstrations. |
category |
phantom, simulation, in-vivo, ex-vivo, or other. |
split |
train, validation, test, or unspecified. |
tags |
Array of unique strings; use [] when there are no tags. |
preview |
Project-relative path or HTTPS image URL; nullable. Clearly label illustrations. |
dimensions |
Three positive integers in axis_order; nullable. |
axis_order |
Three-element array containing x, y, and z once each; nullable. This does not specify anatomical orientation. |
spacing_mm |
Three positive numbers in millimeters, in the same order as axis_order; nullable. |
wavelength_nm |
One positive number in nanometers; nullable. Multiple wavelengths require an explicit Schema extension. |
intensity_unit |
Signal/intensity unit string or bilingual object; nullable. |
data_version |
Data version string for the record; nullable. |
provenance |
Required provenance object containing the four fields below. |
files |
Array of file objects; use [] until resources are integrated. |
provenance must contain four nullable fields: source_id (public source identifier), acquisition (source for the array/acquisition configuration), reconstruction (source for the reconstruction method and version), and standardization_reference (traceable reference to the existing standard and version). source_id remains a string; the other three description fields accept strings or bilingual objects. There is currently no separate array-geometry object. Update the Schema and page together if a structured extension is needed.
Each files element must contain name, format, url, size_bytes, and sha256. name is a nonempty string and format is a format-description string; the last three may be null. A non-null size_bytes is a nonnegative integer, and a non-null sha256 is a 64-character hexadecimal string. The Schema checks only these basic constraints. It does not ensure that URLs work, units are correct, IDs are unique, or publication is authorized.
2. Minimum checks
Identity and grouping
- A stable, unique
id, publictitle, anddescription. category,split, andtagsmust map to actual research definitions. All six images are assigned to Validation by the project owner. Phantom categories follow filenames; in-vivo or ex-vivo status is not inferred.- If real categories or splits fall outside current enumerations, extend the Schema, interface options, and statistics logic first. Do not force unknown categories into existing choices.
- Follow the original study design for data splitting. Check for leakage when the same subject, experiment, or strongly related samples occur across sets.
Scale, units, and coordinates
- Array dimensions state only the element count along each axis. Verify axis order and spatial coordinate conventions as well.
- When
spacing_mmrepresents voxel spacing, ensure its entries correspond to the dimension axes and correctly convert original units to millimeters. - Obtain
wavelength_nmfrom actual records. Do not discard multiwavelength information because a page displays only one number. - Specify physical units, normalization status, and data type. Arbitrary units and physical units are not interchangeable.
- Unknown, not applicable, not provided, and zero have different meanings. Use missing-value representations permitted by the Schema; do not use
0for unknown values.
Acquisition and reconstruction provenance
- Retain publicly shareable sources for array type/geometry, acquisition configuration, reconstruction method, and version.
- Keep relationships between raw signals, reconstructed volumes, and algorithm outputs traceable.
- Provenance fields should identify trusted sources. They must not include participant identities, internal paths, passwords, or tokens.
- Keep nonpublic provenance in controlled records. State its availability clearly on public pages rather than fabricating descriptions.
Versions and resources
- Each resource version must match its associated data release and be updated when content changes.
- Include in
filesonly existing resources authorized for public or controlled distribution. Keep an empty array when no files are available. - Test download URLs in practice. Prefer permanent entry points to expiring signed URLs.
- Compute file size, format information, and SHA-256 from the final release files; do not infer them from filenames.
- Leave placeholder URLs, absent DOIs, and undecided licenses empty. Do not create apparently usable fake links.
3. Recommended mapping record
Maintainers can keep the following record in their own data-processing project. Only an approved public version should enter the website:
| Original field | Original standard version | Website field | Conversion rule | Missing-value handling | Validation method |
|---|---|---|---|---|---|
| To be supplied | To be verified | Follow the Schema | For example, unit conversion or axis reordering | Follow the original standard and Schema | Compare with source records/files |
Record actual operations under “Conversion rule,” especially coordinate reordering, resampling, wavelength selection, and merging versions. If the website only displays data without conversion, explicitly record a direct mapping.
4. Large files, links, and checksums
This repository holds only the website and lightweight public index. Store large scan files, archives, and model weights in a project-approved data repository or object store, then link to them. Controlled data needs a service with access control. Static Pages hosting and hidden buttons cannot protect data.
Compute SHA-256 locally for prepared final files:
# Linux
sha256sum path/to/approved-file
# macOS
shasum -a 256 path/to/approved-fileEnter the actual hash in the corresponding resource field, retaining filename, byte size, and version records. Checksums verify integrity; they do not establish scientific quality or replace publication permission.
Before launch, test public links in a fresh, signed-out browser session. Check the final downloaded files, redirects, access requirements, and checksums. An HTTP 200 response alone is insufficient because the content may be a sign-in or error page.
5. Release boundaries
- The public index is itself public data. It must not contain identifying information or sensitive metadata.
- For human, animal, or other restricted data, first verify approval, consent, licensing, and de-identification requirements.
- Retain DEMO labels and pending-resource states while real data has not been integrated.
- Derive sample counts, total size, and coverage only from a verified real release manifest.
- Update versions, citations, and checksum records whenever files, metadata, or licenses change.
Then work through the pre-release checklist.
6. Bilingual display fields
title, description, intensity_unit, and provenance fields acquisition, reconstruction, and standardization_reference accept a legacy string or { "zh": "中文说明", "en": "English description" }. Nullable fields may still be null. A bilingual object must provide at least one zh or en string; supplying both is recommended. Every supplied title translation must be nonempty. The runtime displays the active language first, falling back to Chinese and then English when valid text is missing.
Keep id, category, split, tags, data_version, provenance.source_id, filenames, URLs, numeric values, and checksums in their original formats. Translations must describe the same facts without changing data, technical identifiers, or demo status. Check that both languages clearly identify synthetic DEMO content and keep missing or unverified research information empty or pending.