Ingestor – Upload and publish a dataset¶
The OpenEM Ingestor guides you through selecting microscope data, extracting metadata and registering the resulting dataset in SciCat. The workflow has four steps:
- Select the data and an extraction method.
- Enter user-specific metadata.
- Review and complete the extracted dataset metadata.
- Check the final dataset record and start the transfer.
Info
Ingestion creates a dataset record in SciCat and initiates the data transfer. If automatic archiving is enabled, archiving starts after the transfer has completed successfully.
Before you begin¶
Make sure that:
- you are connected to your facility network;
- you can sign in to SciCat with your institutional account through eduGAIN or SWITCH edu-ID;
- the files are located in a directory accessible to your facility's Ingestor;
- you have permission to read all files in the directory;
- the data collection is complete and the files will no longer be changed; and
- you know the owner group under which the dataset must be registered.
Choose the dataset boundary carefully. The selected directory and its files are handled as one dataset. A dataset is also the unit that is described, transferred, archived, retrieved and, if applicable, published.
Warning
Do not modify, move or delete source files after starting ingestion. Changes can cause the transfer or subsequent archive job to fail.
Open the Ingestor¶
- Open SciCat and sign in.
- Select the menu icon in the upper-left corner.
- Select Ingestor.
- If prompted, connect or sign in to the Ingestor service.
SciCat normally discovers the Ingestor assigned to your facility. If no Ingestor is shown, verify that you are connected to the facility network. If your facility requires manual configuration, use the URL supplied by your local OpenEM support team. The available facility endpoints are listed under Participating Facilities.
1. Select the data and extraction method¶

- Select the menu icon in the upper-left corner.
- Select Ingestor.
- In the Transfers view, select the + button under New transfer.
- In File Path, select the directory that contains the dataset and under Extraction Method, select the extractor that matches the data.
- Select Next.
The Ingestor can only show paths within the configured facility data collection. Select the directory itself, not an individual representative file, unless your facility's workflow explicitly requires otherwise.
The extraction method determines which files are inspected, which scientific metadata fields are generated and which additional fields appear later in the workflow.
Warning
If no suitable method is available, do not choose one merely because it is listed; contact your local OpenEM support team.
Tip
Before continuing, check that the selected directory contains only the files that belong to this dataset. This avoids transferring unrelated intermediate, temporary or personal files.
2. Enter user-specific metadata¶

The User metadata step contains information that cannot reliably be extracted from the files. Fields marked with an asterisk are mandatory.
- Leave Required Only enabled for a compact view, or disable it to display optional fields as well.
- Open each metadata section and check the pre-filled values.
- Complete every mandatory field.
- Ensure that Creation Location starts with
/. - Select Next.
Common fields include:
| Field | What to enter or verify |
|---|---|
| Owner Group | The group whose members are allowed to access and manage the dataset. Select the correct project or facility group. |
| Is Published | Leave this disabled during ingestion. Publish the archived dataset later with the SciCat publication workflow. |
| Owner | The person responsible for the dataset. This may be pre-filled from your account. |
| Source Folder | The selected source path. Verify it carefully; it is normally read-only. |
| Dataset Name | A concise, meaningful name that helps users identify the dataset in SciCat. |
| Creation Location | The facility or instrument location, written as a path beginning with /, for example /facility/microscope. |
| Principal Investigator | The person responsible for the project or experiment. |
Additional administrative or organisational sections may be shown depending on the selected extractor and facility configuration. Optional metadata makes a dataset easier to find and understand later, so provide it when the information is known.
Warning
The owner group controls access to the dataset. Confirm it before ingestion; do not use a group only because it is the first available option.
3. Review the dataset metadata¶
In the Dataset metadata step, review the information extracted by the selected method. The exact sections and fields depend on the data format and extraction method.
- Review the extracted values for plausibility, including units.
- Correct incomplete or incorrect editable values.
- Fill in any remaining mandatory fields.
- Expand optional sections and add useful scientific context where available.
- Select Next to continue to confirmation.
Pay particular attention to values that cannot be inferred unambiguously from a file, such as the sample identity, experiment description or acquisition context. If extraction fails or important values are missing, go back and verify both the selected path and extraction method.
4. Confirm and start the transfer¶

The Confirm step shows the combined dataset information as JSON. This is the record that will be submitted to SciCat.
- Decide whether Auto Archive should remain enabled. When enabled, the Ingestor starts archiving automatically after the data transfer finishes.
-
Review the generated JSON, especially:
sourceFolderanddatasetName;ownerGroup,ownerandprincipalInvestigator;creationLocation;- the dataset type and data format; and
- extracted scientific metadata and units.
-
Select Back if a value must be corrected.
- When the record is complete, select Ingest.
Do not close the dialog while the submission is being accepted. After successful submission, return to the Ingestor view to follow its progress.
Monitor the transfer¶
The transfer list shows the current state of your ingestion jobs. Refresh the view if the status has not updated yet. Depending on the dataset size and the available facility capacity, transfer and archiving may take some time.
A completed submission should result in a searchable dataset record in SciCat. If automatic archiving was selected, also verify that the archive operation completes successfully before treating the upload as finished.
Do not create another transfer for the same directory merely because processing is still in progress. This can create duplicate dataset records.
Publish the dataset¶
Publishing is a separate SciCat workflow after ingestion and archiving. Only publish data that may be made publicly accessible. The dataset must have the Retrievable status before it can be added to a publication.

- Return to Datasets in SciCat.
- Enable the My data filter.
- Open the Retrievable tab.
-
Select the dataset and select Add to Selection. Repeat this for every dataset that should be published under the same DOI.

-
Open Selection in the upper-right corner.
-
Select Actions.

-
Check that the selection contains the intended datasets.
-
Select Publish.

-
Enter a Title and Abstract.
- Expand Metadata, check any pre-filled values and complete all mandatory publication fields.
- Select Save and Continue.
- Review the publication definition and select Publish to make the data publicly accessible.
Use Save changes instead of Save and Continue if you need to keep the publication as a draft and finish it later.
Troubleshooting¶
The Ingestor cannot be reached¶
- Confirm that you are connected to the facility network or its approved VPN.
- Reload SciCat and sign in again if the session has expired.
- If automatic discovery fails, check the endpoint in Participating Facilities or contact local support.
The source directory is not visible¶
- Confirm that the files are stored below the facility data collection exposed to the Ingestor.
- Check that you have permission to access the directory and all contained files.
- Contact local support if the expected storage area is not available in the path selector.
Next remains disabled¶
- In the data browser select both a valid File Path and Extraction Method.
- In the metadata steps, complete all fields marked with an asterisk.
- Check validation messages and ensure that Creation Location begins with
/.
Metadata is missing or implausible¶
- Return to step 1 and verify the extraction method.
- Confirm that the selected folder contains the expected source files and that the extractor supports their format.
- Correct editable fields before ingestion. For systematic extractor problems, record the extraction method and affected file format when contacting support.
The transfer reports an error¶
- Open the transfer entry and note the displayed error message.
- Check that the source files still exist, have not changed and remain readable.
- Correct the underlying issue before starting a new transfer.
- If the problem persists, provide local support with the dataset path, transfer time, selected extraction method and error message. Do not send confidential data or credentials.
For contact details, see Support.
Useful functions¶
The User metadata and Dataset metadata steps provide tools for reducing the number of displayed fields and reusing metadata.

| Number | Function | Description |
|---|---|---|
| 1 | Required Only | When enabled, the form shows only fields that are required for ingestion. Disable it to view and complete optional metadata as well. |
| 2 | Save or export metadata | Exports selected values from the current form. Use this to save reusable metadata as a template or to download the metadata as JSON. |
| 3 | Load a template | Imports values from a previously saved template into the form. Review all imported values and adapt dataset-specific information before continuing. |
Save or export metadata¶
Select the disk icon (2) to choose which metadata should be exported.

- Export SciCat includes the values from the SciCat Information section.
- Export Organizational includes the organisational metadata.
- Export Sample includes the sample metadata.
- Export All (includes extracted metadata) exports all available sections, including metadata produced by the extractor.
- Export As JSON (not a template format) downloads the selected data as regular JSON instead of an importable Ingestor template.
Select the required options and then select Confirm. To reuse an exported template in another ingestion, select the upload icon (3) and choose the template file.