Frequently asked questions¶
This page provides short answers to common questions about SciCat and OpenEM. Follow the links in each answer for detailed instructions.
SciCat¶
What is SciCat?¶
SciCat is the data catalogue used to register, find, archive, retrieve and publish scientific datasets. The catalogue stores the administrative and scientific metadata; the associated data files are transferred to separate storage and archive systems. OpenEM uses SciCat as its central catalogue and user entry point. See the SciCat Ingestor Manual.
What is a dataset in SciCat?¶
A dataset is a logical collection of files together with its metadata. It is the smallest unit that can be described, transferred, archived, retrieved and published. Choose the boundary carefully: a directory should contain all and only the files that belong to that dataset. See The Concept of Datasets.
Why can I not see or manage a dataset?¶
Access is controlled primarily by the dataset's Owner Group. Only members
of that group can access a non-public dataset. At PSI, proposal groups usually
start with p, while archive groups start with a-. Confirm that the correct
group was selected and that your account is a member of it. Public datasets
can be viewed without group membership.
Where do I get an API token for the SciCat command-line tools?¶
Sign in at discovery.psi.ch, open your user settings
and copy the token. Personal accounts use this token with the --token option;
do not put it into documentation, screenshots or support requests. Functional
accounts may use a different authentication workflow. See
Getting started.
May I modify files after registering a dataset?¶
No. Do not modify, move or delete source files after ingestion has started. Archiving uses the file list created at ingestion time, so later changes can cause the archive job to fail. Ingest only after data collection and processing for that dataset are complete.
How do I archive or retrieve a dataset?¶
An archivable dataset can be sent to long-term storage from SciCat or with the SciCat CLI. Retrieval is a two-step process: first request the archived dataset in SciCat, then copy it from the intermediate cache after retrieval completes. See Archive and Retrieve.
How do I publish datasets and obtain a DOI?¶
Use SciCat's publication workflow to select one or more datasets, complete the publication metadata, and perform the Save, Publish and Register steps. Registration creates the DOI. Confirm that the data may be made public before starting. See Publish.
Which SciCat CLI version should I use?¶
Use SciCat CLI version 3 or newer unless your managed PSI environment provides
the correct version. The former standalone tools are now subcommands of
scicat-cli, and long options require two hyphens. See the
January 2026 upgrade notes.
OpenEM¶
OpenEM connects electron microscopy facilities to SciCat. Its Ingestor registers and transfers datasets, metadata extractors describe them in a structured form, and its Depositor prepares eligible datasets for OneDep.
Log-In & Registration¶
Where do I start using OpenEM?¶
Start at discovery.psi.ch. Sign in there to search for datasets and proposals, start the OpenEM Ingestor and open the Depositor for an eligible dataset. The OpenEM Quick Start shows the main workflow.
Do I need a separate OpenEM account?¶
Usually not. OpenEM uses institutional authentication through eduGAIN and SWITCH edu-ID. You do, however, need an active facility account, the required group memberships and permission from your facility to use its OpenEM services.
How do I sign in?¶
In SciCat, open the profile menu and select Login with Single Sign-On, followed by Single Sign-On with eduGAIN. Select or switch to the appropriate edu-ID and authenticate with your institutional credentials. See Log in.
Why can I sign in to SciCat but not open the Ingestor?¶
The facility Ingestor is normally reachable only from the facility network or its approved VPN. Confirm that you are on that network and that your facility has granted access. If automatic discovery fails, use the Ingestor URL supplied by your facility. See Participating Facilities.
Why is my proposal or owner group missing?¶
First confirm that you are signed in with the correct institutional identity. Then verify your proposal or group membership with the responsible facility or user office. OpenEM cannot assign group membership itself. Do not ingest under another group simply because it is available: the owner group controls access to the dataset.
Ingestor¶
What does the OpenEM Ingestor do?¶
It lets you select microscope data, runs the appropriate metadata extractor, combines extracted and user-provided metadata, creates the dataset record in SciCat and starts the data transfer. If Auto Archive is enabled, archiving starts after a successful transfer.
What should I check before starting an ingestion?¶
Make sure you are connected to the facility network, can access every source file, know the correct owner group and have finished changing the data. The selected directory should contain only the files belonging to one dataset. See Before you begin.
Which file path should I select?¶
Select the dataset directory exposed by your facility's Ingestor, rather than a single representative file or an unrelated parent directory. If the expected directory is absent, check your permissions and confirm that it is below the facility data collection made available to the Ingestor.
Which extraction method should I choose?¶
Choose the method matching the scientific domain, acquisition software and file format. The selection determines which files are inspected and which metadata are produced. If no method matches your data, contact local support instead of choosing an arbitrary extractor. See Available extraction methods.
Why is the Next button disabled?¶
In the first step, both File Path and Extraction Method must be selected.
In later steps, complete every field marked with an asterisk and resolve the
displayed validation messages. In particular, Creation Location must begin
with /.
What does Required Only do?¶
It hides optional metadata fields to provide a shorter form; it does not remove metadata already extracted or change which fields are mandatory. Disable it if you want to add optional context that will make the dataset easier to find and understand.
Should I enable Auto Archive?¶
Enable it when the dataset should be archived automatically after transfer and your facility workflow permits this. Ingestion and archiving are separate operations, so verify that both the transfer and archive job complete before considering the upload finished.
What should I do when a transfer fails or appears stuck?¶
Refresh the transfer view and allow time for large datasets. If an error is shown, record it and check that all source files still exist, are unchanged and remain readable. Do not submit the same directory again merely because it is still processing, as that can create duplicate records. See Ingestor troubleshooting.
Depositor¶
What is the OpenEM Depositor?¶
The Depositor prepares an eligible OpenEM dataset for the wwPDB OneDep system. It converts OSC-EM metadata to mmCIF, combines it with the files you provide and can transfer the result to OneDep or prepare it for manual upload. See the Depositor guide.
Why is the Depositor action not shown for my dataset?¶
Confirm that you are signed in, can access the dataset and that it is registered as an OpenEM dataset with the metadata required by the Depositor. The action may also depend on the configured facility or deployment. Try another known eligible dataset; if the action is still absent, contact local support.
Which files do I need for a deposition?¶
That depends on the experimental method and whether an atomic model is included. A map-only EMDB deposition normally needs a primary map and entry image; a combined PDB/EMDB deposition also needs atomic coordinates. Half maps, masks, FSC curves or other supporting files may also be required. OneDep is the authoritative source for the final required file set. See Decide what to deposit.
Does the Depositor complete the OneDep submission for me?¶
No. After transfer, continue in OneDep, review the processed files, complete any missing mandatory metadata, resolve validation errors, choose the release policy and submit the deposition. Keep the OneDep deposition identifier for future access and support requests.
What should I do if OneDep authorisation fails?¶
Check that you are using the correct OneDep environment and that your token or session has not expired. Sign in again or create a new token if your deployment requires one. Never send tokens, passwords or session cookies to support staff.
Why is the map pixel spacing missing or implausible?¶
The Depositor reads pixel spacing from supported map headers when possible. Verify the source map and its header, compare the value with the reconstruction software and correct it in OneDep if needed. Do not continue with a guessed value. See Depositor troubleshooting.
What should I do when OneDep reports missing or invalid files?¶
Check that every file finished uploading, is readable, uses a supported format and has the correct OneDep file category. Follow the method-specific validation messages in OneDep; it determines the definitive required file set.
Metadata¶
What is the difference between administrative and scientific metadata?¶
Administrative metadata describes ownership and management of the dataset, such as owner group, owner, source folder, creation location and dataset name. Scientific metadata describes the sample, instrument, acquisition and processing. Both types are stored with the SciCat dataset.
Which metadata are filled automatically?¶
The extractor reads supported source files and acquisition metadata. Facility configuration may add instrument identity and other stable settings, while you provide information that cannot be inferred reliably, such as ownership, sample identity and experiment context. Always review automatically generated values. See How extraction works.
Which extraction methods are available?¶
OpenEM provides methods for Life Science and Materials Science data. Support depends on the actual file formats and acquisition software, not only on the scientific discipline. The current inputs and limitations are listed under Available extraction methods.
What should I do if metadata are missing or implausible?¶
Check that you selected the correct dataset directory and extraction method, then compare the result with the acquisition software or source metadata. Fix editable values before ingestion. If the same problem occurs repeatedly for an instrument or format, give local support the extractor name, file format and an example field, but do not send confidential data.
Why does OSC-EM validation fail?¶
Expand all metadata groups and complete mandatory fields, then check data types, allowed values and units. A value can look reasonable while still violating the schema. Record the schema profile, field path and exact validation message if you need support. See Metadata troubleshooting.
Are optional metadata worth completing?¶
Yes. Optional metadata improve discovery, interpretation and reuse, and can reduce manual work during a later OneDep deposition. Provide them when known, but do not guess values merely to fill a field.
Installation¶
Do end users need to install OpenEM?¶
No. End users access SciCat and the facility Ingestor in a web browser. The facility's operators install and maintain the Ingestor, transfer server and related services. End users may need network or VPN access supplied by their facility.
What infrastructure does a facility need?¶
A facility needs a Linux transfer server with access to the raw-data storage, sufficient cache capacity, a supported container runtime and the required network connectivity. Exact sizing depends on data volume and transfer patterns. See Requirements & Infrastructure.
How is the Ingestor installed?¶
Operators deploy it with Docker and Docker Compose, configure the facility URL,
storage paths, extractors and central service endpoints in .env, then start
and verify the containers. Follow the
Ingestor Installation guide rather than
copying settings from another facility.
Is Globus Connect Server required?¶
It is part of the documented facility-to-central transfer setup. Operators must configure the endpoint, network access and identity mapping, then register the domain, endpoint ID and facility name with SciCat Support. End users should not be given access to the service identity. See Globus Connect Server Installation.
Which firewall and domain settings are required?¶
The transfer server needs the inbound and outbound connections documented for the Ingestor, SciCat and Globus services, together with stable facility domain names and TLS. Because the exact rules may change with the deployment, use the current Network Requirements and coordinate changes with SciCat Support.
Why do I receive a 403 Forbidden response during development?¶
Some OpenEM development services accept connections only from approved networks. Use the documented SOCKS5 proxy through an authorised host and configure your browser to use it. This procedure is for developers and operators, not normal end-user access. See Development Proxy.
How do operators update a metadata extractor?¶
The openem-deployment should
contain the latest versions and checksums of the extractors as part of the
ingestor service compose file. Pull the latest openem-deployment release, then
restart the Ingestor (./compose.sh all down && ./compose.sh all up -d).
Contact¶
Who should I contact when I need help?¶
Contact your local facility support team first for access, network, source-data, instrument or facility Ingestor issues. If the facility cannot resolve the problem, contact SciCat Support. See the Support page for the support path.
What information should I include in a support request?¶
Include your facility, the affected component, the approximate time, the dataset PID or transfer/deposition identifier, the extraction method and the exact error message. Explain what you expected and what happened. Screenshots are useful after removing sensitive content.
What must I not include in a support request?¶
Never send passwords, API or OneDep tokens, session cookies, private keys, personal data or confidential research data. Refer to a dataset by its PID and share only the minimum diagnostic information needed.