User guide
Using DeepNEC 2.0
Prepare sequences, choose a prediction level, and interpret enzyme, pathway, EC, motif, and structure results.
Quick start
- Open Prediction and choose Protein or Nucleotide.
- Paste FASTA, upload a file, or retrieve protein accessions. Use Load example to try the supplied protein sequences.
- Select Phase 1, 2, 3, or 4. For the complete hierarchy, choose Phase 4 and Automatic.
- Complete verification, then select Start prediction. Queueing and processing time depend on server load and input size.
- Save the complete private results link. Review the phase tables and download your results.
Sequence input
Protein FASTA
Each record begins with > followed by an identifier, then its sequence on one or more lines. Use a unique first word in every header; descriptions after a space are not part of the identifier. Do not paste alignment gaps, residue numbers, or stop symbols.
The web form accepts the 20 standard amino acids and X, in upper or lower case. X residues are removed before prediction. Protein U, O, B, Z, J, gaps, and * are not accepted by the web form; review the underlying sequence rather than silently substituting uncertain residues.
Nucleotide FASTA
Select Nucleotide before submitting DNA or RNA FASTA. The accepted alphabet is A, C, G, T, U and N. TransDecoder extracts open reading frames and translates them before classification. Results refer to predicted proteins, not directly to the submitted nucleotide records; one record may yield multiple ORFs or none. Short or incomplete transcripts may not yield an ORF. The deployed backend must have TransDecoder available.
Accessions and examples
Choose UniProtKB or NCBI Protein under Accession, enter protein identifiers, and select Retrieve sequences. Review the returned FASTA before submission. Accession retrieval and Load example switch the form to protein input. Invalid, unavailable, or nucleotide-only accessions cannot be retrieved through the protein lookup.
Request limits
The deployment defaults below can be changed by the server administrator. If the live form reports a different limit, follow that message. Limits apply together, not independently.
- Sequences per request
- 10,000
- Total residues or nucleotide bases
- 50,000,000
- Residues or bases per sequence
- 5,000
- Protein accessions per lookup
- 100
- Accession length
- 64 characters
- Request body
- 64 MiB
- Active jobs per client
- 2
For large datasets, we recommend the standalone tool on your computer or cluster.
The four prediction levels
- Phase 1 — Enzyme filter
- Classifies submitted proteins as Enzyme or Non-enzyme. Only predicted enzymes continue.
- Phase 2 — Nitrogen filter
- Classifies those enzymes as Nitrogen or Non-nitrogen, meaning nitrogen-metabolism or non-nitrogen-metabolism enzymes. Only the Nitrogen group continues.
- Phase 3 — Pathway classification
- Assigns each remaining protein to one of the ten classes below.
- Phase 4 — EC assignment
- Assigns an EC-level output within that protein’s Phase 3 route, using a learned classifier or a direct mapping.
Selecting a level runs the preceding gates and stops at that level. A Phase 4 request does not guarantee a Phase 4 row for every submitted sequence. If no proteins pass a gate, later phases may have no table.
Phase 3: ten pathway classes
Combined names denote one shared class, not several independently predicted labels. They do not establish that the source organism carries a complete pathway or operates all of these processes.
View pathway definitions
| Class | Meaning |
|---|---|
| Anammox | Anaerobic ammonium oxidation. |
| Assimilatory | Nitrogen transformations associated with incorporation into cellular material. |
| Denitrification | Reduction steps associated with gaseous nitrogen products. |
| Denitrification / nitrification | A shared functional class associated with both named pathways. |
| Dissimilatory | Nitrogen reduction associated with energy metabolism rather than incorporation into biomass. |
| Dissimilatory / denitrification | A shared functional class associated with dissimilatory reduction and denitrification. |
| Dissimilatory / denitrification / nitrification | A shared functional class associated with all three named contexts. |
| Hydroxylamine reduction | Reduction of hydroxylamine. |
| Nitrification | Oxidative transformations associated with nitrification. |
| Nitrogen fixation | Reduction of molecular nitrogen to ammonia. |
Table headings abbreviate Anammox as Anam., Assimilatory as Assim., Dissimilatory as Diss., Denitrification as Den., Nitrification as Nit., Hydroxylamine reduction as HA red., and Nitrogen fixation as N fix. Slashes join combined pathway names. Downloaded files retain the original machine-readable labels, including underscores.
Phase 4: EC outputs and routing
The current hierarchy reports 21 terminal outputs spanning 26 EC annotations: 16 outputs from five learned classifiers and five direct mappings. It is not a predictor for every EC number.
View all routes and EC outputs
| Phase 3 parent | Assignment | EC outputs |
|---|---|---|
| Anammox | Learned classifier |
|
| Assimilatory | Learned classifier |
|
| Denitrification | Learned classifier |
|
| Denitrification / nitrification | Direct mapping |
|
| Dissimilatory | Learned classifier |
|
| Dissimilatory / denitrification | Direct mapping |
|
| Dissimilatory / denitrification / nitrification | Direct mapping |
|
| Hydroxylamine reduction | Direct mapping |
|
| Nitrification | Learned classifier |
|
| Nitrogen fixation | Direct mapping |
|
Consolidated outputs
1.4.1.13-14+1.4.7.1 represents EC 1.4.1.13, 1.4.1.14 and 1.4.7.1. Similarly, 1.7.1.1-3+1.7.7.2 represents EC 1.7.1.1, 1.7.1.2, 1.7.1.3 and 1.7.7.2. Each group is one model output; it does not resolve an individual member or imply that one protein performs every listed activity.
Automatic or selected pathway
Automatic uses the Phase 3 route for each protein. Selecting one pathway limits Phase 4 processing to proteins that Phase 3 assigns to that pathway. It does not bypass the earlier gates or force other proteins into the selected route. If none match, there may be no Phase 4 output.
A direct mapping assigns the single EC associated with its Phase 3 class. It is not an additional independently evaluated classifier; a deterministic assignment should not be interpreted as 100% biological certainty.
Reading and downloading results
- Sequence ID
- The protein identifier used to join rows across phases. Nucleotide submissions use translated ORF identifiers.
- Prediction / Pathway / EC
- The selected class for that phase. A missing downstream row usually means the protein did not pass an earlier gate or did not match the selected pathway.
- Class percentages
- Softmax scores displayed as percentages for the classes within that task. They are not experimentally verified confidence estimates. Scores from different phases or different Phase 4 classifiers are not interchangeable. Blue, teal, and amber highlights mark the highest, second, and third scores within each row. Tied displayed scores share a rank, with subsequent ranks skipped. Colors represent relative ranking, not calibrated confidence.
- Structure
- Predict or view secondary and 3D structure for a protein. These optional requests do not change its enzyme or EC prediction.
Choose a phase tab, then TSV, CSV, or JSON to download that table. TSV is tab-delimited; CSV opens in spreadsheet software; JSON is convenient for scripts. Exports retain the original field names and values. Keep identifiers intact when joining tables.
The number of rows normally decreases through the hierarchy. This is filtering, not necessarily missing data. Compare each protein’s earlier phase result before interpreting an absent EC assignment.
Secondary and 3D structures
Secondary structure — S4PRED
Select Predict secondary in the results table. The page opens the viewer after prediction finishes. AA is the amino-acid sequence, Pred is the residue assignment, Cart is its graphical representation, and Conf is the predictor’s per-residue confidence score from 0 (low) to 9 (high). H denotes helix, E strand, and C coil. Hover over residues for details. Download the prediction text or export the graphic as PNG or SVG.
Tertiary structure — ESMFold or SWISS-MODEL
The current implementation tries ESMFold for proteins up to 400 residues. Longer proteins use SWISS-MODEL; when configured, SWISS-MODEL also provides a fallback if ESMFold is unavailable. SWISS-MODEL uses template-based modelling and may need additional queueing and processing time. A suitable model is not guaranteed.
Select Predict 3D once. The table shows progress and opens the viewer when coordinates are ready. SWISS-MODEL checks reuse the saved project rather than starting a fresh submission. If you leave and return, use the structure button to resume checking. Completed structures are reused for the same sequence within the same prediction job, including across phase tabs.
In the 3D viewer, drag to rotate and scroll or pinch to zoom. Change the representation, color scheme, or background, highlight supported motifs, or reset the view. Download PDB coordinates for other molecular viewers; PNG and SVG export the displayed view. The SVG contains a rendered image rather than editable molecular geometry.
These are predicted structures, not experimental measurements. Coordinate availability does not establish correct folding, enzyme activity, ligand binding, or pathway membership. Model coverage and quality can vary. Browser WebGL support is required for the interactive 3D display.
Motifs and external database links
The Motifs table counts sequence-pattern matches for Rossmann Fold, NADP Basic, NAD Acidic, Ferredoxin FeS, Mo MGD Binding, Heme Binding CXXCH, and Iron Sulfur CxxC. These are heuristic pattern names, not confirmed domains or cofactor assignments. Zero means no match to that pattern; a nonzero count alone does not establish function. Motif counts are supplemental annotations, not inputs to the deployed ESM-2 classifiers.
Phase 4 links open searches for the EC label in NCBI Protein, UniProt, BRENDA, KEGG, and JGI. They help you inspect reference annotations and related records; they are not evidence that your submitted protein is already present or experimentally characterized in those databases. For a consolidated label, search its individual EC members separately if the combined query is not recognized.
Worked example
On the Prediction page, choose Phase 4 and Automatic, then Load example. The supplied multi-record FASTA starts with AEQ03576.1. Its complete sequence is shown below directly from the website’s example data.
>AEQ03576.1 FYWWSHYPINFVLPSTMIPGALIMDTVMLLTRNWMITALVGGGAFGLLFYPGNWPIFGPTHLPPVAEGVLLSLADYTGFLYVRTGTPEYVRLIEQGSLRTFGGHTTVIAAFFSAFVSMLMFCVWWYFGKLYCTAFYYVKGPRGRVTMKNDVTAYGEEGFPEGN
- Submit the supplied example and save the result link.
- Locate AEQ03576.1 in Phase 1. If it is classified as Enzyme, look for the same ID in Phase 2.
- If Phase 2 labels it Nitrogen, read its Phase 3 pathway. Then inspect its Phase 4 EC output and the relevant class scores.
- Select Predict secondary or Predict 3D in its row. Return to the table using Back to results table.
- Download each phase table you need. Use Sequence ID to follow this protein across them.
This walkthrough describes how to read a run, not a promised label or score. Use the output of the installed model version rather than treating demo sequences as independent validation evidence.
Standalone usage
For large datasets, we recommend using the standalone tool on your computer or cluster. Installation and the source archive are on the Download page.
Protein FASTA
deepnec -i proteins.fasta -od results -l Phase4 -t prot
Nucleotide FASTA
deepnec -i transcripts.fasta -od results -l Phase4 -t nucl
Nucleotide processing requires TransDecoder.LongOrfs on PATH. The translated_proteins.fasta file contains the predicted ORFs used for classification. Use Phase1, Phase2, or Phase3 in place of Phase4 to stop earlier. Run deepnec --help for the options supported by your installed version.
Troubleshooting
Invalid FASTA or unsupported character
Check that every record has a header and sequence, the correct input type is selected, and the alphabet matches the input rules above. Remove alignment formatting and use unique identifiers.
Verification is missing or expired
Allow the verification widget to load and complete it before submission. Refresh if it expires. A configuration error requires administrator assistance.
Request too large or too many jobs
Split the input into smaller batches, wait for existing jobs to finish, or use the standalone tool. All size limits apply together.
No downstream table or EC result
Review Phase 1 and Phase 2 filters and the selected Phase 4 pathway. A protein rejected earlier is not classified at later depths.
No translated proteins
Check the nucleotide sequence and transcript length. TransDecoder may find no suitable ORF; backend installation errors require administrator assistance.
Structure prediction is taking time
Keep the page open for automatic checks. A queued SWISS-MODEL project is not a failed model. After returning to the result link, select the structure control to resume checks.
ESMFold 504, SWISS-MODEL failure, or structure request error
The external service may be unavailable, rate-limited, or unable to generate a model. Check again later; a saved SWISS-MODEL project is reused. A confirmed FAILED project may require support. Classification results remain available even if structure prediction fails.
Blank 3D viewer
Check browser WebGL support or try another browser. You can download the PDB and inspect it in a local molecular viewer.
Results unavailable
Use the complete original bookmark, including the fragment after #. The link may be incomplete, expired, or belong to a different deployment. A job ID alone does not grant access.
Cluster submission or server configuration error
Contact support with the job ID, time, selected level, and exact error text. Do not repeatedly submit the same input while a service problem is unresolved.
Privacy and result retention
Anyone with the complete result link can access that job. Keep the link private; do not post it in public issue reports. Download results before the displayed availability date. The default retention period is 30 days, but administrators can configure it.
Classification runs on the configured compute backend. Requesting tertiary structure sends that protein sequence to ESMFold or SWISS-MODEL. Do not request external structure prediction for confidential sequences unless you are authorized to share them with those services. Their retention and privacy policies are separate from DeepNEC’s local job retention.
Interpretation and limitations
DeepNEC predicts sequence-based functional labels; it does not measure enzyme activity, expression, reaction direction, metabolic flux, or the completeness of an organism’s nitrogen pathways. Its coverage is limited to the documented classes. An unfamiliar or out-of-scope protein can still receive a high score.
Each later phase depends on earlier routing. Accuracy for a pathway-conditioned classifier is not the same as accuracy for the complete hierarchy. Short fragments, ambiguous residues, ORF errors, and sequences unlike the training data can affect predictions. Long proteins are subject to the embedding model’s length handling; accepting a sequence does not guarantee that every residue contributes to the classifier representation.
Interpret predictions together with sequence quality, reference annotations, homology, genomic context, and experimental evidence where available.
Citation, reproducibility, and support
When reporting an analysis, identify DeepNEC 2.0, the webserver URL, access date, selected prediction level and pathway, and input type. For standalone runs, also record the software revision and model version. Keep the downloaded result tables with your analysis.
For the appropriate publication citation or technical assistance, contact Naveen Duhan. Report software issues through the webserver issue tracker or standalone issue tracker. Include the error message and version, but exclude private result tokens, credentials, and confidential sequences.