About AquaVir-KB

Aquatic Invertebrate Virus Knowledge Base

Project Overview

AquaVir-KB is a comprehensive knowledge base dedicated to aquatic invertebrate viruses. It integrates taxonomy, host associations, literature evidence, protein functional annotations, geographic distribution, and phylogenetic analysis across eight host phyla: Arthropoda, Mollusca, Cnidaria, Echinodermata, Porifera, Annelida, Platyhelminthes, and Rotifera.

Version 2.0.0 includes 3531 virus species, 17867 isolates, 29004 proteins, and 341520 evidence-linked records. The database was formerly known as CrustaVirus DB v1 (crustacean virus database) and was expanded in 2026 to cover all aquatic invertebrate viruses.

Team

AquaVir-KB is developed and maintained by a research team focused on aquatic invertebrate virology. Detailed author information will be provided upon publication of the associated manuscript.

Contact: For inquiries regarding the database, data contributions, or collaboration, please reach out via the contact form below.

Bug reports & feedback: Issues and feature requests can be submitted through the database feedback system.

Funding Acknowledgment

AquaVir-KB is a community resource built through collaborative curation.

Maintenance Plan

AquaVir-KB is committed to long-term sustainability. We guarantee:

  • 5-year commitment: The database will remain accessible at the current URL for at least 5 years following publication (through 2031).
  • Quarterly update cycle: Data is refreshed every three months to integrate new GenBank accessions, ICTV taxonomy changes, and literature mining results.
  • Versioned archiving: Each major release is deposited to Zenodo with a versioned DOI, ensuring permanent availability independent of the live website.
  • Source code availability: All curation scripts and frontend code are maintained under an open-source license.

Contact Information

General inquiries: For technical support and data contributions, please reach out via the contact form.

Technical support: the database feedback system

Hosting: Cloud infrastructure with redundant backups.

Data Standards

AquaVir-KB is committed to adhering to community-accepted data standards for virus genome reporting and environmental metadata. The database schema supports:

  • MIUViG (Minimum Information about an Uncultivated Virus Genome) — completeness, contamination, and quality scores via CheckV integration; host prediction metadata.
  • MIxS (Minimum Information about any Sequence) — environmental package fields including salinity, water temperature, pH, dissolved oxygen, sampling depth, water type, culture system, and tissue type.

Compliance statistics are available on the Compliance page.

How to Cite

If you use AquaVir-KB in your research, please cite:

AquaVir-KB Consortium (2026). AquaVir-KB: A Comprehensive Knowledge Base of Aquatic Invertebrate Viruses. Nucleic Acids Research. DOI: (Pending)

A CITATION.cff file is available in the repository root for automated citation management.

Data Standards Compliance

AquaVir-KB is working toward compliance with the following community standards:

  • MIUViG (Minimum Information About an Uncultivated Virus Genome) — genome completeness and quality metadata are being annotated.
  • MIxS (Minimum Information about any (x) Sequence) — sample metadata fields (collection location, date, host) are aligned with MIxS checklists.
  • Darwin Core — host occurrence and geographic distribution data are mapped to Darwin Core terms where applicable.

Full compliance is in progress; please refer to the Help page for current coverage.

License

AquaVir-KB data is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0).

You are free to:

  • Share — copy and redistribute the data
  • Adapt — remix, transform, and build upon the data
  • Use for any purpose, including commercial use

You must: Attribute — give appropriate credit, provide a link to the license, and indicate if changes were made.

Data Availability

All data is freely accessible at https://aquavirdb.com without registration.

Bulk downloads: FASTA sequences, standardized metadata (Excel), host-virus network data (CSV), and reviewed evidence records (Excel) are available from the Download page.

API access: Interactive documentation is available at /docs (Swagger UI) and the OpenAPI schema at /openapi.json.

Permanent archive: Database snapshots are deposited in Zenodo with versioned DOIs.

Current Release

Version: 2.0.0 (2026-06-23)

Virus species: 3531 | Isolates: 17867 | Proteins: 29004 | Evidence records: 341520

Host phyla: 8 | Countries: 64 | Complete genomes: 228

Zenodo DOI: Pending — will be assigned upon publication acceptance

Data Sources

NCBI GenBank
Sequences & Isolates
UniProt
Protein Annotations
ICTV
Virus Taxonomy
Europe PMC
Full-text Literature
InterPro
Protein Domains
GBIF / OBIS
Host Distribution
AlphaFold / ESMFold
3D Protein Structures
KEGG / STRING
Pathways & Interactions

Technology Stack

Backend: Python FastAPI + SQLite

Frontend: Jinja2 + Tailwind CSS + HTMX + ECharts

Data Pipeline: ~60 Python scripts (NCBI / UniProt / InterPro / KEGG / EPMC / SRA)

Deployment: Docker Compose + Nginx + HTTPS

Versioning: Zenodo DOI