About AquaVir-KB
Aquatic Invertebrate Virus Knowledge Base
Project Overview
AquaVir-KB is a comprehensive knowledge base dedicated to aquatic invertebrate viruses. It integrates taxonomy, host associations, literature evidence, protein functional annotations, geographic distribution, and phylogenetic analysis across eight host phyla: Arthropoda, Mollusca, Cnidaria, Echinodermata, Porifera, Annelida, Platyhelminthes, and Rotifera.
Version 2.0.0 includes 3531 virus species, 17867 isolates, 29004 proteins, and 341520 evidence-linked records. The database was formerly known as CrustaVirus DB v1 (crustacean virus database) and was expanded in 2026 to cover all aquatic invertebrate viruses.
Team
AquaVir-KB is developed and maintained by a research team focused on aquatic invertebrate virology. Detailed author information will be provided upon publication of the associated manuscript.
Contact: For inquiries regarding the database, data contributions, or collaboration, please reach out via the contact form below.
Bug reports & feedback: Issues and feature requests can be submitted through the database feedback system.
Funding Acknowledgment
AquaVir-KB is a community resource built through collaborative curation.
Maintenance Plan
AquaVir-KB is committed to long-term sustainability. We guarantee:
- 5-year commitment: The database will remain accessible at the current URL for at least 5 years following publication (through 2031).
- Quarterly update cycle: Data is refreshed every three months to integrate new GenBank accessions, ICTV taxonomy changes, and literature mining results.
- Versioned archiving: Each major release is deposited to Zenodo with a versioned DOI, ensuring permanent availability independent of the live website.
- Source code availability: All curation scripts and frontend code are maintained under an open-source license.
Contact Information
General inquiries: For technical support and data contributions, please reach out via the contact form.
Technical support: the database feedback system
Hosting: Cloud infrastructure with redundant backups.
Data Standards
AquaVir-KB is committed to adhering to community-accepted data standards for virus genome reporting and environmental metadata. The database schema supports:
- MIUViG (Minimum Information about an Uncultivated Virus Genome) — completeness, contamination, and quality scores via CheckV integration; host prediction metadata.
- MIxS (Minimum Information about any Sequence) — environmental package fields including salinity, water temperature, pH, dissolved oxygen, sampling depth, water type, culture system, and tissue type.
Compliance statistics are available on the Compliance page.
How to Cite
If you use AquaVir-KB in your research, please cite:
AquaVir-KB Consortium (2026). AquaVir-KB: A Comprehensive Knowledge Base of Aquatic Invertebrate Viruses. Nucleic Acids Research. DOI: — (Pending)
A CITATION.cff file is available in the repository root for automated citation management.
Data Standards Compliance
AquaVir-KB is working toward compliance with the following community standards:
- MIUViG (Minimum Information About an Uncultivated Virus Genome) — genome completeness and quality metadata are being annotated.
- MIxS (Minimum Information about any (x) Sequence) — sample metadata fields (collection location, date, host) are aligned with MIxS checklists.
- Darwin Core — host occurrence and geographic distribution data are mapped to Darwin Core terms where applicable.
Full compliance is in progress; please refer to the Help page for current coverage.
License
AquaVir-KB data is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0).
You are free to:
- Share — copy and redistribute the data
- Adapt — remix, transform, and build upon the data
- Use for any purpose, including commercial use
You must: Attribute — give appropriate credit, provide a link to the license, and indicate if changes were made.
Data Availability
All data is freely accessible at https://aquavirdb.com without registration.
Bulk downloads: FASTA sequences, standardized metadata (Excel), host-virus network data (CSV), and reviewed evidence records (Excel) are available from the Download page.
API access: Interactive documentation is available at /docs (Swagger UI) and the OpenAPI schema at /openapi.json.
Permanent archive: Database snapshots are deposited in Zenodo with versioned DOIs.
Current Release
Version: 2.0.0 (2026-06-23)
Virus species: 3531 | Isolates: 17867 | Proteins: 29004 | Evidence records: 341520
Host phyla: 8 | Countries: 64 | Complete genomes: 228
Zenodo DOI: Pending — will be assigned upon publication acceptance
Data Sources
Technology Stack
Backend: Python FastAPI + SQLite
Frontend: Jinja2 + Tailwind CSS + HTMX + ECharts
Data Pipeline: ~60 Python scripts (NCBI / UniProt / InterPro / KEGG / EPMC / SRA)
Deployment: Docker Compose + Nginx + HTTPS
Versioning: Zenodo DOI