A data governance and quality control system for viral discovery and contamination tracing

By constructing a full-process data chain for virus discovery (PVDDC), the problem of collecting and identifying contamination information during the virus discovery process was solved, the credibility and reproducibility of virus discovery data were realized, and the data processing efficiency and scientific research credibility were improved.

CN121215039BActive Publication Date: 2026-07-24INST OF PATHOGEN BIOLOGY CHINESE ACADEMY OF MEDICAL SCI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF PATHOGEN BIOLOGY CHINESE ACADEMY OF MEDICAL SCI
Filing Date
2025-08-29
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing virus databases and annotation systems fail to cover the needs for collecting and identifying contamination information related to experimental operations and non-sample sources during the virus discovery process, resulting in problems such as poor data traceability, unidentifiable contamination, and unreproducible research.

Method used

A virus discovery data chain (PVDDC) is constructed. Through structured data modeling, natural language processing, information extraction technology driven by large language models, and a human consistency review mechanism, the core information chain of the entire virus discovery process is recorded, including the source of sampling, reagents and consumables, etc., to improve the traceability and interpretability of the data, and a multi-level quality control mechanism is introduced.

Benefits of technology

It has improved the credibility of virus discovery data, ensured the reproducibility and consistency of data, supported pollution tracing and scientific research analysis, reduced labor costs, and improved data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121215039B_ABST
    Figure CN121215039B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of data management and quality control systems for virus discovery and pollution traceability, the system includes the following modules: module 1: virus discovery whole process data chain module;Module 2: experimental meta-information acquisition and standardization module;Module 3: literature assisted meta-information extraction and large language model module;Module 4: pollution identification support mechanism and data closed loop module;Module 5: the extension and application module of virus discovery whole process data chain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical fields:

[0001] This invention relates to the interdisciplinary fields of viromics data processing, bioinformatics analysis, generative artificial intelligence-assisted information extraction, and high-throughput sequencing data quality control. Background technology:

[0002] This invention is particularly applicable to the application of metagenomic next-generation sequencing (mNGS) technology in emerging pathogen screening, clinical sample analysis, environmental monitoring, and wildlife virus research. It addresses current issues in virus discovery, such as the risk of misjudgment, incomplete data chains, and missing or inconsistent metadata caused by viral nucleic acid contamination from non-sample sources. The system constructs a Panoramic Virus Discovery Data Chain (PVDDC) and introduces a large language model (such as ChatGPT-4o) to assist in information extraction. This enables automatic parsing, semantic standardization, contamination feature annotation, and manual cross-validation of key metadata in virus discovery, thereby improving the reliability and reproducibility of virus discovery results.

[0003] This system has good versatility and scalability, and can be widely used in scientific research institutions, virus resource database construction institutions, medical laboratories, bioinformatics analysis platforms and mNGS technology service companies. It has strong practical application value and commercial transformation potential in virus contamination control, data credibility assessment, quality auditing and AI annotation toolchain integration.

[0004] With the widespread application of metagenomic next-generation sequencing (mNGS) technology in the discovery of emerging pathogens, research on viral diversity, and source tracing analysis, the acquisition and sharing of viral nucleic acid sequences have experienced explosive growth. Current mainstream viral databases (such as NCBI GenBank, GISAID, and EMBL) primarily focus on the storage and classification annotation of viral nucleic acid sequences. While they provide basic metadata fields (such as sampling location, host name, and submitter), they lack any structured recording and collection mechanisms for external influencing factors during sequence generation, such as experimental reagents, consumables, and environmental conditions. Therefore, the source information of a large number of viral records in existing databases is incomplete and unclear, greatly limiting the assessment of data reliability and the tracing of contamination sources.

[0005] In fact, existing research has shown that nucleic acid extraction reagents, sampling consumables, and even the experimental environment widely used in high-throughput sequencing may inadvertently introduce exogenous nucleic acids (such as parvovirus and porcine circovirus) with highly viral sample characteristics. However, since there are currently no databases or information standards specifically designed to identify and label viral sequence contamination information from "non-sample sources," these potential contamination signals are often mistaken for viruses that actually infect the host, leading to systemic biases in cross-species viral transmission, identification of emerging pathogens, and immune-related studies. Many so-called "new discoveries" in virology studies are actually likely "false positive contaminants," and their misleading consequences are becoming increasingly serious.

[0006] On the other hand, existing virus data annotation processes rely heavily on manual entry and subjective judgment, often resulting in the following problems:

[0007] (1) Severe lack of information in the discovery process: Submitters often fail to provide key information such as reagent brand, sampling tube type, experimental procedure and consumables, and the database does not require or support the inclusion of such content; (2) Unidentifiable contamination: Even if viral contamination exists, there is a lack of standardized contamination identification rules and metadata structure to support contamination tracking and source location; (3) Loose metadata structure: Information such as sample type, sampling time and space, host classification, and literature support are mostly in the form of free text, lacking semantic consistency; (4) High subjectivity and repetitive work in manual annotation: Researchers repeatedly consult, copy and compare in literature and databases, which greatly reduces data processing efficiency and results are inconsistent; (5) Lack of quality control: It is impossible to establish a reliable quality audit mechanism for the entire process of viral sequence from discovery to annotation to database entry.

[0008] Although some studies in recent years have attempted to use natural language processing methods to assist in the extraction of structured data, a universally applicable and scalable solution has not yet been formed due to the lack of a unified data chain standard and the difficulty in dealing with abstract reasoning of information such as the semantic diversity of documents and the uses of reagents.

[0009] In summary, existing virus databases and annotation systems fail to cover the needs for collecting and identifying contamination information related to experimental procedures and non-sample sources during virus discovery. This results in numerous problems in practical research, including poor data traceability, unidentifiable contamination, and the inability to reproduce research findings. Therefore, there is an urgent need for a data governance system that covers the entire virus discovery process, capable of recording key metadata across the entire chain, assisting in the identification and tracing of contamination signals, and improving the quality control capabilities and reliability of virome data.

[0010] This invention aims to solve the following core technical problems in the process of virus discovery in high-throughput metagenomic next-generation sequencing (mNGS): Incomplete virus discovery data chain and untraceable information: Existing viral nucleic acid records mostly contain only sequences and limited sample meta-information, and do not cover the entire chain of data information in the virus discovery process, such as sampling, experiment, annotation, literature association, reagents and consumables, and cannot meet the needs of tracing the source of the virus throughout the entire process. (1) Lack of identification basis and information support mechanism for the source of contamination: In viromics research, a large number of contaminated viruses may originate from reagents and consumables, environmental background or human operation, but at present there is a lack of a systematic mechanism for recording the experimental process and the use of materials. Once a contamination signal appears, it is difficult to trace its possible source and cannot form a "chain of evidence" data support. (2) Extraction of literature meta-information depends on manual labor, is inefficient and highly subjective: Viral nucleic acid sequences often lack clear corresponding literature, or their important meta-information (such as sample type, host, reagent use, etc.) is scattered in the main text of the paper, which is difficult to extract in batches, depends on manual reading and comparison, is inefficient, and consistency is difficult to guarantee. (3) Lack of a unified quality control and annotation verification mechanism: The current virus database construction process lacks a double-check, consistency check and anomaly labeling mechanism, which makes it difficult to quantify and assess the credibility of the data, thus limiting its scientific use in the identification of emerging pathogens and scientific research.

[0011] To this end, this invention proposes a data governance and quality control system for virus discovery and contamination tracing. It comprehensively utilizes structured data modeling, natural language processing, information extraction techniques driven by large language models (such as ChatGPT-4o), and a human consistency review mechanism to construct a "panoramic virus discovery data chain." The invention aims to achieve the following objectives: (1) to systematically record the core information chain of the entire virus discovery process, including sampling source, sampling time and space, host classification, reagents and consumables used and their specific uses, submitters and related literature, etc., thereby improving the traceability and interpretability of the data; (2) to construct a standardized data field system and pollution-related tracing framework, and to provide complete and structured information evidence for the identification and tracing of contaminated viruses while retaining key experimental information; (3) to accurately extract key metadata in virus research and store it in a structured manner by using a large language model to assist in the parsing of full-text literature information, thereby reducing manual costs and improving annotation efficiency and consistency; (4) to introduce a multi-level quality control mechanism, including manual double annotation, consistency scoring, abnormal review and verification process, to ensure the scientific nature, accuracy and reproducibility of virus metadata; (5) to realize data sharing and modular output capabilities, support the connection with public virus databases and scientific research analysis platforms, and provide a reliable data foundation for downstream virus evolution analysis, cross-species transmission research, laboratory quality control and other tasks. In summary, this invention constructs a data governance and quality control system that is traceable throughout the entire process, allows for the location of contaminants, and enables the verification of information, based on four dimensions: data structure, information extraction, contamination identification support, and quality control. This provides systematic technical support for improving the accuracy, compliance, and scientific credibility of the virus discovery process, and has significant scientific and practical application value. Summary of the Invention:

[0012] This invention provides a data governance and quality control system for virus discovery and contamination tracing, the system comprising the following modules:

[0013] Module 1: Data chain module for the entire virus discovery process;

[0014] Module 2: Experimental Meta-information Acquisition and Standardization Module;

[0015] Module 3: Document-Assisted Meta-Information Extraction and Large Language Model Module;

[0016] Module 4: Pollution Identification Support Mechanism and Data Closed-Loop Module;

[0017] Module 5: Extension and application module of the virus discovery data chain;

[0018] Module 1 includes the following sub-modules.

[0019] Module 1.1: Sequence and Classification Information Module;

[0020] Module 1.2: Sampling and Host Information Module;

[0021] Module 1.3: Information Module for the Use of Experimental Reagents and Consumables;

[0022] Module 1.4: Literature Relevance and Evidence Sources;

[0023] Module 1.5: Pollution Source Identification Module;

[0024] Module 2 includes the following sub-modules.

[0025] Module 2.1: Virus sequence record retrieval and acquisition module;

[0026] Module 2.2: Field Extraction and Structure Mapping Module;

[0027] Module 2.3: Standardized Rule System and Data Consistency Assurance Module;

[0028] Module 3 includes the following sub-modules.

[0029] Module 3.1: Literature Sources and Association Methods Module;

[0030] Module 3.2: Large Language Model Hint Design and Parsing Process Module;

[0031] Module 3.3: Structured Fields and Data Interface Mapping Module;

[0032] Module 3.4: Manual Review and Consistency Assurance Module;

[0033] Module 3.5: Collaboration mechanism module with the standardization module;

[0034] Module 4 includes the following sub-modules.

[0035] Module 4.1: Pollution Source Identification and Determination Mechanism Module;

[0036] Module 4.2: Pollution Information Sharing and Data Iterative Update Module;

[0037] Module 5 includes the following sub-modules.

[0038] Module 5.1: Epidemiological Data Storage and Association Module;

[0039] Module 5.2: Accuracy and Labeling of Host Source Information;

[0040] Module 5.3: Facilitating Viral Phylogenetic and Genomic Epidemiology Research;

[0041] Module 4.1 further includes the following sub-modules:

[0042] Module 4.1.1: Module for verifying contaminated viruses and related reagents and consumables;

[0043] Module 4.1.2: Potentially Contaminating Viruses and Reagent Association Module;

[0044] Module 4.1.3: Module for recording contaminated viruses.

[0045] The modules mentioned above are connected in sequence to form a data governance and quality control system for the entire virus discovery process. Through modular integration, the system standardizes, identifies contamination, and assesses the credibility of experimentally obtained viral sequences, host information, reagent and consumable backgrounds, and literature support. This allows it to be applied to multiple scenarios, including the identification of emerging pathogens, contamination tracing, clinical sample analysis, and environmental monitoring.

[0046] The system described in this invention is an integrated data governance system for high-throughput viral sequencing data. It aims to comprehensively record and standardize various types of meta-information generated throughout the entire process of virus discovery through modular design, identify and label sources of contamination that may be introduced by experimental reagents or procedures, thereby constructing a traceable, verifiable, and sustainably expandable dataset, and improving the quality control capabilities and credibility of viromics research.

[0047] Based on a database, the system is divided into five functional modules according to the virus discovery process. Each module contains several sub-modules, corresponding to key aspects such as data collection, information standardization, contamination identification, literature support, and support for subsequent scientific research applications. Through unified data field mapping rules, a metadata standard system, and information transmission interfaces between modules, the system integrates scattered experimental information, sequencing data, and literature evidence into structured and standardized data records, forming a data closed loop of "discovery—verification—control—feedback".

[0048] In its actual operation, this system forms a data interface with raw sequencing data (such as FASTQ, contig, etc.), sample registration information, reagent batch information, and literature index results generated in the laboratory. Through automated parsing, high-quality prompt-word-driven large language model-assisted structured annotation, human-machine collaborative review mechanism, and contamination labeling strategy, it gradually accumulates virus record entries with readability and quality control capabilities, forming a highly reliable data resource.

[0049] Furthermore, the system of this invention possesses dynamic updating and open sharing capabilities, enabling continuous access to new viral sequencing data, literature evidence, and contamination source information. It also optimizes the contamination blacklist, prompt word design, and data judgment criteria through a feedback mechanism, supporting various research scenarios such as viral phylogeny, genomic epidemiology, and laboratory tracing. It exhibits good versatility, deployability, and translational potential.

[0050] The "Data Governance and Quality Control System for Virus Discovery and Contamination Tracing" provided by this invention can be deeply integrated with virus-related data obtained during laboratory experiments, supporting two key application directions: (1) Contamination identification and suspected contamination inference.

[0051] The system integrates a large-scale database of known contaminating viruses, their corresponding contamination source labels, and consumable usage backgrounds. Once a new viral sequence generated in the laboratory is imported into the system, it can be compared with existing contamination databases. Simultaneously, the system can automatically identify highly similar contaminating viruses, hosts, sample types, and experimental procedures, thereby inferring potential sources or pathways of contamination. For example, if a virus highly matches a common contaminant found in historical extraction columns from a specific brand, the system will indicate that the sequence is "suspected contamination introduced by this consumable" and suggest verifying the reagent batch or using a control group for validation. This mechanism can improve contamination early warning capabilities during initial screening of emerging viruses, post-library construction evaluation, or multi-center collaboration, avoiding research bias.

[0052] (2) Scientific analysis based on high-quality epidemiological data

[0053] The system's data chain and database provide robust support for research on viral phylogeny and distribution. Viral sequences obtained in the laboratory from clinical or field environments, after being decontaminated, can be incorporated into a unified framework based on sample metadata (sampling time, location, host species, etc.). Combined with other similar sequences and related samples already stored in the system, the system enables visualization of viral distribution and evolutionary path analysis across spatial, host lineage, and temporal dimensions. Furthermore, the system supports integration with data dimensions such as human disease tags and host biological information, serving research on cross-species viral transmission, screening of novel pathogens, and generation of scientific hypotheses.

[0054] Through the two core pathways mentioned above, this system effectively facilitates the transformation of laboratory virus data from "raw output" to "contamination verification" and then to "credible scientific research analysis," thus constructing a credible closed loop for data governance and virus discovery.

[0055] This system automatically integrates virus sequences and their metadata from public databases, combines literature and language models to complete key experimental data, identifies potential sources of contamination, and constructs a standardized, highly reliable data chain for virus discovery, contamination tracing, and data quality control.

[0056] This system relies on a public virus database to acquire virus sequences and metadata. It parses structural fields using standardized rules, and combines a large language model and literature-assisted mechanisms to complete sample information and reagent origins, establishing an experimental traceability chain. Subsequently, it identifies and labels contamination based on information such as nucleic acid extraction methods and co-occurrence relationships, ultimately constructing a high-quality virus discovery dataset for host prediction, contamination identification, and phylogenetic research, applicable to research, clinical, and public health scenarios. For example, if a suspected novel parvovirus sequence is detected in a clinical sample, researchers can use this system to find its record in the public database, automatically extracting contextual information such as sample origin and experimental reagents. The system uses a literature-assisted module and a large language model to complete missing fields. Then, through comparison and analysis with a contamination virus knowledge base, the sequence is identified as a known contaminant originating from a certain brand of nucleic acid extraction kit and automatically labeled. Finally, this information is fed back to the contamination identification closed-loop module to improve database reliability and provide reliable support for virus host inference and evolutionary analysis.

[0057] The system described in this invention has the following functions for each module: Module 1 integrates key metadata such as the background of viral nucleic acid sequence generation, sample source, experimental operation, literature support, reagents and consumables used and their uses, forming a structured data carrier with traceability, verifiability and re-examination, providing basic support for the identification and tracing of viral contamination.

[0058] Module 2's function is to use publicly available viral nucleic acid sequence records in the GenBank database as the starting information source for its data processing flow. By designing automated retrieval, structured field extraction, and standardized rule system, it constructs the initial information architecture of the Virus Discovery Data Chain (PVDDC) and lays the foundation for subsequent literature completion and contamination identification processes.

[0059] Module 3 introduces a literature parsing module based on Large Language Models (LLMs) to extract supplementary information from the full text of scientific research literature in a structured manner. It is mainly used to identify the sampling background of virus samples, the reagents and consumables used in the experimental operation and their specific uses, thereby supplementing the missing key meta-information in the GenBank data records.

[0060] Module 4 provides a contamination identification support mechanism, which aims to identify and mark non-sample source contamination in virus data through an automated contamination source tracking and tracing system. This module can locate potential virus contamination in experimental reagents and consumables, and provide support for the development of subsequent high-throughput quality control algorithms and contamination data exclusion, ensuring the scientific validity and credibility of the final virus discovery data chain.

[0061] Module 5 records genomic data of viral sequences, reagent and consumable information, as well as high-quality epidemiological data and host origin information. By combining global virological data and accurate host identification, PVDDC can provide important support for a wide range of viral phylogenetic and genomic epidemiological studies, and help in the accurate identification and research of emerging pathogens.

[0062] The system described in this invention comprises the following sub-modules: Module 1.1 records the viral nucleic acid sequence ontology and its basic annotation and phylogenetic classification information in the GenBank database, serving as the structural starting point of the virus discovery data chain. This module includes the viral genome sequence, its taxonomic unit (family, genus, species, etc.), the viral strain name, GenBank search number, sequence submission time, submitter's name, submitting institution, sequence molecular type (e.g., ssDNA, dsRNA, etc.), sequence length, and genome structure annotations (e.g., ORF / CDS annotation information). This module provides the core classification and structural foundation for subsequent annotation comparison, literature association, and contamination assessment.

[0063] Module 1.2 is used to describe the source environment and host information of the original virus sample, and is a fundamental component of the analysis of the virus's ecological background and transmission chain. This module includes the sample type (such as throat swabs, blood, tissue, etc.), the host species of the sample (standardized to NCBI Taxonomy terms), the sampling time, the country, province, and specific geographical location of the sampling time, and whether it was a natural field sample (distinguishing it from artificial experimental samples). The standardized processing of this information provides structured support for virus epidemiology, cross-species transmission research, and host consistency review.

[0064] Module 1.3 is a key innovative component that distinguishes this invention from existing virus databases. It is specifically designed to record information on reagents and consumables used in the virus discovery and sequence generation process, as well as the specific use of each consumable in the experimental procedure. This includes the brand, model, action steps, and functions of commercial tools such as nucleic acid extraction kits, library construction reagents (kits), cDNA synthesis reagents, and purification and recovery materials. The retention of this module enables the tracing of the possible source of virus contamination through a data chain when contamination occurs, and is one of the core technical supports for achieving the contamination source tracing function.

[0065] Module 1.4 is used to establish the correspondence between virus sequences and their research literature, which is an important basis for assessing the credibility of virus information and identifying original research content. The system supports extracting the title, author information, and unique identifier (PMID) of the literature associated with the virus sequence. If it has not been formally published, it can record the submitter's institution, preprint information, or technical report information to ensure that each virus record has a clear chain of literature evidence, which helps to determine contamination and review scientific research.

[0066] Module 1.5 is used to label whether the virus or its taxonomic unit has been identified as a potential non-sample source contaminant in previous studies. The contamination identification field is used to help data users quickly assess the source credibility of virus records and can be further combined with reagent information, blank control experiments, systematic retrospective verification and other methods to improve the transparency and efficiency of virus contamination identification.

[0067] Module 2.1 performs the following functions: First, it conducts a targeted or full search of the NCBI GenBank database. The search strategy is flexibly configured based on the viral taxonomic unit of interest in this invention: If the target virus is a specific taxonomic group (e.g., Parvoviridae), a search expression can be directly constructed based on its NCBI taxonomy ID; if it is not a specific taxonomic range (e.g., broadly investigating environmental contamination viruses), a combined advanced search is performed using the Entrez query system. Commonly used fields include: "Viruses [Organism] AND (DNA OR RNA) [Molecule Type] AND 50:30000 [Sequence Length]". All records are downloaded in GenBank flat file format (.gb / .gbff) to ensure that they include original annotations, submitter information, functional labels, and complete sequences. This stage does not perform preliminary exclusion of contaminated records (e.g., "plasmid", "vector", "patent", etc.), but instead uniformly incorporates them into the subsequent quality control and contamination identification process. The download process can be initiated by calling NCBI via a Python script. E-utilities tools (such as esearch+efetch), the Biopython.Entrez module, or the official ncbi-datasets CLI tool can be used to automate batch downloads; all acquired records are retained as intermediate raw datasets for subsequent parsing.

[0068] Module 2.2 uses GenBank files as the "information starting point" and employs the Biopython toolkit for batch structured parsing to extract the following fields: Basic annotation fields: including viral nucleic acid sequence, GenBank accession number, Definition description line, sequence length, molecular type, and submission time; Classification information: parsing the family, genus, and species attribution from / organism and / taxonomy; Virus strain name: preferentially read from the / strain field, and if missing, combined from / isolate or / note; Host and sample information: parsing the / host and / isolation_source fields to form candidate original species descriptions and sample source information; Reference information: including submitter's name, institution, and published or unpublished literature information; The extracted results are uniformly stored in structured JSON or tabular format and input into the standardization module for processing.

[0069] Module 2.3 has the following functions: It establishes a complete set of standardized rules for virus metadata fields: Host standardization: Constructs a host thesaurus that matches NCBI Taxonomy numbers, unifying non-standard descriptions such as "man", "children", and "woman" into "Homo sapiens"; Time format standardization: Sampling times are standardized to ISO format (YYYY-MM-DD), with missing fields placed with "00", such as "2020-00-00"; Virus strain normalization: If both / strain and / isolate fields exist, they are merged into a standardized virus strain name according to priority; Classification system alignment: The latest version of ICTV is compared with the NCBI classification system for hierarchical consistency correction; Submitter / organization cleaning: Standardizes abbreviations, corrects spelling errors, and performs bilingual standardization conversion of Chinese and English names; All standardization processes record the original values, correction rules, and processing results, generating a traceable standardization log file.

[0070] Module 3.1 is designed with a manual and rule-driven document association judgment process to strictly screen whether a document is worth proceeding to the next step of information extraction and processing before LLM parsing. The judgment criteria include, but are not limited to, the following three: 1. Whether the virus sequence is described or explicitly stated for the first time in the article: The system prioritizes judging whether the virus sequence is reported as "newly detected", "newly assembled", or "newly isolated" in the Materials and Methods, Results, or Supplementary Data of the article; if the virus only cites other people's data (such as reusing existing database data), it is not counted as a direct source article; 2. Whether the virus name or GenBank accession appears explicitly in the text or figure captions: If the article contains an explicit accession number (such as "GenBank:ON123456") or a specific strain name (such as "strain"), it is considered a direct source article. If the sequence number is exactly the same as the sequence number recorded in PVDDC (GX2022-H1), it can be confirmed as a direct association; the model will also help identify whether the accession in the figure caption is consistent with the sample label appearing in the phylogenetic tree; 3. Whether it is a duplicate publication or a non-primary document (such as a review, meta-analysis, database index): The system automatically identifies the document type. If it is a review, secondary compilation, systematic review, or other non-original research type, or if there is obvious "secondary citation" behavior, it will be marked as "cannot establish a direct data link association"; the downloaded documents will enter the automated batch processing flow, and the LLM module will perform structured semantic parsing;

[0071] Module 3.2 is designed to call a large language processing interface based on the GPT-4o model and guide the model to extract specific field information from the full text of a research article through specially designed prompts. The LLM-processed output will be returned in JSON or tabular structure and used as a source of metadata for PVDDC.

[0072] Module 3.3 automatically maps the text information extracted by the large language model to the following fields in the PVDDC data chain: sampling host, sample type, sampling time, sampling country, province, and specific geographical information; reagents and consumables used (e.g., "QIAamp Viral RNA Kit", "Trizol", "Illumina DNA Prep") and their uses (e.g., "RNA extraction", "PCR amplification", "cDNA synthesis"); the logical association between the usage steps and the corresponding experimental steps; whether the strain name and accession appearing in the literature match GenBank (for literature validity verification); if there is semantic ambiguity or incomplete structure in the model output, it will be supplemented and confirmed by the manual review module.

[0073] Module 3.4 introduces a "dual manual review mechanism," meaning that each LLM output data must be cross-checked by two bioinformaticians against the original literature. If the consistency is high (field value consistency rate ≥ 90%), it is directly written into the data chain. If there are discrepancies (such as the purpose of a reagent being unclear or the sampling time being ambiguous), the system will record the inconsistent fields, and experts will make the final decision on the value. At the same time, a consistency score log is recorded for each cross-comparison to evaluate the model's generalization ability under different virologic families and sample literature.

[0074] The function of module 3.5 is to send the fields extracted by LLM into the standardization mapping module and map them to the same namespace and data specification as the fields extracted by GenBank, so as to ensure the structural consistency and semantic uniformity of the output data.

[0075] Module 4.1 is designed to categorize contamination assessment into three distinct scenarios, each requiring precise identification and processing based on reagent and consumable information, virus correlation data, and experimental procedures.

[0076] 1) Verified contaminated viruses and related reagents and consumables: When a virus or its closely related viruses have been verified to be contaminated in a specific reagent or consumable, the system will infer from the reagent and consumable information in PVDDC, data in relevant literature, and information on other closely related viruses to further identify possible contaminated viruses and their associated reagents and consumables; this process will expand and add more data chains of contamination sources, providing data support for subsequent contamination source tracing;

[0077] 2) Potential contamination virus and reagent association: For viruses that have not been verified as contaminants, but have a strong brand, type or component association with specific reagent kits or consumables, the system will guide researchers to accurately trace their origin. In this case, the system will generate a "potential contamination risk" warning based on the correlation between the virus and the reagent, and prompt researchers to select relevant reagent kits or conduct more stringent experimental controls to ensure the quality of subsequent data.

[0078] 3) Contaminated virus records: Once a virus is identified as a source of contamination, it will not be excluded from the PVDDC data chain, but will be specially marked as a "contaminated virus" and added to the mNGS virus blacklist. These marked contaminated viruses will serve as "chains of evidence" in experiments to further trace and assist in the discovery of a wider range of contaminated virus groups, providing data support for further analysis of contamination sources. At the same time, these contaminated viruses will be included in the scope of iterative chain of evidence screening to support subsequent correlation analysis with other virus groups and reduce false positives and duplication errors.

[0079] Module 4.2 is designed to provide other researchers with clues for identifying contamination through a data sharing mechanism. It also helps them utilize verified contamination source information, reduce reproducibility issues in experiments, and promote the widespread adoption of virus contamination identification technology.

[0080] The function of module 5.1 is to use the epidemiological data stored in PVDDC, including basic information such as sampling time, location, and sample type, as well as key background information such as host species, infection background, and regional transmission; to conduct virus phylogenetic studies, track virus transmission paths, and assess virus origin tracing.

[0081] Module 5.2 is designed to leverage the accuracy and consistency of PVDDC and combine manual review with AI-assisted analysis to ensure the accurate identification and labeling of virus host information, thus avoiding research misguidance caused by non-standard or incorrect host naming.

[0082] Module 5.3 enables researchers to conduct more accurate phylogenetic analysis and genomic epidemiological studies of viruses by leveraging the high-quality data provided by PVDDC. By analyzing the genomic sequences and host origins of different viruses, scientists can infer the evolutionary history, transmission patterns, and potential cross-species transmission risks of viruses, especially in the monitoring and tracing of emerging pathogens.

[0083] The following provides a detailed description of the system of the present invention and the operation, efficacy, use, and design of each module: 1. Design of the data chain structure for the entire virus discovery process (PVDDC Schema)

[0084] This invention first constructs a standardized data structure for recording multi-source heterogeneous information involved in the entire process of viral nucleic acid sequence "discovery" to "annotation," called the Panoramic Virus Discovery Data Chain (PVDDC). This data chain aims to systematically integrate key metadata such as the background of viral nucleic acid sequence generation, sample source, experimental procedures, literature support, reagents and consumables used, and their applications, forming a structured data carrier with traceability, verifiability, and reproducibility, providing fundamental support for the identification and tracing of viral contamination.

[0085] The PVDDC data chain adopts a modular field system design, covering the following main information dimensions. Field details are shown in Table 1:

[0086] (1) Sequence and Classification Information Module: This module records the viral nucleic acid sequence ontology and its basic annotation and phylogenetic classification information in the GenBank database, serving as the structural starting point of the virus discovery data chain. This module includes the viral genome sequence, its taxonomic unit (family, genus, species, etc.), the viral strain name, GenBank search number, sequence submission time, submitter's name, submitting institution, sequence molecular type (e.g., ssDNA, dsRNA, etc.), sequence length, and genome structure annotations (e.g., ORF / CDS annotation information). This module provides the core classification and structural foundation for subsequent annotation comparisons, literature association, and contamination assessment.

[0087] (2) Sampling and Host Information Module: This module describes the source environment and host information of the original viral sample, forming a fundamental component of viral ecological background and transmission chain analysis. This module includes sample type (e.g., throat swab, blood, tissue, etc.), host species (standardized to NCBI Taxonomy terminology), sampling time, country, province, and specific geographical location of the sampling time, and whether it was a natural field sample (distinguishing it from artificial experimental samples). Standardized processing of this information provides structured support for viral epidemiology, cross-species transmission research, and host consistency review.

[0088] (3) Experimental Reagents and Consumables Information Module: This module is one of the key innovative parts of this invention, distinguishing it from existing virus databases. It is specifically used to structurally record information on the reagents and consumables used in the virus discovery and sequence generation process, as well as the specific use of each consumable in the experimental procedure. This includes the brand, model, action steps, and functions (such as "RNA extraction," "library purification," etc.) of commercial tools such as nucleic acid extraction kits, library construction reagents (kits), cDNA synthesis reagents, and purification and recovery materials. The retention of this module enables the traceability of the possible source of virus contamination through data chain when contamination occurs, and is one of the core technical supports for realizing the contamination source tracing function.

[0089] (4) Literature Association and Evidence Source Module: This module is used to establish the correspondence between virus sequences and their research literature, which is an important basis for assessing the credibility of virus information and identifying original research content. The system supports extracting the title, author information, and unique identifier (PMID) in the PubMed database of literature associated with virus sequences. If it is not formally published, it can record the submitter's affiliation, preprint information, or technical report information to ensure that each virus record has a clear chain of literature evidence, assisting in contamination judgment and scientific research review.

[0090] (5) Contamination Source Identification Module: This module is used to indicate whether the virus or its taxonomic unit has been identified as a potential non-sample source contaminant in previous studies. The contamination identification field helps data users quickly assess the reliability of the source of virus records and can be further combined with reagent information, blank control experiments, systematic retrospective verification, and other methods to improve the transparency and efficiency of virus contamination identification.

[0091] In summary, PVDDC, as the data backbone of this invention, achieves standardized and traceable recording of the entire virus discovery process through the aforementioned five structured information modules. Its technical advantages lie in unifying the virus group data structure, compensating for the shortcomings of existing databases in lacking contamination information, and providing scalable and highly versatile underlying data structure support for the governance, sharing, and contamination identification of large-scale virus data. It is an indispensable key component of this invention.

[0092] Table 1. PVDDC Data Structure and Field Information

[0093]

[0094]

[0095] 2. Experimental Meta-information Collection and Standardization Mechanism

[0096] The data processing flow of this invention takes publicly available viral nucleic acid sequence records in the GenBank database as the starting information source. By designing an automated retrieval, structured field extraction and standardized rule system, it constructs the initial information architecture of the virus discovery data chain (PVDDC) and lays the foundation for subsequent literature supplementation and contamination identification processes.

[0097] (1) Retrieval and Acquisition of Virus Sequence Records

[0098] Virus sequence records were first retrieved through targeted or full searches of the NCBI GenBank database. The search strategy was flexibly configured based on the viral taxonomic units of interest in this invention: if the target virus was a specific taxonomic group (e.g., Parvoviridae), a search query could be directly constructed based on its NCBI taxonomy ID; if it was a non-specific taxonomic range (e.g., a broad search for viruses in environmental contamination), a combined advanced search was performed using the Entrez query system. Commonly used fields included: "Viruses[Organism] AND (DNA OR RNA)[Molecule Type] AND 50:30000[SequenceLength]". All records were downloaded in GenBank flat file format (.gb / .gbff) to ensure that the original annotations, submitter information, functional labels, and complete sequences were included.

[0099] This stage does not involve the initial exclusion of contamination records (such as "plasmid", "vector", "patent", etc.), but rather integrates them into the subsequent quality control and contamination identification process.

[0100] The download process can be automated by using Python scripts to call NCBI E-utilities tools (such as esearch+efetch), the Biopython.Entrez module, or the official ncbi-datasets CLI tool. All acquired records are retained as is as an intermediate raw dataset for subsequent parsing.

[0101] (2) Field extraction and structure mapping

[0102] The obtained GenBank files contain blocks such as LOCUS, DEFINITION, FEATURES, and REFERENCE. Although they are the main archiving format for viral nucleic acid sequences, they have serious deficiencies in their metadata structure. These deficiencies are mainly manifested in the following ways: the lack of standard fields to record key information such as experimental reagents, consumables, and library construction procedures used in the virus discovery process; metadata such as host, sample type, and sampling time exists in free text form, with non-standard content and loose structure; and most records are not accompanied by formal literature or are outdated, lacking verifiable evidence chains.

[0103] Therefore, this invention uses GenBank files as the "information starting point" and employs the Biopython toolkit for batch structured parsing to extract the following fields: Basic annotation fields: including viral nucleic acid sequence, GenBank accession number, Definition description line, sequence length, molecular type, and submission time; Classification information: parsing the family, genus, and species attribution information from / organism and / taxonomy; Virus strain name: preferentially read from the / strain field, and if missing, combined from / isolate or / note; Host and sample information: parsing the / host and / isolation_source fields to form candidate original species descriptions and sample source information; Reference information: including the submitter's name, institution, and published or unpublished literature information.

[0104] The extracted results are uniformly stored in structured JSON or table format and input into the standardized module for processing.

[0105] (3) Standardized rule system and data consistency guarantee

[0106] To eliminate information redundancy and semantic discrepancies caused by differences in submitter information in the database, this invention establishes a complete set of standardization rules for virus meta-information fields: Host standardization: Constructing a host thesaurus that matches NCBI Taxonomy numbers, unifying non-standard descriptions such as "man," "children," and "woman" into "Homo sapiens"; Time format standardization: Sampling times are all standardized to ISO format (YYYY-MM-DD), with missing fields placed with "00," such as "2020-00-00"; Virus strain normalization: If both / strain and / isolate fields exist, they are merged into a standardized virus strain name according to priority; Classification system alignment: Using the latest version of ICTV and the NCBI classification system for hierarchical consistency correction; Submitter / organization cleaning: Unifying abbreviations, correcting spelling errors, and performing bilingual standardization conversion on Chinese and English names. All standardization processes record the original values, correction rules, and processing results, generating a traceable standardization log file.

[0107] 3. Document-Assisted Meta-Information Extraction and Large Language Model Module

[0108] Because the GenBank database itself lacks fields for recording reagents, experimental procedures, and data generation background, and the information provided by submitters is often incomplete, non-standard, or untraceable, relying solely on information extracted from the database cannot meet the requirements of the Virus Discovery Data Chain (PVDDC) for high-quality, verifiable metadata. To address this issue, this invention introduces a literature parsing module based on Large Language Models (LLMs) to extract supplementary information from full-text scientific literature in a structured manner. This module is primarily used to identify the sampling background of virus samples, the reagents and consumables used in the experimental procedures, and their specific applications, thereby supplementing the missing key metadata in the GenBank data records.

[0109] (1) Literature sources and association methods

[0110] In the previous module, the system extracted the literature information corresponding to each virus sequence (such as PubMed ID, title, author, and submitting institution). This module uses this information to retrieve and download the full-text PDF file of the literature from public databases (such as PubMed Central and CNKI), or, if formally published articles cannot be obtained, attempts to perform semantic searches based on combinations of author name, institution name, and virus strain name using public search engines (such as Google, Bing, and Baidu) (see the literature association process for details).

[0111] To ensure a clear and verifiable correspondence between retrieved literature and viral sequence records, this invention designs a manual + rule-driven literature association determination process to rigorously screen whether a document is "worth proceeding to the next step of information extraction and processing" before LLM parsing. The judgment criteria include, but are not limited to, the following three: 1. Whether the viral sequence is described or explicitly stated for the first time in the article: The system prioritizes determining whether the virus sequence is reported as "newly detected," "newly assembled," or "newly isolated" in the Materials and Methods, Results, or Supplementary Data sections of the literature. If the virus merely cites data from others (e.g., reusing existing database data), it is not considered a direct source; 2. Whether the virus name or GenBank accession explicitly appears in the text or figure captions: If a clear accession number (e.g., "GenBank:ON123456") or specific strain name (e.g., "strain GX2022-H1") appears in the literature and is completely consistent with the sequence number recorded in PVDDC, a direct association can be confirmed. The model also assists in identifying whether the accession in the figure caption matches the sample label appearing in the phylogenetic tree; 3. Whether it is a duplicate publication or a non-primary document (such as a review, meta-analysis, or database index): The system automatically identifies the document type. If it is a review, secondary compilation, systematic review, or other non-original research type, or if there is obvious "secondary citation" behavior, it will be marked as "cannot be directly linked to establish a data chain". The downloaded documents will enter an automated batch processing flow, where the LLM module will perform structured semantic parsing.

[0112] (2) Large Language Model Hint Design and Parsing Process

[0113] This system calls a large language processing interface built on the GPT-4o model, guiding the model to extract specific field information from the full text of a research article through specially designed prompts. The following is a standard prompt example: "Imagine you are an expert in virology and molecular biology. Based on the extracted full text of this academic article about viruses, please extract the following information in a structured format:"

[0114] –Virus host origin

[0115] –Type of biological samples collected

[0116] –Sample ing date(year / month / day if available)

[0117] –Country and location of sampling

[0118] –Reagents and consumables used in each stage of the virus discovery process (eg, nucleic acid extraction, PCR, library prep), along with their intended purpose.

[0119] Please ensure that all information is accurate, traceable to the text, and described clearly.”

[0120] The output after LLM processing will be returned in JSON or tabular structure and used as a source for PVDDC metadata completion.

[0121] (3) Mapping structured fields to data interfaces

[0122] The text information extracted by the large language model will be automatically mapped to the following fields in the PVDDC data chain: sampling host, sample type, sampling time, sampling country, province, and specific geographical information; reagents and consumables used (e.g., "QIAamp ViralRNA Kit", "Trizol", "Illumina DNA Prep") and their uses (e.g., "RNA extraction", "PCR amplification", "cDNA synthesis"); the logical relationship between the usage steps and the corresponding experimental procedures; and whether the strain name and accession appearing in the literature match GenBank (for literature validity verification). If there is semantic ambiguity or incomplete structure in the model output, it will be supplemented and confirmed by the manual review module.

[0123] (4) Manual review and consistency assurance

[0124] To ensure the accuracy and consistency of the output results of the large language model, this invention introduces a "dual manual review mechanism," whereby each LLM output data must be cross-checked by two bioinformaticians against the original literature: if the consistency is high (field value consistency rate ≥ 90%), it is directly written into the data chain; if there are deviations (such as the purpose of a reagent cannot be determined, or the sampling time is ambiguous), the system will record the inconsistent fields, and experts will make the final decision on the value; at the same time, a consistency score log is recorded for each cross-comparison, which is used to evaluate the model's generalization ability under different virologic families and sample literature.

[0125] (5) Collaboration mechanism with standardization modules

[0126] All fields extracted via LLM will be sent to the standardization mapping module and mapped to the same namespace and data specification as the fields extracted from GenBank, ensuring structural consistency and semantic uniformity of the output data. For example, "QIAamp" → standard consumable name, "blood plasma" → "plasma", "children" → "Homo sapiens", sampling date → "YYYY-MM-DD", etc.

[0127] 4. Pollution Identification Support Mechanism and Data Closed-Loop Module

[0128] To ensure the integrity of the data chain and reduce errors caused by external contamination, this invention designs a contamination identification support mechanism. This mechanism aims to identify and label non-sample source contamination in virus data through automated contamination source tracking and tracing systems. This module can locate potential virus contamination in experimental reagents and consumables, and provides support for subsequent high-throughput quality control algorithm development and contamination data exclusion, ensuring the scientific validity and reliability of the final virus discovery data chain.

[0129] (1) Pollution source identification and determination mechanism

[0130] In the construction of viral sequence data chains, contamination source identification is a crucial step. Therefore, this invention categorizes contamination assessment into three different scenarios, each based on reagent and consumable information, viral correlation data, and experimental procedures for precise determination and processing:

[0131] 1) Confirmed contamination of viruses and related reagents and consumables.

[0132] When a virus or its closely related viruses are confirmed to contaminate specific reagents and consumables, the system will infer from the reagent and consumable information in the PVDDC, data in relevant literature, and information on other closely related viruses to further identify possible contaminating viruses and their associated reagents and consumables. This process will expand and add more data links to the contamination source chain, providing data support for subsequent contamination source tracing.

[0133] 2) Potentially contaminated viruses are associated with reagents.

[0134] For viruses that are not verified as contaminants, but have a strong brand, type, or component association with specific reagent kits or consumables, the system will guide researchers to accurately trace their origin. In this case, the system will generate a "potential contamination risk" warning based on the correlation between the virus and the reagent, and prompt researchers to select relevant reagent kits or implement more stringent experimental controls to ensure the quality of subsequent data.

[0135] 3) Records identified as contaminated with viruses.

[0136] Once a virus is identified as a source of contamination, it is not excluded from the PVDDC data chain. Instead, it is specially marked as a "contaminated virus" and added to the mNGS virus blacklist. These marked contaminated viruses will serve as "chains of evidence" in experiments for further investigation, assisting in the discovery of a wider range of contaminated virus groups and providing data support for further analysis of the contamination source. Simultaneously, these contaminated viruses will be included in iterative chain of evidence screening, supporting subsequent correlation analysis with other virus groups and reducing false positives and duplication errors.

[0137] (2) Pollution information sharing and data iterative updates

[0138] The identification of contamination sources is not only used for processing current experimental data, but also for building a long-term data update and iteration platform. Through the globally open platform PVDDC, contamination information is continuously shared and updated, and all contamination data is publicly available and fed back in real time within the global scientific community. This data-sharing mechanism not only provides other researchers with clues for contamination identification, but also helps them utilize verified contamination source information, reducing reproducibility issues in experiments and promoting the widespread adoption of virus contamination identification technology.

[0139] With the continuous updating and expansion of PVDDC, global viral contamination information will be systematically accumulated and shared. PVDDC will record the correlation between various reagents and consumables and viruses, constructing a contamination source information database. This platform not only provides strong support for academic research but also offers valuable resources for subsequent contamination detection and data quality control.

[0140] 5PVDDC Extensions and Applications: Supporting High-Quality Epidemiological Data and Host Origin Information

[0141] The PVDDC not only records genomic data of viral sequences and reagent and consumable information, but also aggregates high-quality epidemiological data and host origin information. By combining global virological data and precise host identification, the PVDDC can provide important support for a wide range of viral phylogenetic and genomic epidemiological studies, and help in the accurate identification and research of emerging pathogens.

[0142] (1) Epidemiological Data Storage and Correlation: The epidemiological data stored in PVDDC includes not only basic information such as sampling time, location, and sample type, but also key background information such as host species, infection background, and geographical transmission. This data is the foundation for conducting viral phylogenetic studies, tracing viral transmission routes, and assessing viral origin tracing. High-quality recording and standardized processing of epidemiological information helps to conduct cross-regional and cross-species comparative analyses of different research results, promoting global virus surveillance and control research.

[0143] (2) Accuracy and Labeling of Host Origin Information: PVDDC places particular emphasis on accuracy and consistency in labeling viral host origins. By combining manual verification with AI-assisted analysis, this invention ensures the accurate identification and labeling of viral host information, avoiding research misguidance caused by non-standard or incorrect host naming. In PVDDC, each viral sequence is associated with precise host information, and standardized host naming (such as Homo sapiens, Canis lupus, etc.) ensures data consistency and reusability.

[0144] (3) Facilitating Viral Phylogenetic and Genomic Epidemiological Studies: With the high-quality data provided by PVDDC, researchers can conduct more accurate phylogenetic analyses and genomic epidemiological studies of viruses. By analyzing the genome sequences and host origins of different viruses, scientists can infer the evolutionary history, transmission patterns, and potential cross-species transmission risks of viruses. Especially in the monitoring and tracing of emerging pathogens, the data provided by PVDDC can serve as an important foundation for virus tracing, pathogen transmission analysis, and viral vaccine development.

[0145] The beneficial effects of the present invention: The data governance and quality control system (PVDDC) based on the whole process of virus discovery tracking and contamination identification proposed in this invention has significant advantages over the existing technology and can solve the current technical problems in virus sequence data management, contamination source tracing, and data quality control. Specifically, it has the following beneficial effects: (1) Improves the data quality and reliability in the virus discovery process.

[0146] Existing virus discovery technologies, when using high-throughput sequencing (mNGS) for viromics analysis, often overlook the issue of contamination of experimental reagents and consumables, leading to false positives in viral sequence results, especially for low-abundance viruses or viruses suspected of cross-species transmission. By introducing PVDDC, this invention can comprehensively record the usage information of all key reagents and consumables in the virus discovery process and link it with viral sequence data, forming a traceable experimental data chain. This systematic traceability and contamination source identification mechanism ensures the reliability and high quality of the data, avoiding erroneous conclusions caused by experimental contamination.

[0147] (2) Achieving comprehensive tracing and precise location of virus contamination sources: Traditional virus contamination identification methods are usually limited to some basic screening rules (such as manually removing known contaminating sequences), lacking a systematic record of contamination sources at multiple stages of the entire experimental process. This invention provides researchers, data analysts, or reagent manufacturers with a traceable data chain by systematically storing key traceability information such as reagents, consumables, and their uses throughout the virus discovery process in PVDDC. Users can perform automated correlation calculations based on this structured information in an external analysis environment, such as comparing the reagent and consumable information of a verified contaminated virus with the metadata of other closely related viruses, thereby inferring more potential contamination sources and assisting in tracing. PVDDC does not directly perform inference calculations, but provides a complete and standardized original evidence chain and contamination labeling, ensuring that users can quickly and accurately locate contamination sources and formulate corresponding prevention and control strategies in different application scenarios.

[0148] (3) Automated quality control and error removal for high-throughput data: The automated contamination identification and data quality control mechanism of this invention can monitor, detect, and remove erroneous data caused by contamination in real time. Traditional high-throughput virome data processing often relies on manual screening and post-processing, which is not only time-consuming but also prone to overlooking subtle contamination effects. By integrating efficient quality control algorithms and contamination tracing mechanisms, PVDDC can automatically remove contaminated records during the data generation stage and ensure the accuracy of downstream analysis. Even when faced with a large number of samples and complex viral populations, PVDDC can efficiently complete contamination removal and reduce the burden of manual review.

[0149] (4) Improve the accuracy and coverage of viral sequence and literature association: Existing viral databases (such as...)

[0150] GenBank has limitations in linking viral sequence records to relevant literature, especially for uncited or unpublished papers, where effective linking mechanisms are often lacking. This invention, using a method combining literature-aided information extraction and Large Language Model (LLM), allows PVDDC to accurately extract key information about viral sequences from research literature and automatically match it with viral records, significantly improving the relevance and completeness of the link between literature and viral sequences. Furthermore, by combining public search engines with metadata retrieval, PVDDC's literature linking capabilities are significantly enhanced, covering more informal literature and multilingual materials, thus overcoming the limited coverage of traditional methods.

[0151] (5) Promoting Global Viromics Data Sharing and Collaborative Research: PVDDC provides a powerful support platform for tracing viral contamination sources, sharing viral sequence data, and facilitating scientific collaboration. Through integration with global open data platforms (such as ParvoDB), PVDDC not only achieves global data sharing but also supports other researchers in reducing errors and reproducibility issues in experiments based on validated contamination source information. With iterative data updates, researchers worldwide can utilize this shared contamination identification and tracing data to further improve the efficiency and quality of viromics research.

[0152] (6) Providing a standardized data framework for future virology and high-throughput sequencing research: The standardized data structure, contamination labeling mechanism, and literature association verification system provided by this invention lay the foundation for future virology research and the popularization of high-throughput sequencing (mNGS) technology. As the cost of mNGS continues to decrease and the technology matures, the standardized data chain and contamination identification mechanism provided by PVDDC will become indispensable tools for researchers when conducting cross-species and cross-domain virus detection and discovery, promoting the comprehensive development of virology research from data cleaning and quality control to precise tracing and cross-research data sharing.

[0153] In summary, this invention brings significant innovative breakthroughs to viromics research through its unique data tracing mechanism, literature association method, automated quality control function, and powerful data sharing platform. It not only provides an efficient and reliable mechanism for contamination identification and data cleaning in current viromics research, but also lays a foundational framework for future high-throughput virus detection and cross-disciplinary collaboration, greatly improving the accuracy, efficiency, and data quality of virus research. Attached image description:

[0154] Figure 1 This invention provides a technical roadmap and system functions for constructing PVDDC.

[0155] Figure 2 ParvoDB Data Resource Platform Homepage Overview Detailed implementation method:

[0156] The method of using the system described in this invention includes the following steps:

[0157] Obtain nucleic acid sequences and metadata related to virus discovery; perform field parsing and structure mapping through standardized rules; combine literature and large language models to help extract and complete key information; identify and label potential sources of contamination; and use the data results for reliable association between the virus and the host, the disease, contamination tracing, and quality control.

[0158] Specifically, this may include the following steps: (1) Data acquisition: Obtain viral nucleic acid sequences and their metadata from public databases, collect literature and experimental reagent usage records, and construct a data chain for the entire process of virus discovery; (2) Data standardization processing: Parse and map the viral sequence fields and their associated metadata using preset rules to achieve data field uniformity and semantic consistency; (3) Auxiliary information completion and verification: Combine literature retrieval, large language model parsing, and manual review mechanisms to complete and cross-verify metadata such as missing experimental steps, sample information, and reagent and consumable sources in the process of virus discovery; (4) Identification and labeling of contamination sources: Based on nucleic acid... Extraction methods, detection frequency, co-occurrence relationships and other information to identify possible sources of contamination in virus sequences, establish a contamination risk identification mechanism and label it; (5) Quality control and correlation inference: Based on the completed data chain, conduct correlation analysis between virus-host-disease, identify high-risk data or contamination interference, and output traceable, assessable and quality-controllable virus discovery datasets; (6) Application of results: Use the processed data in scenarios such as identification of emerging pathogens, tracing the source of virus contamination, inference of host affiliation and improvement of public database data quality, to provide decision-making basis and data support for scientific research institutions, public health departments, clinical units and reagent manufacturers.

[0159] This system can be applied to real-world scenarios involving virus contamination tracing. For example, if an unknown parvovirus sequence is detected in a clinical sample, researchers can use this system to first search for the sequence's public database information within the entire virus discovery data chain to obtain its experimental context. Then, the system uses a large language model module and a literature support module to complete the sample source and reagent information. The system automatically compares the type of extraction kit used with existing records of contaminated viruses to determine if it is a source of contamination. If a contamination risk exists, the sequence will be marked in the contamination identification module, and feedback will be generated in the data loop mechanism to drive updates to the original database. Ultimately, this information can also be used for host inference, infection route analysis, and virus evolution research.

[0160] The present invention will be further described below through specific embodiments.

[0161] Example 1: Constructing a Virus Discovery Data Chain (PVDDC) and its Data Governance and Quality Control System, using parvovirus as an example.

[0162] (I) Selection of Parvoviridae Model and Data Scale

[0163] In the prototype verification process of this invention, the Parvoviridae family (NCBITaxonomy ID: 10780) was selected as the research model. On the one hand, this family has a broad host spectrum and ecological distribution, containing representative members that infect various vertebrates and invertebrates. Furthermore, it has been repeatedly reported as a significant case of reagent contamination in viral contamination tracing studies (such as the NIH-CQV virus), thus making it representative for contamination identification and data chain construction. On the other hand, there is ample data on this family in public databases, providing the conditions for large-scale data chain construction. Based on this Taxonomy ID, this invention retrieved 34,908 initial sequence records from the NCBI GenBank database. After rigorous quality control and standardization, 30,271 high-quality records were ultimately retained as the foundation dataset for subsequent literature association, reagent and consumable information matching, and contamination labeling.

[0164] (II) Sequence Retrieval and Quality Control

[0165] 2.1 Sequence Record Retrieval and Download: In this invention, the nucleotide sequence record retrieval strategy needs to be specifically designed according to the pathogen category being studied (such as bacteria, fungi, viruses, or parasites). For pathogen groups with a clear taxonomic affiliation (such as specific viral families or genera), the optimal method is to directly limit the search criteria to their unique taxonomic identifier (Taxonomy ID, taxid) in the NCBI Taxonomy database and batch retrieve all GenBank records under that taxonomic unit. This method ensures the taxonomic specificity of the search results and facilitates the batch download of relevant GenBank flat files.

[0166] If the research subject does not belong to a clearly defined taxonomic group (e.g., newly discovered or unclassified pathogens), then it is necessary to use the NCBINucleotide advanced search interface ( https: / / www.ncbi.nlm.nih.gov / nuccore / advanced This method allows for the development of more flexible search strategies. It can combine keywords, gene names, host metadata, sequence length ranges, and publication date intervals to maximize the completeness of search results.

[0167] 2.2 Metadata Extraction and Standardization: The downloaded GenBank files were parsed using the Biopython toolkit in a Python 3 environment. Biopython is maintained by the international open-source community and covers a variety of bioinformatics computational tasks. Detailed documentation can be found here. https: / / biopython.org / Before extracting metadata, each GenBank file must be checked for integrity and structure to ensure that no information is lost during the parsing process.

[0168] The main metadata fields that can be extracted using BioPython include: virus strain-related information (such as descriptive fields and taxonomic information), spatiotemporal sampling metadata, related literature and submitter / institution information, sample type description, and host biological information. The extracted metadata is stored in a structured Excel spreadsheet for subsequent manual review and bioinformatics annotation.

[0169] Because the metadata collected in GenBank comes from submissions by researchers worldwide, it often suffers from incomplete information, outdated updates, or inconsistent structures. To ensure data quality and interoperability, this invention standardizes the metadata, including:

[0170] The classification information is mapped according to the NCBI Taxonomy system, and the ICTV (International Committee on Taxonomy of Viruses) classification is referenced and corrected when necessary; the strain and isolate fields are merged into a unified virus strain name; host information is standardized, and vague or non-standard descriptions (such as "man", "woman", "children") are unified into the standard species name "Homo sapiens"; the sampling date adopts the ISO standard format (YYYY-MM-DD), and "00" is used as a placeholder when the date is incomplete (such as "2022-00-00" to indicate only the year is known); the country and region information is standardized according to the ISO 3166 standard.

[0171] 2.3 Quality Control of Low-Quality Records: To ensure the biological relevance, accuracy, and subsequent usability of viral sequence data, this invention implements rigorous quality control on all records at the initial stage of biological cataloging to systematically identify and remove low-quality or biologically irrelevant entries. These entries include synthetic constructs, patent-related sequences, plasmid products, misclassified records, and extremely short sequences.

[0172] 2.4 Removal of Patent and Plasmid Source Records: Viral nucleotide sequences associated with patents or laboratory constructs (such as expression vectors and cloning plasmids) are systematically removed because these records do not represent natural viral diversity and may introduce misleading signals in virome analysis. Such sequences typically originate from genetic engineering, molecular cloning, or the development of commercial reagents. For example, GenBank entries BD393208 and BD393209 are vector constructs, while FB295589 and FB295590 are patent submission sequences.

[0173] During the metadata parsing process, this invention implements a multi-layered filtering strategy: scanning fields such as DEFINITION, COMMENT, REFERENCE, and FEATURES, and matching the following case-insensitive keywords:

[0174] "patent", "pat.", "proprietary", "plasmid", "vector", "construct", "recombinant", "expression cassette", "shuttle", "cosmid".

[0175] Furthermore, if the source feature in the FEATURES section contains the qualifier / plasmid=, the record is directly marked as a plasmid source and removed. Descriptions of known patent institutions or "Patent No." in the submitter information and references are also considered exclusion criteria. To maintain consistency, this invention maintains a manually created blacklist of patent and plasmid-related terms, which is updated regularly, and a case-insensitive substring matching method is uniformly applied across all relevant metadata fields.

[0176] All records excluded for the above reasons are recorded in a structured exclusion log, which includes the retrieval number, the reason for exclusion, and the triggering keywords or fields to ensure the transparency of subsequent audits and updates.

[0177] 2.5 Sequence Length-Based Filtering: To eliminate records lacking biological interpretability, this invention sets a sequence length threshold: nucleotide sequences shorter than 50 bp are considered unusable for viral genome analysis and are eliminated (e.g., MZ681471). These extremely short records typically originate from sequencing artifacts, primer dimers, or non-functional fragments, lacking value for classification and functional analysis.

[0178] 2.6 Classification Errors and Manual Verification: Since the classification of public virus databases largely relies on metadata provided by submitters, misclassification occasionally occurs, such as misclassifying non-parvoviruses into the Parvoviridae family. To address this, this invention performs genomic structure re-annotation and classification verification on all candidate parvovirus records, including verifying whether they possess typical genomic characteristics (such as the NS1 gene), genomic organization structure, and consistency with the ICTV classification system. Entries confirmed to be misclassified during the verification process (such as JN857331) are removed from the dataset.

[0179] (III) Literature Retrieval and Association

[0180] 3.1 General Overview and Technical Necessity: To systematically identify, verify, and associate peer-reviewed research papers or technical reports with viral sequence records, this invention establishes a standardized process for literature collation and association annotation to enhance the traceability and interpretability of database entries. In the NCBI GenBank database, some viral sequence records include a PubMed ID provided by the submitter, which directly points to published research papers. However, in practice, these links are often incomplete, expired, or even completely missing. In such cases, even if the viral sequence has been described in a paper, the GenBank metadata may not be updated, preventing users from directly tracing relevant literature from the database. To address this issue, this invention develops a multi-step literature retrieval and matching process, combining structured retrieval from biomedical databases with the fuzzy matching function of public search engines to maximize the comprehensiveness and accuracy of literature association.

[0181] 3.2 Metadata Extraction and Database Verification: Before linking documents, this invention prioritizes extracting key information for matching from virus sequence records, including PubMedID (if it already exists), sequence title, submitter's name, affiliated institution, and virus strain name. When a record contains a valid PubMed ID or a clear paper title, it will be directly verified in the PubMed database, and marked as "linked document" after confirmation. If there is no direct link, it is necessary to determine whether the sequence is unpublished or published but the linking information has not been updated. For the latter case, based on the extracted metadata, targeted searches will be conducted in databases such as PubMed, Google Scholar, and CNKI using various combinations of virus strain name, submitter's name, submitting institution, or part of the title.

[0182] 3.3 Assisted Use of Public Search Engines: When no matching results are found in professional literature databases, this invention further constructs keyword combinations and searches on public search engines such as Google, Bing, and Baidu. These combinations typically include formats such as "virus strain name + submitter name," "accession number + submitting institution," and "virus name + research keywords," aiming to supplement the search coverage, especially to obtain literature from non-English journals, local research reports, or literature not standardized and included in international biomedical databases. This step utilizes the more lenient fuzzy search algorithms and the indexing advantages of multilingual content in public search engines to improve the detection rate of potentially related literature.

[0183] 3.4 Manual Review and Association Criteria: All candidate literature retrieved through any channel must undergo manual review. The criteria for confirming association include: the literature must report or describe the virus sequence for the first time in the Materials and Methods or Results section; the virus strain name and / or GenBank accession number in the literature must be completely consistent with the target record (e.g., clearly appearing in the phylogenetic tree, figure captions, or data tables); and the literature must be an original research report rather than a secondary citation or derivative description. Literature with duplicate records, irrelevant studies, or only indirect mentions will be excluded. The finally confirmed associated literature will be recorded in detail in the database, including complete citation information, matching reasons, search path, and literature type, thus forming a structured, traceable, and transparent literature association system, providing solid support for the scientific application of PVDDC.

[0184] (iv) Information extraction based on large language models

[0185] 4.1 Objectives and Background: To improve the efficiency and depth of metadata extraction from relevant scientific literature, this invention introduces a Large Language Model (LLM), specifically the ChatGPT-4o model, into the data processing workflow for semi-automatic extraction of structured information from full-text academic papers. Traditional manual annotation is not only time-consuming but also easily affected by individual subjective differences. Leveraging the understanding and inductive abilities of LLM for natural language, standardization and accelerated processing can be achieved when capturing key virology-related information (such as host species, sample type, geographical background, and experimental methods), making it particularly suitable for literature with irregular text structures, non-English publications, or containing a large number of domain-specific terms.

[0186] 4.2 Full-text Acquisition and Preprocessing: Once a document has been confirmed to be associated with a specific viral sequence record using the methods described above, the full-text acquisition stage begins. Depending on access permissions, PDF files can be downloaded from institutional subscriptions, open access resources, or publisher websites. A standardized file naming system and hierarchical directory structure are established according to GenBank accession numbers or PVDDC record numbers to facilitate batch management and retrieval.

[0187] 4.3 Prompt Term Design and API Batch Processing: In the batch processing stage, this invention utilizes the ChatGPT-4o API to achieve programmatic interaction with the LLM. Standardized prompt terms are constructed for each paper, clearly defining the extraction scope and focus. For example, the model is required to act as an expert with a strong virology background, analyzing and accurately extracting from the literature text the host origin of the virus sample, the type of sample collected, the collection time, the country and specific location of collection, as well as all reagents and consumables used in the research process and their uses. Subsequently, a Python script is used to extract the PDF text, which is then submitted to the ChatGPT-4o API along with the aforementioned prompt terms. In cases where the PDF contains complex layouts or scanned images, text recognition and cleaning are performed first using an OCR tool (such as Tesseract) to ensure readability. The output returned by the model is stored in JSON or tabular format according to downstream integration requirements.

[0188] The prompt words are as follows: "Imagine you are an expert with a strong background invirology. Based on the academic article about parvovirus provided in the extracted text from the PDF, thoroughly analyze and extract the following information as described in the article: the host origin of the virus

[0189] samples,type of samples collected,sampling date,country of sampling,

[0190] sampling location, and all reagents and consumables used in theresearch process along with their purposes. Ensure that the extracted information is detailed and accurately reflects the content of the article."

[0191] 4.4 Output Integration and Manual Verification

[0192] The model output is mapped to standardized fields of PVDDC, including host, sample type, collection date, collection country, collection location, and reagents used. Each field is then checked by a bioinformatics coordinator to ensure complete consistency with the original literature descriptions, paying particular attention to resolving ambiguous expressions (such as "tissue" or generic terms like "avian host"), standardizing place names, and mapping proprietary reagent names to their corresponding functional categories. Any inconsistencies are flagged and corrected through further literature review.

[0193] 4.5 Accuracy Assessment and Quality Assurance

[0194] To verify the reliability of the metadata extracted by LLM, this invention selected 50 peer-reviewed papers as the evaluation set, and data was extracted independently by both manual and ChatGPT-4o batch processing. A total of 762 metadata fields were obtained in the evaluation, covering categories such as host origin, sample type, collection time, collection location, reagents and consumables used and their uses. The consistency between the model and manual results at the field level reached 90.0% (685 / 762). Almost all of the 76 inconsistencies stemmed from the model omitting vaguely described or incidentally mentioned highly specialized reagents or consumables, and all were corrected during the manual review stage. Notably, these differences did not involve important epidemiological or taxonomic metadata. Furthermore, this method reduced the manual workload of initial data extraction by more than 70%, significantly improving processing efficiency without compromising data quality.

[0195] (V) Verification of Epidemiological Information

[0196] 5.1 Objectives and Scope: This module aims to validate and standardize key epidemiological metadata extracted from literature using Large Language Modeling (LLM). This metadata includes core information such as host species, sample type, collection date, and geographical location. These elements are crucial for analyzing viral ecological characteristics, transmission patterns, and spatiotemporal dynamics, and their accuracy directly affects the scientific validity and effectiveness of downstream analyses.

[0197] 5.2 Review and Processing Process: Epidemiological data extracted by LLM typically includes specific species information of the virus host, sample type, collection date (potentially accurate to a specific date, month, year, or time interval), and geographic location information at different levels of granularity (such as country, province / state, region, or even precise coordinates). The processor must compare the extracted results with the full text of the original literature, focusing on the "Materials and Methods" and "Results" sections to confirm the accuracy of the information sources and descriptions. Any discrepancies, ambiguities, or omissions must be manually corrected in conjunction with the original context to ensure consistency with the original report. The corrected epidemiological fields will undergo uniform formatting across records to support comprehensive querying and geospatial analysis within the PVDDC framework.

[0198] (vi) Reagents and consumables and their uses

[0199] 6.1 Objectives and Importance: This module aims to systematically record the experimental reagents and consumables used in the virus sequencing and analysis procedures described in relevant literature, and to clarify their specific uses in the research. Complete documentation of this information not only enhances methodological transparency but also provides crucial data support for identifying potential sources of contamination, replicating metagenomic experiments, and optimizing subsequent analytical procedures.

[0200] 6.2 Verification and Supplementation: For reagent and consumable information related to LLM extraction (such as nucleic acid extraction kits, reverse transcriptase, PCR reaction reagents, etc.), the compiler needs to verify each item against the full text of the original literature, especially the "Materials and Methods" section. When the literature does not clearly state the specific purpose of a reagent or consumable, the compiler should supplement the information by consulting the manufacturer's product manual, official website, or peer-reviewed experimental method literature. For example, if the literature only mentions "RNeasy MiniKit" without specifying its function, it needs to be supplemented by stating its purpose as "for total RNA extraction from low-volume biological samples." The reagent and consumable data, supplemented manually and with clearly labeled uses, not only provide a basis for experimental protocol reproduction but also support reagent-based specific data screening and contamination signal identification in the PVDDC contamination tracing module.

[0201] (vii) Host taxonomic standardization

[0202] 7.1 Objectives and Necessity: Standardizing host names is crucial for ensuring consistency in host-virus association records and enables accurate taxonomic retrieval within the PVDDC dataset. Given the prevalence of colloquial, non-standardized, or obsolete names in literature and GenBank records, standardizing host names is essential for bioinformatics organization.

[0203] 7.2 Taxonomic Resolution and Alignment: All host names extracted by LLM and manually reviewed by the compiler undergo rigorous scrutiny in terms of scientific validity, accuracy, and consistency. Non-standard descriptions such as "man," "children," "swine," and "duckling" should be mapped to the corresponding standard species names and taxonomic identifiers (TaxIDs) based on the NCBI Taxonomy database, ensuring consistency with their complete taxonomic phylogenetic information to achieve machine-readable, standardized records. This standardization process provides a reliable data foundation for cross-species, cross-study host range analysis and zoonotic disease surveillance.

[0204] (viii) Data verification and cross-checking mechanism

[0205] This section defines a standardized verification and cross-checking process for organized metadata under the PVDDC framework to ensure the scientific accuracy, consistency, and reproducibility of all organized records. Verification must be conducted independently by at least two organized individuals with relevant domain qualifications, comprehensively evaluating the metadata, including taxonomic identifiers, host information, sampling data, literature links, and reagent annotations. The independent consistency of the organized results should reach 90%–95%. If it falls below this threshold, it must be flagged and enter a discrepancy review process. During the discrepancy review, the relevant organized individuals must collaborate and discuss based on the original data source, re-examine the evidence, and resolve disagreements through consensus. If no consensus can be reached, the dispute will be submitted to a third-party organized individual or senior domain expert for adjudication. Their decision will be based on scientific evidence and established standards and will serve as the final conclusion.

[0206] To minimize processing bias and improve data objectivity and protocol compliance, the PVDDC framework implements a multi-level cross-checking mechanism, including independent annotation, system sampling, and quantitative consistency assessment. Each data record must be independently annotated by two processors during the initial processing phase, with their results remaining invisible to each other to avoid subjective influence. After completion, the annotation results are automatically or manually compared to identify consistency and discrepancies in fields such as taxonomy, host information, sampling data, and reagent records. High consistency indicates a clear protocol and unified processing standards; discrepancies are addressed through targeted review according to the validation process.

[0207] All data organizers must strictly adhere to the standardized data organization guidelines outlined in this invention for constructing the Parvovirus PVDDC, including metadata field definitions, acceptable vocabularies, formatting specifications, and document verification procedures. To further reduce subjective differences, at least 10% of records in each batch will be randomly selected for review by senior data organizers to check their compliance with the protocol, thereby identifying potential standard deviations, systematic misunderstandings, and weaknesses in training or documentation.

[0208] After batch processing is completed, the system will quantitatively assess the consistency among the processors and calculate the consistency rate of each field. If the overall consistency rate is below 90%, the batch of data will be marked as needing reassessment, and a review process will be initiated to re-examine and coordinate corrections of discrepancies. If necessary, the field definitions or operational guidelines in the processing agreement will be optimized. Simultaneously, the cross-check results will be fed back to the processors to promote continuous improvement in annotation consistency and data quality.

[0209] (ix) Description of Parvovirus PVDDC Data Quality Improvement

[0210] In the parvovirus PVDDC (n=30,271) constructed in this example, the data quality was significantly improved compared to the original GenBank records. The literature correlation rate increased from 42.3% to 70.9%, providing a more reliable evidence base for metadata annotation and subsequent analysis; the coverage of virus sampling date information increased from 78.7% to 92.1%, the coverage of sampling country information increased from 89.2% to 98.5%, and the coverage of specific sampling locations increased from 15.8% to 56%, significantly expanding the refined distribution information of global sampling points; the coverage of sample type information increased from 49.8% to 95%, and 18,344 reagent and consumable records related to the virus discovery process were extracted from 1,545 documents; the virus classification completeness increased from 9.1% to 100%, the standardization rate of host or source information increased from 84.9% to 98.8%, and the matching rate with NCBI taxonomic records increased from 40.7% to 98.8%; at the same time, 211 new host species were added, expanding the coverage of standardized host groups to 591. Through the aforementioned quality improvement measures, this invention achieves high coverage, high consistency, and high standardization of parvovirus metadata processing, significantly enhancing the scientific credibility and reproducibility of the data in application scenarios such as virus discovery, host spectrum analysis, spatiotemporal distribution research, and tracing of experimental reagent contamination. It ensures the structured, consistent, and machine-readable nature of the virus-host relationship network, providing an operable technical path for the accurate identification of emerging pathogens and reducing the risk of misjudgment. It can be widely applied in fields such as metagenomics, virology, and pathogen monitoring.

[0211] (X) Based on parvovirus PVDDC data, the uncertainty of the parvovirus host origin is revealed.

[0212] Previous studies have demonstrated that parvoviruses can exist on the diatomaceous earth membrane of the centrifuge column in nucleic acid extraction kits, thus contaminating high-throughput sequencing results and causing misleading information in the identification and judgment of emerging pathogens. Against this backdrop, this invention systematically analyzes the nucleic acid extraction methods used in the parvovirus discovery process based on the reagent and consumable dataset in PVDDC, categorizing them into silica membrane centrifuge column method, magnetic bead method, and other methods (including phenol-chloroform extraction and TRIzol reagent method). The results show that the silica membrane centrifuge column method was the most frequently used, accounting for 15,373 records (50.78%), followed by records with no clearly defined extraction method (n=10,902, 36.01%), magnetic bead method records (1,921 records, 6.35%), and other methods (2,075 records, 6.85%). Since 2008, the growth trend of the silica membrane centrifuge column method and the undefined method is highly consistent, suggesting that a large number of records with undefined methods likely also used the centrifuge column method. Of the 188 parvoviruses identified by ICTV, 149 were first reported after 2008. The vast majority of these discoveries relied on centrifugation column methods, and 75 viruses were discovered solely using this method. To further assess the reliability of virus-host associations, the PVDDC dataset also integrated multiple types of evidence beyond molecular detection, including virus isolation and culture, animal inoculation, in situ hybridization, microscopic observation, immunological identification, and research findings satisfying Koch's postulates. Viruses possessing any one of these supplementary pieces of evidence were considered species with strong host association evidence, while viruses with only disease association information were classified as species with weaker host association support. The results showed that only 66 parvoviruses exhibited clear disease associations (e.g., Penstylhamaparvovirus decapod1 caused infectious subcutaneous and hematopoietic tissue necrosis in shrimp, and Protocolaparvovirus ungulate1 caused reproductive disorders in pigs), and only 58 viruses had strong evidence of host association (e.g., Amdoparvovirus carnivoran1 was associated with Aleutian disease in mink, supported by multiple pieces of evidence including virus isolation and culture, animal inoculation, microscopic observation, and immunological identification). However, even these viruses with strong evidence of host association generally exhibited multi-host phenomena and uncertain host spectrums. For example, Iteradensovirus lepidopteran2 was associated with hosts from 5 different classes, Penstylhamaparvovirus decapod1 involved hosts from 4 classes, and Protocolaparvovirus carnivoran1 involved hosts from 3 classes and 65 species.In summary, this case study reveals the widespread use of the silica membrane centrifugation column method in parvovirus discovery, the lack of diverse host evidence for verification, and the high degree of diffusion of some host spectrum. These factors, together with our existing experimental verification results, indicate that there is significant uncertainty in the host origin of parvovirus, and to some extent, they increase the risk of misjudgment in the identification of emerging pathogens.

[0213] (xi) ParvoDB based on PVDDC promotes the correct association between parvovirus, host, and disease.

[0214] This invention integrates the PVDDC high-quality parvovirus data chain into a publicly accessible online platform—the Parvovirus Database (ParvoDB). http: / / web3.mgc.ac.cn:8080 / parvodb / ), homepage as Figure 2 As shown, four core functional modules have been constructed: a PVDDC strain retrieval module, a human parvovirus module, a parvovirus monitoring system based on reagent and consumable traceability, and a virus-host relationship network. This platform can be used not only by researchers in virology and pathogenic biology for basic research and clinical monitoring, but also provides reagent kit manufacturers and quality control institutions with authoritative data references that can be directly applied to product development, batch testing, and contamination traceability.

[0215] In practical applications, ParvoDB supports the following workflows: First, users can quickly assess whether suspected zoonotic parvoviruses have been detected in laboratory reagents or consumables in previous studies through a parvovirus monitoring system that tracks reagent and consumable traceability. This provides the first line of screening for scientific experimental design, interpretation of clinical test results, and contamination risk assessment in reagent production processes. Second, using the PVDDC strain retrieval function, users can accurately obtain the entire chain record of strains highly related to the target virus genus or species in terms of laboratory background, sampling information, and quality control, especially identifying potential reagent contamination signals. Researchers can adjust experimental protocols accordingly, and reagent manufacturers can conduct raw material traceability and process optimization.

[0216] If no potential source of contamination is found, users can further utilize the human parvovirus module and virus-host relationship network to systematically analyze whether the relevant virus has established a known association with humans or animals at the same or higher taxonomic level (such as order), and search for published supporting evidence. For associations lacking evidence beyond molecular testing, ParvoDB recommends combining non-molecular biological methods such as electron microscopy, in situ hybridization, and immunological detection to conduct supplementary verification, ensuring the scientific validity and traceability of the host-virus-disease association.

[0217] After establishing a reliable virus-host-disease association, ParvoDB can integrate spatiotemporal sampling information, experimental methods, and reagent usage records from PVDDC, providing data support for pathogen tracing, monitoring strategy development, and antiviral measure research. Simultaneously, reagent kit manufacturers and quality control institutions can use it as a contamination risk database for background noise detection in production batches, contamination early warning in the supply chain, and continuous product quality improvement, thereby generating direct economic and social value in reducing false positives, mitigating market recall risks, and enhancing product reputation. This platform achieves seamless integration of scientific research findings and industrial applications, demonstrating significant translational potential and widespread application value.

[0218] (XII) Summary of Invention Examples

[0219] This embodiment addresses the accuracy issues of parvovirus discovery, tracing, and host association. It constructs a full-process data governance and quality control system—PVDDC—based on multi-source metadata integration, in-depth literature mining, and large language model-assisted analysis, and integrates it into the publicly accessible ParvoDB platform. Through systematic data retrieval, metadata standardization, literature completion, LLM-driven high-throughput information extraction, dual-curator cross-validation, and reagent and consumable correlation tracking, this invention achieves full-link tracing and high-precision data integration of the parvovirus discovery process, significantly improving the completeness and reliability of virus classification, host information, sampling spatiotemporal information, and experimental methods.

[0220] Compared with existing technologies, this invention introduces a strategy of multidimensional data chain construction and large language model co-curation in the field of virology for the first time. It can not only identify and label potential reagent or consumable contamination records, but also provide systematic decision support for the scientific determination of host-virus-disease associations. This technology effectively solves long-standing pain points such as uncertainty of virus origin, missing metadata, host misattribution, and insufficient identification of contamination signals, and realizes closed-loop management from data collection, quality assessment, standardization to result publication.

[0221] The innovations of this invention are mainly reflected in the following aspects: First, it deeply integrates multi-source database records with full-text literature evidence to form a high-quality virus data chain covering sampling background, experimental process and reagent traceability; second, it introduces a multimodal large language model to achieve accurate extraction and automated processing of meta-information across languages ​​and fields; and third, it establishes a scalable and shareable online platform, ParvoDB, which enables researchers, clinical testing institutions and reagent manufacturers and quality control providers to directly call and apply the results, promoting the interconnection of basic research, clinical monitoring and industrial quality control.

[0222] In summary, this invention is not only original in its technical approach and data governance model, but also provides new solutions in virological research methodology and quality control of pathogen discovery, possessing significant scientific research value and industrial transformation potential.

[0223] Key technical points of this invention: 1. Construction and standardized governance mechanism of the entire virus discovery data chain.

[0224] For the first time, a standardized record of the entire chain has been achieved, from virus sequence and classification information, sampling and host information, to experimental reagent and consumable usage information, literature association information, and contamination labeling. Through literature information extraction and semantic association analysis assisted by a large language model, the missing reagent and consumable metadata, sampling details, and experimental background information in public databases such as GenBank are automatically completed, significantly improving data integrity and traceability.

[0225] 2. Pollution labeling and source tracing system based on traceable evidence chains

[0226] This system innovatively introduces structured fields for reagents, consumables, and their uses into the data chain and uniquely binds them to the virus sequence, forming a traceable chain of evidence for the source of contamination. Combined with the cross-document association and reasoning capabilities of LLM, it achieves automated association and labeling of verified or highly suspected contaminating viruses with multi-dimensional information such as reagents, consumables, brands, and ingredients, supporting researchers, testing institutions, and reagent manufacturers in conducting accurate contamination tracing and risk assessment.

[0227] 3. Multi-source information fusion and scalable data governance platform

[0228] Establish a high-quality data governance system for research institutions, public health departments, testing laboratories, and reagent manufacturers, providing standardized data interfaces and contamination information labeling services. The data structure supports integration with external systems to achieve contamination source tracing, blacklist virus management, and subsequent iterative updates. It also has the capability to expand to more virus families and unknown viruses, providing long-term, evolvable technical support for virus identification and quality control in mNGS.

Claims

1. A data governance and quality control system for virus discovery and contamination tracing, characterized in that, The system includes the following modules: Module 1: Data chain module for the entire virus discovery process; Module 2: Experimental Meta-information Acquisition and Standardization Module; Module 3: Document-Assisted Meta-Information Extraction and Large Language Model Module; Module 4: Pollution Identification Support Mechanism and Data Closed-Loop Module; Module 5: Extension and application module of the virus discovery data chain; in Module 1 includes the following sub-modules Module 1.1: Sequence and Classification Information Module; Module 1.2: Sampling and Host Information Module; Module 1.3: Information Module for the Use of Experimental Reagents and Consumables; Module 1.4: Literature Relevance and Evidence Sources; Module 1.5: Pollution Source Identification Module; in Module 2 includes the following sub-modules Module 2.1: Virus sequence record retrieval and acquisition module; Module 2.2: Field Extraction and Structure Mapping Module; Module 2.3: Standardized Rule System and Data Consistency Assurance Module; in Module 3 includes the following sub-modules Module 3.1: Literature Sources and Association Methods Module; Module 3.2: Large Language Model Hint Design and Parsing Process Module; Module 3.3: Structured Fields and Data Interface Mapping Module; Module 3.4: Manual Review and Consistency Assurance Module; Module 3.5: Collaboration mechanism module with the standardization module; in Module 4 includes the following sub-modules Module 4.1: Pollution Source Identification and Determination Mechanism Module; Module 4.2: Pollution Information Sharing and Data Iterative Update Module; in Module 5 includes the following sub-modules Module 5.1: Epidemiological Data Storage and Association Module; Module 5.2: Accuracy and Labeling of Host Source Information; Module 5.3: Facilitating Viral Phylogenetic and Genomic Epidemiology Research; in, Module 4.1 further includes the following sub-modules: Module 4.1.1: Module for verifying contaminated viruses and related reagents and consumables; Module 4.1.2: Potentially Contaminating Viruses and Reagent Association Module; Module 4.1.3: Module for recording contaminated viruses The modules are connected sequentially through data interfaces and functional logic to form a computer-based data management system. The system runs on a local server, cloud platform, or hybrid architecture and can interface and integrate with raw sequencing data, sample information, and experimental process data obtained in the laboratory. This enables standardized and traceable management of the entire virus discovery process. The system is suitable for screening emerging pathogens, analyzing clinical samples, identifying viruses in environmental samples, and studying wildlife viruses. It can effectively identify potential sources of contamination and improve the accuracy and reliability of analysis results. Module 1 is designed to integrate the background of viral nucleic acid sequence generation, sample source, experimental operation, literature support, reagents and consumables used and their uses into a structured data carrier that is traceable, verifiable and retrievable, providing basic support for the identification and tracing of viral contamination. Module 2's function is to use publicly available viral nucleic acid sequence records in the GenBank database as the starting information source for its data processing flow. By designing automated retrieval, structured field extraction, and standardized rule system, it constructs the initial information architecture of the data chain for the entire virus discovery process and lays the foundation for subsequent literature supplementation and contamination identification processes. Module 3 introduces a literature parsing module based on a large language model to extract supplementary information from the full text of scientific literature in a structured manner. It is mainly used to identify the sampling background of virus samples, the reagents and consumables used in the experimental operation and their specific uses, thereby supplementing the missing key meta-information in the GenBank data records. Module 4 provides a contamination identification support mechanism, which aims to identify and mark non-sample source contamination in virus data through an automated contamination source tracking and tracing system. This module can locate virus contamination in experimental reagents and consumables, and provide support for the development of subsequent high-throughput quality control algorithms and contamination data exclusion, ensuring the scientific validity and credibility of the final virus discovery data chain. Module 5 records genomic data of viral sequences, reagent and consumable information, as well as high-quality epidemiological data and host origin information. By combining global virological data and accurate host identification, PVDDC provides important support for a wide range of viral phylogenetic and genomic epidemiological studies, and helps in the accurate identification and research of emerging pathogens.

2. The system according to claim 1, characterized in that, in, Module 1.1 records the viral nucleic acid sequence ontology and its basic annotation and phylogenetic classification information in the GenBank database, serving as the structural starting point of the virus discovery data chain. This module includes the viral genome sequence, its family, genus, species, and strain name, as well as the GenBank search number, sequence submission time, submitter's name, submitting institution, sequence molecular type, sequence length, and genome structure annotations. This module provides the core classification and structural foundation for subsequent annotation comparisons, literature association, and contamination assessments. Module 1.2 is used to describe the source environment and host information of the original virus sample, and is a fundamental component of the analysis of the virus's ecological background and transmission chain. This module includes descriptions of sample type, host species, sampling time, country, province, specific geographical location, and whether it was a natural field sample. The standardized processing of this information provides structured support for virus epidemiology, cross-species transmission research, and host consistency review. Module 1.3 is a key innovative component that distinguishes it from existing virus databases. It is specifically designed to record information on reagents and consumables used in the virus discovery and sequence generation process, as well as the specific use of each consumable in the experimental procedure. This includes the brand, model, action steps, and functions of nucleic acid extraction kits, library construction reagents, cDNA synthesis reagents, and purification and recovery materials. The retention of this module enables the traceability of virus sources through data chains when contamination occurs, and is one of the core technical supports for achieving contamination source tracing. Module 1.4 is used to establish the correspondence between virus sequences and their research literature, which is an important basis for assessing the credibility of virus information and identifying original research content. The system supports extracting the title, author information, and unique identifier in the PubMed database of the literature associated with the virus sequence. If it has not been formally published, it records the submitter's institution, preprint information, or technical report information to ensure that each virus record has a clear chain of literature evidence, which helps in the judgment of contamination and scientific research review. Module 1.5 is used to label whether the virus or its taxonomic unit has been identified as a potential non-sample source contaminant in previous studies. The contamination identification field is used to help data users quickly assess the credibility of the source of virus records and further combine it with reagent information, blank control experiments, and systematic retrospective verification to improve the transparency and efficiency of virus contamination identification.

3. The system according to claim 1, Its features are, in, The function of module 2.1 is to first perform targeted or full search through the NCBI GenBank database. The search strategy is flexibly configured according to the virus taxonomy unit of interest: if the target virus is a specific taxonomy group, the search query is directly constructed based on its NCBI taxonomy ID. For non-specific classification ranges, the Entrez query system is used for combined advanced searches. Commonly used fields include: "Viruses[Organism] AND (DNA OR RNA)[MoleculeType] AND 50:30000[Sequence Length]". All records are downloaded in GenBank flat file format to ensure the inclusion of original annotations, submitter information, functional labels, and complete sequences. No preliminary exclusion of contaminated records is performed at this stage; instead, they are uniformly incorporated into subsequent quality control and contamination identification processes. The download process utilizes Python scripts to call the NCBIE-utilities tool, the Biopython.Entrez module, or the official ncbi-datasets CLI tool for automated batch downloading. All acquired records are retained as is as intermediate raw datasets for subsequent analysis. Module 2.2 uses GenBank files as the "information starting point" and leverages the Biopython toolkit for batch structured parsing to extract the following fields: Basic annotation fields: including viral nucleic acid sequence, GenBank accession number, Definition description line, sequence length, molecular type, and submission time; Classification information: parsing the family, genus, and species attribution from / organism and / taxonomy; Virus strain name: preferentially read from the / strain field, and if missing, combined from / isolate or / note; Host and sample information: parsing the / host and / isolation_source fields to form candidate original species descriptions and sample source information; Reference information: including submitter's name, institution, and published or unpublished literature information; The extracted results are uniformly stored in structured JSON or tabular format and input into the standardization module for processing. Module 2.3 defines a set of standardized rules for virus metadata fields: Host standardization: Constructs a host vocabulary that matches NCBI Taxonomy numbers, unifying non-standard descriptions of "man", "children", and "woman" as "Homo sapiens"; Time format standardization: Sampling times are standardized to ISO format, with missing fields placed with "00"; Virus strain normalization: If both / strain and / isolate fields exist, they are merged into a standardized virus strain name according to priority. Classification system alignment: Hierarchical consistency correction was performed by comparing the latest version of ICTV with the NCBI classification system; Submitter / organization cleansing: Abbreviations were standardized, spelling errors were corrected, and bilingual standard conversions were performed on Chinese and English names; All standardization processes recorded the original values, correction rules, and processing results, generating a backtracking standardization log file.

4. The system according to claim 1, characterized in that, in, Module 3.1's function is to design a manual + rule-driven literature association determination process to rigorously screen whether a document is "worth proceeding to the next step of information extraction and processing" before LLM parsing. The judgment criteria include, but are not limited to, the following three:

1. Whether the viral sequence is described or explicitly stated for the first time in the article: The system prioritizes determining whether the virus sequence is reported as "newly detected," "newly assembled," or "newly isolated" in the Materials and Methods, Results, or Supplementary Data of the article; if the virus only cites other people's data, it is not counted as a direct source article; 2. Whether the virus name or GenBank accession explicitly appears in the text or figure captions: If a clear accession number or specific strain name appears in the article and is completely consistent with the sequence number recorded in PVDDC, it can be confirmed as a direct association; the model will also assist in identifying the accession in the figure captions.

3. Whether it is a duplicate publication or a non-primary document: The system automatically identifies the document type. If it is a review, secondary compilation, systematic review, or if there is obvious "recitation" behavior, it will be marked as "cannot be directly linked to data"; The downloaded documents will enter the automated batch processing flow, and the LLM module will perform structured semantic parsing. Module 3.2 is designed to call a large language processing interface based on the GPT-4o model, which uses specially designed prompts to guide the model to extract specific field information from the full text of the research article. The LLM-processed output will be returned in JSON or tabular structure and used as a source of metadata for PVDDC. Module 3.3 automatically maps the text information extracted by the large language model to the following fields in the PVDDC data chain: sampling host, sample type, sampling time, sampling country, province and specific geographical information; reagents and consumables used and their uses; logical association between usage steps and corresponding experimental steps; whether the strain name and accession appearing in the literature match GenBank; if there is semantic ambiguity or incomplete structure in the model output, it will be supplemented and confirmed by the manual review module. Module 3.4 introduces a "dual manual review mechanism," meaning that each LLM output data must be cross-checked by two bioinformaticians against the original literature. If the consistency is high, it is directly written into the data chain; if there is a discrepancy, the system will record the inconsistent field, and the final value will be determined by experts. At the same time, each cross-comparison will record a consistency score log, which is used to evaluate the model's generalization ability under different virologic families and sample literature. The function of module 3.5 is to send the fields extracted by LLM into the standardization mapping module and map them to the same namespace and data specification as the fields extracted by GenBank, so as to ensure the structural consistency and semantic uniformity of the output data.

5. The system according to claim 1, Its features are, Module 4.1 is designed to categorize contamination assessment into three distinct scenarios, each requiring precise identification and processing based on reagent and consumable information, virus correlation data, and experimental procedures. 1) Confirmed contaminated viruses and related reagents and consumables When a virus or its closely related viruses have been verified to contaminate a specific reagent or consumable, the system will infer the contaminating virus and its associated reagent or consumable by using the reagent or consumable information in PVDDC, data in relevant literature, and information on other closely related viruses. This process will expand and add more data chains of contamination sources, providing data support for subsequent contamination source tracing. 2) Potentially contaminated viruses are associated with reagents. For viruses that have not been verified as contaminants, but have a strong brand, type, or component association with specific reagent kits or consumables, the system will guide researchers to accurately trace their origin. In this case, the system will generate a "potential contamination risk" warning based on the correlation between the virus and the reagent, and prompt researchers to select relevant reagent kits or conduct more stringent experimental controls to ensure the quality of subsequent data. 3) Records identified as contaminated viruses Once a virus is identified as a source of contamination, it will not be excluded from the PVDDC data chain. Instead, it will be specially marked as a "contaminated virus" and added to the mNGS virus blacklist. These marked contaminated viruses will serve as "chains of evidence" in experiments to further trace and assist in the discovery of a wider range of contaminated virus groups, providing data support for further analysis of the source of contamination. At the same time, these contaminated viruses will be included in the scope of iterative chain of evidence screening to support subsequent correlation analysis with other virus groups, reducing false positives and repetitive errors. Module 4.2 is designed to provide other researchers with clues for identifying contamination through a data sharing mechanism. It also helps them utilize verified contamination source information, reduce reproducibility issues in experiments, and promote the widespread adoption of virus contamination identification technology.

6. The system according to claim 1, characterized in that, in, The function of module 5.1 is to utilize the epidemiological data stored in PVDDC, including sampling time, location, sample type, host species, infection background, and geographical transmission, to conduct virus phylogenetic studies, track virus transmission paths, and assess virus origin tracing. Module 5.2 is designed to leverage the accuracy and consistency of PVDDC and combine manual review with AI-assisted analysis to ensure the accurate identification and labeling of virus host information, thus avoiding research misguidance caused by non-standard or incorrect host naming. Module 5.3 is designed to enable researchers to conduct more accurate phylogenetic analysis and genomic epidemiological studies of viruses by leveraging high-quality data provided by PVDDC; to infer the evolutionary history, transmission patterns, and potential cross-species transmission risks of viruses by analyzing the genomic sequences and host origins of different viruses; and to monitor and trace emerging pathogens.

7. The method of using the system according to claim 1, characterized in that, Includes the following steps: Obtain nucleic acid sequences and metadata related to virus discovery; perform field parsing and structure mapping through standardized rules; combine literature and large language models to help extract and complete key information; identify and label potential sources of contamination; and use the data results for reliable association between the virus and the host, the disease, contamination tracing, and quality control.

8. The method of use according to claim 7, Its features are, Includes the following steps: (1) Data acquisition: Obtain viral nucleic acid sequences and their meta-information from public databases, collect literature and experimental reagent usage records, and construct a data chain for the entire process of virus discovery; (2) Data standardization processing: The virus sequence field and its associated metadata are parsed and mapped in structure through preset rules to achieve data field unification and semantic consistency; (3) Supplementing and verifying auxiliary information: Combining literature retrieval, large language model analysis and manual review mechanism, the missing experimental steps, sample information and reagent and consumable sources in the process of virus discovery are supplemented and cross-verified; (4) Identification and labeling of contamination sources: Based on nucleic acid extraction methods, detection frequency, and co-occurrence relationships, identify the contamination sources of viral sequences, establish a contamination risk identification mechanism, and label them; (5) Quality control and correlation analysis: Based on the completed data chain, conduct correlation analysis between virus-host-disease, identify high-risk data or contamination interference, and output a traceable, assessable and quality-controllable virus discovery dataset; (6) Application of results: The processed data will be used for the identification of emerging pathogens, the tracing of viral contamination, the inference of host affiliation, and the improvement of the quality of public database data, providing decision-making basis and data support for research institutions, public health departments, clinical units and reagent manufacturers.

9. The method of use according to claim 7, characterized in that, The steps are as follows: An unknown parvovirus sequence was detected in a clinical sample. Researchers used this system to first search for the sequence's public database information in the entire virus discovery data chain to obtain its experimental context. Then, the large language model module and literature assistance module supplemented the sample source and reagent information. The system automatically compared the type of extraction kit used with existing records of contaminated viruses to determine whether it was a source of contamination. If there was a risk of contamination, the sequence would be marked in the contamination identification module, and feedback would be generated in the data closed-loop mechanism to drive the update of the original database. Finally, this information was also used for host inference, infection route analysis, and virus evolution research.