Classification annotation method for adeno-associated virus data on exogenous comprehensive database based on next-generation sequencing

Through multi-stage classification, full-genome high-depth region + t test and resistance gene database dual verification, the problems of inefficiency, high false positive rate and cumbersome detection in AAV data analysis were solved, and efficient and accurate identification of AAV data components and recombinant and resistance gene detection were achieved.

CN120220822APending Publication Date: 2025-06-27SHANGHAI WEIKE BIOTECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510277693.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has a trade-off between reference genome richness and analysis efficiency in the classification method of AAV data, resulting in a significant increase in the analysis time; at the same time, the false positive rate in recombinant detection is high, and the resistance gene detection method is cumbersome and inefficient.

Method used

The multi-stage classification method was adopted, first using the prior reference genome for the first stage classification, and then using the exogenous comprehensive database for the second stage classification; the recombinant regions were screened through the whole genome search for high-depth regions + t test; a resistance gene database was established to conduct dual tests of DNA and protein.

Benefits of technology

On the premise of ensuring the correctness of classification results, the analysis efficiency is significantly improved, the false positive rate of recombinant detection is reduced, and the sensitivity and accuracy of resistance gene detection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220822A_ABST
    Figure CN120220822A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of bioinformatics, and provides a next-generation sequencing-based adeno-associated virus data classification annotation method, which comprises the steps of multi-stage classification, recombinant examination and resistance gene dual verification. Through multi-stage comparison of a priori reference genome and an exogenous comprehensive database, efficient and accurate qualitative and quantitative component analysis is realized. Whole genome depth distribution is combined with t detection, so that the false positive rate of recombinant detection is effectively reduced. Meanwhile, a resistance gene database of DNA and protein levels is constructed, and the accuracy and efficiency of resistance gene detection are improved through dual verification. The method is suitable for antiviral drug target screening, drug resistance research and drug effect evaluation in drug research and development, strong technical support is provided for the field of gene therapy, and the efficiency and accuracy of AAV data analysis are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of second-generation sequencing data analysis in bioinformatics, and specifically discloses a method for classifying and annotating adeno-associated virus data based on second-generation sequencing on an exogenous comprehensive database, which is applied to the classification of the main components of recombinant adeno-associated virus (rAAV) vector samples, the secondary classification of exogenous contaminant components, and the annotation of cancer-related gene regions. Background Art

[0002] Recombinant adeno-associated virus (rAAV) is a small virus that is widely used in the field of gene therapy. Due to its simple structure, excellent safety, and low immunogenicity, rAAV is regarded as an important vector in gene therapy. Adeno-associated virus (AAV) belongs to the genus Dependovirus of the Parvoviridae family and has the characteristics of being non-enveloped and single-stranded DNA defective. It can infect human hosts without causing diseases, so it shows great potential in the development of gene therapy vectors. Identifying the effective components of AAV data and mining key gene-related data are important research topics in this field currently.

[0003] In the application of AAV as a gene vector, it is often necessary to modify its genome to insert the target gene. The production process of rAAV vectors involves multiple complex steps, such as the cloning of viral genes, the transfection of host cells, and the collection and purification of viral particles. Therefore, during the preparation of rAAV, the viral vector may be affected by potential contaminants, resulting in the mixing of impurity components, which brings challenges to subsequent quality control and analysis. When identifying and classifying the components of AAV samples, it is usually necessary to rely on a large exogenous database to identify and distinguish various pollution sources. However, with the continuous expansion of the exogenous database, the increase in data volume often leads to a significant extension of the analysis time, and there is an obvious positive correlation between the analysis efficiency and the database complexity. Therefore, how to improve the analysis efficiency on the premise of ensuring the accuracy of the component identification results is one of the problems that need to be solved urgently in the current AAV analysis field.

[0004] In gene therapy, it is crucial to detect and monitor the presence of recombinants when using vectors such as recombinant adeno-associated virus (rAAV). As a virus that widely exists in nature, AAV has been found in humans and other primates. Due to the large number of homologous sequences between AAV and the human genome, AAV may recombine with the human genome to form recombinants when analyzing human-derived genomic data.

[0005] A widely used recombinant detection tool is Manta

[0006] (https: / / github.com / Illumina / manta), developed by multiple research institutions such as Pacific Biosciences and the Broad Institute. Manta uses a graph-based method for recombinant detection, and its algorithm features are as follows: Manta will build a graph model for each potential variant region and detect abnormal structures in the genome by comparing with the reference genome; Manta pays special attention to paired-end reads and split reads, which are important information sources for detecting structural variations. By analyzing paired short reads and their abnormal positions, recombinant detection is achieved. However, due to the large number of homologous sequences between AAV and the human genome, a large number of false positives are included in the Manta detection results, and it is impossible to ensure the detection of normal recombinants.

[0007] During the preparation of recombinant adeno-associated virus (rAAV), there may be a risk of recombination with resistance genes carried by pathogens in the environment, which will provide a drug-resistant environment for bacteria during the production and transportation of rAAV, and may then lead to serious clinical adverse events. Therefore, identifying resistance genes in the recombinant rAAV product genome can reduce this risk and avoid unnecessary contamination during the subsequent production and transportation of rAAV. The existing methods for detecting antibiotic resistance genes (such as patent application number: CN202410508532.X) mainly use qPCR to achieve result detection. Such methods can detect resistance genes from a single DNA level, but they cannot detect multiple resistance genes simultaneously, and are also interfered by non-specific binding of primers. They are characterized by relatively cumbersome operations and the risk of false positives, and cannot meet the requirements for time efficiency and accuracy in rAAV production.

[0008] To sum up, for the identification of effective components of AAV data and the mining of key gene-related data, there are currently the following three main problems:

[0009] 1) At present, there is a large trade-off problem between the richness of the reference genome and the analysis efficiency of the classification method for AAV data. To confirm the sources of all sequencing data as comprehensively as possible, it is usually necessary to rely on a more detailed external database. However, due to the large size of the external database, increasing the richness of the database often leads to a significant increase in the analysis time. Currently, there is a lack of a method to ensure the correctness of the results and improve the analysis efficiency when identifying the components of AAV data.

[0010] 2) Detecting the integration of host-derived genes into rAAV in rAAV products is one of the key steps in rAAV quality control. However, due to the large number of homologous sequences between AAV and the human genome, the current conventional recombinant detection algorithms only judge the recombinant sequences, which often leads to a large number of false positive results.

[0011] 3) The current resistance gene detection methods mainly rely on qPCR to achieve result detection. When the primers have non-specific binding, there is a lack of multiple verification and it is prone to the risk of false positives. Moreover, when using the qPCR method to detect multiple resistance genes simultaneously, there are problems such as cumbersome operation, long time consumption, and low analysis efficiency. Summary of the Invention

[0012] In view of the above problems actually encountered in the quality control of rAAV: 1) The identification efficiency of unknown components in the recombinant adeno-associated virus (rAAV) vector is low, 2) The false positive rate of recombinant sequence detection is relatively high, 3) The detection sensitivity of resistance genes is insufficient. The present invention respectively proposes the following solutions: including multi-stage classification of sequencing data; screening the recombinant sequences detected by the conventional algorithm by searching for high-depth regions in the whole genome and combining with t-test; in addition, establishing a resistance gene database to conduct double DNA and protein inspections on the sequencing results and detect multiple resistance genes simultaneously. Using the above methods, the above problems encountered in the quality control of rAAV in the actual scenario are solved.

[0013] The present invention includes the following specific technical solutions:

[0014] A method for classifying and annotating adeno-associated virus data based on next-generation sequencing on an exogenous comprehensive database, including one or more of the following steps 1)-3):

[0015] 1) Multi-stage classification:

[0016] Align the original sequencing data with the prior reference (prior reference genome), and classify it into known source sequences and unknown source sequences through the homologous sequence information file;

[0017] Perform a secondary alignment of the unknown source sequences with the exogenous comprehensive database to complete the qualitative and quantitative analysis of the components;

[0018] 2) Recombinant detection:

[0019] Extract the whole genome depth distribution from the host-derived sequences to identify potential recombinant regions;

[0020] Perform a t-test on the front and back control regions of each recombinant region to screen out the recombinant detection regions with a depth significantly higher than the control group;

[0021] 3) Dual verification of resistance genes:

[0022] Construct a resistance gene database at the DNA and protein levels;

[0023] Align the virus-derived sequences with the DNA database and the protein database respectively, and determine the detection of resistance genes through the intersection alignment results.

[0024] Furthermore, in the above method for classifying and annotating adeno-associated virus data based on next-generation sequencing on an exogenous comprehensive database, the original sequencing data in step 1) is a paired-end Fastq file.

[0025] Aiming at the problem that the analysis time is too long due to the large size of the exogenous database when identifying the components of AAV data, we propose a multi-stage classification method for sequencing data. This method first uses a priori reference genomes to perform the first-stage classification on the sequencing data, and then uses an exogenous comprehensive database to perform the second-stage classification on the sequences that could not be classified to a definite source in the first stage. On the premise of ensuring the correctness of the classification results, the analysis efficiency is improved. The specific scheme is as follows:

[0026] Step 1) includes the following specific steps:

[0027] S1. First-stage classification:

[0028] S11. Align the original paired-end Fastq file with the a priori reference genome to obtain multiple original BAM files;

[0029] S12. Introduce a homologous sequence information file, classify the sequences in the original BAM file by source to obtain a known-source sequence BAM file and an unknown-source sequence BAM file;

[0030] S2. Second-stage classification:

[0031] S21. For the unknown-source sequence BAM file, extract the corresponding sequences from the original paired-end Fastq file to obtain an unknown-source sequence Fastq file;

[0032] S22. Introduce an exogenous comprehensive database, align the unknown-source sequence Fastq file to obtain a second-stage classification result BAM file;

[0033] S3. Qualitative and quantitative classification results:

[0034] S31. For the second-stage classification result BAM file, count its sequence types and quantities to obtain the qualitative and quantitative results of the components of the second-stage classification.

[0035] Since the prior reference genome can cover most of the sequence sources for the current analysis, after the first-stage classification of the sequencing data using the prior reference genome, the number of position-source sequences that need to be classified in the second stage can be effectively reduced, thereby shortening the calculation time and improving the analysis efficiency. Finally, accurate and efficient qualitative and quantitative analysis of the sample components is achieved.

[0036] Secondly, to address the problem of a relatively high false positive rate in the detection results of recombinant bodies in human-derived sequences, we propose a method for detecting recombinant bodies that combines genome-wide search for high-depth regions and t-tests. First, use a conventional method for detecting recombinant bodies to capture potential recombinant regions across the entire genome and calculate the depth of each region; for each depth region obtained in the previous step, extract control regions at a certain distance before and after; then further divide the extracted regions into multiple small regions and compare them through t-tests to achieve the detection of recombinant regions. The specific scheme is as follows:

[0037] Step 2) includes the following specific steps:

[0038] S4. Extract potential recombinant data:

[0039] S41. Classify the source of the sample sequencing sequences to obtain a host BAM file, and the sequences it contains are human-derived sequences;

[0040] S42. For each chromosome data in the host BAM file, respectively count the depth of each n-base length window on each chromosome to obtain the depth distribution of each chromosome of the host genome. These regions are regarded as potential recombinant regions, where n is a natural number and the value range is: (100 ≤ n ≤ 1000);

[0041] S5. Use t-tests to compare and detect recombinant regions:

[0042] S51. For each region x in the depth distribution of each chromosome of the host genome, respectively extract two other regions at a distance of l before and after region x as control groups xL and xR, where the length of l is: n-base length ≤ l ≤ 10n-base length;

[0043] S52. For region x, control group xL, and control group xR, further divide each region into smaller regions;

[0044] S53. Conduct a t-test on the subdivided regions of region x and control group xL to confirm whether the depth of region x is significantly higher than that of control group xL;

[0045] S54. Conduct a t-test on the subdivided regions of region x and control group xR to confirm whether the depth of region x is significantly higher than that of control group xR;

[0046] S55. If the depth of region x is significantly higher than that of the control group xL and significantly higher than that of the control group xR, region x is determined to be a recombinant detection region.

[0047] At this point, the detection of the recombinant region is complete. Since AAV and the human genome have homologous sequences, the recombinant regions identified by conventional recombinant detection methods cannot effectively filter out the sequences of the human genome itself. We first detected sequences of human origin as potential recombinant data, and then identified regions with significantly high depth as detected recombinant regions, which improved the accuracy of the results and reduced false positives.

[0048] Finally, in order to solve the problem of cumbersome operation in the current qPCR-based resistance gene detection, we proposed a resistance gene database DNA + protein double test method. By pre-building a database and combining bioinformatics analysis methods, we realized the parallel analysis of large quantities of data, which greatly improved the accuracy of the analysis results in complex biological systems and the analysis efficiency. The specific plan is as follows:

[0049] The step 3) comprises the following specific steps:

[0050] S6. Construction of resistance gene database:

[0051] S61, collect resistance genes and corresponding protein reference databases at the protein level and DNA level;

[0052] S62, integrating reference data from different sources into a single DNA input file and a single protein input file to obtain a resistance gene DNA database and a resistance gene protein database, respectively;

[0053] S7. Resistance gene detection:

[0054] S71, according to the virus BAM file, backtrack and extract its corresponding virus Fastq file;

[0055] S72, using the virus Fastq file as an input file, and aligning it with the integrated resistance gene and protein reference database using the same alignment parameters at the protein level and DNA level, respectively, to obtain the original DNA alignment standard result and protein standard alignment result;

[0056] S73, removing redundancy from the original DNA alignment results, removing the sequences that were not aligned successfully, and obtaining a complete single sequence alignment result of the resistance gene;

[0057] S74, screening the original protein alignment results, and retaining only the alignment sequences whose similarity exceeds a certain threshold;

[0058] S75. Compare the detection results at the DNA level and the protein level, and extract the intersection of the detected resistance genes to obtain the detection results.

[0059] The present invention also discloses a hardware device for executing the above classification and annotation method, including:

[0060] A memory for storing the original sequencing data, the prior reference genome, the exogenous comprehensive database, the resistance gene DNA database, and the resistance gene protein database;

[0061] A processor configured to execute the method steps described in any one of claims 1 to 5;

[0062] An input interface for receiving the original sequencing data;

[0063] An output interface for outputting the classification and quality control results and generating a classification and annotation report.

[0064] The present invention also discloses a computer-readable storage medium having stored thereon a computer program, which when executed by a processor, implements the above classification and annotation method.

[0065] The present invention also discloses a server system for remotely executing the above classification and annotation method, including:

[0066] At least one central processing unit;

[0067] At least one graphics processing unit, optionally used for accelerating data processing;

[0068] At least one memory unit for storing data and programs;

[0069] A network interface for communicating with a client device, receiving sequencing data, and returning analysis results;

[0070] An operating system configured to support the collaborative work of the above hardware resources and execute the classification and annotation method described in any one of claims 1 to 5.

[0071] The present invention also discloses a cloud-based service platform for providing the service of the above classification and annotation method, including:

[0072] A cloud server cluster for storing and processing a large amount of sequencing data;

[0073] A user interface allowing a user to upload sequencing data, view the analysis progress, and obtain the final results;

[0074] A security mechanism to ensure the security of data transmission and storage;

[0075] A billing system for charging users according to the service usage.

[0076] The present invention also discloses the application of the above classification annotation method in drug research and development, and the application includes one or more of 1)-3):

[0077] 1) Screening of antiviral drug targets: screening potential drug targets through recombinant detection and resistance gene identification;

[0078] 2) Research on drug resistance: studying the drug resistance mechanism of the virus through double verification of resistance genes;

[0079] 3) Evaluation of drug effects: evaluating the impact of antiviral drugs on genomic variations; evaluating the effects and safety of antiviral drugs through virus genome analysis.

[0080] Compared with the prior art, the present invention has the following outstanding beneficial effects:

[0081] 1) This method uses a multi-stage classification method. First, the sequencing data is classified in the first stage using a specialized small genome obtained from prior knowledge. The sequences that have completed classification in the first stage will not enter the second stage of classification, thereby obtaining a smaller sequencing data of unknown origin and classifying it using an exogenous database. Through two-stage classification, this method can greatly reduce the analysis time while ensuring the correctness of classification, and improve the analysis efficiency.

[0082] 2) Sequences of human origin may be mixed into AAV to form recombinants. Currently, the conventional method for detecting recombinants generally directly detects structural variations in the sequences. However, due to the high homology between AAV data and the human genome, there are significant problems in dealing with false positives in this method. This method uses a scheme of searching for high-depth regions in the whole genome + depth subdivision for t-test comparison, which solves the problem of false positive filtering of the recombinant results detected by the conventional method.

[0083] 3) This method proposes a method of double verification of DNA + protein, which realizes efficient and accurate detection of resistance genes. Compared with the qPCR detection method, when the primers have non-specific binding, the double verification of DNA + protein reduces the risk of false positives, improves the specificity of the detection results, and the process-based analysis operation process is more convenient and efficient. It can perform parallel analysis of data from multiple samples, thereby effectively reducing the analysis time. Brief Description of the Drawings

[0084] Figure 1 Schematic diagram of the multi-stage classification method;

[0085] Figure 2 Comparison of the time used by two classification methods;

[0086] Figure 3Schematic diagram of the t-test comparison method for high-depth regions + depth subdivision

[0087] Figure 4 Box plot of the depth distribution of the xi region in a certain area and the control regions xL and xR

[0088] Figure 5 Schematic diagram of the detection method for resistance genes by DNA + protein dual detection Detailed implementation manners

[0089] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0090] The specific meanings of the terms used in the present invention are as follows.

[0091] AAV: Adeno-Associated Virus, adeno-associated virus

[0092] Resistance gene: Resistance genes refer to genes encoding the ability of microorganisms to resist antibiotics, disinfectants or other antibacterial agents.

[0093] reference: Reference genome. A reference genome is a genome that has been widely studied and annotated, representing the typical or representative genome of the species or population, and is used to compare and annotate the genomic data of other individuals.

[0094] FASTA file: A text-based file format for representing nucleotide sequences or amino acid sequences.

[0095] Exogenous comprehensive database: An integrated database of exogenous reference genomes that integrates representative collections of microorganisms and viruses, and is filtered and integrated to obtain a comprehensive database. The data is sourced from the public database NCBI Nucleotide Sequence Database.

[0096] Resistance gene DNA database: A reference database for resistance genes at the DNA level established by this method. The data mainly comes from the public databases CARD and SARG.

[0097] Resistance gene protein database: A reference database for resistance genes at the protein level established by this method. The data mainly comes from the public database NCBI Non-Redundant Protein Database.

[0098] Homologous information file: A file that stores homologous sequence information, generated by performing homology comparison and analysis on the sequences of two reference genomes. The stored content includes the positions of homologous sequences on different genomes, homology scores, etc.

[0099] Fastq file: A file format that stores biological sequences (usually nucleic acid sequences) and their sequencing quality score information.

[0100] depth: Refers to the number of sequences obtained by sequencing technology at a specific position or region. For example, if a genomic position is covered by 10 different sequences, the depth at this position is 10. Depth is an important indicator for measuring sequencing quality and accuracy, and directly affects the accuracy and reliability of variant detection, genome assembly, and subsequent analysis.

[0101] BAM file: Binary Alignment / Map, a file format that stores sequence alignment results.

[0102] VCF file: Variant Call Format, a file format used to store genomic variant site data.

[0103] Example 1

[0104] To solve the problems described in the background art, we proposed 3 corresponding solutions, namely, a multi-stage classification method for sequencing data, a method for detecting recombinants by searching for high-depth regions in the whole genome + t-test, and a method for double-checking DNA + protein in the resistance gene database. The above methods are collectively referred to as a classification and annotation method for adeno-associated virus data based on next-generation sequencing on an exogenous comprehensive database.

[0105] This Example 1 generally discloses a classification and annotation method for adeno-associated virus data based on next-generation sequencing on an exogenous comprehensive database, including one or more of the following steps 1)-3):

[0106] 1) Multi-stage classification

[0107] Align the original sequencing data with the prior reference genome, and classify it into known-source sequences and unknown-source sequences through the homologous information file;

[0108] Perform a secondary alignment of the unknown-source sequences with the exogenous comprehensive database to complete the qualitative and quantitative analysis of the components;

[0109] 2) Recombinant detection

[0110] Extract the whole-genome depth depth distribution from the host-source sequences to identify potential recombinant regions;

[0111] Perform a t-test on the pre- and post-control regions for each region, and screen out the recombinant detection regions with significantly higher depths than the control group;

[0112] 3) Dual verification of resistance genes

[0113] Construct a resistance gene database at the DNA and protein levels;

[0114] Align the virus-derived sequences with the DNA database and the protein database respectively, and determine the detection of resistance genes through the intersection alignment results.

[0115] Example 2

[0116] In Example 2, the paired-end Fastq files of AAV data are qualitatively and quantitatively analyzed for components using two methods respectively, and the time required for the two methods is compared. It is a specific refinement corresponding to step 1) in Example 1, aiming to verify that using the multi-stage classification method can improve the analysis efficiency during the alignment process with the external database, and achieve short-time and high-precision classification of unknown sequence data. Figure 1 It is a schematic diagram of the multi-stage classification method. The paired-end Fastq files of the sample will first be aligned and classified using the prior reference (prior reference genome). Sequences that cannot be classified to a clear source in the first stage will be aligned and classified using the external comprehensive database in the second stage, and sequences that already have a clear source in the first stage will not be additionally classified in the second stage;

[0117] In this example, the input data to be prepared includes:

[0118] (1) 1 example of AAV sample, after extracting the cell fluid containing AAV and sequencing it respectively, 1 pair of paired-end Fastq files is obtained;

[0119] (2) The prior reference genome in FASTA file format, including virus genome FASTA file, host genome FASTA file, and impurity genome FASTA file in this example;

[0120] (3) Homologous sequence information file: A file storing homologous sequence information, generated by performing a homology comparison analysis on the sequences of two reference genomes, and the stored content includes the positions of homologous sequences on different genomes, homology scores, etc.

[0121] (4) Comprehensive external database, an external reference genome database integrating representative collections of microorganisms and viruses.

[0122] Method 1: Perform qualitative and quantitative analysis of sample components through the multi-stage classification method, and statistically analyze the time.

[0123] The specific operations are as follows:

[0124] 1. First, perform the first-stage classification. Using alignment software such as BWA, Bowtie2, SOAP, etc., align the paired-end Fastq files to the prior reference genome in FASTA file format to obtain multiple raw BAM files.

[0125] 2. Introduce the information in the file of the positions of homologous sequences on the genome, and precisely classify the sequences in the raw BAM files so that each sequence is classified and only classified to a unique source. In this embodiment, multiple BAM files of known source sequences and 1 BAM file of unknown source sequences are obtained. Thus, each sequence appears and only appears in 1 BAM file.

[0126] 3. Then, perform the second-stage classification. For the BAM file of unknown source sequences obtained from the first-stage classification, according to the sequence information it contains, extract the corresponding sequences from the paired-end Fastq files to obtain 2 paired-end Fastq files of unknown source sequences. Compared with the original paired-end Fastq files, the file size of the paired-end Fastq files of unknown source sequences has been significantly reduced.

[0127] 4. Use the alignment tool again to align the paired-end Fastq files of unknown source sequences to the exogenous comprehensive database to obtain 1 BAM file of exogenous sequences.

[0128] 5. Statistically analyze the sequence information in the BAM files of known source sequences and BAM files of exogenous sequences obtained in the above steps to obtain the qualitative and quantitative results of the composition of the sample.

[0129] 6. In this analysis, the paired-end Fastq files contain 5,398,008 sequences. After processing with the prior reference genome, there are still 298,734 sequences of unknown source, reducing the number of sequences aligned to the exogenous comprehensive database in the second stage to 6% of the original number of sequences.

[0130] 7. Finally, statistically analyze the time used in the two-stage classification method.

[0131] Method 2: Without performing classification in multiple stages, directly align and classify all sample sequences to the exogenous comprehensive database and statistically analyze the time consumption.

[0132] 1. Use alignment tools such as BWA, Bowtie2, SOAP, etc., to align the paired-end Fastq files to the exogenous comprehensive database. The parameter information used in the alignment analysis is the same as that used in the second stage of Method 1 to obtain 1 BAM file.

[0133] 2. Statistically analyze the chromosomal source information in the BAM file to obtain the qualitative and quantitative results of the composition of the sample.

[0134] 3. In this analysis, the paired - end Fastq files contained 5,398,008 sequences, and all of them were aligned with the exogenous comprehensive database.

[0135] 4. Finally, the time taken by the single - stage alignment classification method was counted.

[0136] So far, the analysis times of the two classification methods have been obtained and plotted as a bar chart ( Figure 2 , using the same pair of original paired - end Fastq files to be classified by the single - stage classification method and the multi - stage classification method respectively, and their analysis times were counted separately. The single - stage classification method directly compared and classified the original paired - end Fastq files with the exogenous database, taking about 50 minutes; the multi - stage classification method first compared and classified the original paired - end Fastq files with the prior reference, and then compared and classified the sequences that could not be classified to a clear source in the first stage with the exogenous database in the second stage, with a total time of about 26 minutes). As can be seen from Figure 2 , for the same data, when using the multi - stage classification method, compared with the single - stage classification method, the total analysis time was reduced from 50 minutes to 26 minutes, which was 52% of the time taken by the single - stage classification method, proving that the multi - stage classification method can effectively reduce the analysis time and achieve efficient identification and analysis of AAV data components.

[0137] Example 3

[0138] Example 3 aims to screen the results of conventional tests by implementing a method of genome - wide search for high - depth regions + depth subdivision for t - test comparison on a sample. It is a specific refinement of step 2) in Example 1. By comparing the change in the number of candidate recombinant regions before and after filtering, it is verified that this method can effectively reduce the false positives of the results of conventional recombinant detection methods. Figure 3 It is a schematic diagram of the method of high - depth region + depth subdivision for t - test comparison. First, the depth of the sequenced and aligned genome is statistically analyzed, that is, the depth of each region of the whole genome is obtained. The depth of a region x is compared with the depth of the control regions at a certain distance to its left and right, and it is judged whether the depth of region x is significantly higher than that of the control group through t - test. This example will follow the Figure 3 described method for corresponding data analysis.

[0139] In this example, the input data to be prepared includes:

[0140] (1) 1 AAV sample. After extracting the cell fluid containing AAV and sequencing it, 1 pair of paired - end Fastq files was obtained.

[0141] In this example, the size of window n is 500 bases, and l = 10n = 5000 bases;

[0142] The specific analysis steps are as follows:

[0143] 1. First, classify the paired-end Fastq files through the same method as in Example 1 to obtain the host-source BAM file of the sample, and the sequences it contains are human-derived sequences.

[0144] 2. Then, for each chromosome data in the host BAM file, count the depth of each 500-base length window on each chromosome to obtain the depth distribution of each chromosome of the host genome, and these regions are regarded as potential recombinant regions.

[0145] 3. Next, detect the high-depth regions through a t-test. For each region x in the depth distribution of each chromosome of the host genome, take another 2 regions 5000 bases away from the front and back of region x as the control groups xL and xR respectively.

[0146] 4. Further subdivide regions x, xL, and xR. Divide each region into small windows with a length of 50 bases, and count the depth of each small window to obtain the subdivided data of region x, the subdivided data of region xL, and the subdivided data of region xR.

[0147] 5. Conduct a t-test on the subdivided data of region x and the subdivided data of region xL to obtain the corresponding test results tL and pL; conduct a t-test on the subdivided data of region x and the subdivided data of region xR to obtain the test results tR and pR.

[0148] 6. Judge whether the depth of region x is significantly higher than the depth of the control groups xL and xR through the p-value and t-value. The specific judgment logic is as follows:

[0149] (1) If the results simultaneously meet the conditions: pL < 0.05, pR < 0.05, tL > 0, tR > 0. Then it indicates that the depth of region x is significantly higher than that of the control group, and thus region x is the detected recombinant region;

[0150] (2) If the results do not meet the conditions in (1), then it is considered that region x is not the detected recombinant region.

[0151] 7. Use the subdivided data of region x, the subdivided data of region xL, and the subdivided data of region xR to draw a box plot ( Figure 4 ), the vertical axis of this box plot represents the depth value, the horizontal axis is the 3 groups of regions xL, x, and xR respectively, the short lines above and below each box represent the maximum and minimum values of the data respectively, the horizontal line inside the box represents the median, and the length of the box can reflect the dispersion of the data. It can be seen from the figure that the distributions of the 3 groups of data are relatively concentrated, and the box of region x is significantly higher than the boxes of the control group regions.

[0152] 8. The data recorded in Table 1 are the results obtained by using the method of t - test comparison with high depth region + depth subdivision to detect the recombinant regions in the sample. As can be seen from the table, the depth of region x is significantly higher compared with the control groups xL and xR, meeting the design expectations and proving that this method can effectively detect the recombinant regions in the host - derived sequences.

[0153] Results of t - test of region x with control group xL and control group xR respectively in Table 1

[0154] p-value t-value x_vs_xL 1.96E-06 7.5657 x_vs_xR 1.58E-06 8.3803

[0155] 9. For the sequences classified as human from AAV data, their actual sources are of two types, actual recombinants or ordinary human genomic contamination. In this case, the recombinant regions are more likely to show abnormal depth. Therefore, we first detected the human - derived sequences as potential recombinant regions, and then used the depth + t - test method to identify the regions with significantly high depth as the recombinant detection regions. As can be seen from Table 2, in this example, we used the conventional method to extract 9062 potential recombinant regions from chromosome 21 of the sample, and finally detected 3537 recombinant regions, reducing the false positive rate by 60.97%, which proves that this method can effectively improve the accuracy of the detection results.

[0156] Table 2 Detection results of recombinant regions of the sample on chromosome 21

[0157]

[0158] Example 4

[0159] Example 4 aims to demonstrate a method of double verification by DNA + protein to solve the problems of cumbersome operation and long time - consuming of the qPCR detection method, and high false positive rate of the detection results when the primers have non - specific binding, so as to improve the efficiency and reduce the false positive rate in the detection of resistance genes. It is a specific refinement corresponding to step 3) in Example 1. Figure 5 It is a schematic diagram of the method of double verification by DNA + protein.

[0160] In this example, the input data to be prepared include:

[0161] (1) 1 AAV sample, after extracting the cell fluid containing AAV and sequencing it respectively, 1 pair of paired - end Fastq files are obtained.

[0162] 1. First, collect resistance genes and corresponding protein reference databases from the DNA level and protein level respectively, and integrate reference data from different sources into a single DNA input and a single protein input file. So far, we have obtained a DNA level resistance gene database and a protein level resistance gene database.

[0163] 2. Then, using the same multi-stage classification method as in Example 1, the virus-derived sequence is taken out from the double-ended Fastq file to obtain a virus double-ended Fastq file.

[0164] 3. Next, compare the virus double-ended Fastq file to the database. The specific operation method is: use a comparison tool, such as bwa, to compare the double-ended Fastq file to the DNA level resistance gene database to obtain the original DNA comparison result; use a comparison tool, such as diamond, to compare the double-ended Fastq file to the protein level resistance gene database to obtain the original protein comparison result.

[0165] 4. Filter the original DNA alignment results, remove the sequences that have not been aligned successfully and the sequences that have been aligned to multiple sites, and obtain the DNA alignment result that uniquely aligns to the resistance gene.

[0166] 5. Filter the original protein comparison results, remove the results with similarity less than 80%, and obtain the protein comparison results.

[0167] 6. Compare the DNA comparison results and protein comparison results of one sample, and take the intersection of the detected resistance genes to obtain the resistance gene detection results. From the DNA comparison results, you can obtain information such as the comparison length and comparison position of the sequence, and from the protein comparison results, you can obtain the sequence similarity between the query sequence and the reference sequence.

[0168] 7. As shown in Table 3, the same sample was compared using bwa and diamond, and the comparison results were filtered through the above steps. The DNA comparison results and protein comparison results were compared. It was found that the resistance genes detected at the DNA level and protein level had an 8.08% intersection, which reduced the false positive rate by 91.92%. It was verified that the sensitivity and accuracy of the detection results can be improved by the double comparison method of DNA and protein.

[0169] Table 3 Results obtained by double comparison of DNA and protein

[0170]

[0171] 8. The DNA+protein dual verification method used in this embodiment, as a process-based bioinformatics analysis method, can perform parallel analysis on multiple resistance genes of multiple samples, that is, the increase in data volume will not affect the analysis time length. Compared with the limitation of the qPCR method that can only detect a single resistance gene each time, the analysis efficiency of the DNA+protein dual verification method has been greatly improved.

[0172] As can be seen from the above embodiments, for the problems actually encountered in the quality control of rAAV, such as 1) low identification efficiency of unknown components of recombinant adeno-associated virus (rAAV) vector data, 2) high false positive rate of recombinant detection, and 3) low detection efficiency of resistance genes, we have proposed the following solutions, including multi-stage classification of sequencing data, screening of recombinants detected by conventional algorithms through searching for high-depth regions in the whole genome + t-test, and additionally establishing a resistance gene database to perform DNA+protein dual verification on the sequencing results and detect multiple resistance genes simultaneously. Using the above methods, the above problems encountered in the quality control of rAAV in the actual scenario have been solved.

[0173] The above are only a limited number of preferred embodiments of the present invention, and the description is relatively specific and detailed, but it should not be construed as a limitation to the scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention.

Claims

1. A classification and annotation method for adeno-associated virus data based on second-generation sequencing on an exogenous comprehensive database, characterized in that: The method comprises at least one of the following steps 1)-3): 1) Multi-stage classification: The raw sequencing data are compared with the a priori reference genome and classified into known source sequences and unknown source sequences through the homologous sequence information file; Performing a secondary comparison of the unknown source sequence with an exogenous comprehensive database to complete qualitative and quantitative analysis of the components; 2) Recombinant detection: Extract the whole genome depth distribution from the host source sequence to identify potential recombinant regions; A t-test of the before and after control areas was performed on each recombinant region to screen out the recombinant detection areas with a depth significantly higher than that of the control group; 3) Double verification of resistance genes: Construct a resistance gene database at the DNA and protein levels; The virus-derived sequences are compared with the DNA database and the protein database respectively, and the detection of the resistance gene is determined by the intersection comparison results.

2. The classification annotation method according to claim 1, characterized in that: The original sequencing data in step 1) is a double-end Fastq file.

3. The classification annotation method according to claim 1, characterized in that: The step 1) comprises the following specific steps: S1. First stage classification: S11, aligning the original double-end Fastq file with the a priori reference genome to obtain multiple original BAM files; S12, introducing the homologous sequence information file, classifying the sources of the sequences in the original BAM file, and obtaining a BAM file of known source sequences and a BAM file of unknown source sequences; S2, second stage classification: S21, for the unknown source sequence BAM file, extract the corresponding sequence from the original double-end Fastq file to obtain the unknown source sequence Fastq file; S22, introduce an external comprehensive database, compare the Fastq file of the unknown source sequence, and obtain the BAM file of the second-stage classification result; S3. Qualitative and quantitative classification results: S31. For the BAM file of the second-stage classification result, count the sequence types and quantities to obtain the qualitative and quantitative results of the components of the second-stage classification.

4. The classification annotation method according to claim 1, characterized in that: The step 2) comprises the following specific steps: S4. Extract potential recombinant data: S41, classifying the source of the sample sequencing sequence to obtain a host BAM file, wherein the sequence contained in the file is a sequence of human origin; S42. For each chromosome data in the host BAM file, the depth of each n-base length window on each chromosome is counted to obtain the depth distribution of each chromosome of the host genome. These regions are regarded as potential recombinant regions, where n is a natural number in the range of 100≤n≤1000. S5. t-test to detect the recombinant region: S51, for each region x in the depth distribution of each chromosome of the host genome, two other regions with a distance l before and after the region x are taken as control groups xL and xR, respectively, where the length of l is: n bases ≤ l ≤ 10n bases; S52, for region x, control group xL and control group xR, further subdividing each region into smaller regions; S53, performing a t-test on the subdivided areas of region x and control group xL to confirm whether the depth of region x is significantly higher than that of control group xL; S54, performing a t-test on the subdivided regions of region x and control group xR to confirm whether the depth of region x is significantly higher than that of control group xR; S55. If the depth of region x is significantly higher than that of the control group xL and significantly higher than that of the control group xR, region x is determined to be a recombinant detection region.

5. The classification annotation method according to claim 1, characterized in that: The step 3) comprises the following specific steps: S6. Construction of resistance gene database: S61, collect resistance genes and corresponding protein reference databases at the protein level and DNA level; S62, integrating reference data from different sources into a single DNA input file and a single protein input file to obtain a resistance gene DNA database and a resistance gene protein database, respectively; S7. Resistance gene detection: S71, according to the virus BAM file, backtrack and extract its corresponding virus Fastq file; S72, using the virus Fastq file as an input file, and aligning it with the integrated resistance gene and protein reference database using the same alignment parameters at the protein level and DNA level, respectively, to obtain the original DNA alignment standard result and protein standard alignment result; S73, removing redundancy from the original DNA alignment results, removing the sequences that were not aligned successfully, and obtaining a complete single sequence alignment result of the resistance gene; S74, screening the original protein alignment results, and retaining only the alignment sequences whose similarity exceeds a certain threshold; S75. Compare the detection results at the DNA level and the protein level, take out the intersection of the detected resistance genes, and obtain the detection result.

6. A hardware device for executing the classification annotation method according to any one of claims 1 to 5, characterized in that: include: A memory device for storing raw sequencing data, a priori reference genomes, an exogenous comprehensive database, a resistance gene DNA database, and a resistance gene protein database; A processor configured to perform the method steps as claimed in any one of claims 1 to 5; An input interface for receiving raw sequencing data; Output interface, used to output classification and quality control results and generate classification annotation reports.

7. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, the classification annotation method according to any one of claims 1 to 5 is implemented.

8. A server system for remotely executing the classification annotation method according to any one of claims 1 to 5, characterized in that: include: at least one central processing unit; at least one graphics processing unit, optionally for accelerating data processing; at least one memory unit for storing data and programs; A network interface for communicating with a client device, receiving sequencing data and returning analysis results; An operating system configured to support the collaborative work of the above-mentioned hardware resources and to execute the classification annotation method as described in any one of claims 1 to 5.

9. A cloud-based service platform for providing the service of the classification annotation method according to any one of claims 1 to 5, characterized in that: include: Cloud server clusters for storing and processing large amounts of sequencing data; A user interface that allows users to upload sequencing data, view analysis progress, and obtain final results; Security mechanism to ensure the security of data transmission and storage; A billing system that charges users based on service usage.

10. Application of the classification annotation method according to any one of claims 1 to 5 in drug development, characterized in that: The application includes one or more of 1)-3): 1) Antiviral drug target screening; 2) Drug resistance research; 3) Evaluation of drug efficacy.

Citation Information

Patent Citations

  • Antibiotic resistance gene detection method in livestock and poultry wastewater struvite phosphorus recovery process

    CN118291594A