A method for processing viral data based on high-throughput sequencing data

Through identification of models and linkage relationship analysis, viral lineages and mutation sites are determined, and viral haplotype sequences are constructed, which solves the problem of inefficient processing of high-throughput sequencing data and achieves improved efficiency and accuracy of viral data processing.

CN119446257BActive Publication Date: 2025-05-27SUZHOU INST OF SYST MEDICINE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510032740.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-27
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

The prior art is inefficient in virus data processing based on high-throughput sequencing data, and it is difficult to fully reveal the full picture and variant pattern of the quasi-species of the host.

Method used

By reading the original sequencing data, using preset identification models to determine the virus lineage, determine the mutation site and linkage relationship, screen specific mutation sites, construct viral haplotype sequences, and improve the efficiency and accuracy of viral data processing.

Benefits of technology

The efficiency and accuracy of virus data processing for high-throughput sequencing data are improved, and the genetic relationship and mutation characteristics of viral quasi-species in the host can be more comprehensively analyzed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119446257B_ABST
    Figure CN119446257B_ABST
Patent Text Reader

Abstract

This specification discloses a virus data processing method based on high-throughput sequencing data. The server can analyze the original sequencing data through a neural network model to utilize the information contained in the original sequencing data, determine the virus lineages to which each virus contained in the target sample belongs, and then, based on the virus lineages to which each virus contained in the target sample belongs output by the neural network model, perform virus data processing on the target sample based on high-throughput sequencing data to improve the efficiency and accuracy of virus data processing based on high-throughput sequencing data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of bioinformatics analysis technology, and particularly to a method for processing virus data based on high-throughput sequencing data. Background Art

[0002] Currently, when a virus replicates in a host, due to factors such as errors during replication and pressure from the host immune system, genetic variations may occur, thus forming a population composed of multiple related but slightly different virus strains, namely the so-called quasispecies. The composition and evolutionary dynamics of the virus quasispecies in the host are closely related to clinical manifestations such as the severity of the disease, the ability of the virus to evade the host immune response, and the emergence of antiviral drug resistance. Therefore, analyzing virus quasispecies to study their characteristics is crucial for deeply understanding the virus pathogenesis mechanism and optimizing treatment strategies.

[0003] Generally, although a large amount of raw sequencing data can be obtained at one time through high-throughput sequencing technology when processing virus data based on high-throughput sequencing data, it is relatively difficult to extract and analyze meaningful quasispecies characteristics from the vast amount of data, which results in low efficiency of virus data processing based on high-throughput sequencing data.

[0004] Therefore, how to improve the efficiency of virus data processing based on high-throughput sequencing data is an urgent problem to be solved. Summary of the Invention

[0005] This specification provides a method for processing virus data based on high-throughput sequencing data to partially solve the above problems existing in the prior art.

[0006] This specification adopts the following technical solutions:

[0007] This specification provides a method for processing virus data based on high-throughput sequencing data, including:

[0008] Reading the raw sequencing data of a target sample, where the raw sequencing data is used to characterize the nucleic acid sequence of the target sample;

[0009] Inputting the raw sequencing data into a preset recognition model to determine the virus lineages to which each virus contained in the target sample belongs through the recognition model; and

[0010] Determining each mutation site included in the raw sequencing data and the mutation frequency of each mutation site, and determining the linkage relationship between the mutation sites, where the linkage relationship is used to reflect the genetic pattern in which some of the mutation sites are passed on to offspring as a whole;

[0011] For each virus lineage, identify the variant sites that match the preset specific variant sites corresponding to the virus lineage from among the various variant sites, as the variant sites belonging to the virus lineage;

[0012] Based on the virus lineage to which each variant site belongs, the mutation frequency of each variant site, and the linkage relationship between the various variant sites, determine the genetic relationship between the variant sites, so as to determine the various virus haplotype sequences contained in the target sample according to the genetic relationship, and perform tasks according to the various virus haplotype sequences. The virus haplotype sequence is used to reflect the combination of at least some variant sites with a close genetic relationship among the various variant sites.

[0013] Optionally, determining the various variant sites contained in the original sequencing data specifically includes:

[0014] Perform preprocessing on the original sequencing data to obtain preprocessed original sequencing data. The preprocessing includes at least one of quality assessment, trimming, and deduplication;

[0015] Determine the various variant sites contained in the preprocessed original sequencing data.

[0016] Optionally, the original sequencing data is composed of each segment data. Among them, for each segment data, the segment data is used to represent a short sequence segment in the nucleic acid sequence of the target sample;

[0017] Determining the various variant sites contained in the original sequencing data specifically includes:

[0018] For each segment data, align the segment data with a preset virus reference genome to determine the corresponding number of the segment data. The number is used to represent the position of the segment data on the virus reference genome;

[0019] According to the number corresponding to each segment data, sort and recombine each segment data into each segment data sequence as each virus sequence, and detect the difference points between each virus sequence and the virus reference genome as the various variant sites contained in the original sequencing data.

[0020] Optionally, determining the linkage relationship between the various variant sites specifically includes:

[0021] For each set of variant sites, according to the distribution of each variant site in each segmented data, determine the co-occurrence frequency between the variant sites in the set of variant sites, and according to the co-occurrence frequency, determine the linkage relationship between the variant sites in the set of variant sites, where the set of variant sites contains at least two variant sites, and the co-occurrence frequency is used to reflect the frequency of each variant site in the variant site combination appearing in the same segmented data.

[0022] Optionally, for each virus lineage, determine the variant sites that match the specific variant sites corresponding to the preset virus lineage from the variant sites as the variant sites belonging to the virus lineage, specifically including:

[0023] Obtain the expected frequency of each virus lineage;

[0024] For each virus lineage, determine the variant sites that match the preset specific variant sites corresponding to the virus lineage from the variant sites as the candidate variant sites belonging to the virus lineage;

[0025] According to the expected frequency of each virus lineage, determine the reference variant frequency of the candidate variant sites;

[0026] If it is determined that the difference between the reference variant frequency of the candidate variant site and the variant frequency of the candidate variant site is less than the preset difference threshold, then determine that the candidate variant site is a variant site belonging to the virus lineage.

[0027] Optionally, input the original sequencing data into a preset recognition model to determine the virus lineages to which the viruses contained in the target sample belong through the recognition model, specifically including:

[0028] Input the original sequencing data into a preset recognition model to determine the virus lineages to which the viruses contained in the target sample belong through the recognition model, and the credibility of each virus lineage;

[0029] According to the virus lineage to which each variant site belongs, the variant frequency of each variant site, and the linkage relationship between the variant sites, determine the genetic relationship between the variant sites, specifically including:

[0030] According to the credibility of each virus lineage, screen out the target virus lineage from the virus lineages;

[0031] According to the target virus lineage, the variant frequency of each variant site belonging to the target virus lineage, and the linkage relationship between the variant sites, determine the genetic relationship between the variant sites.

[0032] Optionally, task execution is performed according to each of the virus haplotype sequences, specifically including:

[0033] For each virus haplotype sequence, determine the frequency value of the occurrence of the virus haplotype sequence in the original sequencing data as the abundance value of the virus haplotype sequence;

[0034] According to the abundance values of each virus haplotype sequence, screen out the dominant haplotype sequences from each of the virus haplotype sequences, and perform task execution according to the dominant haplotype sequences.

[0035] This specification provides a virus data processing device based on high-throughput sequencing data, including:

[0036] An acquisition module, configured to read the original sequencing data of a target sample, where the original sequencing data is used to characterize the nucleic acid sequence of the target sample;

[0037] An identification module, configured to input the original sequencing data into a preset identification model to determine, through the identification model, each virus lineage to which each virus contained in the target sample belongs;

[0038] A determination module, configured to determine each mutation site included in the original sequencing data and the mutation frequency of each mutation site, and determine the linkage relationship between the mutation sites, where the linkage relationship is used to reflect the genetic pattern in which some of the mutation sites among the mutation sites are passed on to offspring as a whole;

[0039] A matching module, configured to, for each virus lineage, determine, from the mutation sites, the mutation sites that match the preset specific mutation sites corresponding to the virus lineage as the mutation sites belonging to the virus lineage;

[0040] An execution module, configured to determine the genetic relationship between the mutation sites according to the virus lineage to which each mutation site belongs, the mutation frequency of each mutation site, and the linkage relationship between the mutation sites, and determine each virus haplotype sequence included in the target sample according to the genetic relationship, and perform task execution according to each virus haplotype sequence, where the virus haplotype sequence is used to reflect the combination of at least some of the mutation sites with a close genetic relationship among the mutation sites.

[0041] The above at least one technical solution adopted in this specification can achieve the following beneficial effects:

[0042] In the method for processing virus data based on high-throughput sequencing data provided in this specification, first, the original sequencing data of the target sample is read. The original sequencing data is used to characterize the nucleic acid sequence of the target sample. Then, the original sequencing data is input into a preset recognition model to determine, through the recognition model, the virus lineages to which the viruses contained in the target sample belong, the mutation sites contained in the original sequencing data and the mutation frequency of each mutation site, and the linkage relationship between the mutation sites. Here, the linkage relationship is used to reflect the relative positions of the mutation sites in the original sequencing data and the genetic pattern in which some of the mutation sites are passed on to offspring as a whole. For each virus lineage, mutation sites that match the specific mutation sites corresponding to the preset virus lineage are determined from the mutation sites as the mutation sites belonging to the virus lineage. According to the virus lineage to which each mutation site belongs, the mutation frequency of each mutation site, and the linkage relationship between the mutation sites, the genetic relationship between the mutation sites is determined to determine the virus haplotype sequences contained in the target sample according to the genetic relationship, and tasks are executed according to the virus haplotype sequences. The virus haplotype sequences are used to reflect the combination of at least some mutation sites with a close genetic relationship among the mutation sites.

[0043] As can be seen from the above method, a neural network model can be used to analyze the original sequencing data to utilize the information contained in the original sequencing data to determine the virus lineages to which the viruses contained in the target sample belong. Furthermore, based on the virus lineages to which the viruses contained in the target sample belong output by the neural network model, virus data processing based on high-throughput sequencing data can be performed on the target sample to improve the efficiency and accuracy of virus data processing based on high-throughput sequencing data. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The drawings described herein are used to provide a further understanding of this specification and form a part of this specification. The schematic embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:

[0045] Figure 1 It is a schematic flowchart of a method for processing virus data based on high-throughput sequencing data provided in this specification;

[0046] Figure 2 It is a schematic diagram of a target sample provided in this specification;

[0047] Figure 3 It is a schematic diagram of a device for processing virus data based on high-throughput sequencing data provided in this specification;

[0048] Figure 4 This specification provides a corresponding toFigure 1 Schematic diagram of electronic equipment. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.

[0050] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.

[0051] During the replication process of the virus in the host, due to replication errors and immune pressure, genetic mutations may occur, resulting in a series of related mutants. These mutants constitute a viral quasispecies. These mutants belonging to the same viral quasispecies are genetically similar to the original virus strain, but may differ in certain characteristics, such as replication speed, transmission ability, or sensitivity to drugs. The formation and evolution of viral quasispecies are of great significance for understanding the clinical manifestations of viral infections such as infection severity, immune escape, and drug resistance.

[0052] At present, virus data processing based on high-throughput sequencing data mainly relies on sequencing a small number of isolated strains, which makes it difficult to fully reveal the full picture of quasispecies within the host. Although high-throughput sequencing technology can obtain a large number of virus sequences at one time, the process from raw sequencing data obtained through high-throughput sequencing technology to quasispecies variation analysis is relatively difficult, and it is difficult to accurately identify the variation pattern within the host in massive data. In addition, the efficiency of processing massive data is low, which makes it difficult to meet the needs of large-scale population research.

[0053] Figure 1 A schematic diagram of a method for processing viral data based on high-throughput sequencing data provided in this specification, comprising the following steps:

[0054] S101: Reading original sequencing data of a target sample, wherein the original sequencing data is used to characterize the nucleic acid sequence of the target sample.

[0055] In this specification, the business platform can use samples that need to undergo virus data processing based on high-throughput sequencing data as target samples, so that it can read the original sequencing data obtained after the target sample is sequenced in advance by a high-throughput sequencing platform, and then perform virus data processing based on high-throughput sequencing data based on the read original sequencing data of the target sample.

[0056] Among them, the above-mentioned target sample can be a sample containing viral genetic information. For example: body fluids (such as blood, saliva, urine, cerebrospinal fluid, etc.), tissue samples, and cell cultures of organisms infected with ribonucleic acid (RNA) viruses (such as humans, animals, or plants). Another example: virus particles extracted from the environment, etc.

[0057] The above-mentioned raw sequencing data can be the sequence information of deoxyribonucleic acid (DNA) fragments or RNA fragments obtained by performing high-throughput sequencing on the above-mentioned target sample (in the case where the virus contained in the above-mentioned target sample is an RNA virus, it is usually necessary to first reverse-transcribe RNA into DNA). Here, the DNA fragments or RNA fragments can be obtained by splitting the nucleic acid sequence of the virus in the target sample. Among them, the sequence information of each DNA fragment or RNA fragment is a segmented data that makes up the raw sequencing data. The sequence information of each DNA fragment or RNA fragment can be presented in the form of bases (A, T, C, G), and each sequence fragment has a unique identifier corresponding to it (this unique identifier is automatically generated during the sequencing process). Each of the above-mentioned DNA fragments or RNA fragments is a short sequence fragment.

[0058] It should be noted that after obtaining the raw sequencing data of the target sample by performing high-throughput sequencing on the above-mentioned target sample, the raw sequencing data can be saved to a database, so that when the business platform needs to perform virus data processing based on the high-throughput sequencing data of the target sample, it can read the raw sequencing data of the target sample for analysis.

[0059] In this specification, the execution subject for implementing the virus data processing method based on high-throughput sequencing data can refer to a designated device such as a server set up in the business platform, or it can also refer to terminal devices such as desktop computers and laptop computers. For the convenience of description, hereinafter, only the case where the server is the execution subject will be used as an example to illustrate the virus data processing method provided in this specification.

[0060] S102: Input the raw sequencing data into a preset recognition model to determine, through the recognition model, the virus lineages to which the various viruses contained in the target sample belong.

[0061] In this specification, after the server reads the raw sequencing data of the target sample, it can input the raw sequencing data into a preset recognition model to identify the raw sequencing data through the recognition model, so as to determine the virus lineages to which the various viruses contained in the target sample belong.

[0062] Among them, the above-mentioned recognition model can be set according to actual needs. Preferably, the above-mentioned recognition model can be the Bidirectional Encoder Representations from Transformers (BERT).

[0063] The above recognition model needs to be trained before it can be deployed to the server for the recognition of raw sequencing data. The method for training the recognition model can be as follows: for each virus lineage, select the virus species belonging to that virus lineage as the sample virus species. After splitting the nucleic acid sequence data of the sample virus species, obtain each sample segment data. Input each sample segment data into the initial recognition model so that the initial recognition model recognizes each sample segment data, and obtain the recognition results of the virus lineages to which each sample segment data belongs output by the initial recognition model. Furthermore, the pre-training loss can be determined based on the deviation between the recognition results output by the initial recognition model and the recognition results of the virus lineages to which each sample segment data actually belongs (the pre-training loss is positively correlated with the above deviation). Furthermore, the initial recognition model can be trained with the goal of minimizing the pre-training loss to obtain the pre-trained recognition model.

[0064] Furthermore, the server can obtain the sample sequencing data and input the sample sequencing data into the pre-trained recognition model to determine, through the pre-trained recognition model, the virus lineages to which each virus contained in the training sample corresponding to the sample sequencing data belongs. Furthermore, the recognition loss can be determined based on the deviation between the virus lineages to which each virus contained in the training sample corresponding to the sample sequencing data belongs determined by the pre-trained recognition model and the label of the training sample corresponding to the sample sequencing data. Furthermore, the model parameters of the pre-trained recognition model can be fine-tuned with the goal of minimizing the recognition loss to obtain the trained recognition model.

[0065] Among them, the label of the above training sample is used to represent the virus lineages to which each virus actually contained in the training sample belongs.

[0066] The above recognition loss is positively correlated with the deviation between the virus lineages to which each virus contained in the training sample corresponding to the sample sequencing data belongs determined by the recognition model and the label of the training sample corresponding to the sample sequencing data.

[0067] It should be noted that the above virus lineage can refer to different evolutionary branches formed by viruses through genetic variation in the host or population. Each virus lineage represents a group of virus species with similar genetic characteristics. The virus lineage reflects the changes of the virus over time, especially the adaptive changes in the face of selection pressures (such as host immune response, drug treatment, etc.).

[0068] In addition, before the server inputs the raw sequencing data into the recognition model, the server can also preprocess the raw sequencing data to obtain the preprocessed raw sequencing data, and then input the preprocessed raw sequencing data into the recognition model to determine, through the recognition model, the virus lineages to which each virus contained in the target sample belongs.

[0069] Among them, the above-mentioned preprocessing includes at least one of quality assessment, shearing, and deduplication.

[0070] S103: Determine each variant site included in the original sequencing data and the variant frequency of each variant site, and determine the linkage relationship between the variant sites. The linkage relationship is used to reflect the relative positions of the variant sites in the original sequencing data and the genetic pattern in which some of the variant sites are transmitted to offspring as a whole.

[0071] Furthermore, the server can determine each variant site included in the original sequencing data and the variant frequency of each variant site, and determine the linkage relationship between the variant sites.

[0072] In an actual application scenario, during the high-throughput sequencing of a target sample, it is necessary to cut the nucleic acid sequence of the target sample to obtain each segmented data. However, during this process, factors such as sequencing errors and random cutting of short sequence fragments may affect the generation and position information of short sequence fragments, thereby affecting each segmented data. Therefore, after obtaining the original sequencing data, and before determining each variant site included in the original sequencing data, the variant frequency of each variant site, and the linkage relationship between the variant sites, it is also necessary to align each segmented data to determine the position of each segmented data on the viral reference genome, so as to determine the relative position of each segmented data and restore the complete nucleic acid sequence.

[0073] Based on this, the server can align each segmented data included in the preprocessed original sequencing data with a preset viral reference genome to determine the corresponding number of the segmented data. Furthermore, according to the numbers corresponding to each segmented data, each segmented data can be sorted and recombined into each segmented data sequence as each viral sequence, and the difference points between each viral sequence and the viral reference genome can be detected as each variant site included in the original sequencing data.

[0074] Among them, the above-mentioned number is used to represent the position of the segmented data on the viral reference genome.

[0075] Furthermore, the server can determine, for each variant site, through a preset detection tool, the frequency at which the variant site appears in the original sequencing data as the variant frequency of the variant site.

[0076] The detection tools in the above content can be tools such as Samtools and VarScan for monitoring variants in genomic sequences.

[0077] Furthermore, the server can determine the linkage relationship between each mutation site based on the mutation frequency of each mutation site and the positions of the respective mutation sites.

[0078] Among them, the above-mentioned linkage relationship is used to reflect the genetic pattern in which some of the mutation sites among the respective mutation sites are passed on to offspring as a whole.

[0079] For example: If the positions of two mutation sites are very close, then the probability of recombination (i.e., chromosomal segment exchange) occurring between them is low. Therefore, the mutations at these two sites are more likely to be replicated and transmitted together. So, there is a strong linkage relationship between these two mutation sites.

[0080] Another example: If another mutation site often appears simultaneously when one mutation site appears, it may be due to selection pressure or other biological mechanisms that cause these two mutation sites to tend to appear together (e.g., when mutation site a and mutation site b appear simultaneously, it may confer higher drug resistance to the virus, thus making the combination of ectopic site a and mutation site b more common in the virus population). So, there is a strong linkage relationship between these two mutation sites.

[0081] Among them, the method by which the server determines the linkage relationship between each mutation site based on the mutation frequency of each mutation site and the positions of the respective mutation sites can be that the server can, for each set of mutation sites, determine the co-occurrence frequency between the mutation sites in the set according to the distribution of each mutation site in each segment of data, and determine the linkage relationship between the mutation sites in the set according to the co-occurrence frequency between the mutation sites in the set.

[0082] Among them, for each set of mutation sites, the set of mutation sites contains at least two mutation sites.

[0083] The above-mentioned co-occurrence frequency is used to reflect the frequency at which the mutation sites in the combination of mutation sites appear in the same segment of data.

[0084] The method by which the server determines the linkage relationship between the mutation sites in the set according to the co-occurrence frequency between the mutation sites in the set can be that the server determines the expected co-occurrence frequency of the mutation sites in the set according to the mutation frequency of each mutation site in the set, and then can determine whether there is a strong linkage relationship between the mutation sites in the set according to whether the ratio between the expected co-occurrence frequency of the mutation sites in the set and the co-occurrence frequency between the mutation sites in the set exceeds a preset ratio threshold.

[0085] For ease of understanding, the following describes in detail the method by which the above server determines the linkage relationship between each mutation site based on the mutation frequency of each mutation site and the positions of each mutation site in conjunction with an embodiment.

[0086] For example: If the total number of each segmented data contained in the original sequencing data is 10,000, the number of segmented data containing mutation site A is 3,000, the number of segmented data containing mutation site B is 2,000, the number of segmented data containing mutation site C is 4,000, the number of segmented data containing mutation site D is 1,000, the number of segmented data containing both mutation site A and mutation site B is 1,500, and the number of segmented data containing both mutation site A and mutation site B is 400.

[0087] At this time, the mutation frequency of mutation site A is 30%, the mutation frequency of mutation site B is 20%, the frequency of mutation site C is 40%, the frequency of mutation site D is 10%, the co-occurrence frequency of mutation site A and mutation site B appearing together is 15%, the co-occurrence frequency of mutation site C and mutation site D appearing together is 4%, and based on the mutation frequency of mutation site A and the mutation frequency of mutation site B, it can be deduced that the expected co-occurrence frequency between mutation site A and mutation site B is 30% * 20%, that is, 6%. In fact, the co-occurrence frequency of mutation site A and mutation site B appearing together is 15%. From this, it can be determined that the co-occurrence frequency between mutation site A and mutation site B is significantly different from the co-occurrence frequency of random distribution, indicating that there may be some association between them (that is, a strong linkage relationship).

[0088] Similarly, the expected co-occurrence frequency between mutation site C and mutation site D is 40% * 10%, that is, 4%. In fact, the co-occurrence frequency of mutation site A and mutation site B appearing together is 4%. From this, it can be determined that the co-occurrence frequency between mutation site A and mutation site B is the same as the co-occurrence frequency of random distribution, indicating that there may be a weak association between them (that is, a weak linkage relationship).

[0089] S104: For each virus lineage, determine the mutation sites that match the preset specific mutation sites corresponding to the virus lineage from the above-mentioned mutation sites as the mutation sites belonging to the virus lineage.

[0090] In this specification, the server can determine, for each virus lineage, the mutation sites that match the preset specific mutation sites corresponding to the virus lineage from the mutation sites as the mutation sites belonging to the virus lineage.

[0091] Among them, for each virus lineage, the specific variant sites of the virus lineage can refer to the variant sites that frequently appear in the virus lineage but rarely or do not appear in other lineages, and these specific variants can be used to distinguish different virus lineages.

[0092] In the above content, for each variant site, if the server determines that there is the above-mentioned linkage relationship between the variant site and the specific variant site corresponding to a virus lineage, it can be considered that the variant site matches the specific variant site of this virus lineage.

[0093] Furthermore, the server can also pre-determine the expected frequency of each virus lineage in advance, so that for each virus lineage, the variant sites determined from each variant site that match the preset specific variant sites corresponding to the virus lineage can be used as candidate variant sites belonging to the virus lineage. Furthermore, according to the expected frequency of each virus lineage, the reference variant frequency of the candidate variant sites belonging to the virus lineage can be determined. If it is determined that the difference between the reference variant frequency of the candidate variant site and the variant frequency of the candidate variant site is less than the preset difference threshold, it can be determined that the candidate variant site is a variant site belonging to the virus lineage.

[0094] Among them, the expected frequency of each virus lineage mentioned above can be set in advance. Of course, it can also be predicted from the original sequencing data of the target sample through a preset prediction model.

[0095] It should be noted that since the variant sites belonging to a virus lineage are often inherited as a whole, the variant frequencies of the variant sites belonging to a virus lineage should be similar. That is, when the expected frequency of a virus lineage is 30%, the reference variant frequency of each variant site belonging to this virus lineage should be 30%. Therefore, the server can determine the reference variant frequency of the candidate variant sites belonging to the virus lineage according to the expected frequency of each virus lineage.

[0096] In an actual application scenario, there are some variant sites that belong to more than two virus lineages at the same time. At this time, the reference variant frequency of this variant site can be the sum of the expected frequencies of the virus lineages to which this variant site belongs.

[0097] For example: If a variant site a belongs to virus lineage A and virus lineage B at the same time, and the expected frequency of virus lineage A is 25% and the expected frequency of virus lineage B is 30%, then the reference variant frequency of this variant site should be 55%.

[0098] Since the reference lineage frequencies of each virus lineage represent the proportions of all virus particles in the sample in different lineages, the sum of the frequencies of the virus lineages to which the viruses belonging to each sample belong should be approximately equal to 100%. Specifically, such asFigure 2 as shown

[0099] Figure 2 It is a schematic diagram of a target sample provided in this specification.

[0100] Combined with Figure 2 It can be seen that when a sample contains viruses of two virus lineages (i.e., virus lineage A and virus lineage B), when the reference lineage frequency of virus lineage A is 75%, the reference lineage frequency of virus lineage B should be about 25%, that is, the sum of the reference lineage frequencies of virus lineage A and virus lineage B should be about 100%.

[0101] Based on this, the server can also determine the frequency of each virus lineage according to each mutation site belonging to each virus lineage, and then can judge whether the sum of the frequencies of each virus lineage is approximately equal to 100%. If so, for each virus lineage, the candidate mutation sites belonging to the virus lineage can be used as the mutation sites belonging to the virus lineage.

[0102] S105: Determine the genetic relationship between each mutation site according to the virus lineage to which each mutation site belongs, the mutation frequency of each mutation site, and the linkage relationship between each mutation site, so as to determine each virus haplotype sequence included in the target sample according to the genetic relationship, and perform tasks according to each virus haplotype sequence. The virus haplotype sequence is used to reflect the combination of at least some mutation sites with a close genetic relationship among each mutation site.

[0103] In this specification, the server can determine the genetic relationship between each mutation site according to the virus lineage to which each mutation site belongs, the mutation frequency of each mutation site, and the linkage relationship between each mutation site, so as to determine each virus haplotype sequence included in the target sample according to the genetic relationship between each mutation site.

[0104] Specifically, the server can screen out each mutation site with a mutation frequency higher than a preset mutation frequency threshold from each mutation site as each high-frequency mutation site, and then can divide each high-frequency mutation site into different high-frequency mutation site combinations according to the linkage relationship between each mutation site and the virus lineage to which each high-frequency mutation site belongs.

[0105] Among them, for each high-frequency mutation site combination, it can be considered that there is a close genetic relationship between each high-frequency mutation site included in the high-frequency mutation site combination, and then each virus haplotype sequence included in the target sample can be determined according to the genetic relationship between each high-frequency mutation site.

[0106] In addition, the server can also screen out target virus lineages from each virus lineage according to the credibility of each virus lineage, and then can determine the genetic relationship between each mutation site based on the target virus lineage, the mutation frequency of each mutation site belonging to the target virus lineage, and the linkage relationship between each mutation site, so as to determine each virus haplotype sequence contained in the target sample according to the genetic relationship.

[0107] Among them, the credibility of each virus lineage above can be obtained through an identification model.

[0108] Furthermore, the server can determine the frequency value of each virus haplotype sequence appearing in the original sequencing data as the abundance value of the virus haplotype sequence, and then can screen out the dominant haplotype sequence from each virus haplotype sequence according to the abundance value of each virus haplotype sequence, and perform tasks according to the dominant haplotype sequence.

[0109] It should be noted that the tasks to be executed above can be determined according to actual needs. For example, the server can display the determined dominant haplotype sequence to the user so that the user can develop a vaccine based on the dominant haplotype sequence contained in the target sample.

[0110] For another example: the server can determine the relationship between the dominant haplotype contained in the target sample and drug metabolism, so that it can recommend a list of drugs that can be used to the user based on the relationship between the dominant haplotype and drug metabolism.

[0111] It can be seen from the above content that the server can analyze the original sequencing data through a neural network model to utilize the information contained in the original sequencing data to determine each virus lineage to which each virus contained in the target sample belongs. Furthermore, based on each virus lineage to which each virus contained in the target sample belongs output by the neural network model, virus data processing based on high-throughput sequencing data can be performed on the target sample to improve the efficiency and accuracy of virus data processing based on high-throughput sequencing data.

[0112] The above is one or more embodiments of the method for processing virus data based on high-throughput sequencing data in this specification. Based on the same idea, this specification also provides a corresponding device for processing virus data based on high-throughput sequencing data, as Figure 3 shown.

[0113] Figure 3 is a schematic diagram of a device for processing virus data based on high-throughput sequencing data provided by this specification, including:

[0114] An acquisition module 301, configured to read the original sequencing data of the target sample, where the original sequencing data is used to characterize the nucleic acid sequence of the target sample;

[0115] An identification module 302, configured to input the original sequencing data into a preset identification model, so as to determine, through the identification model, each virus lineage to which each virus contained in the target sample belongs;

[0116] A determination module 303, configured to determine each mutation site included in the original sequencing data and the mutation frequency of each mutation site, and determine the linkage relationship between the mutation sites, where the linkage relationship is used to reflect the genetic pattern in which some of the mutation sites are transmitted to offspring as a whole;

[0117] A matching module 304, configured to, for each virus lineage, determine, from the mutation sites, the mutation sites that match the preset specific mutation sites corresponding to the virus lineage as the mutation sites belonging to the virus lineage;

[0118] An execution module 305, configured to determine the genetic relationship between the mutation sites according to the virus lineage to which each mutation site belongs, the mutation frequency of each mutation site, and the linkage relationship between the mutation sites, so as to determine each virus haplotype sequence included in the target sample according to the genetic relationship, and perform a task according to the virus haplotype sequences, where the virus haplotype sequence is used to reflect the combination of at least some mutation sites with a close genetic relationship among the mutation sites.

[0119] Optionally, the determination module 303 is specifically configured to preprocess the original sequencing data to obtain preprocessed original sequencing data, where the preprocessing includes at least one of quality assessment, clipping, and deduplication; and determine each mutation site included in the preprocessed original sequencing data.

[0120] Optionally, the original sequencing data is composed of segmented data, where, for each segmented data, the segmented data is used to represent a short sequence segment in the nucleic acid sequence of the target sample;

[0121] The determination module 303 is specifically configured to, for each segmented data, compare the segmented data with a preset virus reference genome to determine the number corresponding to the segmented data, where the number is used to represent the position of the segmented data on the virus reference genome; sort and recombine the segmented data into segmented data sequences as virus sequences according to the number corresponding to each segmented data, and detect the difference points between the virus sequences and the virus reference genome as each mutation site included in the original sequencing data.

[0122] Optionally, the determining module 303 is specifically configured to, for each set of mutation sites, determine the co-occurrence frequency between the mutation sites in the set of mutation sites according to the distribution of the mutation sites in each segment of data, and determine the linkage relationship between the mutation sites in the set of mutation sites according to the co-occurrence frequency, where the set of mutation sites contains at least two mutation sites, and the co-occurrence frequency is used to reflect the frequency of the mutation sites in the combination of mutation sites appearing in the same segment of data.

[0123] Optionally, the matching module 304 is specifically configured to obtain the expected frequency of each virus lineage; for each virus lineage, determine, from the mutation sites, the mutation sites that match the preset specific mutation sites corresponding to the virus lineage as the candidate mutation sites belonging to the virus lineage; determine the reference mutation frequency of the candidate mutation sites according to the expected frequency of each virus lineage; if it is determined that the difference between the reference mutation frequency of the candidate mutation sites and the mutation frequency of the candidate mutation sites is less than a preset difference threshold, determine that the candidate mutation sites are the mutation sites belonging to the virus lineage.

[0124] Optionally, the identifying module 302 is specifically configured to input the original sequencing data into a preset identification model to determine, through the identification model, the virus lineages to which the viruses contained in the target sample belong, and the credibility of each virus lineage;

[0125] The determining module 303 is specifically configured to screen out the target virus lineages from the virus lineages according to the credibility of each virus lineage; determine the genetic relationship between the mutation sites according to the target virus lineages, the mutation frequency of each mutation site belonging to the target virus lineage, and the linkage relationship between the mutation sites.

[0126] Optionally, the execution module 305 is specifically configured to, for each virus haplotype sequence, determine the frequency value of the occurrence of the virus haplotype sequence in the original sequencing data as the abundance value of the virus haplotype sequence; screen out the dominant haplotype sequences from the virus haplotype sequences according to the abundance value of each virus haplotype sequence, and perform tasks according to the dominant haplotype sequences.

[0127] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above Figure 1 A virus data processing method based on high-throughput sequencing data provided.

[0128] This specification also provides Figure 4 The schematic structural diagram of an electronic device corresponding to Figure 1 as shown. Figure 4As described above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 method for processing virus data based on high-throughput sequencing data. Of course, in addition to the software implementation method, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.

[0129] The improvement of a technology can be clearly distinguished as a hardware improvement (for example, the improvement of circuit structures such as diodes, transistors, switches, etc.) or a software improvement (the improvement of a method flow). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always program the improved method flow into the hardware circuit to obtain the corresponding hardware circuit structure. Therefore, it cannot be said that an improvement in a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device PLD (such as a field programmable gate array FPGA) is such an integrated circuit, and its logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system on a piece of PLD without having to ask a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a hardware description language HDL. And there are not only one kind of HDL, but many kinds, such as ABEL, AHDL, Confluence, CUPL, HDCal, JHDL, Lava, Lola, MyHDL, PALASM, RHDL, etc. Currently, the most commonly used are VHDL and Verilog. Those skilled in the art should also be clear that as long as the method flow is slightly logically programmed with the above several hardware description languages and programmed into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0130] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.

[0131] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0132] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0133] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0134] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0135] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0136] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0137] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0138] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0139] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media such as modulated data signals and carrier waves.

[0140] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0141] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0142] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0143] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the description of the method embodiment.

[0144] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A virus data processing method based on high-throughput sequencing data, characterized in that: include: Reading raw sequencing data of the target sample, wherein the raw sequencing data is used to characterize the nucleic acid sequence of the target sample; Inputting the raw sequencing data into a preset recognition model to determine the virus lineages to which the viruses contained in the target sample belong, and the credibility of each virus lineage through the recognition model; as well as Determine each variant site contained in the original sequencing data and the variant frequency of each variant site, and determine the linkage relationship between the variant sites, wherein the linkage relationship is used to reflect the inheritance pattern of whether some of the variant sites in each variant site are transmitted to the offspring as a whole; For each viral lineage, a mutation site that matches a preset specific mutation site corresponding to the viral lineage is determined from the mutation sites as a mutation site belonging to the viral lineage; According to the credibility of each virus lineage, the target virus lineage is screened out from each virus lineage; According to the target virus lineage, the mutation frequency of each mutation site belonging to the target virus lineage, and the linkage relationship between the mutation sites, the genetic relationship between the mutation sites is determined, so as to determine the haplotype sequences of the viruses contained in the target sample based on the genetic relationship, and the task is performed based on the haplotype sequences of the viruses. The haplotype sequences of the viruses are used to reflect the combination of at least some mutation sites that have a close genetic relationship among the mutation sites.

2. The method according to claim 1, characterized in that Determining each variant site contained in the original sequencing data specifically includes: Preprocessing the raw sequencing data to obtain preprocessed raw sequencing data, wherein the preprocessing includes: at least one of quality assessment, shearing, and deduplication; Determine each variant site contained in the raw sequencing data after the preprocessing.

3. The method according to claim 1, characterized in that The original sequencing data is composed of segmented data, wherein for each segmented data, the segmented data is used to characterize a short sequence fragment in the nucleic acid sequence of the target sample; Determining each variant site contained in the original sequencing data specifically includes: For each segmented data, the segmented data is compared with a preset viral reference genome to determine a number corresponding to the segmented data, wherein the number is used to characterize the position of the segmented data on the viral reference genome; According to the number corresponding to each segmented data, the segmented data are sorted and reorganized into segmented data sequences as each virus sequence, and the difference points between the each virus sequence and the virus reference genome are detected as each mutation site contained in the original sequencing data.

4. The method according to claim 3, characterized in that Determining the linkage relationship between the variant sites specifically includes: For each variant site set, the co-occurrence frequency between the variant sites in the variant site set is determined according to the distribution of the variant sites in each segmented data, and the linkage relationship between the variant sites in the variant site set is determined according to the co-occurrence frequency, wherein the variant site set contains at least two variant sites, and the co-occurrence frequency is used to reflect the frequency of each variant site in the variant site combination appearing in the same segmented data.

5. The method according to claim 1, characterized in that For each virus lineage, a mutation site that matches a preset specific mutation site corresponding to the virus lineage is determined from the mutation sites as the mutation site belonging to the virus lineage, specifically including: Obtain the expected frequency of each viral lineage; For each viral lineage, a mutation site that matches a preset specific mutation site corresponding to the viral lineage is determined from the mutation sites as a candidate mutation site belonging to the viral lineage; Determining the reference variation frequency of the candidate variation site according to the expected frequency of each viral lineage; If it is determined that the difference between the reference mutation frequency of the candidate mutation site and the mutation frequency of the candidate mutation site is less than a preset difference threshold, the candidate mutation site is determined to be a mutation site belonging to the viral lineage.

6. The method according to claim 1, characterized in that The tasks are executed according to the haplotype sequences of each virus, specifically including: For each viral haplotype sequence, determining the frequency value of the viral haplotype sequence in the original sequencing data as the abundance value of the viral haplotype sequence; According to the abundance value of each viral haplotype sequence, a dominant haplotype sequence is screened out from the viral haplotype sequences, and the task is executed according to the dominant haplotype sequence.

7. A virus data processing device based on high-throughput sequencing data, characterized in that: include: An acquisition module is used to read the original sequencing data of the target sample, wherein the original sequencing data is used to characterize the nucleic acid sequence of the target sample; An identification module, used to input the raw sequencing data into a preset identification model, so as to determine the virus lineages to which the viruses contained in the target sample belong, and the credibility of each virus lineage through the identification model; A determination module, used to determine each variant site contained in the original sequencing data and the variant frequency of each variant site, and determine the linkage relationship between the variant sites, wherein the linkage relationship is used to reflect the inheritance mode of whether some variant sites in each variant site are transmitted to the offspring as a whole; A matching module, for determining, for each viral lineage, from the various mutation sites, a mutation site that matches a preset specific mutation site corresponding to the viral lineage as a mutation site belonging to the viral lineage; An execution module is used to screen out a target virus lineage from each virus lineage according to the credibility of each virus lineage; determine the genetic relationship between each variant site according to the target virus lineage, the mutation frequency of each variant site belonging to the target virus lineage, and the linkage relationship between the variant sites, so as to determine each virus haplotype sequence contained in the target sample according to the genetic relationship, and perform tasks according to each virus haplotype sequence, wherein the virus haplotype sequence is used to reflect the combination of at least part of the variant sites that have a close genetic relationship among the variant sites.

8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Copy number variations predictive of risk of schizophrenia

    CA2729856A1

  • Conditionally active chimeric antigen receptors for modified t-cells

    CN107922473A