Construction method of segmented virus gene reassortment host adaptability prediction model
By constructing a host adaptive prediction model for segmented virus gene reassignment, using key codon regions to simulate reassignment and neural network training, the problem of difficult host adaptability after gene reassignment of influenza virus is solved, and early identification and prevention of high-risk reassignment results are achieved.
Patent Information
- Application Number
- CN202510518031.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The prior art cannot effectively predict the host adaptability after gene reassorption of influenza viruses, especially the viral reassorption process for multigene fragments, making it difficult to identify variant strains that are at high risk to humans in the early stage.
By constructing a host adaptability prediction model for segmented virus gene reassignment, key codon regions are used to simulate reassignment, and combined with neural network training, the host adaptability after virus reassignment is predicted, and high-risk reassignment results are identified.
It has achieved early identification of human adaptability after virus reassorption, provided targeted preventive measures, and improved the efficiency and accuracy of influenza virus mutation monitoring.
Smart Images

Figure CN120356513A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medicine, and more specifically, to a method, device, medium and program product for constructing a segmented virus gene reassortment host adaptability prediction model. Background Art
[0002] The influenza virus genome is segmented. When different subtype virus strains undergo genetic material recombination in the same host cell, progeny viruses with new antigenicity, transmissibility and pathogenicity may be produced, and this process is called gene reassortment. Gene reassortment is not only one of the important mechanisms for the evolution and variation of influenza viruses, but also a key factor in triggering human-to-human or cross-species transmission. Therefore, studying influenza virus reassortment helps to deeply understand the virus transmission and variation mechanisms, and can provide important scientific basis for vaccine research and development, antiviral drug development, epidemic prediction and prevention and control strategy formulation.
[0003] On the one hand, in traditional biological research, co-infection methods and genetic engineering techniques can be used for virus reassortment, but often require high experimental skills and equipment conditions, and new virus strains with higher pathogenicity or drug resistance may be produced during the virus reassortment process, posing a threat to laboratory biosafety. On the other hand, artificial intelligence (AI) methods such as machine learning (ML) and deep learning (DL) have achieved remarkable results in the biological field in recent years, and also shown amazing effects in virus host specificity and adaptability prediction. Babayan et al. can predict the vector host based on virus gene codons and the amino acids they encode, etc. The machine learning model established by Li et al. can accurately predict the adaptive host of influenza A virus (IAV) based on the dinucleotide (DNT) characteristics of genes, and reveal the host specificity of coronavirus genes based on the gene DNT cluster characteristics. However, the above AI methods can only conduct prediction research on the host adaptability of a single gene segment, and do not involve virus reassortment prediction-related content.
[0004] Meanwhile, virological evidence emphasizes the polygenic nature of IAV adaptation and adaptive reassortment. Two envelope glycoproteins, hemagglutinin (HA) and neuraminidase (NA), which form the basis of virus serotyping, mediate virus binding and release. To date, only influenza virus types H1N1, H2N2, and H3N2 can specifically bind to human influenza virus receptors through their HA and cause influenza pandemics. Adaptive mutations in HA only result in fluctuations in receptor affinity, rather than a shift in avian-to-human binding specificity. However, the compatibility and constraints between RNA polymerase-related genes are more complex and are crucial for IAV adaptation and adaptive reassortment. RNA polymerase-related genes such as PB2, PB1, PA, and NP are responsible for influenza virus proliferation and packaging and are prone to reassortment between influenza viruses. Therefore, the present invention conducts research on polymerase-related genes. Summary of the Invention
[0005] In view of the above problems, the present invention provides a method for constructing a segmented virus gene reassortment host adaptability prediction model, which uses the critical codon region to simulate reassortment to achieve the prediction of the host adaptability of newly simulated reassorted viruses, thereby facilitating early monitoring and response to human high-risk infection mutations.
[0006] The first aspect of the present application first discloses a method for ranking the importance of codons of virus genes, the method comprising:
[0007] S1: Obtain a test data set representing virus genes and host labels, where each gene data in the test data set is a data representation of each codon of the gene;
[0008] S2: Based on the test data set, perform virus gene host adaptability prediction to obtain predicted host labels. After comparing the predicted host labels with the host labels, obtain the first prediction accuracy rate of the test set;
[0009] S3: Use the sliding window method to mask the data representation of the codons in the first window of the test data set to obtain the first masked test data set. Based on the first masked test data set, perform virus gene host adaptability prediction to obtain the predicted host labels of the masked first window. After comparing the predicted host labels of the masked first window with the host labels, obtain the prediction accuracy rate of the masked first window;
[0010] S4: Slide the window to sequentially mask the data representation of the codons in the first to the nth windows and then repeat the steps of S3 to obtain the second prediction accuracy rate, where the second prediction accuracy rate includes the prediction accuracy rates of masking the first to the nth windows, and n is the number of codons of the virus gene;
[0011] S5: Based on the difference between the second prediction accuracy rate and the first prediction accuracy rate, obtain the regional importance of masking the first to the nth windows;
[0012] S6: Average the regional importance of the window containing each codon to obtain the codon importance.
[0013] Furthermore, slide the window mask on the data of the test set after connecting the gene codons end to end, where the end refers to the digital representation of the last codon;
[0014] Optionally, the method for digital representation of codons is as follows: Connect the gene sequences of the single gene end to end and use the sliding window method to traverse the gene sequence. The frequencies of 64 different codons within the window form the window codon frequency vector, and use the window codon frequency vector to represent the central codon of the window. As the window slides, each codon of the gene sequence is sequentially represented as the window codon frequency vector when the codon is the central codon;
[0015] Optionally, the gene data is an encoding feature matrix composed of the window codon frequency vectors obtained by sliding window traversal;
[0016] Optionally, use the virus gene host adaptability prediction model to predict the virus gene host adaptability, and the virus gene host adaptability prediction model is obtained by training on the training data set of the virus gene;
[0017] Optionally, S3 further includes: obtaining the Bayesian posterior accuracy of the masked first window based on the prediction accuracy of the masked first window and the probability of each masked window; repeating the S3 step to obtain the third prediction accuracy, where the third prediction accuracy includes the Bayesian posterior accuracies of the masked first window to the nth window, and obtaining the Bayesian regional importance of the first to the nth window based on the difference between the third prediction accuracy and the first prediction accuracy; averaging the Bayesian regional importance of each codon to obtain the Bayesian posterior codon importance;
[0018] Optionally, obtain the optimal codon importance based on the weighted average of the codon importance and the Bayesian posterior codon importance.
[0019] The second aspect of the present application discloses a method for constructing a segmented virus gene reassortment host adaptability prediction model, and the method includes:
[0020] Obtain a virus data set and host labels;
[0021] For different genes of the viruses in the virus data set, filter the codon importance based on the codon importance threshold of different genes to obtain the key codons of different genes, and combine the digital representations of the key codons of different genes to obtain the reassortment result of the virus data set;
[0022] Input the reassortment results of the virus dataset and the host labels into a neural network for iterative training to obtain a prediction model for the host adaptability of viral gene reassortment.
[0023] Further, the method includes:
[0024] The way to combine the digital characterizations of the key codons of different genes is to merge the digital characterizations of the key codons of the genes in the order of the genes in the virus.
[0025] Optionally, the codon frequency vectors of the key codons of different genes are merged in the order of the genes in the virus to obtain a simulated sequence, and the simulated sequence is a matrix.
[0026] The third aspect of the present application discloses a method for predicting the adaptability of simulated reassortment of a viral genome, including:
[0027] Obtain a virus dataset, where the virus dataset is a digital characterization set of the virus;
[0028] For different genes of the viruses in the virus dataset, after filtering the codon importance based on the codon importance threshold of different genes, obtain the key codon regions of different genes, and randomly replace the key codons in the key codon regions of 1, 2, or 3 genes respectively to obtain a simulated reassortment dataset of the virus dataset;
[0029] Input the simulated reassortment results of the virus dataset into the prediction model for the host adaptability of viral gene reassortment according to any one of the above to obtain the host adaptability corresponding to the simulated reassortment results.
[0030] Further, the method for randomly replacing the key codons in the key codon regions is:
[0031] According to the codon importance of different genes, respectively intercept the digital characterizations of the key codons of different genes constituting the virus, and use the digital characterizations of the key codons of 1, 2, or 3 corresponding genes from another strain of virus for random replacement to obtain the digital characterization of the reassortment simulation sequence;
[0032] Optionally, the different viruses are viruses of different subtypes;
[0033] Optionally, randomly replace any several key codon short fragments in the gene fragments of the influenza virus from humans with the corresponding key codon short fragments of the corresponding genes of the avian virus subtype to obtain the simulated reassortment results.
[0034] The fourth aspect of the present application discloses a system for ranking the codon importance of viral genes, including:
[0035] The first acquisition module: used to acquire a test data set representing viral genes and host labels, where each gene data in the test data set is a digital representation of each codon of the gene;
[0036] The first prediction module: used to predict the host adaptability of viral genes based on the test data set to obtain a predicted host label, and obtain the first prediction accuracy of the test set after comparing the predicted host label with the host label;
[0037] The sliding mask prediction module: used to mask the digital representation of the codons in the first window of the test data set using the sliding window method to obtain the first masked test data set, predict the host adaptability of viral genes based on the first masked test data set to obtain the host label predicted by masking the first window, and obtain the prediction accuracy of masking the first window after comparing the host label predicted by masking the first window with the host label; The sliding window sequentially masks the digital representations of the codons in the first to the nth windows and repeats the above steps of sliding mask prediction to obtain the second prediction accuracy, and the second prediction accuracy includes the prediction accuracies of masking the first to the nth windows, where n is the number of codons of the viral gene;
[0038] The region importance calculation module: used to obtain the region importance of masking the first to the nth windows based on the difference between the second prediction accuracy and the first prediction accuracy;
[0039] The codon importance calculation module: obtains the codon importance by averaging the region importance of the windows containing each codon.
[0040] This application (fifth aspect) discloses a construction system for a segmented viral gene reassortment host adaptability prediction model, including:
[0041] The acquisition module: used to acquire a viral data set and host labels;
[0042] The reassortment module: used to filter the codon importance of different genes in the viral data set based on the codon importance threshold of different genes to obtain the key codons of different genes, and combine the digital representations of the key codons of different genes to obtain the reassortment result of the viral data set;
[0043] The training module: inputs the reassortment result of the viral data set and the host labels into a neural network for iterative training to obtain a viral gene reassortment host adaptability prediction model.
[0044] This application (sixth aspect) also discloses a segmented viral gene reassortment host adaptability prediction system, including:
[0045] The third acquisition module: used to acquire a viral data set, and the viral data set is a digital representation set of the virus;
[0046] The simulation reassortment module: used to filter the codon importance of different genes in the virus dataset based on the codon importance threshold of different genes to obtain the critical codon regions of different genes, and randomly replace the critical codons in the critical codon regions to obtain the simulation reassortment results of the virus dataset;
[0047] The third prediction module: used to input the simulation reassortment results of the virus dataset into the virus gene reassortment host adaptability prediction model to obtain the host adaptability corresponding to the simulation reassortment results.
[0048] The seventh aspect of the present application discloses a computer device, which includes: a memory and a processor; the memory is used to store program instructions; the processor is used to call the program instructions, and when the program instructions are executed, it is used to execute the steps of the above method.
[0049] The eighth aspect of the present application discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the above method.
[0050] The ninth aspect of the present application discloses a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the above method.
[0051] The present application has the following beneficial effects:
[0052] (1) The present application determines the critical region of codons by sorting the codon importance, and trains the virus gene reassortment model based on the critical codons to obtain a model that can predict the human adaptability after virus reassortment;
[0053] (2) Based on the critical codon region, the simulation reassortment of the virus is realized, and the human adaptability of the reassorted sequence is predicted based on the virus reassortment adaptability prediction model. The present application can timely identify the reassortment results that pose a greater threat to humans from multiple reassortment possibilities earlier, facilitating targeted research and prevention of the reassortment results that pose a greater threat to humans as early as possible. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0055] Figure 1 It is a schematic flowchart of the method provided in the second aspect of the embodiment of the present invention;
[0056] Figure 2It is a schematic diagram of a program product provided in the fifth aspect of the embodiments of the present invention;
[0057] Figure 3 It is a schematic diagram of a computer device provided in the embodiments of the present invention;
[0058] Figure 4 It is a schematic diagram of the architecture of an exemplary computing device provided in the embodiments of the present invention;
[0059] Figure 5 It is a schematic diagram of a storage medium provided in the embodiments of the present invention;
[0060] Figure 6 It is a flowchart of Codon2Vec sliding window characterization provided in the embodiments of the present invention;
[0061] Figure 7 It is a flowchart of Codon2Vec preprocessing and characterization evaluation provided in the embodiments of the present invention;
[0062] Figure 8 It is a framework diagram of a ResNet residual network provided in the embodiments of the present invention;
[0063] Figure 9 It is a flowchart of calculating the importance degree of codons provided in the embodiments of the present invention;
[0064] Figure 10 It is a schematic diagram of calculating the importance degree of codons by a Bayesian model provided in the embodiments of the present invention;
[0065] Figure 11 It is a schematic diagram of a flowchart of an ablation experiment provided in the embodiments of the present invention;
[0066] Figure 12 It is a schematic diagram of generating simulated reassortant data provided in the embodiments of the present invention;
[0067] Figure 13 It is a flowchart of constructing and predicting a reassortment adaptability model provided in the embodiments of the present invention. Detailed implementation manners
[0068] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0069] In some processes described in the specification, claims, and the above-mentioned drawings of the present invention, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as S101, S102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., do not represent a sequence, and do not limit that "first" and "second" are of different types.
[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0071] Figure 1 It is a schematic flowchart of a method for constructing a prediction model for the host adaptability of virus gene reassortment provided by an embodiment of the present invention. Specifically, the method includes the following steps:
[0072] S101: Obtain a virus data set and host labels;
[0073] S102: For different genes of the viruses in the virus data set, key codons of different genes are obtained after filtering the codon importance based on the codon importance thresholds of different genes, and the digitalized representations of the key codons of different genes are combined to obtain a simulated sequence of the virus data set;
[0074] S103: Input the simulated sequence of the virus data set and host labels into a neural network for iterative training to obtain a prediction model for the host adaptability of virus gene reassortment.
[0075] A gene segment of a virus refers to a nucleic acid sequence region with specific functional or structural characteristics in the virus genome, usually corresponding to encoding structural proteins of the virus (such as capsid proteins, envelope proteins) or non-structural proteins (such as replicases, regulatory proteins). Different viruses may contain multiple functionally independent gene segments according to their genomic characteristics (such as segmented or non-segmented, single-stranded or double-stranded, RNA or DNA).
[0076] Some examples of virus gene segments:
[0077] NP (Nucleoprotein) gene segment of avian influenza virus (such as H5N1): Encodes the nucleoprotein, which is responsible for encapsulating viral RNA to form ribonucleoprotein complexes (RNPs) and is a key component for viral genome replication and transcription. This gene is highly conserved and is commonly used for influenza virus typing and molecular diagnosis (such as RT-PCR detection).
[0078] HA (Hemagglutinin) and NA (Neuraminidase) gene segments of influenza virus: HA is responsible for host cell receptor binding and membrane fusion, and NA assists in virus release. These two are the core basis for influenza virus subtype classification (such as H1N1, H5N1).
[0079] gag, pol, and env gene segments of HIV: Encodes the viral capsid protein, reverse transcriptase / integrase, and envelope glycoprotein respectively, which are key targets for antiviral drug (such as reverse transcriptase inhibitors) and vaccine research.
[0080] S (Spike protein) gene segment of SARS-CoV-2: Encodes the viral spike protein, which mediates the binding to the host cell ACE2 receptor and is the core target for the development of COVID-19 vaccines (such as mRNA vaccines).
[0081] The research on gene segments provides a molecular basis for virus tracing (such as analyzing the avian influenza host source through the NP gene), diagnostic reagent development (such as antigen detection based on HA / NA), vaccine design (such as screening S protein antigenic epitopes), and antiviral drug research and development.
[0082] In some embodiments, the research aim is to provide a method for constructing a reassortment adaptability prediction model of IAV (Influenza A Virus), and use this model to predict the human adaptability of avian influenza virus (AIV) simulated reassortment sequences.
[0083] The present invention first adopts the self-created Codon2Vec characterization method, which performs encoding characterization based on sliding window codon frequency statistics and pre-training of the transformer encoding layer;
[0084] Then, based on the ResNet residual network, single-gene adaptability prediction models for four segments of PB2, PB1, PA, and NP are established, and the key codon regions in each gene are calculated through ablation experiments;
[0085] Finally, according to the calculated key codon regions, the simulation reassortment generation of avian influenza virus sequences and human influenza virus sequences is realized, and a reassortment host adaptability prediction model of IAV is constructed based on ResNet, and finally it is applied to the reassortment host adaptability prediction of simulated reassortment sequences.
[0086] First, the 8 segments of the influenza virus genome, HA, NA, M, NS, NP, PB1, PB2, and PA, encode 10 proteins: hemagglutinin (HA), neuraminidase (NA), matrix proteins M2 and M1, non-structural proteins NS1 and NS2, nucleocapsid, and three polymerases. The NP protein encoded by the 5th segment, together with the polymerase trimer PB1, PB2, PA, and viral RNA, constitutes the viral ribonucleoprotein (vRNP) of the influenza virus. The RNA of the influenza virus is transcribed and replicated in the nucleus, and one of the key processes is the nucleocytoplasmic shuttling of the ribonucleoprotein. In the early infection stage when the influenza virus invades cells, the ribonucleoprotein of the influenza virus quickly enters the nucleus to start transcription. As the main component of the vRNP complex, NP mediates the nuclear import of the vRNP complex through its nuclear localization signals (NLSs). The NP protein enters the nucleus after being translated and expressed from the viral mRNA in the cytoplasm, promoting the replication and transcription processes of the viral genome. In summary, the influenza virus NP protein is an important viral protein that determines the host range and pathogenicity of influenza A virus.
[0087] Currently, the present invention mainly conducts research on polymerase-related genes, and the single-gene host adaptability and reassortment host adaptability prediction models trained are also for these 4 genes. However, the research method in the present invention can be extended to other gene segments or even other viruses, that is, this method can be used to perform sequence characterization on other genes or viruses and train corresponding adaptability prediction models. For other gene segments of IAV, they can be generated by simulating reassortment together with polymerase-related genes, and an AI model suitable for predicting the adaptability of more gene simulations and reassortments can be trained.
[0088] The processing flow and specific operation steps of the present invention are as follows:
[0089] I. Codon2Vec Characterization Method
[0090] 1.1 (Step 1): Input the cleaned csv file containing the IAV gene sequence:
[0091] The file contains multiple lines, and each line includes the IAV virus sequence and its corresponding host label;
[0092] Avian influenza virus (Influenza A Virus, IAV) is the general term for influenza A viruses and belongs to the family Orthomyxoviridae. Its subtype classification is based on two glycoprotein antigens on the virus surface: Hemagglutinin (HA) and Neuraminidase (NA). Currently, 18 subtypes of HA (H1 - H18) and 11 subtypes of NA (N1 - N11) are known. Different combinations of HA and NA form various subtypes of IAV (such as H1N1, H3N2, etc.).
[0093] Therefore, the IAV gene sequence file contains the gene sequences of IAV of different subtypes.
[0094] The cleaned IAV gene sequence file contains the gene sequences of N virus fragments and their corresponding host tags.
[0095] N = 8, including fragments HA, NA, M, NS, NP, PB1, PB2, PA;
[0096] Since the adaptability needs to be predicted separately for different fragments, different fragments are processed separately.
[0097] First, for the first fragment, such as the NP fragment, use the method of a sliding window (window size (WS1)=128) to traverse each row of the NP fragment gene sequence, and encode the central codon within the current window using the codon frequencies in the window. Since there are 64 types of codons, each central codon is represented as a one - dimensional vector with 64 digital features.
[0098] As the window slides, encoding and representing each codon in the sequence can obtain a two - dimensional matrix of the number of codons × 64, as Figure 6 shown.
[0099] Figure 6 The gene sequence of a certain row shown as: ATG1 GAA2 AGA3 ATA4 AAA5GAA6 CTA7GAT8……TAGn
[0100] The size of the sliding window is S. When S contains 128 codons, S = 384;
[0101] Based on the current window, calculate the frequency of 64 - bit codons, count the frequency of each codon within the current window, and the calculation is as follows:
[0102]
[0103] When the window size is 128 codons, then S = 384, and the above formula corresponds to:
[0104]
[0105] condon j represents 64 codons;
[0106] Thus, the vector representation of codon frequency is obtained as:
[0107]
[0108] represents the codon frequency vector within the window, with a dimension of 64*1; thus, the central codon is encoded into a 1×64 vector,
[0109] For example, the codons contained in the 64th window are ATG1 GAA2 AGA3 ATA4 AAA5 GAA6CTA7 GAT8…GUC 63 AUA 64 AUC 65 …UUC 128 ;
[0110] The central codon in the current window is AUA 64 ,
[0111] After the above operations, as the window slides, each codon in the sequence is encoded and characterized, that is, the central codon is characterized by the codon frequency vector, and a two-dimensional matrix of the number of codons contained in the gene sequence × 64 can be obtained. That is, assuming that the number of codons contained in the gene sequence of an NP fragment is n, the dimension of the encoding feature matrix is n×64, as Figure 6 shown.
[0112] Since we need to encode each codon, therefore, when the central codon of the window is ATG1, the sequence of the window is obtained by splicing the codons at the end of the sequence, and then the codon frequency vector is calculated. When the central codon is the last codon, the window sequence is supplemented by intercepting the codons in the first half of the sequence and then the codon frequency is calculated.
[0113] The determination of the sliding window is shown in the following pseudocode:
[0114] Table 1 Logic of Encoding Codons with a Sliding Window
[0115]
[0116]
[0117] Due to the input size requirements of the subsequent ResNet framework, zero vector padding is performed at the end of the feature matrix. For example, the size of the feature matrix on the PB2 gene is 760×64, while the input parameters of the subsequent ResNet model framework are 128×128×3. Therefore, the size of the feature matrix on the PB2 gene is adjusted to 768×64, that is, an 8×64 zero vector is added at the end of the original feature matrix. Finally, the characterization matrices of the same gene of multiple different subtype virus sequences are stored in an npy file.
[0118]
[0119] Among them, represents the codon frequency vector when the first codon in the sequence is the central codon; ORF represents the gene fragment sequence, and A(ORF) represents the encoded feature matrix after encoding.
[0120] In some embodiments, for effective encoding, that is, to reduce the 0 values in the codon frequency vector, it is recommended that the window width be greater than 64.
[0121] In some embodiments, the window width of the sliding window is 128 codons.
[0122] In some embodiments, the window width of the sliding window is 100 codons.
[0123] In some embodiments, the prediction network model used replaces the RestNet network with any one of the following: DenseNet, ResNeXt, Inception-ResNet, MobileNetV2, EfficientNet, Wide ResidualNetworks (WRN). When applying, the dimension of the transformation matrix is changed based on the requirements of the network for the input, or the window width of the sliding window is changed.
[0124] 1.2 (Step 2): Input the npy file obtained in Step 1, and use the transformer encoding layer to perform pre-training on it to obtain the optimized feature encoding (with the same size as in Step 1), and store it in another npy file. Among them, the Transformer pre-training performs an autoencoding task to reconstruct the input data using unsupervised training to obtain the optimized feature encoding, Figure 7 The A’(ORF) shown.
[0125] 1.3 (Step 3): Input the two npy files obtained in Step 1 and Step 2. Use two methods, UMAP dimensionality reduction and hierarchical clustering in unsupervised learning, to evaluate the representation effect of the representation matrices in these two npy files (i.e., evaluate the representation effect of the pre-training of the transformer encoding layer), and use external evaluation indicators of clustering, the adjusted rand index (ARI) and the normalized mutual information (NMI), to evaluate the clustering effect. The final evaluation result is the dimensionality reduction visualization result and the clustering effect score of the two representation matrices.
[0126] According to the evaluation results, store the optimal npy file of the corresponding encoding matrix of the training set, that is, the optimal encoding feature matrix (if the pre-training effect of the transformer encoding layer is not good on this segment, the original representation matrix is used for representation, that is, the npy file obtained in Step 1; otherwise, the representation matrix after the pre-training of the transformer encoding layer is used, that is, the npy file obtained in Step 2). The processing flow of Steps 2 - 3 is as Figure 7 shown.
[0127] In some embodiments, only the clustering evaluation index is used to select the optimal encoding feature matrix.
[0128] In some embodiments, only the dimensionality reduction visualization result is used to select the optimal encoding feature matrix.
[0129] In some embodiments, the optimal encoding feature matrix is selected by synthesizing multiple evaluation results. The selection process is as Figure 7 shown. Update the value of c according to different indicators, and select A (ORF) or A' (ORF) as the optimal encoding feature matrix according to the value of c.
[0130] In some embodiments, the following subsequent steps are performed based on the encoding feature matrix obtained by the sliding window.
[0131] In some embodiments, the following subsequent steps are performed using the encoding feature matrix after Transformer auto-encoding.
[0132] In some embodiments, the following subsequent steps are performed using the optimal encoding feature matrix.
[0133] Unless otherwise specified, the following encoding feature matrix A (ORF) is any one of the following: the encoding feature matrix, the encoding feature matrix output by the Transformer, and the optimal encoding feature matrix.
[0134] II. Establish a ResNet single-segment adaptability prediction model
[0135] 2.1 (Step 4): Use the ResNet residual network framework (as shown in Figure 8 ) to train the single-gene host adaptability prediction model. Among them, the label used for training is the original host label of the sequence. If its host is avian, the model label is "0"; if its host is human, the model label is "1". Save the trained model (i.e., the single-fragment host adaptability prediction model), which can be used to predict the host adaptability of IAV with unknown host.
[0136] That is, for different fragments, repeat Steps 1-4 to obtain the single-fragment adaptability prediction models of different fragments respectively.
[0137] In some embodiments, for 4 gene sequences (PB2, PB1, PA, and NP), Steps 1-4 need to be run independently, and finally 4 single-gene adaptability prediction models are obtained.
[0138] In some embodiments, for 8 gene sequences, run Steps 1-4 independently, and finally 8 single-gene adaptability prediction models are obtained.
[0139] III. Ablation experiment / Bayesian model calculates the importance order of codons in each gene fragment
[0140] 3.1 (Step 5): Use the trained single-fragment host adaptability prediction model to predict the encoded feature matrix of the test set (the genes of the test set are processed in the same way to obtain the encoded feature matrix of the test set), and obtain the current prediction accuracy (acc1), which will be used later to observe the impact of the mask on the accuracy in the ablation experiment (Step 6).
[0141] Step 6: To evaluate the importance of each codon in the gene fragment, perform a sliding window mask (window size = 80) along the row direction of the encoded feature matrix of the test set (for example, if the gene fragment contains 700 codons, the dimension of the encoded feature matrix is: 700×64), that is, cover the encoded sequences of the corresponding 80 codons within the sliding window, that is, mask the codon frequency vector representing the codon within the sliding window area as a zero vector, because the codon frequency vector within the window is a 0 vector, and the codon mask is as shown in the black mask in the matrix in Figure 10 or as shown in the masked encoded feature matrix in Figure 11 , and some of its vectors are represented as 0 vectors; moreover, when performing a sliding window mask on the encoded feature matrix, the judgment of the corresponding masked area is shown in Table 2.
[0142] Table 2 Judgment logic (pseudo-code) for sliding window masking the encoded feature matrix
[0143]
[0144]
[0145] The codon feature vector within the first window of the mask is a zero vector, obtaining the encoded feature matrix A(IRF) of the first window of the mask mask1 ,
[0146] For the second window, obtain A(ORF) mask2 ,
[0147] For the nth window, obtain A(ORF) maskn ,
[0148] Finally, for a gene fragment containing n codons, n masked encoded feature matrices are obtained. After inputting these n masked encoded feature matrices into the pre-trained Transformer representation matrix, optimized feature encoding is obtained. After obtaining the optimal feature matrix in the manner of step 3, it is input into the trained single-gene / fragment host adaptability model for prediction, and the accuracy rate (acc2, where acc2 is a list) corresponding to each masked feature matrix is obtained. By comparing it with acc1, the difference between the two can reflect the importance degree of the codon vector in the current masked region. If the difference between acc2 and acc1 is large, it indicates that the masking of 80 codons in this part of the vector has a greater impact on the accuracy rate, that is, the importance degree of the codon vector representation in this part is high
[0149] Then, slide the window. As the window slides, according to the difference between acc2 and acc1 predicted after masking the vectors in the corresponding region, the importance degree of each block of region vectors is calculated in turn. Since the current obtained importance degree is for the vectors within an entire masked region, the importance degree of the codon frequency vectors within the region is averaged according to the window size (80) to obtain the importance degree (Vector importance, VI) of each vector. Then, according to the sliding window size (128) in the representation process of step 1, the importance degree of each vector is averaged to obtain the importance degree (Codon importance, CI) of each codon in the sequence - CI1. The specific details are as Figure 9 , Table 3 shows
[0150] Table 3 Calculation process of codon importance degree
[0151]
[0152]
[0153]
[0154] where acc nRepresents the accuracy rate of the nth sliding window position, VI n Represents the importance degree of the nth vector, CI n Represents the importance degree of the nth codon.
[0155] 3.2 (Step 7:) Use the Bayesian model to calculate the importance degree (CI2) of each codon:
[0156] Such as Figure 10 As shown, after encoding the feature matrix and inputting it into the single-fragment host adaptability prediction model, the accuracy rate acc1 is obtained;
[0157] After performing a sliding window mask on the encoded feature matrix, the encoded feature matrix of the i-th window of the mask is obtained in sequence and then input into the single-fragment host adaptability prediction model to obtain the accuracy rate acc2;
[0158] Among them, acc2 is a list, and the elements in the list respectively represent the accuracy rates obtained when masking the i-th window:
[0159] acc2 = [acc21, acc22, …, acc2 i , …, acc2 n
[0160] Among them, acc2 i Represents the accuracy rate obtained by inputting the encoded feature matrix of the i-th window of the mask into the single-fragment host adaptability prediction model; substituting acc2 into the Bayesian model to obtain the Bayesian accuracy rate acc3, that is, P(B i |A)
[0161]
[0162] Among them, P(B i |A) represents the probability of B i occurring based on condition A, where event A represents the prediction accuracy rate; B i represents the event of the encoded feature matrix of the i-th window of the mask;
[0163] P(B i ) represents the probability of the encoded feature matrix of the i-th window of the mask appearing. In our scenario, since we use a sliding window mask with a step size of 1 codon, therefore, n is the number of codons;
[0164] A represents the prediction accuracy rate, A|B i represents the prediction accuracy rate obtained when masking the i-th window, P(A|B i ) represents the probability of the feature A matrix under the condition of event B i , which is acc2 i here;
[0165] Therefore, the posterior probability of acc2 (post-acc2), i.e., acc3, is expressed as:
[0166]
[0167] acc2 i represents the accuracy obtained by inputting the encoded feature matrix of the i-th window of the mask into the single-fragment host adaptability prediction model;
[0168] acc3 i represents the posterior accuracy obtained by inputting the encoded feature matrix of the i-th window of the mask into the single-fragment host adaptability prediction model;
[0169] acc3 = [acc31, acc32, …, acc3 i , …, acc3 n
[0170] The posterior importance of the mask window is obtained based on the difference between acc3 and acc1: that is, the acc-list in Table 3 is replaced with acc3 - acc1, and then CI2 is calculated based on the codon importance calculation process shown in Table 3.
[0171] Step 8: Take the average of the codon importance levels calculated in Steps 6 and 7. As shown in Formula 1, obtain the importance of each final codon. As Figure 10 shown, normalize it and save it to a csv file.
[0172]
[0173] For the ablation experiment of 4 gene fragments, Steps 5 - 8 need to be run independently, and finally the codon importance results on the 4 gene fragments are obtained.
[0174] In some embodiments, the codon importance is obtained only using the CI calculated in Step 6.
[0175] In some embodiments, the codon importance is obtained using the Bayesian-corrected CI2.
[0176] In some embodiments, the codon importance is obtained by taking the mean or weighted combination of the codon importance level (CI1) in Step 6 and the Bayesian-corrected CI2.
[0177] IV. Key Codon Combinations
[0178] Step 9: According to the evaluation results in Step 3, input the optimal encoded feature matrix of the 4-fragment training set at the same time. According to the csv file of the codon importance results obtained in Step 8, intercept and combine the key codons that are important for adaptability in the 4 fragments.
[0179] That is, according to the codon importance calculated from the 4 genes, compare it with the key codons that can be found in the literature. On the premise of including as many key codons found in the literature as possible, select the codon importance threshold, intercept the vector representation of the codons greater than the importance threshold from the characterization matrix, and arrange them in sequence according to their order in this fragment to obtain the characterizations of the key codons of the 4 genes respectively. Finally, merge the 4 fragments in the order of PB2, PB1, PA, and NP. Different importance thresholds can be selected as needed, but after the threshold is selected, there is only one possible splicing result.
[0180] According to the requirements of the subsequent model size, perform size filling on the above merged characterization matrix (the same as step 1) to obtain the coding matrix of the key codons of the 4 genes representing the training set, as Figure 12 shown in the upper part: It can be seen that for the key codons of the 4 genes (PB2, PB1, PA, NP), the different shades of green in the figure represent different codon importance thresholds. Taking the strictest codon importance threshold as an example,'simultaed ORF’ represents the simulated sequence. It can be seen that the simulated sequence is obtained by concatenating the key codon regions of PB2, PB1, PA, and NP.
[0181] Based on this characterization and the ResNet residual network framework in step 4, train the reassortment adaptability prediction model (here, the self-concatenation form of the 4-gene fragment characterizations of the same strain virus in the training set is adopted, and its label is the same as the original strain), as Figure 13 shown, for predicting the reassortment host adaptability of avian influenza viruses after 2020.
[0182] Step 10: Since the present invention predicts the reassortment host adaptability of reassortment sequences that do not exist in nature at present, it is necessary to generate the characterizations of all possible simulated reassortment sequences of avian influenza according to the existing avian influenza virus sequences.
[0183] In some embodiments, input the 4-fragment coding feature matrix of the prediction set. According to the codon importance result csv file obtained in step 8, respectively intercept the key codon characterizations of the 4 fragments, and randomly replace the key codon coding matrix of the 4 sequence fragments of the selected human influenza virus reference strain, and then generate the feature representation of the reassortment simulated sequence, as Figure 12 shown, and save it to an npy file.
[0184] In some embodiments, Figure 12The simulated reassortments show a schematic diagram of simulated reassortment, where the green represents the IAV gene sequences detected in birds or pigs in the training set; the sky blue represents the gene sequences of human H1N1, and the dark blue represents the gene sequences of H3N2 with human hosts;
[0185] PB2’ is a short fragment of key codons that uses the key codons calculated in the PB2 gene to replace the entire PB2 gene fragment of the virus strain;
[0186] PB1’ is a short fragment of key codons that uses the key codons calculated in the PB1 gene to replace the entire PB1 gene fragment of the virus strain;
[0187] PA’ is a short fragment of key codons that uses the key codons calculated in the PA gene to replace the entire PA gene fragment of the virus strain;
[0188] NP’ is a short fragment of key codons that uses the key codons calculated in the NP gene to replace the entire NP gene fragment of the virus strain;
[0189] After random replacement, the gene sequences of simulated reassortment are obtained.
[0190] In some embodiments, the possibility of virus reassortment is expressed as:
[0191]
[0192] Where R n represents the number of possibilities of reassortment, in which 4 represents 4 segments for reassortment, represents replacing 1 gene of the human prototype virus strain, represents replacing 2 genes of the human prototype virus strain, represents replacing 3 genes of the human prototype virus strain.
[0193] For different genes of the viruses in the virus dataset, after filtering the codon importance based on the codon importance threshold of different genes, the key codon regions of different genes are obtained. After randomly replacing the key codons in the key codon regions, the simulated reassortment results of the virus dataset are obtained; the method of randomly replacing the key codons in the key codon regions is:
[0194] According to the codon importance of different genes, the digital characterizations of the key codons of different genes constituting the virus are intercepted respectively, and the digital characterization of the reassortment simulation sequence is obtained by randomly replacing them with the digital characterizations of the key codons of the corresponding genes from different viruses; the corresponding genes here refer to: according to the codon importance calculated on each gene, intercept the digital characterizations of the key codons on the corresponding gene. For example, if codons 1, 2, and 3 are important on the PB2 gene, then intercept the characterizations of codons 1, 2, and 3, and if codons 4, 5, and 6 are important on the PB1 gene, then intercept the characterizations of codons 4, 5, and 6. Randomly replace the important codons of 1, 2, or 3 genes, that is, use the characterizations of codons 1, 2, and 3 on the PB2 gene of virus A to replace the characterizations of codons 1, 2, and 3 on the PB2 gene of virus B, which is to randomly replace the key codons of 1 gene.
[0195] In some embodiments, the steps of simulating reassortment are as follows: Randomly replace 1, 2, or 3 short fragments of key codons of four gene segments (PB2, PB1, PA, NP) from human H1N1 or H3N2 prototype virus with the corresponding short fragments of key codons of PB2, PB1, PA, NP of avian virus subtypes to obtain the gene sequence of simulated reassortment.
[0196] In some embodiments, select the key codon regions of PB2, PB1, PA, NP, and randomly combine the screened key codons to obtain the result of simulated reassortment.
[0197] The above simulated reassortment is based on the study of 4 gene segments in this research. When the research involves 8 genes, randomly replace any 1-7 short fragments of key codons of 8 gene segments (HA, NA, M, NS, NP, PB1, PB2, PA) from human H1N1 or H3N2 prototype virus with the corresponding short fragments of key codons of the corresponding gene segments of avian virus subtypes to obtain the gene sequence of simulated reassortment.
[0198] Step 11: Input the npy file obtained in step 10, and use the reassortment adaptability prediction model trained in step 9 to perform simulated reassortment host adaptability prediction, and obtain the adaptability prediction result of simulated reassortment, that is, the predicted host label (adapt to humans or adapt to birds) of the simulated reassortment sequence ( Figure 13 ). Analyze the prediction results of each simulated reassortment, and it can be concluded which fragments of the avian influenza virus have a higher risk of reassortment and adaptation to humans, and which subtypes have a higher risk of simulated reassortment and adaptation to humans, providing a direction for the prevention and control of influenza A virus. At the same time, this method can be applied to the reassortment adaptability prediction problems of other viruses, or extended to the reassortment monitoring of more gene segments.
[0199] In the reassortment adaptability prediction problems of some other viruses, if the segments they contain are different from those of IAV, after collecting data, the above-mentioned method is used to predict the adaptability of the segments they contain, determine the importance ranking of segment codons, screen out the key codons, and then train the reassortment adaptability prediction model of the virus, and simulate the adaptability of different reassortment results after reassortment, and carry out targeted prevention.
[0200] Figure 3 is a schematic diagram of a computer device provided by an embodiment of the present invention, as Figure 3 shown, the device 2000 may include: one or more processors 2010, and one or more memories 2020; wherein, computer-readable code is stored in the memory, and when the computer-readable code is run by the one or more processors, the above-mentioned method can be executed.
[0201] The processor in this embodiment may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, operations and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc., and may be of the X86 architecture or the ARM architecture.
[0202] Generally speaking, the various example embodiments of the present disclosure may be implemented in hardware or special circuits, software, firmware, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices. When the aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, technologies, or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, special circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.
[0203] For example, the method or device according to the embodiments of the present disclosure may also be implemented by means of Figure 4 the architecture of the computing device 3000 shown. As Figure 4As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. Storage devices in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for the processing and / or communication of the methods provided by the present disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 4 the architecture shown is merely exemplary, and when implementing different devices, one or more components in the shown Figure 4 computing device may be omitted according to actual needs.
[0204] An embodiment of the present invention also provides a computer-readable storage medium, such as Figure 5 shown, which is a schematic diagram of the storage medium 4000 provided by an embodiment of the present invention. Computer-readable instructions 4010 are stored on the computer storage medium 4020. When the computer-readable instructions 4010 are run by a processor, the methods according to the embodiments of the present disclosure described with reference to the above figures may be executed. The computer-readable storage medium in the embodiments of the present disclosure may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memories for the methods described herein are intended to include, but are not limited to, these and any other suitable types of memories. It should be noted that the memories for the methods described herein are intended to include, but are not limited to, these and any other suitable types of memories.
[0205] An embodiment of the present disclosure also provides a computer program product or a computer program, which, when executed by a processor, implements the steps of the above method, such as Figure 2 shown, the computer program product or the computer program includes:
[0206] Acquisition module 201: Acquire a virus dataset and host labels;
[0207] Reconfiguration module 202: For different genes of the viruses in the virus dataset, after filtering the codon importance based on the codon importance threshold of different genes, obtain the key codons of different genes, and combine the digital representations of the key codons of different genes to obtain the reconfiguration result of the virus dataset;
[0208] Training module 203: Input the reconfiguration result of the virus dataset and the host labels into a neural network for iterative training to obtain a virus gene reconfiguration host adaptability prediction model.
[0209] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0210] Generally speaking, the various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.
[0211] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0212] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0213] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0214] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0215] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art should understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A method for ranking the codon importance of a viral gene, characterized in that, The method includes: S1: Obtain a test data set representing viral genes and host labels, where each gene data in the test data set is a digital representation of each codon of the gene; S2: Perform viral gene host adaptability prediction based on the test data set to obtain a predicted host label. After comparing the predicted host label with the host label, obtain the first prediction accuracy rate of the test set; S3: Use the sliding window method to mask the digital representation of the codons in the first window of the test data set to obtain a first masked test data set. Perform viral gene host adaptability prediction based on the first masked test data set to obtain the predicted host label of the masked first window. After comparing the predicted host label of the masked first window with the host label, obtain the prediction accuracy rate of the masked first window; S4: Slide the window to sequentially mask the digital representations of the codons in the first to the nth windows, and then repeat the steps of S3 to obtain a second prediction accuracy rate. The second prediction accuracy rate includes the prediction accuracy rates of masking the first to the nth windows, where n is the number of codons of the viral gene; S5: Obtain the regional importance of masking the first to the nth windows based on the difference between the second prediction accuracy rate and the first prediction accuracy rate; S6: Average the regional importance of the windows containing each codon to obtain the codon importance.
2. The method for ranking the codon importance of a viral gene according to claim 1, characterized in that Mask the data of the test set by sliding the window after connecting the gene codons end to end. Here, the end refers to the digital representation of the last codon; Optionally, the method for digital representation of codons is: Connect the gene sequences of a single gene end to end and use the sliding window method to traverse the gene sequence. The frequencies of 64 different codons in the window form a window codon frequency vector, and use the window codon frequency vector to represent the central codon of the window. As the window slides, sequentially represent each codon of the gene sequence as the window codon frequency vector when this codon is the central codon; Optionally, the gene data is a coding feature matrix composed of the window codon frequency vectors obtained by sliding window traversal; Optionally, use a viral gene host adaptability prediction model to perform viral gene host adaptability prediction. The viral gene host adaptability prediction model is trained by a training data set of viral genes; Optionally, S3 further includes: Obtain the Bayesian posterior accuracy rate of the masked first window based on the prediction accuracy rate of the masked first window and the probability of masking each window; Repeat the steps of S3 to obtain a third prediction accuracy rate. The third prediction accuracy rate includes the Bayesian posterior accuracy rates of masking the first window to the nth windows. Obtain the Bayesian regional importance of the first to the nth windows based on the difference between the third prediction accuracy rate and the first prediction accuracy rate; Average the Bayesian regional importance of the windows containing each codon to obtain the Bayesian posterior codon importance; Optionally, obtain the optimal codon importance based on the weighted average of the codon importance and the Bayesian posterior codon importance.
3. A method for constructing a prediction model of host adaptability for segmented virus gene reassortment, characterized in that, The method includes: Obtain a viral data set and host labels; For different genes of the viruses in the virus dataset, after filtering the codon importance based on the codon importance thresholds of different genes, the key codons of different genes are obtained, and after combining the digital characterizations of the key codons of different genes, the reassortment result of the virus dataset is obtained; wherein the codon importance is obtained based on the method described in any one of claims 1-2. The reassortment result of the virus dataset and the host label are input into a neural network for iterative training to obtain a virus gene reassortment host adaptability prediction model.
4. A method for constructing a prediction model for the host adaptability of segmented virus gene reassortment, characterized in that, The method includes: Obtain a virus dataset and a host label; For different genes of the viruses in the virus dataset, after filtering the codon importance based on the codon importance thresholds of different genes, the key codons of different genes are obtained, and after combining the digital characterizations of the key codons of different genes, a simulated sequence of the virus dataset is obtained; The simulated sequence of the virus dataset and the host label are input into a neural network for iterative training to obtain a virus gene reassortment host adaptability prediction model.
5. The construction method of the segmented virus gene reassortment host adaptability prediction model according to claim 4, characterized in that, The method includes: The way of combining the digital characterizations of the key codons of different genes is to merge the digital characterizations of the key codons of the genes in the order of the genes in the virus. Optionally, the codon frequency vectors of the key codons of different genes are merged in the order of the genes in the virus to obtain a simulated sequence, and the simulated sequence is a matrix.
6. A method for predicting the adaptability of segmented virus genome simulation reassortment, characterized in that, The method includes: Obtain a virus dataset, where the virus dataset is a digital characterization set of viruses; For different genes of the viruses in the virus dataset, after filtering the codon importance based on the codon importance thresholds of different genes, the key codon regions of different genes are obtained, and after randomly replacing the key codons in the key codon regions of K genes respectively, a simulated reassortment dataset of the virus dataset is obtained; The simulated reassortment result of the virus dataset is input into the virus gene reassortment host adaptability prediction model described in any one of claims 3-5 to obtain the host adaptability corresponding to the simulated reassortment result.
7. The method for predicting the adaptability of segmented virus genome simulation reassortment according to claim 6, characterized in that The method of randomly replacing the key codons in the key codon regions is: According to the codon importance of different genes, the digital characterizations of the key codons of K different genes constituting the virus are respectively intercepted, and the digital characterizations of the key codons of K corresponding genes from different viruses are used for random replacement to obtain the digital characterization of the reassortment simulated sequence; Optionally, the different viruses are viruses of different subtypes; Optionally, the key codon short fragments in the K gene fragments of the influenza virus from humans are randomly replaced with the corresponding key codon short fragments of the corresponding genes of the avian virus subtype to obtain a simulated reassortment result; Optionally, the value range of K is 1-7.
8. A computer device, characterized in that, The device includes: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by the processor, the steps of the method described in any one of claims 1-7 are implemented.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Non-evolutionary tree-dependent segmented RNA virus reconfiguration method
CN115910377A
Model training and gene expression optimization method, device, equipment and medium
CN117976048A
Virus host prediction method and system based on neural network, and storage medium
CN117995267A
Virus host prediction method and system based on fuzzy control optimization and medium
CN118173164A
Construction method of virus single-gene host adaptability prediction model, equipment and medium
CN120356519A