Construction method of virus single-gene host adaptability prediction model, equipment and medium

By constructing a single-gene host adaptive prediction model for viruses, using sliding window method and neural network technology, the problem of host adaptive prediction of influenza viruses is solved, efficient adaptive analysis of single-gene influenza viruses and identification of key codons is achieved, and influenza virus prevention and control is supported.

CN120356519AActive Publication Date: 2025-07-22ACADEMY OF MILITARY MEDICAL SCIENCES

Patent Information

Application Number
CN202510518036.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-22
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively monitor and predict the adaptability of influenza viruses to different hosts, especially the adaptive mutation and transmission ability of highly pathogenic avian influenza viruses to humans, affecting the study of influenza virus prevention and control.

Method used

By constructing a single-gene host adaptability prediction model for viruses, using sliding window method to characterize the codon frequency of gene sequences, combining neural network model training, predicting the host adaptability of single-gene viruses, and optimizing the feature matrix through self-coding learning and dimensionality reduction technology to screen out key codons.

Benefits of technology

It realizes efficient adaptability analysis of single genes of influenza viruses, can identify key codons and predict their host adaptability, and provides a scientific basis for influenza virus prevention and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356519A_ABST
    Figure CN120356519A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of intelligent medical treatment, and particularly relates to a construction method of a virus single-gene host adaptability prediction model, equipment and a medium. The method comprises the steps that a training set virus single gene data set and a host tag are obtained, single genes in the virus single gene data set are sequentially encoded to obtain an encoding feature matrix set, and encoding comprises the steps that a sliding window method is used for traversing a gene sequence, frequencies of 64 different codons in a window form a window codon frequency vector, and the window codon frequency vector is used for encoding the host tag; representing a central codon of a window by using the window codon frequency vector, and representing each codon of a gene sequence as a window codon frequency vector in sequence; and inputting the coding feature matrix set into a residual network model to obtain a prediction label, and obtaining a single-gene host adaptability prediction model based on comparison iteration training of the prediction label and a host label. According to the method, the model capable of predicting the gene adaptability is obtained by training after the codon is represented, and the gene codon adaptability can be analyzed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent medicine, and more particularly, to a method, device, medium, and program product for constructing a prediction model for the host adaptability of a single gene of a virus. Background Art

[0002] Influenza virus, the representative species of the Orthomyxoviridae family, is abbreviated as the flu virus, including human influenza virus with humans as the host and animal influenza virus with animals as the host. Influenza viruses are divided into four types: A, B, C, and D. Among them, influenza A and B viruses are the pathogens of influenza, which can cause acute respiratory infectious diseases. The main symptoms after infection are cough, high fever, runny nose, muscle pain, etc. A small number are accompanied by severe pneumonia, and severe cases can lead to multiple organ failures such as the heart and kidneys, inducing death. The fatality rate caused by highly pathogenic avian influenza virus infection is very high. The antigenicity of influenza A virus is prone to mutation, causing worldwide pandemics many times. The segmented genome structure of influenza virus makes it easy for different viruses to undergo genomic segment reassortment during co-infection. At the same time, the surface proteins of influenza virus are under the selective pressure of the population immunity of the host, resulting in antigenic drift, causing the repeated epidemics of influenza A or B viruses. The host adaptability of influenza virus is a key scientific issue in its prevention and control research.

[0003] The 8 segments of the influenza virus genome, HA, NA, M, NS, NP, PB1, PB2, and PA, respectively encode 10 proteins: hemagglutinin (HA), neuraminidase (NA), matrix proteins M2 and M1, non-structural proteins NS1 and NS2, nucleocapsid, and three polymerases. In the past two decades, in the research on virus adaptability, the definition of adaptability at least includes that the virus can infect, cause disease, and be easily transmitted to a certain type of host. High adaptability means that the virus can spread more effectively in the host population. The high infectivity, high transmissibility, and low pathogenicity of the virus match its high adaptability, and vice versa. So far, the IAVs adapted to humans mainly include H3N2 and H1N1 subtypes. Avian influenza virus subtypes such as H5N1 and H7N9 occasionally infect humans, but cannot be transmitted between humans and do not have human adaptability. The three swine influenza virus subtypes H1N1 (classical and avian-like H1N1 in pigs), H3N2 (human-like and reassortant H3N2), and H1N2 that are widely prevalent in the world can also occasionally infect humans, but do not have the ability to spread effectively among the population. After continuous passage of avian influenza virus (AIV) through mammalian cells or mammals, the phenotypes such as replication ability and polymerase activity are significantly enhanced, improving its adaptability to mammals (cells). Summary of the Invention

[0004] In view of the above problems, the present invention provides a method for constructing a virus single-gene host adaptability prediction model. After codon characterization of the gene, a training set is constructed with a feature coding matrix to train the single-gene host adaptability prediction model, effectively monitoring the variability of a single gene in the virus.

[0005] This application (in the first aspect) discloses a method for constructing a virus single-gene host adaptability prediction model, including:

[0006] Obtaining a training set of virus single-gene data sets and host labels,

[0007] Encoding the single genes in the virus single-gene data set in sequence to obtain a set of encoded feature matrices. The encoding is as follows: After connecting the head and tail of the gene sequence of the single gene, the sliding window method is used to traverse the gene sequence. The frequencies of 64 different codons within the window form a window codon frequency vector, and the central codon of the window is characterized by the window codon frequency vector. As the window slides, each codon of the gene sequence is sequentially characterized as the window codon frequency vector when this codon is the central codon. After the traversal ends, an encoded feature matrix composed of window codon frequency vectors is obtained;

[0008] Inputting the set of encoded feature matrices into a neural network model to obtain predicted labels, and iteratively training based on the comparison between the predicted labels and the host labels to obtain a virus single-gene host adaptability prediction model.

[0009] Further, the method includes: The central codon of the sliding window represents the codon located at the center position of the sliding window;

[0010] Optionally, the width of the sliding window is S codons. If S is an odd number, the central codon is the (S + 1) / 2-th codon in the window. If S is an even number, the central codon is the S / 2-th codon or the S / 2-th codon in the window;

[0011] Optionally, when the index of the central codon in the gene sequence is less than S / 2 or, the codons missing from the sliding window are supplemented with the codons at the end of the gene sequence. When the index of the central codon in the gene sequence is greater than n - S, the codons missing from the sliding window are supplemented with the codons at the starting position of the gene sequence, where n represents that the gene sequence contains n codons in total;

[0012] The 64 codons are the 64 codons of the standard genetic code;

[0013] Optionally, the dimension of the encoded feature matrix is n×64, where n represents that the gene sequence contains n codons in total;

[0014] Optionally, the first row of the encoded feature matrix is the window codon frequency vector when the first codon of the gene sequence is the central codon, and the last row is the window codon frequency vector when the last codon of the gene sequence is the central codon;

[0015] Optionally, the step size of the sliding window method is 1 codon.

[0016] Furthermore, perform autoencoding learning on the set of encoded feature matrices to obtain a second set of encoded feature matrices, correct the encoded feature matrices in the second set of encoded feature matrices to the same dimension as the encoded feature matrices by padding with 0s to obtain a third set of encoded feature matrices, and input the third set of encoded feature matrices and the host labels into the neural network model for iterative training to obtain a virus single-gene host adaptability prediction model;

[0017] Optionally, perform autoencoding learning based on one or more of the following models: LTSM, Transformer, Transformer-XL, MLP, MLP-Mixer.

[0018] Optionally, perform autoencoding learning on the set of encoded feature matrices to obtain a second set of encoded feature matrices and simultaneously obtain a trained autoencoder;

[0019] Optionally, if the encoded feature matrix does not match the input dimension of the autoencoding model, pad with 0s at the end to meet the requirements;

[0020] Optionally, evaluate the representation effect of the matrices in the set of encoded feature matrices and the third set of encoded feature matrices, and select the matrix with good representation effect as the optimal encoded feature matrix. The optimal encoded feature matrices of the training set virus single genes constitute the optimal encoded feature matrix set, and input the optimal encoded feature matrices and the host labels into the neural network model for iterative training to obtain a virus single-gene host adaptability prediction model;

[0021] Optionally, the representation effect evaluation is one or more of the following: clustering evaluation, umap dimensionality reduction evaluation;

[0022] Furthermore, the training set virus single genes are any one of the following genes: HA, NA, M, NS, NP, PB1, PB2, PA, and the corresponding virus single-gene host adaptability prediction models are obtained after training.

[0023] This application (second aspect) discloses a method for predicting the host adaptability of virus single genes, including:

[0024] Obtain the encoded feature matrix of the virus single gene;

[0025] Input the encoded feature matrix into the viral single-gene host adaptability prediction model constructed by using the construction method of any one of the viral single-gene host adaptability prediction models to obtain the host adaptability of the viral single-gene.

[0026] This application (third aspect) discloses a method for ranking the importance of codons in a viral gene sequence, including:

[0027] S1: Obtain the encoded feature matrix set and host labels of the test set viral single genes;

[0028] S2: Input the encoded feature matrix set of the test set into the single-gene host adaptability prediction model constructed by using the construction method of the viral single-gene host adaptability prediction model to obtain the predicted host labels. After comparing the predicted host labels with the host labels of the test set, obtain the first prediction accuracy rate;

[0029] S3: Use the sliding window method to mask the first window of the encoded feature matrix set of the test set to obtain the masked first window encoded feature matrix set. The central row of the first window is the first row of the encoded feature matrix set. Input the masked first window encoded feature matrix set into the same single-gene host adaptability prediction model and calculate the prediction accuracy rate of the masked first window. After the sliding window slides and traverses sequentially along the rows of the encoded feature matrix, repeat S3 to obtain the second prediction accuracy rate. The second prediction accuracy rate includes the prediction accuracy rates of the masked first window to the nth window, where n represents the number of codons of the single gene;

[0030] S4: Obtain the regional vector importance of the first to the nth windows based on the difference between the second prediction accuracy rate and the first prediction accuracy rate;

[0031] S5: Average the regional vector importance of each codon to obtain the importance of each codon.

[0032] Furthermore, the encoded feature matrix set of the test set is the optimized encoded feature matrix set obtained by inputting the encoded feature set of the test set into the autoencoder constructed by the autoencoding learning of the encoded feature matrix set of the training set;

[0033] Optionally, the encoded feature matrix set of the test set is a set composed of the optimal encoded feature matrices obtained by evaluating the characterization effects of the matrices in the encoded feature matrix set and the optimized encoded feature matrix set;

[0034] Optionally, S3 further includes: obtaining the Bayesian posterior accuracy of the first masked window based on the prediction accuracy of the first masked window and the probability of each masked window; repeating step S3 to obtain the third prediction accuracy, where the third prediction accuracy includes the Bayesian posterior accuracies of the first to the nth masked windows, and obtaining the Bayesian region vector importance of the first to the nth windows based on the difference between the third prediction accuracy and the first prediction accuracy; averaging the Bayesian region vector importance of each codon to obtain the Bayesian posterior codon importance of each codon;

[0035] Optionally, obtaining the optimal codon importance vector based on the weighted average of the codon importance and the Bayesian posterior codon importance;

[0036] Optionally, screening the key codons of each fragment based on the set importance threshold;

[0037] Optionally, the Bayesian posterior accuracy of the first masked window obtained based on the prediction accuracy of the masked ith window and the probability of each masked window is expressed as:

[0038]

[0039] where P(B i |A) represents the Bayesian posterior accuracy of the ith masked window; B i represents the event of masking the ith window; P(B i ) represents the probability of masking the ith window; P(A|B i ) represents the prediction accuracy of the ith masked window.

[0040] The fourth aspect of the present application discloses a method and system for constructing a virus single-gene host adaptability prediction model, including:

[0041] An acquisition module: used to acquire a training set of virus single-gene data sets and host labels;

[0042] A feature encoding module: used to encode the single genes in the virus single-gene data set in sequence to obtain an encoded feature matrix set. The encoding is as follows: concatenating the head and tail of the gene sequence of the single gene and traversing the gene sequence using the sliding window method. The frequencies of 64 different codons within the window form a window codon frequency vector, and the central codon of the window is characterized by the window codon frequency vector. As the window slides, each codon of the gene sequence is sequentially characterized as the window codon frequency vector when the codon is the central codon. After the traversal, an encoded feature matrix composed of window codon frequency vectors is obtained;

[0043] Training module: configured to input the set of encoded feature matrices into a neural network model to obtain predicted labels, and iteratively train based on the comparison between the predicted labels and the host labels to obtain a virus single-gene host adaptability prediction model.

[0044] This application (fifth aspect) discloses a prediction system for virus single-gene host adaptability, including:

[0045] Second acquisition module: configured to acquire the encoded feature matrix of the virus single gene;

[0046] Second prediction module: configured to input the encoded feature matrix into a virus single-gene host adaptability prediction model constructed by using the construction method of the virus single-gene host adaptability prediction model described in any one of the above to obtain the host adaptability of the virus single gene.

[0047] This application (sixth aspect) discloses a method for ranking the importance of codons in a virus gene sequence, including:

[0048] Third acquisition module: configured to acquire the set of encoded feature matrices and host labels of the test set virus single gene;

[0049] First accuracy prediction module: configured to input the set of encoded feature matrices of the test set into a single-gene host adaptability prediction model constructed by using the construction method of the virus single-gene host adaptability prediction model to obtain the predicted host labels, and obtain the first prediction accuracy after comparing the predicted host labels with the test set host labels;

[0050] Sliding window mask prediction module: configured to use the sliding window method to mask the first window of the set of encoded feature matrices of the test set to obtain the masked first window encoded feature matrix set, the central row of the first window is the first row of the set of encoded feature matrices, input the masked first window encoded feature matrix set into the same single-gene host adaptability prediction model and calculate the predicted accuracy of the masked first window, and repeat after the sliding window slides along the rows of the encoded feature matrix in turn to obtain the second prediction accuracy, the second prediction accuracy includes the predicted accuracies of the masked first window to the nth window, where n represents the number of codons of the single gene;

[0051] Region vector importance calculation module: configured to obtain the region vector importance of the first to nth windows based on the difference between the second prediction accuracy and the first prediction accuracy;

[0052] Codon importance calculation module: configured to average the region vector importance of each codon to obtain the importance of each codon.

[0053] The seventh aspect of the present application discloses a computer device, which includes: a memory and a processor; the memory is used to store program instructions; the processor is used to call the program instructions, and when the program instructions are executed, it is used to execute the steps of the above-mentioned method.

[0054] The eighth aspect of the present application discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the above-mentioned method.

[0055] The ninth aspect of the present application discloses a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the above-mentioned method.

[0056] The present application has the following beneficial effects:

[0057] (1) By using the model trained after characterizing codons, the present application can analyze the adaptability of gene codons;

[0058] (2) By using the single-gene host adaptability prediction model to analyze the codon importance of genes, the present application can be used to screen adaptive gene loci and identify the genetic causal relationships of the existing influenza virus host adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0060] Figure 1 It is a schematic flowchart of the method provided in the first aspect of the embodiments of the present invention;

[0061] Figure 2 It is a schematic diagram of the program product provided in the fourth aspect of the embodiments of the present invention;

[0062] Figure 3 It is a schematic diagram of the computer device provided in the embodiments of the present invention;

[0063] Figure 4 It is a schematic diagram of the architecture of an exemplary computing device provided in the embodiments of the present invention;

[0064] Figure 5 It is a schematic diagram of the storage medium provided in the embodiments of the present invention;

[0065] Figure 6 It is a schematic flowchart of a Codon2Vec sliding window characterization provided in the embodiments of the present invention;

[0066] Figure 7 is a flowchart of Codon2Vec preprocessing and characterization evaluation provided by an embodiment of the present invention;

[0067] Figure 8 is a framework diagram of the ResNet residual network provided by an embodiment of the present invention;

[0068] Figure 9 is a flowchart of calculating the importance degree of codons provided by an embodiment of the present invention;

[0069] Figure 10 is a schematic diagram of calculating the importance degree of codons by a Bayesian model provided by an embodiment of the present invention;

[0070] Figure 11 is a schematic diagram of a flowchart of an ablation experiment provided by an embodiment of the present invention;

[0071] Figure 12 is a schematic diagram of generating simulated reassortant data provided by an embodiment of the present invention;

[0072] Figure 13 is a flowchart of constructing and predicting a reassortment adaptation model provided by an embodiment of the present invention. Detailed implementation manners

[0073] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0074] In some processes described in the specification, claims and above-mentioned drawings of the present invention, a plurality of operations appear in a specific order, but it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as S101, S102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0075] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0076] Figure 1 It is a schematic flowchart of a method for constructing a virus single-gene host adaptability prediction model provided by an embodiment of the present invention, including:

[0077] S101: Obtain a training set of virus single-gene data sets and host labels.

[0078] S102: Encode the single genes in the virus single-gene data set in sequence to obtain a set of encoded feature matrices. The encoding is as follows: After connecting the head and tail of the gene sequence of the single gene, use the sliding window method to traverse the gene sequence. The frequencies of 64 different codons within the window form a window codon frequency vector, and the window codon frequency vector is used to represent the central codon of the window. As the window slides, each codon of the gene sequence is sequentially represented as the window codon frequency vector when the codon is the central codon. After the traversal, an encoded feature matrix composed of window codon frequency vectors is obtained.

[0079] S103: Input the set of encoded feature matrices into a neural network model to obtain predicted labels, and iteratively train based on the comparison between the predicted labels and the host labels to obtain a virus single-gene host adaptability prediction model.

[0080] A gene fragment of a virus refers to a nucleic acid sequence region in the virus genome with specific functional or structural characteristics, and usually corresponds to encoding a structural protein (such as a capsid protein, an envelope protein) or a non-structural protein (such as a replicase, a regulatory protein) of the virus. Different viruses may contain multiple functionally independent gene fragments according to their genomic characteristics (such as segmented or non-segmented, single-stranded or double-stranded, RNA or DNA).

[0081] Some examples of virus gene fragments:

[0082] The NP (nucleocapsid protein) gene fragment of the avian influenza virus (such as H5N1): Encodes the nucleocapsid protein, which is responsible for wrapping the viral RNA to form a ribonucleoprotein complex (RNP), and is a key component for viral genome replication and transcription. This gene is highly conserved and is commonly used for influenza virus typing and molecular diagnosis (such as RT-PCR detection).

[0083] The HA (hemagglutinin) and NA (neuraminidase) gene fragments of the influenza virus: HA is responsible for host cell receptor binding and membrane fusion, and NA assists in virus release. Both are the core basis for influenza virus subtype classification (such as H1N1, H5N1).

[0084] The gag, pol, and env gene fragments of HIV: Encoding the viral capsid protein, reverse transcriptase / integrase, and envelope glycoprotein respectively, are key targets for antiviral drugs (such as reverse transcriptase inhibitors) and vaccine research.

[0085] The S (spike protein) gene fragment of SARS-CoV-2: Encodes the viral spike protein, mediates the binding of the host cell ACE2 receptor, and is the core target for the development of COVID-19 vaccines (such as mRNA vaccines).

[0086] The research on gene fragments provides a molecular basis for virus origin tracing (such as analyzing the host origin of avian influenza through NP gene analysis), the development of diagnostic reagents (such as antigen detection based on HA / NA), vaccine design (such as screening S protein antigenic epitopes), and the research and development of antiviral drugs.

[0087] In some embodiments, the research aim is to provide a method for constructing a reassortment adaptability prediction model of IAV (Influenza A virus), and use this model to predict the human adaptability of the simulated reassortment sequences of avian influenza virus (AIV).

[0088] The present invention first adopts the self-created Codon2Vec characterization method, and performs encoding characterization based on the statistical frequency of sliding window codons and pre-training of the transformer encoding layer;

[0089] Then, based on the ResNet residual network, single-gene adaptability prediction models for four fragments, PB2, PB1, PA, and NP, are established, and the key codon regions in each gene are calculated through ablation experiments;

[0090] Finally, according to the calculated key codon regions, the simulation reassortment generation of avian influenza virus sequences and human influenza virus sequences is realized, and a reassortment host adaptability prediction model of IAV is constructed based on ResNet, and finally it is applied to the reassortment host adaptability prediction of the simulated reassortment sequences.

[0091] First, the 8 segments of the influenza virus genome, HA, NA, M, NS, NP, PB1, PB2, and PA, encode 10 proteins: hemagglutinin (HA), neuraminidase (NA), matrix proteins M2 and M1, non-structural proteins NS1 and NS2, nucleocapsid, and three polymerases. The NP protein encoded by segment 5 forms the viral ribonucleoprotein (vRNP) together with the polymerase trimer PB1, PB2, PA, and viral RNA. The RNA of the influenza virus is transcribed and replicated in the nucleus, and one of the key processes is the nucleocytoplasmic shuttle of the ribonucleoprotein. In the early infection stage when the influenza virus invades cells, the ribonucleoprotein of the influenza virus quickly enters the nucleus to start transcription. As the main component of the vRNP complex, NP mediates the nuclear import of the vRNP complex through its nuclear localization signals (NLSs). The NP protein enters the nucleus after being translated and expressed from the viral mRNA in the cytoplasm, promoting the replication and transcription processes of the viral genome. In summary, the influenza virus NP protein is an important viral protein that determines the host range and pathogenicity of influenza A virus.

[0092] Currently, the present invention mainly conducts research on polymerase-related genes, and the single-gene host adaptability and reassortment host adaptability prediction models trained are also for these 4 genes. However, the research method in the present invention can be extended to other gene segments or even other viruses, that is, this method can be used to perform sequence characterization on other genes or viruses and train corresponding adaptability prediction models. For other gene segments of IAV, they can be generated by simulating reassortment with polymerase-related genes together, and an AI model suitable for predicting the adaptability of more gene simulations and reassortments can be trained.

[0093] The processing flow and specific operation steps of the present invention are as follows:

[0094] I. Codon2Vec Characterization Method

[0095] 1.1 (Step 1): Input the cleaned csv file containing the IAV gene sequence:

[0096] The file contains multiple lines, and each line includes the IAV virus sequence and its corresponding host label;

[0097] The cleaned IAV gene sequence file contains the gene sequences of N virus segments and their corresponding host labels.

[0098] N = 8, including segments HA, NA, M, NS, NP, PB1, PB2, PA;

[0099] Since the adaptability needs to be predicted separately for different segments, therefore, different segments are processed separately.

[0100] First, for the first fragment, such as the NP fragment, the method of using a sliding window (window size (WS1) = 128) is used to traverse the NP fragment gene sequence of each line, and the codon frequency in the window is used to encode the central codon within the current window. Since there are 64 types of codons, each central codon is represented as a one-dimensional vector with 64 digital features.

[0101] As the window slides, each codon in the sequence is encoded and characterized, and a two-dimensional matrix of codon number × 64 can be obtained, as Figure 6 shown.

[0102] Figure 6 For example, the gene sequence of a certain line shown as: ATG1 GAA2 AGA3 ATA4 AAA5 GAA6CTA7GAT8……TAG n

[0103] The size of the sliding window is S. When S contains 128 codons, S = 384;

[0104] Based on the current window, calculate the frequency of 64-bit codons, count the frequency of each codon within the current window, and the calculation is as follows:

[0105]

[0106] When the window size is 128 codons, then S = 384, and the above formula corresponds to:

[0107]

[0108] condon j represents 64 types of codons;

[0109] Thus, the vector representation of the codon frequency is obtained as:

[0110]

[0111] represents the codon frequency vector within the window, with a dimension of 64*1; thus, the central codon is encoded into a 1×64 vector,

[0112] For example, the codons contained in the 64th window are ATG1 GAA2 AGA3 ATA4 AAA5 GAA6CTA7 GAT8…GUC 63 AUA 64 AUC 65 …UUC 128 ;

[0113] The central codon in the current window is AUA 64 ,

[0114] After the above operations, as the window slides, each codon in the sequence is encoded and characterized. That is, the central codon is characterized by a codon frequency vector, and a two-dimensional matrix of the number of codons contained in the gene sequence × 64 can be obtained. That is, assuming that the number of codons contained in the gene sequence of an NP fragment is n, the dimension of the encoding feature matrix is n × 64, as Figure 6 shown.

[0115] Since we need to encode each codon, when the central codon of the window is ATG1, the sequence of the window is obtained by splicing the codons at the end of the sequence, and then the codon frequency vector is calculated. When the central codon is the last codon, the window sequence is supplemented by intercepting the codons in the first half of the sequence and then the codon frequency is calculated.

[0116] The determination of the sliding window is shown in the following pseudocode:

[0117] Table 1 Logic of Encoding Codons with a Sliding Window

[0118]

[0119] Due to the input size requirements of the subsequent ResNet framework, zero vector padding is performed at the end of the feature matrix. For example, the size of the feature matrix on the PB2 gene is 760 × 64, while the input parameter of the subsequent ResNet model framework is 128 × 128 × 3. Therefore, the size of the feature matrix on the PB2 gene is adjusted to 768 × 64, that is, an 8 × 64 zero vector is added at the end of the original feature matrix. Finally, the characterization matrices of the same gene of multiple virus sequences are stored in an npy file.

[0120]

[0121] Among them, represents the codon frequency vector when the first codon in the sequence is the central codon; ORF represents the gene fragment sequence, and A(ORF) represents the encoded encoding feature matrix.

[0122] In some embodiments, for effective encoding, that is, to reduce the 0 values in the codon frequency vector, it is recommended that the window width be greater than 64.

[0123] In some embodiments, the window width of the sliding window is 128 codons.

[0124] In some embodiments, the window width of the sliding window is 100 codons.

[0125] In some embodiments, any of the following can be used to replace the RestNet network in the prediction network model: DenseNet, ResNeXt, Inception-ResNet, MobileNetV2, EfficientNet, Wide Residual Networks (WRN). When applying, the dimension of the transformation matrix is changed based on the requirements of the network for the input, or the window width of the sliding window is changed.

[0126] 1.2 (Step 2): Input the npy file obtained in Step 1, and use the Transformer encoding layer to pre-train it to obtain the optimized feature encoding (with the same size as in Step 1), and store it in another npy file. Among them, the Transformer pre-training performs an auto-encoding task using unsupervised training to reconstruct the input data to obtain the optimized feature encoding. Figure 7 The A’ (ORF) shown.

[0127] 1.3 (Step 3): Input the 2 npy files obtained in Step 1 and Step 2, and use two methods of umap dimensionality reduction and hierarchical clustering in unsupervised learning to evaluate the representation effect of the representation matrices in these two npy files (that is, evaluate the representation effect of the pre-training of the Transformer encoding layer), and use external evaluation indicators of clustering, the adjusted Rand index (ARI) and the normalized mutual information (NMI), to evaluate the clustering effect. The final evaluation result is the dimensionality reduction visualization result and the clustering effect score of the two representation matrices.

[0128] According to the evaluation results, store the optimal npy file of the corresponding encoding matrix of the training set, that is, the optimal encoding feature matrix (if the pre-training effect of the Transformer encoding layer is not good on this segment, then use the original representation matrix for representation, that is, the npy file obtained in Step 1; otherwise, use the representation matrix after the pre-training of the Transformer encoding layer, that is, the npy file obtained in Step 2). The processing flow of Steps 2-3 is as Figure 7 shown.

[0129] In some embodiments, only the clustering evaluation index is used to select the optimal encoding feature matrix.

[0130] In some embodiments, only the dimensionality reduction visualization result is used to select the optimal encoding feature matrix.

[0131] In some embodiments, the optimal encoding feature matrix is selected by comprehensively considering various evaluation results. The selection process is as Figure 7As shown, the c value is updated according to different metrics, and A (ORF) or A' (ORF) is selected as the optimal coding feature matrix according to the c value.

[0132] In some embodiments, the following subsequent steps are performed based on the coding feature matrix obtained by the sliding window.

[0133] In some embodiments, the following subsequent steps are performed using the coding feature matrix after Transformer autoencoding.

[0134] In some embodiments, the following subsequent steps are performed using the optimal coding feature matrix.

[0135] Unless otherwise specified, the following coding feature matrix A (ORF) is any one of the following: the coding feature matrix, the coding feature matrix output by Transformer, and the optimal coding feature matrix.

[0136] II. Establish a ResNet single - fragment adaptability prediction model

[0137] 2.1 (Step 4): Use the ResNet residual network framework (as Figure 8 shown) to train a single - gene host adaptability prediction model. Among them, the label used for training is the original host label of the sequence. If its host is avian, the model label is "0"; if its host is human, the model label is "1". Save the trained model (i.e., the single - fragment host adaptability prediction model), which can be used for predicting the host adaptability of IAV with unknown host.

[0138] That is, for different fragments, repeat Steps 1 - 4 to obtain single - fragment adaptability prediction models for different fragments respectively.

[0139] In some embodiments, for 4 gene sequences (PB2, PB1, PA, and NP), Steps 1 - 4 need to be run independently, and finally 4 single - gene adaptability prediction models are obtained.

[0140] In some embodiments, for 8 gene sequences, run Steps 1 - 4 independently, and finally 8 single - gene adaptability prediction models are obtained.

[0141] III. Ablation experiment / Bayesian model calculates the codon importance order in each gene fragment

[0142] 3.1 (Step 5): Use the trained single - fragment host adaptability prediction model to predict the coding feature matrix of the test set (the genes of the test set are processed in the same way to obtain the coding feature matrix of the test set), and obtain the current prediction accuracy (acc1), which will be used to observe the impact of the mask on the accuracy in the ablation experiment (Step 6).

[0143] Step 6: To evaluate the importance of each codon in the gene fragment, a sliding window mask (window size = 80) is applied to the encoded feature matrix of the test set along the row direction of the matrix. That is, the encoded sequences of the corresponding 80 codons within the sliding window are covered, and the codon frequency vectors representing the codons within the sliding window region are masked as zero vectors. Since the codon frequency vectors within the window are zero vectors, the codon mask is as shown in the black mask in the matrix in Figure 10 or as shown in the masked encoded feature matrix in Figure 11 , and some of its vectors are represented as zero vectors. Moreover, when applying the sliding window mask to the encoded feature matrix, the judgment of the corresponding masked region is shown in Table 2.

[0144] Table 2 Judgment Logic (Pseudocode) for Sliding Window Mask Encoded Feature Matrix

[0145]

[0146] The codon feature vector within the first masked window is a zero vector, obtaining the encoded feature matrix A(ORF) of the first masked window mask1 ,

[0147] For the second window, obtain A(ORF) mask2 ,

[0148] For the nth window, obtain A(ORF) maskn ,

[0149] Finally, for a gene fragment containing n codons, n masked encoded feature matrices are obtained. After inputting these n masked encoded feature matrices into the pre-trained Transformer representation matrix, optimized feature encodings are obtained. After obtaining the optimal feature matrix in the manner of Step 3, it is input into the trained single-fragment host adaptability model for prediction, and the accuracy rate (acc2, where acc2 is a list) corresponding to each masked feature matrix is obtained and compared with acc1. The difference between the two can reflect the importance of the codon vectors in the current masked region. If the difference between acc2 and acc1 is large, it indicates that the masking of the 80 codons in this part of the vector has a great impact on the accuracy rate, that is, the importance represented by this part of the codon vector is high.

[0150] Then, slide the window. As the window slides, calculate the importance of each block of regional vectors in turn according to the difference between acc2 and acc1 predicted after masking by the corresponding regional vector mask. Since the importance obtained currently is for the vectors within an entire region of the mask, the importance of the codon frequency vectors within the region is averaged according to the window size (80) to obtain the importance of each vector (Vector importance, VI). Then, according to the window size (128) in the characterization process of step 1, the importance of each vector is averaged to obtain the importance of each codon in the sequence (Codon importance, CI) - CI1. The specific details are as Figure 9 , Table 3 shows:

[0151] Table 3 Calculation process of codon importance

[0152]

[0153]

[0154] where acc n represents the accuracy at the nth sliding window position, VI n represents the importance of the nth vector, and CI n represents the importance of the nth codon.

[0155] 3.2 (Step 7:) Use the Bayesian model to calculate the importance of each codon (CI2):

[0156] Regarding the distribution difference problem of 64 codons at the specified virus gene and specified gene locus for human and avian influenza viruses, the human and avian host priors (P(H)) of the virus and the joint distribution of the virus host and a certain codon (P(H,C)) can be calculated, and then the host conditional probability of this codon can be calculated: (P(C|H) = P(H,C) / P(H).

[0157] As Figure 10 shown, after encoding the feature matrix and inputting it into the single - fragment host adaptability prediction model, the accuracy acc1 is obtained;

[0158] After masking the encoding feature matrix with a sliding window, the encoding feature matrix of the i - th window of the mask is obtained in turn and input into the single - fragment host adaptability prediction model to obtain the accuracy acc2;

[0159] Among them, acc2 is a list, and the elements in the list respectively represent the accuracy obtained when masking the i - th window of the mask:

[0160] acc2 = [acc21, acc22, …, acc2 i,…,acc2 n

[0161] Among them, acc2 i represents the accuracy rate obtained by inputting the encoded feature matrix of the i-th window of the mask into the single-fragment host adaptability prediction model; substituting acc2 into the Bayesian model to obtain the Bayesian accuracy rate acc3, that is, P(B i |A)

[0162]

[0163] Among them, P(B i |A) represents B based on condition A i occurrence probability, where event A represents the prediction accuracy rate; B i represents the encoded feature matrix of the i-th window of the event mask;

[0164] P(B i ) represents the probability of the encoded feature matrix of the i-th window of the mask appearing. In our scenario, since we use a sliding window mask with a step size of 1 codon, therefore, n is the number of codons;

[0165] A represents the prediction accuracy rate, A|B i represents the prediction accuracy rate obtained when masking the i-th window, P(A|B i ) represents the probability of feature A matrix under event B i condition, which is acc2 here i ;

[0166] Therefore, the posterior probability (post-acc2) of acc2, that is, acc3, is expressed as:

[0167]

[0168] acc2 i represents the accuracy rate obtained by inputting the encoded feature matrix of the i-th window of the mask into the single-fragment host adaptability prediction model;

[0169] acc3 i represents the posterior accuracy rate obtained by inputting the encoded feature matrix of the i-th window of the mask into the single-fragment host adaptability prediction model;

[0170] acc3 = [acc31, acc32, …, acc3 i , …, acc3 n

[0171] ​​Obtain the posterior importance of the mask window based on the difference between acc3 and acc1: that is, replace the acc-list in Table 3 with acc3 - acc1, and then calculate CI2 based on the codon importance calculation process shown in Table 3.

[0172] Step 8: Take the average of the codon importance levels calculated in Steps 6 and 7. As shown in Formula 1, obtain the importance of each codon finally. Figure 10 As shown, normalize it and save it to a csv file.

[0173]

[0174] For the ablation experiment of 4 gene fragments, Steps 5 - 8 need to be run independently, and finally obtain the codon importance results on 4 gene fragments.

[0175] In some embodiments, obtain the codon importance only using the CI calculated in Step 6.

[0176] In some embodiments, obtain the codon importance with the Bayesian - corrected CI2.

[0177] In some embodiments, obtain the codon importance with the mean or weighted combination of the codon importance level (CI1) in Step 6 and the Bayesian - corrected CI2.

[0178] IV. Key Codon Combinations

[0179] Step 9: According to the evaluation results in Step 3, input the optimal coding feature matrix of the 4 - fragment training set simultaneously. According to the csv file of the codon importance results obtained in Step 8, intercept and combine the key codons that are important for fitness in the 4 fragments.

[0180] That is: compare the codon importance calculated in the 4 genes with the key codons that can be found in the literature. On the premise of including as many key codons found in the literature as possible, select the codon importance threshold, intercept the vector representations of the codons greater than the importance threshold from the characterization matrix, arrange them in the order in which they appear in the fragment, respectively obtain the characterizations of the key codons of the 4 genes, and finally combine the 4 fragments in the order of PB2, PB1, PA, NP. Different importance thresholds can be selected according to needs, but after the threshold is selected, there is only one possible splicing result.

[0181] According to the requirements of the subsequent model size, perform size filling on the above - combined characterization matrix (same as Step 1) to obtain the coding matrix of the key codons of the 4 genes representing the training set, as Figure 12 shown in the upper part.

[0182] Based on this representation and the ResNet residual network framework in Step 4, train a reassortment adaptability prediction model (here, the self-concatenation form of the 4-gene fragment representation of the same strain virus in the training set is used, and its label is the same as the original strain), as Figure 13 shown, for predicting the reassortment host adaptability of avian influenza viruses after 2020.

[0183] Step 10: Since the present invention predicts the reassortment host adaptability of reassortment sequences that do not exist in nature currently, it is necessary to generate the representations of all possible simulated reassortment sequences of avian influenza viruses based on the existing avian influenza virus sequences. Input the 4-fragment encoded feature matrix of the prediction set. According to the codon importance result csv file obtained in Step 8, intercept the key codon representations of the 4 fragments respectively, and randomly replace the key codon encoding matrix of the 4 sequence fragments of the selected human influenza virus prototype strain, and then generate the feature representation of the reassortment simulation sequence, as Figure 13 shown, and save it to an npy file.

[0184] Step 11: Input the npy file obtained in Step 10, and use the reassortment adaptability prediction model trained in Step 9 to perform simulated reassortment host adaptability prediction, and obtain the adaptability prediction result of the simulated reassortment, that is, the predicted host label (adapt to humans or adapt to birds) of the simulated reassortment sequence. Analyze the prediction results of each simulated reassortment, and it can be concluded which fragments of the avian influenza virus have a higher risk of reassortment and adaptation to humans, and which subtypes have a higher risk of simulated reassortment and adaptation to humans, providing a direction for the prevention and control of influenza A viruses. At the same time, this method can be applied to the reassortment adaptability prediction problems of other viruses, or extended to the reassortment monitoring of more gene fragments.

[0185] In the reassortment adaptability prediction problems of some other viruses, if the fragments it contains are different from those of IAV, after collecting data, use the above method to perform fragment adaptability prediction on the fragments it contains, determine the importance ranking of fragment codons, screen out the key codons, then train the reassortment adaptability prediction model of this virus, and output the adaptability of different reassortment results after simulation, and conduct targeted prevention.

[0186] Figure 3 is a schematic diagram of a computer device provided by an embodiment of the present invention, as Figure 3 shown. The device 2000 may include: one or more processors 2010 and one or more memories 2020; wherein, computer-readable code is stored in the memory, and when the computer-readable code is run by the one or more processors, the above-mentioned method can be executed.

[0187] The processor in this embodiment may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, operations, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc., and may be of the X86 architecture or the ARM architecture.

[0188] Generally speaking, the various exemplary embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices. When the aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, technologies, or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0189] For example, the method or device according to the embodiments of the present disclosure may also be implemented by means of Figure 4 the architecture of the computing device 3000 shown. As Figure 4 shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for the processing and / or communication of the method provided by the present disclosure and the program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 4 the architecture shown is only exemplary, and when implementing different devices, one or more components shown in the Figure 4 computing device may be omitted according to actual needs.

[0190] The embodiment of the present invention also provides a computer-readable storage medium, such as Figure 5As shown, it is a schematic diagram of the storage medium 4000 provided by an embodiment of the present invention. Computer-readable instructions 4010 are stored on the computer storage medium 4020. When the computer-readable instructions 4010 are run by a processor, the methods according to the embodiments of the present disclosure described with reference to the above drawings can be executed. The computer-readable storage medium in the embodiments of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memories for the methods described herein are intended to include but are not limited to these and any other suitable types of memories. It should be noted that the memories for the methods described herein are intended to include but are not limited to these and any other suitable types of memories.

[0191] The embodiments of the present disclosure also provide a computer program product or a computer program. When the computer program is executed by a processor, the steps of the above method are implemented, as Figure 2 shown. The computer program product or the computer program includes:

[0192] An acquisition module 201: configured to acquire a training set of virus single-gene data sets and host labels;

[0193] A feature encoding module 202: configured to encode the single genes in the virus single-gene data set in sequence to obtain a set of encoded feature matrices. The encoding is as follows: After connecting the gene sequences of the single genes end to end, the sliding window method is used to traverse the gene sequences. The frequencies of 64 different codons within the window form a window codon frequency vector, and the window codon frequency vector is used to represent the central codon of the window. As the window slides, each codon of the gene sequence is sequentially represented as the window codon frequency vector when the codon is the central codon. After the traversal is completed, an encoded feature matrix composed of window codon frequency vectors is obtained;

[0194] Training module 203: configured to input the set of encoded feature matrices into a neural network model to obtain predicted labels, and iteratively train based on the comparison between the predicted labels and the host labels to obtain a single-gene host adaptability prediction model.

[0195] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0196] Generally speaking, the various example embodiments of the present disclosure can be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, technologies, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.

[0197] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0198] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0199] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0200] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0201] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art should understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.

Claims

1. A method for constructing a prediction model for the host adaptability of a single gene of a virus, characterized in that, The method includes: Obtaining a training set of viral single-gene data sets and host labels; Successively encoding the single genes in the viral single-gene data set to obtain a set of encoded feature matrices. The encoding is as follows: After connecting the head and tail of the gene sequence of the single gene, the gene sequence is traversed using a sliding window method. The frequencies of 64 different codons within the window form a window codon frequency vector, and the central codon of the window is characterized by the window codon frequency vector. As the window slides, each codon in the gene sequence is successively characterized as the window codon frequency vector when the codon is the central codon. After the traversal, an encoded feature matrix composed of window codon frequency vectors is obtained; Inputting the set of encoded feature matrices into a neural network model to obtain predicted labels, and iteratively training based on the comparison between the predicted labels and the host labels to obtain a viral single-gene host adaptability prediction model.

2. The method for constructing a viral single-gene host adaptability prediction model according to claim 1, wherein In the method, the central codon of the sliding window represents the codon located at the central position of the sliding window; Optionally, the width of the sliding window is S codons. If S is odd, the central codon is the (S + 1) / 2-th codon in the window. If S is even, the central codon is the S / 2-th codon or the S / 2-th codon in the window; Optionally, when the index of the central codon in the gene sequence is less than S / 2 or [the description seems incomplete here], the codons missing in the sliding window are supplemented with the codons at the end of the gene sequence. When the index of the central codon in the gene sequence is greater than n - S, the codons missing in the sliding window are supplemented with the codons at the starting position of the gene sequence, where n represents that the gene sequence contains n codons in total; The 64 different codons are the 64 codons of the standard genetic code; Optionally, the dimension of the encoded feature matrix is n×64, where n represents that the gene sequence contains n codons in total; Optionally, the first row of the encoded feature matrix is the window codon frequency vector when the first codon of the gene sequence is the central codon, and the last row is the window codon frequency vector when the last codon of the gene sequence is the central codon; Optionally, the step size of the sliding window method is 1 codon.

3. The method for constructing a viral single-gene host adaptability prediction model according to claim 1, wherein Performing auto-encoding learning on the set of encoded feature matrices to obtain a second set of encoded feature matrices, correcting the encoded feature matrices in the second set of encoded feature matrices to the same dimension as the encoded feature matrices by padding with 0 to obtain a third set of encoded feature matrices, and inputting the third set of encoded feature matrices and the host labels into the neural network model for iterative training to obtain a viral single-gene host adaptability prediction model; Optionally, performing auto-encoding learning based on one or more of the following models: LTSM, Transformer, Transformer-XL, MLP, MLP-Mixer; Optionally, performing auto-encoding learning on the set of encoded feature matrices to obtain a second set of encoded feature matrices and simultaneously obtaining a trained auto-encoder; Optionally, if the dimension of the encoded feature matrix does not match the input dimension of the auto-encoding model, padding with 0 at the end to meet the requirements; Optionally, after evaluating the representation effects of the matrices in the set of encoded feature matrices and the third set of encoded feature matrices, select the matrices with good representation effects as the optimal encoded feature matrices. The optimal encoded feature matrices of the viral single genes in the training set form an optimal encoded feature matrix set. Input the optimal encoded feature matrices and host labels into the neural network model for iterative training to obtain a prediction model for the host adaptability of viral single genes; Optionally, the representation effect evaluation is one or more of the following: clustering evaluation, UMAP dimensionality reduction evaluation.

4. The method for constructing a viral single-gene host adaptability prediction model according to claim 1, wherein The training set of viral single genes is any one of the following genes: HA, NA, M, NS, NP, PB1, PB2, PA. After training, the corresponding prediction model for the host adaptability of viral single genes is obtained.

5. A method for predicting the host adaptability of viral single genes, characterized in that Obtain the encoded feature matrix of the viral single gene; Input the encoded feature matrix into the prediction model for the host adaptability of single genes constructed by the method according to any one of claims 1-4 to obtain the host adaptability of the viral single gene.

6. A method for ranking the importance of codons in a viral gene sequence, characterized in that, The method includes: S1: Obtain the set of encoded feature matrices and host labels of the viral single gene in the test set; S2: Input the set of encoded feature matrices of the test set into the prediction model for the host adaptability of viral single genes constructed by the method according to any one of claims 1-4 to obtain the predicted host labels. After comparing the predicted host labels with the host labels in the test set, obtain the first prediction accuracy rate; S3: Use the sliding window method to mask the first window of the set of encoded feature matrices of the test set to obtain the masked first window encoded feature matrix set. The central row of the first window is the first row of the set of encoded feature matrices. Input the masked first window encoded feature matrix set into the same prediction model for the host adaptability of single genes and calculate the prediction accuracy rate of the masked first window; S4: Slide the window along the rows of the encoded feature matrix in turn and repeat S3 to obtain the second prediction accuracy rate. The second prediction accuracy rate includes the prediction accuracy rates of the masked first window to the nth window, where n represents the number of codons of the single gene; S5: Obtain the regional vector importance of the first to nth windows based on the difference between the second prediction accuracy rate and the first prediction accuracy rate; S6: Average the regional vector importance of each codon to obtain the importance of each codon.

7. The method for sorting the importance of codons in a viral gene sequence according to claim 6, characterized in that The set of encoded feature matrices of the test set is the optimized set of encoded feature matrices obtained by inputting the encoded feature set of the test set into the autoencoder constructed by the autoencoding learning of the set of encoded feature matrices of the training set; Optionally, the set of encoded feature matrices of the test set is a set composed of the optimal encoded feature matrices obtained after evaluating the representation effects of the matrices in the set of encoded feature matrices and the optimized set of encoded feature matrices; Optionally, S3 further includes: obtaining the Bayesian posterior accuracy of the first masked window based on the prediction accuracy of the first masked window and the probability of each masked window; repeating step S3 to obtain the third prediction accuracy, where the third prediction accuracy includes the Bayesian posterior accuracies of the first to the nth masked windows, and obtaining the Bayesian regional vector importance of the first to the nth windows based on the difference between the third prediction accuracy and the first prediction accuracy; averaging the Bayesian regional vector importance of each codon to obtain the Bayesian posterior codon importance of each codon; Optionally, the optimal codon importance vector is obtained based on the weighted average of the codon importance and the Bayesian posterior codon importance; Optionally, the key codons of each fragment are screened based on the set importance threshold; Optionally, the Bayesian posterior accuracy of the first masked window obtained based on the prediction accuracy of the ith masked window and the probability of each masked window is expressed as: Among them, P(B i |A) represents the Bayesian posterior accuracy of the i-th window of the mask; B i represents the i-th window of the event mask; P(B i ) represents the probability of the i-th window of the mask; P(A|B i ) represents the prediction accuracy of the i-th window of the mask.

8. A computer device, characterized in that, The device includes: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Virus gene recognition and host prediction method and system

    CN114512182A

  • Virus detection method and system for tumor RNA sequencing data

    CN117746985A

  • Virus host prediction method and system based on neural network, and storage medium

    CN117995267A

  • Virus host prediction method and system based on fuzzy control optimization and medium

    CN118173164A

  • Processing method and device for tRNA (transfer ribonucleic acid) adaptive index weight prediction model

    CN119207560A

Cited By

  • Construction method of segmented virus gene reassortment host adaptability prediction model

    CN120356513A

  • Method for constructing segmented virus gene reassortant host adaptation prediction model

    CN120356513B