A method and apparatus for predicting key mutation sites in RNA viruses

By constructing a viral evolutionary branch tree and predictive models to identify key mutation sites in viral gene sequences, the problems of low efficiency and high cost in existing technologies have been solved, enabling efficient viral gene sequence analysis and promoting the development of vaccines and drugs.

CN119479802BActive Publication Date: 2025-10-28SUZHOU INST OF SYST MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510032739.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-10-28
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

Existing technologies are inefficient and costly in identifying key mutation sites in viral gene sequences, which may lead to new and unknown mutations in the virus, affecting drug and vaccine development and disease control.

Method used

By constructing a viral evolutionary branch tree, the first and second mutation probabilities of nucleotide sites are determined. Using a pre-trained site prediction model or a weighted summation method, key mutation sites of the target virus are screened out.

Benefits of technology

It has improved the efficiency of discovering key mutation sites in viral gene sequences, saved time and resources, promoted the research and development of vaccines or drugs, and safeguarded public health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479802B_ABST
    Figure CN119479802B_ABST
Patent Text Reader

Abstract

This specification discloses a data prediction method and apparatus for key mutation sites in RNA viruses. Specifically, it includes: constructing a viral evolutionary branch tree corresponding to the target virus based on acquired historical viral evolutionary data; determining the first mutation probability and second mutation probability for each nucleotide site in the target virus gene sequence based on the viral evolutionary branch tree; and identifying key mutation sites in the target virus gene sequence based on the first and second mutation probabilities for each nucleotide site. This method effectively improves the efficiency of discovering key mutation sites in viral gene sequences, saving time and resources, and indirectly significantly improving the efficiency of subsequent vaccine or drug development targeting key mutation sites, thus safeguarding public health and safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of biotechnology, and in particular to a data prediction method and apparatus for key mutation sites of RNA viruses. Background Technology

[0002] Viruses are a unique type of microorganism in nature. They cannot survive and reproduce independently and must rely on host cells to complete their life activities. Currently, the parasitism of most viruses usually has some negative impact on the host, such as causing dysfunction of the host's physiological system and triggering serious diseases.

[0003] During transmission and adaptation between different hosts, viruses can undergo various genetic mutations depending on their survival environment and conditions to ensure their stable survival within host cells. Therefore, analyzing mutation sites in viral genes and designing drugs is of great help in solving viral diseases.

[0004] Current traditional techniques primarily involve conducting extensive and repeated laboratory experiments on viruses to determine potential mutation sites in their gene sequences. While this method can identify some important mutation sites, the overall process is time-consuming, inefficient, and extremely costly. This not only hinders the development of drugs and vaccines against the virus but also risks the emergence of new, unknown mutations due to the prolonged testing period, potentially leading to more serious disease problems.

[0005] Therefore, it is crucial to efficiently and accurately identify key variant gene sites in viral gene sequences. Summary of the Invention

[0006] This specification provides a data prediction method and apparatus for key mutation sites of RNA viruses, in order to partially solve the aforementioned problems existing in the prior art.

[0007] The following technical solution is adopted in this specification:

[0008] This specification provides a data prediction method for key mutation sites in RNA viruses, including:

[0009] Obtain historical viral evolution data corresponding to the target virus, wherein the historical viral evolution data includes at least: viral gene sequence data and host information of the target virus, and viral gene sequence data and host information of each historically evolved virus obtained by the target virus in the course of historical evolution.

[0010] Based on the historical virus evolution data, construct the virus evolution branch tree corresponding to the target virus;

[0011] For each nucleotide site of the target virus, based on the viral evolutionary branch tree, a first mutation probability and a second mutation probability corresponding to the nucleotide site are determined. The first mutation probability is used to represent the probability of a fixed mutation occurring at the nucleotide site, and the second mutation probability is used to represent the probability of a parallel mutation occurring at the nucleotide site.

[0012] Based on the first mutation probability and the second mutation probability corresponding to each nucleotide site of the target virus, the key mutation sites of the target virus are determined.

[0013] Optionally, based on the historical virus evolution data, a virus evolution branch tree corresponding to the target virus is constructed, specifically including:

[0014] Based on the viral gene sequence data of the target virus in the historical viral evolution data, and the viral gene sequence data of each historically evolved virus corresponding to the target virus, a basic evolutionary tree corresponding to the target virus is constructed.

[0015] Based on the host information of the target virus and the host information of each historical evolutionary virus corresponding to the target virus, the basic phylogenetic tree is subjected to host origin classification and coloring processing, and the colored basic phylogenetic tree is used as the viral evolutionary branch tree corresponding to the target virus.

[0016] Optionally, based on the viral evolutionary branching tree, for each nucleotide site of the target virus, the first mutation probability corresponding to that nucleotide site is determined, specifically including:

[0017] Based on a preset fixed mutation cycle duration, the historical evolution duration of the target virus is divided into cycles to determine each fixed mutation cycle corresponding to the target virus.

[0018] For each fixed mutation cycle corresponding to the target virus, at least one viral evolutionary branch corresponding to the fixed mutation cycle is determined according to the viral evolutionary branch tree, and the fixed mutation evaluation result corresponding to the target virus within the fixed mutation cycle is determined according to the at least one viral evolutionary branch corresponding to the fixed mutation cycle.

[0019] Based on the fixed mutation assessment results corresponding to each fixed mutation cycle, the first mutation probability corresponding to the nucleotide site is determined.

[0020] Optionally, based on the first mutation probability and the second mutation probability corresponding to each nucleotide site of the target virus, the key mutation sites of the target virus are determined, specifically including:

[0021] For each nucleotide site of the target virus, the first mutation probability and the second mutation probability corresponding to the nucleotide site are input into a pre-trained site prediction model, so that the site prediction model can determine the comprehensive site feature data corresponding to the nucleotide site based on the first mutation probability and the second mutation probability corresponding to the nucleotide site.

[0022] Based on the comprehensive site feature data corresponding to each nucleotide site of the target virus, key mutation sites are identified from each nucleotide site of the target virus.

[0023] Optional, training site prediction models, specifically including:

[0024] Obtain the first mutation probability and the second mutation probability corresponding to each nucleotide site of the sample virus;

[0025] For each nucleotide site of the sample virus, the first mutation probability and the second mutation probability corresponding to the nucleotide site are input into the site prediction model to be trained. This allows the site prediction model to determine the comprehensive site feature data corresponding to the nucleotide site based on the first mutation probability and the second mutation probability. Based on the comprehensive site feature data and the standard comprehensive site feature data corresponding to the nucleotide site, the model determines the loss value of the site prediction model for the nucleotide site. The magnitude of the loss value is negatively correlated with the feature similarity between the comprehensive site feature data and the comprehensive standard site feature data corresponding to the nucleotide site.

[0026] Based on the loss values ​​of the site prediction model to be trained for each nucleotide site of the sample virus, the total loss value corresponding to the site prediction model to be trained is determined, and the site prediction model to be trained is trained based on the total loss value.

[0027] Optionally, based on the first mutation probability and the second mutation probability corresponding to each nucleotide site of the target virus, the key mutation sites of the target virus are determined, specifically including:

[0028] For each nucleotide site of the target virus, the first mutation probability and the second mutation probability corresponding to the nucleotide site are weighted and summed according to the preset fixed mutation weight and parallel mutation weight to obtain the comprehensive probability of the mutation site corresponding to the nucleotide site.

[0029] Based on the preset comprehensive probability threshold of mutation sites and the comprehensive probability of mutation sites corresponding to each nucleotide site of the target virus, the key mutation sites corresponding to the target virus are determined from each nucleotide site of the target virus.

[0030] Optionally, based on the first mutation probability and the second mutation probability corresponding to each nucleotide site of the target virus, the key mutation sites of the target virus are determined, specifically including:

[0031] Based on a preset site screening strategy, the key mutation sites corresponding to the target virus are determined from each nucleotide site of the target virus according to the first mutation probability and the second mutation probability of each nucleotide site of the target virus.

[0032] This specification provides a data prediction device for key mutation sites in RNA viruses, including:

[0033] The acquisition module is used to acquire historical viral evolution data corresponding to the target virus. The historical viral evolution data includes at least: viral gene sequence data and host information of the target virus, and viral gene sequence data and host information of each historically evolved virus obtained by the target virus in the course of historical evolution.

[0034] A construction module is used to construct a virus evolution branch tree corresponding to the target virus based on the historical virus evolution data;

[0035] The probability determination module is used to determine, for each nucleotide site of the target virus, a first mutation probability and a second mutation probability corresponding to that nucleotide site based on the viral evolutionary branch tree. The first mutation probability is used to represent the probability of a fixed mutation occurring at that nucleotide site, and the second mutation probability is used to represent the probability of a parallel mutation occurring at that nucleotide site.

[0036] The site determination module is used to determine the key mutation sites of the target virus based on the first mutation probability and the second mutation probability corresponding to each nucleotide site of the target virus.

[0037] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described data prediction method for key mutation sites of RNA viruses.

[0038] This specification provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described data prediction method for key mutation sites of RNA viruses.

[0039] The above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects:

[0040] As can be seen from the above method, the data prediction method for key mutation sites of RNA viruses provided in this specification can construct a viral evolutionary branch tree corresponding to the target virus based on the obtained historical viral evolutionary data. Then, for each nucleotide site in the target virus gene sequence, the first mutation probability and the second mutation probability corresponding to that nucleotide site are determined based on the viral evolutionary branch tree. Finally, based on the first mutation probability and the second mutation probability corresponding to each nucleotide site of the target virus, the key mutation sites in the target virus gene sequence are determined.

[0041] As can be seen from the above, the data prediction method for key mutation sites in RNA viruses provided in this specification can identify key mutation sites in the target virus gene sequence that may have a serious impact on viral gene mutations based on the historical viral evolution data corresponding to the target virus. This method can effectively save the time and resources spent on discovering and identifying key mutation sites in viral gene sequences, significantly improving the efficiency of discovering key mutation sites. This indirectly leads to a significant improvement in the efficiency of subsequent vaccine or drug development targeting key mutation sites in gene sequences, thus safeguarding the health and safety of the public. Attached Figure Description

[0042] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings:

[0043] Figure 1 This is a flowchart illustrating a data prediction method for key mutation sites in RNA viruses provided in this specification.

[0044] Figure 2 This is a schematic diagram of the overall workflow framework for a data prediction method for key mutation sites of RNA viruses provided in this specification.

[0045] Figure 3 This is a schematic diagram of a data prediction device for key mutation sites of RNA viruses provided in this specification.

[0046] Figure 4 The one provided in this specification corresponds to Figure 1 A schematic diagram of the structure of an electronic device. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.

[0048] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0049] Figure 1 This is a flowchart illustrating a data prediction method for key variant sites in RNA viruses provided in this specification, including the following steps:

[0050] S101: Obtain historical virus evolution data corresponding to the target virus.

[0051] In the current natural world, viruses, as a relatively unique type of microorganism, inevitably have some impact on the physiological activities of their hosts due to their unique mode of survival, which relies solely on parasitism. During transmission, viruses undergo genetic mutations to adapt to different living environments and conditions, thus helping them to survive stably within the host. Therefore, analyzing mutation sites in viral genes is crucial for developing drugs and vaccines to overcome viral diseases.

[0052] Current traditional methods mostly rely on extensive and repeated experiments and tests in laboratories, which are time-consuming and costly. Furthermore, delays can lead to new and unknown mutations in the virus, increasing the risk of more serious problems. Therefore, how to efficiently identify key mutant gene sites from viral gene sequences is a pressing issue that needs to be addressed.

[0053] Therefore, this specification provides a data prediction method for key variant sites of RNA viruses. The execution subject of this method can be a terminal device such as a desktop computer or laptop computer, or a server. Alternatively, the execution subject can be software, such as a client installed on a terminal device. For ease of explanation, this specification will only use a terminal device as the execution subject to describe the provided data prediction method for key variant sites of RNA viruses.

[0054] Based on this, a terminal device using the data prediction method for key mutation sites of RNA viruses provided in this specification can determine, based on the historical viral evolution data corresponding to the target virus, a first mutation probability representing the probability of fixed mutations occurring at each nucleotide site in the target virus, and a second mutation probability representing the probability of parallel mutations occurring at each nucleotide site in the target virus. Then, the terminal device can determine the key mutation sites from the gene sequence of the target virus based on the first and second mutation probabilities corresponding to each nucleotide site of the target virus.

[0055] The specific application scenarios for terminal devices using the methods described in this manual to predict key mutation sites in target viruses can be determined according to actual needs. For example, in the scenario of monitoring an epidemic, terminal devices using the methods described in this manual can identify key mutation sites in the gene sequence of the target virus that may mutate during future transmission, based on historical viral evolution data corresponding to the target virus. The relevant information data of the identified key mutation sites, such as protein structure information of nucleotide sites and host information of the mutated virus, can be presented to relevant medical professionals for data reference regarding disease transmission and the possible future mutation trends of the virus. Accurate prediction of key mutation sites in the target virus can play a highly beneficial role in early warning of epidemics and in the rational control of epidemic spread and outbreaks.

[0056] For example, in the field of vaccine development targeting diseases caused by a target virus, the terminal device using the methods described in this manual can identify key mutation sites in the gene sequence of the target virus that cause severe illness, based on historical viral evolution data corresponding to the target virus. By recommending the information on these identified key mutation sites to medical researchers developing vaccines against variants of the target virus, this can provide effective assistance. This will accelerate the development of vaccines and drugs to some extent, thereby contributing to solving the disease problems caused by viral transmission and protecting the health and safety of the public.

[0057] In this specification, the terminal device can obtain historical viral evolution data corresponding to the target virus. The historical viral evolution data may include at least: the viral gene sequence data of the target virus and the host information corresponding to the target virus, as well as the viral gene sequences of each historically evolved virus and the host information corresponding to each historically evolved virus during the historical evolution process of the target virus.

[0058] It should be noted that the historical viral evolution data corresponding to the target virus in this specification can be updated in real time. Over time, the viral gene sequence may undergo various unpredictable gene mutations under different living environments, changing living conditions, and diverse parasitic hosts. By updating the historical viral evolution data in real time, the various data identified in subsequent processes and the ultimately predicted key mutation sites are guaranteed to have strong real-time effectiveness and practicality.

[0059] In addition to the viral gene sequence data and host information of the target virus and its various historical evolutionary viruses mentioned above, the historical viral evolution data corresponding to the target virus may also include relevant data such as the isolation determination time and location of the target virus and its various historical evolutionary viruses, and the protein structure corresponding to each nucleotide site in the gene sequence. This type of data can play a useful role in the subsequent prediction of key mutation sites of the target gene, making the identified key mutation sites more accurate and the information more comprehensive.

[0060] S102: Based on the historical virus evolution data, construct the virus evolution branch tree corresponding to the target virus.

[0061] S103: For each nucleotide site of the target virus, based on the viral evolutionary branch tree, determine the first mutation probability and the second mutation probability corresponding to that nucleotide site.

[0062] In this specification, the terminal device can construct a virus evolution branch tree corresponding to the target virus based on the historical virus evolution data obtained through the above steps.

[0063] Specifically, the terminal device can construct a basic phylogenetic tree for the target virus based on the viral gene sequence data of the target virus in historical viral evolution data, as well as the viral gene sequence data of each historically evolved virus corresponding to the target virus. Next, the terminal device can color each viral branch in the basic phylogenetic tree according to the host origin corresponding to the host information, based on the host information of the target virus and each historically evolved virus. The terminal device can then use the colorized basic phylogenetic tree as the viral evolutionary branch tree corresponding to the target virus.

[0064] Furthermore, for each nucleotide site in the gene sequence of the target virus, the terminal device can determine the first mutation probability and the second mutation probability corresponding to each nucleotide site based on the viral evolutionary branching tree constructed above. Specifically, the first mutation probability represents the probability of a fixed mutation occurring at that nucleotide site, while the second mutation probability represents the probability of a parallel mutation occurring at that nucleotide site.

[0065] The specific process for determining the first mutation probability corresponding to each nucleotide site of the target virus is as follows: First, the terminal device can divide the historical evolution time of the target virus into periods based on a pre-set fixed mutation period, thereby determining each fixed mutation period corresponding to the target virus.

[0066] Next, the terminal device can, for each fixed mutation cycle of the target virus, determine at least one viral evolutionary branch corresponding to the target virus within that fixed mutation cycle, based on the viral evolutionary branch tree corresponding to the target virus constructed above. The terminal device can then determine the fixed mutation assessment result for the target virus within that fixed mutation cycle based on the relevant data corresponding to at least one viral evolutionary branch. Specifically, the fixed mutation assessment result can be used to indicate which nucleotide sites in the gene sequence of the target virus may have undergone fixed mutations within that fixed mutation cycle.

[0067] Then, the terminal device can determine the first mutation probability corresponding to each nucleotide site in the target virus, i.e. the probability of a fixed mutation occurring at each nucleotide site, based on the fixed mutation evaluation results of the target virus in each fixed mutation cycle.

[0068] In the current field of viral evolution, viruses can gradually adapt to changing environments within host cells by undergoing fixed mutations. These fixed mutations can fundamentally alter the viral gene sequence, ensuring that offspring inherit the resulting mutations at birth. Therefore, reasonable research into viral fixed mutations can greatly benefit the subsequent development of vaccines and drugs targeting specific viruses.

[0069] The first mutation probability determined through the above steps effectively reflects the fixed mutation probability of each nucleotide site in the target virus throughout its evolutionary history. A clear indication of this fixed mutation trend effectively aids in the subsequent identification of key mutation sites, improving the accuracy and practicality of this identification. The specific process for determining the first mutation probability can be found using the following formula:

[0070]

[0071] Among them, targeting each nucleotide site in the target virus , Used to represent this nucleotide site The corresponding first mutation frequency. Used to indicate the historical evolution time of the target virus. This is used to represent the duration of historical evolution based on a preset fixed mutation cycle. The fixed mutation cycles are divided into different periods. Used to represent nucleotide sites In a fixed mutation cycle The fixed number of mutations within a given period, i.e., the fixed mutation cycle mentioned above. This nucleotide site is located in at least one viral evolutionary branch corresponding to the target virus. The corresponding fixed mutation assessment results, This indicates a fixed mutation cycle. The total number of mutations of the target virus.

[0072] The above formula is used for each fixed mutation cycle. Each nucleotide site in the target virus The corresponding fixed mutation probability can be calculated by the terminal device by summing and averaging to obtain the nucleotide site in the target virus. Corresponding first mutation frequency .

[0073] While determining the first mutation probability for each nucleotide site of the target virus, the terminal device can also determine the second mutation probability for each nucleotide site based on the viral evolutionary branching tree constructed above. The second mutation probability specifically represents the probability of parallel mutations occurring at nucleotide sites in the target virus.

[0074] It's important to note that the determination process for the second mutation probability differs from that of the first. The second mutation probability specifically reflects the occurrence of parallel mutations, focusing on each evolutionary branch of the target virus, rather than targeting each mutation cycle as in the first mutation probability determination process. In the current field of viral evolution research, parallel mutations, along with the aforementioned fixed mutations, are crucial aspects. Parallel mutations occur when viral gene sequences undergo identical variations across different host populations due to similar selective pressures. Parallel mutations exist in almost all populations and evolutionary branches, but not necessarily in all viral individuals; however, they may occur when facing similar or identical survival environments and pressures. Fixed mutations, on the other hand, are gene mutations present in a single population, carried by all viral individuals and tending towards stable inheritance.

[0075] The occurrence of parallel mutations indicates that the virus is adapting to its current survival environment and selective pressures, thereby better parasitizing host cells. Therefore, studying and determining the probability of parallel mutations at each nucleotide site can effectively improve the accuracy of the process for identifying key mutation sites. The specific process for determining the probability of the aforementioned second mutation can be referenced using the following formula:

[0076]

[0077] Among them, targeting each nucleotide site in the target virus , Used to represent nucleotide sites The corresponding second mutation probability. Used to represent the number of evolutionary branches in the historical evolutionary process of a target virus. This indicates the number of evolutionary branches. Each evolutionary branch in it. Used to represent nucleotide sites In evolutionary branches The number of parallel mutations that occur in the process, and This is used to indicate the evolutionary branch of the target virus. The total number of mutation events that occurred in the process.

[0078] Using the formula above, the terminal device can determine each nucleotide site in the target virus. In each evolutionary branch The corresponding parallel mutation probabilities. Then, by summing and averaging the probabilities of each parallel mutation, the probability of each nucleotide site is obtained. The corresponding second mutation probability. The determined second mutation probability can effectively reflect the actual situation of parallel mutations occurring at each nucleotide site in the target virus during its historical evolution. Combining it with the first mutation probability determined in the above process takes into account both the impact of fixed mutations on the evolution of the target virus and integrates relevant data on parallel mutations, effectively improving the accuracy of the subsequent process for determining key mutation sites in the target virus.

[0079] S104: Determine the key mutation sites of the target virus based on the first mutation probability and the second mutation probability corresponding to each nucleotide site of the target virus.

[0080] In this specification, the terminal device can screen out key mutation sites of the target virus from each nucleotide site of the target virus based on the first mutation probability and the second mutation probability corresponding to each nucleotide site of the target virus determined through the above steps.

[0081] Specifically, for each nucleotide site in the target viral gene sequence, the terminal device can input the first mutation probability and the second mutation probability corresponding to the nucleotide site into a pre-trained site prediction model, so that the site prediction model can determine the comprehensive site feature data corresponding to the nucleotide site based on the first mutation probability and the second mutation probability corresponding to the nucleotide site.

[0082] Then, the terminal device can screen out key variant sites from the target virus's nucleotide sites based on the comprehensive site feature data corresponding to each nucleotide site in the target virus's gene sequence. The process for determining the comprehensive site feature data corresponding to nucleotide sites and the specific data format can be referenced in the following formula:

[0083]

[0084] in, Used to represent each nucleotide site in the target virus The corresponding comprehensive site feature data can be specifically represented as a vector matrix as shown in the formula above. and These are used to represent the nucleotide sites determined through the above steps. The corresponding first mutation probability and second mutation probability.

[0085] Using the above formula, the terminal device can detect nucleotide sites. Corresponding first mutation probability Second mutation probability These are input together into the pre-trained site prediction model, enabling the site prediction model to predict based on nucleotide sites. The corresponding first mutation probability Second mutation probability After performing feature processing separately, the feature data are fused to obtain nucleotide sites. Corresponding comprehensive site feature data .

[0086] It should be noted that, as mentioned in the above steps, the historical viral evolution data corresponding to the target virus may also include relevant data such as the isolation and determination time and location of each historical evolutionary virus corresponding to the target virus, and the protein structure corresponding to each nucleotide site in the gene sequence of the target virus and each historical evolutionary virus.

[0087] In this specification, in addition to inputting the first mutation probability and the second mutation probability corresponding to each nucleotide site of the target virus into the site prediction model to determine the comprehensive site feature data corresponding to each nucleotide site, the terminal device can also input other data from the aforementioned historical viral evolution data into the site prediction model. This allows the site prediction model to determine the comprehensive site feature data corresponding to each nucleotide site not only based on the first mutation probability and the second mutation probability corresponding to each nucleotide site, but also based on relevant data such as the isolation determination time and location of each historically evolved virus corresponding to the target virus, and the protein structure corresponding to each nucleotide site in the gene sequences of the target virus and each historically evolved virus.

[0088] The comprehensive site feature data determined by this method involves richer information and data, thus enabling more accurate prediction of key mutation sites in the target virus. The determination process and specific data format of the comprehensive site feature data corresponding to this method can be referenced in the following formula:

[0089]

[0090] in, , as well as This has the same meaning as the formula described above for determining the comprehensive site feature data based solely on the first and second mutation probabilities. Used to represent each nucleotide site in the target virus The corresponding comprehensive site feature data can be specifically represented as a vector matrix as shown in the formula above, and and These are used to represent the nucleotide sites determined through the above steps. The corresponding first mutation probability and second mutation probability.

[0091] And in the above formula Then used to represent nucleotide sites The corresponding genotype characteristic data are specifically represented by nucleotide sites. The corresponding gene composition or protein structure. Then used to represent nucleotide sites The corresponding geographic location feature data is specifically represented by nucleotide sites. Geographical distribution of virus samples successfully isolated when mutations occur. Then used to represent nucleotide sites The corresponding host type characteristic data is specifically represented by nucleotide sites. The host information of the historical evolutionary virus corresponding to the target virus after mutation, such as whether the host is human or animal. As for... Then used to represent nucleotide sites Structural feature data in the gene sequence of the target virus can specifically be represented by the location of nucleotide sites in the protein structure of the target virus and the spatial domain information of their location.

[0092] Through the calculation using the above formula, the terminal device can use the site prediction model to extract the first and second mutation probabilities corresponding to each nucleotide site of the target virus, as well as related data such as genotype feature data, geographical location feature data, host type feature data, and structural feature data. Then, it performs feature matrix fusion to obtain the comprehensive site feature data corresponding to each nucleotide site.

[0093] After determining the comprehensive site feature data corresponding to each nucleotide site in the gene sequence of the target virus, the terminal device can screen out the key variant sites corresponding to the target virus from each nucleotide site of the target virus based on the comprehensive site feature data corresponding to each nucleotide site and the preset key variant site feature evaluation criteria.

[0094] It should be noted that the specific model category of the site prediction model mentioned in the above method is not strictly limited in this specification. It can be a mathematical model with the ability to extract features based on data, such as a deep neural network (DNN). The model category of the site prediction model can be flexibly selected and set according to the actual application scenario and needs.

[0095] Furthermore, the site prediction model mentioned above is a pre-trained mathematical model. Specifically, the terminal device can obtain the first mutation probability and the second mutation probability corresponding to each nucleotide site of the sample virus. Then, for each nucleotide site of the sample virus, the terminal device can input the first mutation probability and the second mutation probability corresponding to that nucleotide site into the site prediction model to be trained. This allows the site prediction model to determine the comprehensive site feature data corresponding to that nucleotide site based on the first mutation probability and the second mutation probability.

[0096] Next, the terminal device can determine the loss value of the site prediction model to be trained for that nucleotide site based on the comprehensive site feature data and the standard comprehensive site feature data corresponding to that nucleotide site. The magnitude of the loss value is negatively correlated with the similarity between the comprehensive site feature data and the standard comprehensive site feature data. Then, the terminal device can determine the total loss value of the site prediction model to be trained and the total loss value of the sample virus based on the loss values ​​corresponding to each nucleotide site. Finally, the terminal device can train the site prediction model based on the total loss value. The trained site prediction model is then used to implement the aforementioned data prediction method for key variant sites of RNA viruses.

[0097] It should also be noted that this specification does not strictly limit the specific training and optimization process for the site prediction model to be trained. For example, in this specification, the optimization training process for the site prediction model can utilize the Gradient Boosting Decision Tree (GBDT) machine learning algorithm, as shown in the following formula:

[0098]

[0099] Specifically, for each nucleotide site in the sample virus, Used to represent nucleotide sites Corresponding comprehensive site feature data, Used to represent the site prediction model to be trained. A decision tree, This is the learning rate set for the site prediction model to be trained. By calculating and summing using this formula, the terminal device can determine each nucleotide site in the sample virus. Corresponding loss value By optimizing all nucleotide sites Corresponding loss value This is done to improve the capabilities of the site prediction model, thereby obtaining the trained site prediction model.

[0100] In addition to the training methods mentioned above, this specification also allows for the training and optimization of the site prediction model using a random forest model. The specific training process can be referenced in the following formula:

[0101]

[0102] As shown in the formula above, the process of determining the overall loss value is similar to that of gradient boosting decision trees, but the preset learning rate is removed. Instead of using the default settings, we chose to perform an averaging process, averaging the nucleotide sites in each decision tree corresponding to the site prediction model. The average of the loss values ​​is used as the nucleotide site. The loss value during the overall training process.

[0103] In this specification, in addition to the method described above for determining key mutation sites in the target virus using a site prediction model, the terminal device can also directly perform a weighted summation of the first and second mutation probabilities at each nucleotide site in the target virus to determine the comprehensive probability of the mutation site corresponding to each nucleotide site. Then, the terminal device can determine the key mutation sites corresponding to the target virus from each nucleotide site of the target virus based on the comprehensive probability of the mutation site corresponding to each nucleotide site and a preset comprehensive probability threshold for mutation sites.

[0104] The specific process for determining the overall probability of the mutation site corresponding to each nucleotide site in the target virus mentioned above can refer to the following formula:

[0105]

[0106] Specifically, for each nucleotide site in the target virus, and These are respectively represented as nucleotide sites. The corresponding first mutation probability and second mutation probability. and Each is targeted at a nucleotide site The corresponding first mutation probability Second mutation probability The probability weight parameters are set, and their sum is 1, that is... .

[0107] The terminal device can use the above calculation formula to determine the first mutation probability at each nucleotide site. Second mutation probability A weighted summation process is performed to obtain the overall probability of the variant site corresponding to each nucleotide site. The terminal device can then use this probability, based on a preset threshold, to identify nucleotide sites with an overall probability greater than the threshold as the key variant sites corresponding to the target virus.

[0108] In addition to determining the key mutation sites corresponding to the target virus through site prediction models and weighted summation, this manual also allows the terminal device to directly screen sites based on the first mutation probability and the second mutation probability corresponding to each nucleotide site, according to a preset key mutation site evaluation strategy, thereby determining the key mutation sites in the target virus gene sequence.

[0109] The specific critical variant site assessment strategy is not strictly limited in this specification. For example, it could involve setting corresponding probability thresholds for the first and second mutation probabilities for each nucleotide site. If a mutation probability exceeds a certain threshold, the nucleotide site is considered a critical variant site. The specific content of the critical variant site assessment strategy can be set according to the actual application scenario and requirements.

[0110] It should be noted that the above three methods for determining the key mutation sites of the target virus based on the first and second mutation probabilities corresponding to each nucleotide site are: site prediction models, weighted summation of mutation probabilities, and screening based on key mutation site evaluation strategies. In practical applications, one or more methods can be arbitrarily selected according to the actual application scenario and needs. The specific combination and selection of methods are not strictly limited in this specification. The key mutation sites corresponding to the target virus obtained through multiple methods can be used as evaluation and verification standards for the prediction accuracy of each method. This results in higher accuracy and practicality of the finally determined key mutation sites.

[0111] At this point, the data prediction process for key mutation sites of RNA viruses using terminal devices is complete. To facilitate understanding of the overall prediction process, a flowchart illustrating the overall workflow of a data prediction method for key mutation sites of RNA viruses is provided below. Figure 2 As shown in the image.

[0112] Figure 2 This is a schematic diagram of the overall workflow framework for a data prediction method for key mutation sites of RNA viruses provided in this specification.

[0113] like Figure 2 As shown, the terminal device can construct a viral evolutionary branch tree corresponding to the target virus based on the acquired historical viral evolutionary data. Then, based on the viral evolutionary branch tree, the terminal device can determine the first mutation probability and the second mutation probability corresponding to each nucleotide site in the target virus gene sequence. Finally, based on the first mutation probability and the second mutation probability corresponding to each nucleotide site in the target virus, the terminal device can screen out the corresponding key mutation sites from the gene sequence of the target virus.

[0114] As can be seen from the above, the data prediction method for key mutation sites in RNA viruses provided in this specification can identify key mutation sites in the target virus gene sequence that may have a significant impact on viral gene mutations based on the historical viral evolution data corresponding to the target virus. This method can effectively improve the efficiency of discovering key mutation sites in viral gene sequences, greatly saving time and resources, and consequently significantly improving the efficiency of subsequent vaccine or drug development targeting key mutation sites in gene sequences.

[0115] The above describes one or more implementations of the methods described in this specification. Based on the same approach, this specification also provides corresponding data prediction devices for key mutation sites in RNA viruses, such as... Figure 3 As shown.

[0116] Figure 3 This is a schematic diagram of a data prediction device for key mutation sites of RNA viruses provided in this specification, including:

[0117] The acquisition module 301 is used to acquire historical virus evolution data corresponding to the target virus. The historical virus evolution data includes at least: viral gene sequence data and host information of the target virus, and viral gene sequence data and host information of each historically evolved virus obtained by the target virus in the course of historical evolution.

[0118] Construction module 302 is used to construct a virus evolution branch tree corresponding to the target virus based on the historical virus evolution data;

[0119] The probability determination module 303 is used to determine, based on the viral evolutionary branch tree, a first mutation probability and a second mutation probability corresponding to each nucleotide site of the target virus. The first mutation probability is used to represent the probability of a fixed mutation occurring at the nucleotide site, and the second mutation probability is used to represent the probability of a parallel mutation occurring at the nucleotide site.

[0120] The site determination module 304 is used to determine the key mutation sites of the target virus based on the first mutation probability and the second mutation probability corresponding to each nucleotide site of the target virus.

[0121] Optionally, the construction module 302 is specifically used to: construct a basic evolutionary tree corresponding to the target virus based on the viral gene sequence data of the target virus in the historical viral evolution data and the viral gene sequence data of each historical evolutionary virus corresponding to the target virus; perform host source classification and coloring processing on the basic evolutionary tree based on the host information of the target virus and the host information of each historical evolutionary virus corresponding to the target virus; and use the colored basic evolutionary tree as the viral evolutionary branch tree corresponding to the target virus.

[0122] Optionally, the probability determination module 303 is specifically used to: divide the historical evolution time of the target virus into periods according to a preset fixed mutation period duration, and determine each fixed mutation period corresponding to the target virus; for each fixed mutation period corresponding to the target virus, determine at least one viral evolutionary branch corresponding to the fixed mutation period according to the viral evolutionary branch tree, and determine the fixed mutation evaluation result corresponding to the target virus within the fixed mutation period according to the at least one viral evolutionary branch corresponding to the fixed mutation period; and determine the first mutation probability corresponding to the nucleotide site according to the fixed mutation evaluation result corresponding to each fixed mutation period.

[0123] Optionally, the site determination module 304 is specifically used to input the first mutation probability and the second mutation probability corresponding to each nucleotide site of the target virus into a pre-trained site prediction model, so that the site prediction model determines the comprehensive site feature data corresponding to the nucleotide site based on the first mutation probability and the second mutation probability corresponding to the nucleotide site; and determines the key mutation sites from each nucleotide site of the target virus based on the comprehensive site feature data corresponding to each nucleotide site of the target virus.

[0124] Optionally, the site determination module 304 is specifically used to: obtain the first mutation probability and the second mutation probability corresponding to each nucleotide site of the sample virus; for each nucleotide site of the sample virus, input the first mutation probability and the second mutation probability corresponding to the nucleotide site into the site prediction model to be trained, so that the site prediction model to be trained determines the comprehensive site feature data corresponding to the nucleotide site based on the first mutation probability and the second mutation probability corresponding to the nucleotide site, and determines the loss value of the site prediction model to be trained for the nucleotide site based on the comprehensive site feature data corresponding to the nucleotide site and the standard comprehensive site feature data corresponding to the nucleotide site, wherein the magnitude of the loss value is negatively correlated with the feature similarity between the comprehensive site feature data corresponding to the nucleotide site and the comprehensive standard site feature data corresponding to the nucleotide site; determine the total loss value corresponding to the site prediction model to be trained based on the loss value of the site prediction model to be trained for each nucleotide site of the sample virus, and train the site prediction model to be trained based on the total loss value.

[0125] Optionally, the site determination module 304 is specifically used to: for each nucleotide site of the target virus, perform a weighted summation of the first mutation probability and the second mutation probability corresponding to the nucleotide site according to a preset fixed mutation weight and parallel mutation weight, to obtain the comprehensive probability of the variant site corresponding to the nucleotide site; and determine the key variant site corresponding to the target virus from each nucleotide site of the target virus according to a preset comprehensive probability threshold of variant sites and the comprehensive probability of variant sites corresponding to each nucleotide site of the target virus.

[0126] Optionally, the site determination module 304 is specifically used to determine the key mutation sites corresponding to the target virus from each nucleotide site of the target virus based on a preset site screening strategy and according to the first mutation probability and the second mutation probability of each nucleotide site of the target virus.

[0127] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 The provided method is a data prediction method for key mutation sites in RNA viruses.

[0128] This instruction manual also provides Figure 4 The one shown corresponds to Figure 1 A schematic diagram of the structure of an electronic device. (e.g.) Figure 4As shown, at the hardware level, this electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to achieve the above. Figure 1 The data prediction method shown is for key mutation sites in RNA viruses.

[0129] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0130] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0131] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0132] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware components.

[0133] Those skilled in the art will understand that embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0134] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0135] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0136] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0137] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0138] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0139] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0140] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0141] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0142] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0143] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0144] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.

Claims

1. A data prediction method for key mutation sites in RNA viruses, characterized in that, include: Obtain historical viral evolution data corresponding to the target virus, wherein the historical viral evolution data includes at least: viral gene sequence data and host information of the target virus, and viral gene sequence data and host information of each historically evolved virus obtained by the target virus in the course of historical evolution. Based on the historical virus evolution data, construct the virus evolution branch tree corresponding to the target virus; For each nucleotide site of the target virus, based on the viral evolutionary branch tree, a first mutation probability and a second mutation probability corresponding to the nucleotide site are determined. The first mutation probability is used to represent the probability of a fixed mutation occurring at the nucleotide site, and the second mutation probability is used to represent the probability of a parallel mutation occurring at the nucleotide site. Based on the first mutation probability and the second mutation probability corresponding to each nucleotide site of the target virus, the key mutation sites of the target virus are determined; Specifically, based on the viral evolutionary branching tree, for each nucleotide site of the target virus, the first mutation probability corresponding to that nucleotide site is determined, including: Based on a preset fixed mutation cycle duration, the historical evolution duration of the target virus is divided into cycles to determine each fixed mutation cycle corresponding to the target virus. For each fixed mutation cycle corresponding to the target virus, at least one viral evolutionary branch corresponding to the fixed mutation cycle is determined according to the viral evolutionary branch tree, and the fixed mutation evaluation result corresponding to the target virus within the fixed mutation cycle is determined according to the at least one viral evolutionary branch corresponding to the fixed mutation cycle. Based on the fixed mutation evaluation results corresponding to each fixed mutation cycle, the first mutation probability corresponding to the nucleotide site is determined; The formula for determining the first mutation probability is: Among them, targeting each nucleotide site in the target virus , Used to represent this nucleotide site The corresponding first mutation frequency, Used to indicate the historical evolution time of the target virus. This is used to represent the duration of historical evolution based on a preset fixed mutation cycle. The fixed mutation cycles are divided into several groups. Used to represent nucleotide sites In a fixed mutation cycle The fixed number of mutations within, Indicates a fixed mutation cycle The total number of mutations in the target virus; The formula for determining the second mutation probability is: Among them, targeting each nucleotide site in the target virus , Used to represent nucleotide sites The corresponding second mutation probability, Used to represent the number of evolutionary branches in the historical evolutionary process of a target virus. This indicates the number of evolutionary branches. Each evolutionary branch in it, Used to represent nucleotide sites In evolutionary branches The number of parallel mutations that occurred in the process. This is used to indicate the evolutionary branch of the target virus. The total number of mutation events that occurred in the process.

2. The method as described in claim 1, characterized in that, Based on the historical virus evolution data, a virus evolution branch tree corresponding to the target virus is constructed, specifically including: Based on the viral gene sequence data of the target virus in the historical viral evolution data, and the viral gene sequence data of each historically evolved virus corresponding to the target virus, a basic evolutionary tree corresponding to the target virus is constructed. Based on the host information of the target virus and the host information of each historical evolutionary virus corresponding to the target virus, the basic phylogenetic tree is subjected to host origin classification and coloring processing, and the colored basic phylogenetic tree is used as the viral evolutionary branch tree corresponding to the target virus.

3. The method as described in claim 1, characterized in that, Based on the first mutation probability and the second mutation probability corresponding to each nucleotide site of the target virus, the key mutation sites of the target virus are determined, specifically including: For each nucleotide site of the target virus, the first mutation probability and the second mutation probability corresponding to the nucleotide site are input into a pre-trained site prediction model, so that the site prediction model can determine the comprehensive site feature data corresponding to the nucleotide site based on the first mutation probability and the second mutation probability corresponding to the nucleotide site. Based on the comprehensive site feature data corresponding to each nucleotide site of the target virus, key mutation sites are identified from each nucleotide site of the target virus.

4. The method as described in claim 3, characterized in that, The training site prediction model specifically includes: Obtain the first mutation probability and the second mutation probability corresponding to each nucleotide site of the sample virus; For each nucleotide site of the sample virus, the first mutation probability and the second mutation probability corresponding to the nucleotide site are input into the site prediction model to be trained. This allows the site prediction model to determine the comprehensive site feature data corresponding to the nucleotide site based on the first mutation probability and the second mutation probability. Based on the comprehensive site feature data and the standard comprehensive site feature data corresponding to the nucleotide site, the model determines the loss value of the site prediction model for the nucleotide site. The magnitude of the loss value is negatively correlated with the feature similarity between the comprehensive site feature data and the comprehensive standard site feature data corresponding to the nucleotide site. Based on the loss values ​​of the site prediction model to be trained for each nucleotide site of the sample virus, the total loss value corresponding to the site prediction model to be trained is determined, and the site prediction model to be trained is trained based on the total loss value.

5. The method as described in claim 1, characterized in that, Based on the first mutation probability and the second mutation probability corresponding to each nucleotide site of the target virus, the key mutation sites of the target virus are determined, specifically including: For each nucleotide site of the target virus, the first mutation probability and the second mutation probability corresponding to the nucleotide site are weighted and summed according to the preset fixed mutation weight and parallel mutation weight to obtain the comprehensive probability of the mutation site corresponding to the nucleotide site. Based on the preset comprehensive probability threshold of mutation sites and the comprehensive probability of mutation sites corresponding to each nucleotide site of the target virus, the key mutation sites corresponding to the target virus are determined from each nucleotide site of the target virus.

6. The method as described in claim 1, characterized in that, Based on the first mutation probability and the second mutation probability corresponding to each nucleotide site of the target virus, the key mutation sites of the target virus are determined, specifically including: Based on a preset site screening strategy, the key mutation sites corresponding to the target virus are determined from each nucleotide site of the target virus according to the first mutation probability and the second mutation probability of each nucleotide site of the target virus.

7. A data prediction device for key mutation sites in RNA viruses, characterized in that, include: The acquisition module is used to acquire historical viral evolution data corresponding to the target virus. The historical viral evolution data includes at least: viral gene sequence data and host information of the target virus, and viral gene sequence data and host information of each historically evolved virus obtained by the target virus in the course of historical evolution. A construction module is used to construct a virus evolution branch tree corresponding to the target virus based on the historical virus evolution data; The probability determination module is used to determine, for each nucleotide site of the target virus, a first mutation probability and a second mutation probability corresponding to that nucleotide site based on the viral evolutionary branch tree. The first mutation probability is used to represent the probability of a fixed mutation occurring at that nucleotide site, and the second mutation probability is used to represent the probability of a parallel mutation occurring at that nucleotide site. Specifically, based on the viral evolutionary branching tree, for each nucleotide site of the target virus, the first mutation probability corresponding to that nucleotide site is determined, including: Based on a preset fixed mutation cycle duration, the historical evolution duration of the target virus is divided into cycles to determine each fixed mutation cycle corresponding to the target virus. For each fixed mutation cycle corresponding to the target virus, at least one viral evolutionary branch corresponding to the fixed mutation cycle is determined according to the viral evolutionary branch tree, and the fixed mutation evaluation result corresponding to the target virus within the fixed mutation cycle is determined according to the at least one viral evolutionary branch corresponding to the fixed mutation cycle. Based on the fixed mutation evaluation results corresponding to each fixed mutation cycle, the first mutation probability corresponding to the nucleotide site is determined; The formula for determining the first mutation probability is: Among them, targeting each nucleotide site in the target virus , Used to represent this nucleotide site The corresponding first mutation frequency, Used to indicate the historical evolution time of the target virus. This is used to represent the duration of historical evolution based on a preset fixed mutation cycle. The fixed mutation cycles are divided into several groups. Used to represent nucleotide sites In a fixed mutation cycle The fixed number of mutations within, Indicates a fixed mutation cycle The total number of mutations in the target virus; The formula for determining the second mutation probability is: Among them, targeting each nucleotide site in the target virus , Used to represent nucleotide sites The corresponding second mutation probability, Used to represent the number of evolutionary branches in the historical evolutionary process of a target virus. This indicates the number of evolutionary branches. Each evolutionary branch in it, Used to represent nucleotide sites In evolutionary branches The number of parallel mutations that occurred in the process. This is used to indicate the evolutionary branch of the target virus. The total number of mutation events that occurred in the process; The site determination module is used to determine the key mutation sites of the target virus based on the first mutation probability and the second mutation probability corresponding to each nucleotide site of the target virus.

8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1 to 6.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 6.