Genetic variation pathogenicity prediction method, device, storage medium and computer equipment

By designing a genetic variant pathogenic prediction model integrated with multiple target prediction sub-models, the problem that the existing technology cannot predict all mutation types is solved, and a higher accuracy of pathogenic prediction is achieved.

CN114300036BActive Publication Date: 2025-05-06BGI GENOMICS CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111638768.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-05-06
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

The prior art cannot predict all mutation types, resulting in low accuracy of pathogenic prediction.

Method used

A genetic variant pathogenic prediction method is designed to make predictions by obtaining characteristic data of gene variant sites and inputting them into the genetic variant pathogenic prediction model obtained by integrating multiple target predictor models in advance. Each target predictor model uses the characteristic data of the training gene variant site classified by its corresponding pathogenicity rating as the training sample, and trains the pathogenicity rating classification of the training gene variant site as the sample label.

Benefits of technology

This method can be applied to all mutation types and improves the accuracy of predicting pathogenicity of genetic variants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114300036B_ABST
    Figure CN114300036B_ABST
Patent Text Reader

Abstract

The genetic variation pathogenicity prediction method, device, storage medium and computer equipment provided in the present application first determine the characteristic data corresponding to the gene variation site to be predicted, then predict the characteristic data through the genetic variation pathogenicity prediction model, and obtain the genetic variation pathogenicity prediction result; because the genetic variation pathogenicity prediction model is obtained by integrating multiple target prediction sub-models designed according to different pathogenicity level classifications, and each target prediction sub-model is trained using the characteristic data of the training gene variation site of its corresponding pathogenicity level classification as a training sample, and the pathogenicity level classification of the training gene variation site is used as a sample label for training, and the training sample contains all mutation types. Therefore, the genetic variation pathogenicity prediction model is applicable to all mutation types, and can obtain accurate genetic variation pathogenicity prediction results based on the input characteristic data of the gene variation site.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of gene detection technology, and in particular to a method, device, storage medium and computer equipment for predicting the pathogenicity of genetic variation. Background Art

[0002] As high-throughput sequencing technology matures and costs continue to decline, sequencing has become an important means of clinical research and diagnosis. After years of clinical application, major public genome databases have collected a large amount of variant site information. Accurately interpreting these variant sites is the key to achieving human precision medicine.

[0003] In 2015, the American College of Medical Genetics and Genomics (ACMG) released the classification and interpretation standards for gene variant sites, which mainly divided gene variant sites into five categories: pathogenic, likely pathogenic, clinically unclear, likely benign, and benign, to evaluate the pathogenicity and potential risks of variant sites. However, the currently commonly used genetic variation pathogenicity prediction methods mainly target missense mutations or non-synonymous mutations, cannot predict all mutation types, and do not consider the detailed classification of gene variant sites, resulting in low prediction accuracy. Summary of the invention

[0004] The purpose of the present application is to solve at least one of the above-mentioned technical defects, especially the technical defect that the prior art cannot predict all mutation types, resulting in a low prediction accuracy.

[0005] The present application provides a method for predicting pathogenicity of genetic variation, the method comprising:

[0006] Obtaining the gene variation site to be predicted;

[0007] Determining feature data corresponding to the gene variation site to be predicted;

[0008] Inputting the characteristic data into a pre-configured genetic variation pathogenicity prediction model to obtain a genetic variation pathogenicity prediction result output by the genetic variation pathogenicity prediction model;

[0009] Among them, the genetic variation pathogenicity prediction model is obtained by integrating multiple target prediction sub-models designed according to different pathogenicity level classifications, and each target prediction sub-model uses the characteristic data of the training gene variation site of its corresponding pathogenicity level classification as a training sample, and is trained using the pathogenicity level classification of the training gene variation site as a sample label.

[0010] Optionally, the determining of feature data corresponding to the gene variation site to be predicted includes:

[0011] Calling a local annotation file, wherein the local annotation file is an annotation file downloaded in advance from a variable site functional annotation library;

[0012] Searching the local annotation file for feature data corresponding to the gene variation site to be predicted;

[0013] If the feature data corresponding to the gene variation site to be predicted is not found in the local annotation file, the webpage API interface is called to obtain the feature data corresponding to the gene variation site through the webpage API interface.

[0014] Optionally, before inputting the feature data into a preconfigured genetic variation pathogenicity prediction model, the method further comprises:

[0015] Adding derived variables to the feature data according to the missing condition of the feature data;

[0016] The derived variables are displayed as missing feature data for completion;

[0017] Normalize the completed feature data.

[0018] Optionally, the training process of each target prediction sub-model includes:

[0019] Determining an initial prediction sub-model, and obtaining a pathogenicity grade classification corresponding to the initial prediction sub-model, and characteristic data of a training gene variation site of the pathogenicity grade classification;

[0020] Using a k-fold cross-validation method, the characteristic data of the training gene mutation site is divided into k data sets, and k rounds of model training are performed;

[0021] In each round of model training, one of the data sets is selected as a validation set, and the remaining data sets are selected as training sets. The initial prediction sub-model is trained using the feature data in the training set, and the trained initial prediction sub-model is verified using the feature data in the validation set to obtain a verification result.

[0022] According to the pathogenicity grade classification corresponding to the initial prediction sub-model, the model with the highest prediction accuracy in the verification results of k rounds of model training is selected as the target prediction sub-model.

[0023] Optionally, the genetic variation pathogenicity prediction model includes a main direction prediction layer, a direction correction layer, a degree prediction layer and a mapping layer;

[0024] Inputting the characteristic data into a preconfigured genetic variation pathogenicity prediction model to obtain a genetic variation pathogenicity prediction result output by the genetic variation pathogenicity prediction model, including:

[0025] Inputting the characteristic data into the main direction prediction layer, predicting the main direction of genetic variation of the characteristic data, and obtaining a prediction result of the main direction of genetic variation;

[0026] Correcting the main direction of the genetic variation of the main direction prediction result of the genetic variation by the direction correction layer to obtain a main direction correction result of the genetic variation;

[0027] Using the degree prediction layer to predict the probability value of the correction result of the main direction of the genetic variation, to obtain a probability value prediction result;

[0028] The probability value prediction result is mapped to the corresponding mapping interval through the mapping layer to obtain a final mapping score, and the mapping score is used as the genetic variation pathogenicity prediction result.

[0029] Optionally, the target prediction sub-model includes a three-classification model and a two-classification model;

[0030] The three-class model is applied to the main direction prediction layer and the direction correction layer, and the two-class model is applied to the direction correction layer and the degree prediction layer.

[0031] Optionally, the three-classification model applied to the main direction prediction layer includes a first classification model, the first classification model is used to predict the main direction of the genetic variation of the feature data, the main direction of the genetic variation includes pathogenic tendency, unclear clinical significance and benign tendency;

[0032] The three-classification model applied to the direction correction layer includes a second classification model and a third classification model, wherein the second classification model is used to correct pathogenicity, possible pathogenicity, and unclear clinical significance in the main direction of the genetic variation, and the third classification model is used to correct benignity, possible benignity, and unclear clinical significance in the main direction of the genetic variation;

[0033] The binary classification model applied to the direction correction layer includes a fourth classification model, and the fourth classification model is used to correct the pathogenic tendency and benign tendency in the main direction of the genetic variation;

[0034] The binary classification models applied to the degree prediction layer include a fifth classification model, a sixth classification model, a seventh classification model and an eighth classification model;

[0035] Among them, the fifth classification model is used to predict the probability value of pathogenic or possibly pathogenic in the main direction correction results of the genetic variation, the sixth classification model is used to predict the probability value of possibly pathogenic or clinically unclear in the main direction correction results of the genetic variation, the seventh classification model is used to predict the probability value of benign or possibly benign in the main direction correction results of the genetic variation, and the eighth classification model is used to predict the probability value of possibly benign or clinically unclear in the main direction correction results of the genetic variation.

[0036] The present application also provides a genetic variation pathogenicity prediction device, comprising:

[0037] A site acquisition module is used to obtain the gene variation site to be predicted;

[0038] A feature determination module, used to determine feature data corresponding to the gene variation site to be predicted;

[0039] A pathogenicity prediction module, used to input the characteristic data into a pre-configured genetic variation pathogenicity prediction model to obtain a genetic variation pathogenicity prediction result output by the genetic variation pathogenicity prediction model;

[0040] Among them, the genetic variation pathogenicity prediction model is obtained by integrating multiple target prediction sub-models designed according to different pathogenicity level classifications, and each target prediction sub-model uses the characteristic data of the training gene variation site of its corresponding pathogenicity level classification as a training sample, and is trained using the pathogenicity level classification of the training gene variation site as a sample label.

[0041] The present application also provides a storage medium, characterized in that: the storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the genetic variation pathogenicity prediction method as described in any of the above embodiments.

[0042] The present application also provides a computer device, characterized in that it includes: one or more processors, and a memory;

[0043] The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the method for predicting the pathogenicity of genetic variations as described in any one of the above embodiments is executed.

[0044] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0045] The genetic variation pathogenicity prediction method, device, storage medium and computer equipment provided in the present application, when predicting the pathogenicity of the gene variation site to be predicted, first determine the characteristic data corresponding to the gene variation site to be predicted, then predict the characteristic data through a pre-configured genetic variation pathogenicity prediction model, and obtain the genetic variation pathogenicity prediction result; in the present application, since the genetic variation pathogenicity prediction model is obtained by integrating multiple target prediction sub-models designed according to different pathogenicity level classifications, and each target prediction sub-model is trained using the characteristic data of the training gene variation site classified by its corresponding pathogenicity level as a training sample, and the pathogenicity level classification of the training gene variation site is used as a sample label for training, compared with the network model directly trained according to a single mutation type, the present application uses the characteristic data of the training gene variation site classified by different pathogenicity levels as a training sample, so that the training sample contains all mutation types. Therefore, the genetic variation pathogenicity prediction model is applicable to all mutation types, and can obtain accurate genetic variation pathogenicity prediction results according to the input characteristic data of the gene variation site, thereby effectively improving the prediction accuracy of genetic variation pathogenicity. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0047] Figure 1 A schematic diagram of a method for predicting pathogenicity of genetic variation provided in an embodiment of the present application;

[0048] Figure 2 A schematic diagram of the training architecture of each target prediction sub-model provided in the embodiments of the present application;

[0049] Figure 3 A schematic diagram of the structure of training parameters in the model training phase provided in an embodiment of the present application;

[0050] Figure 4 A network architecture diagram of a genetic variation pathogenicity prediction model provided in an embodiment of the present application;

[0051] Figure 5 A schematic diagram of the structure of a genetic variation pathogenicity prediction device provided in an embodiment of the present application;

[0052] Figure 6 A schematic diagram of the internal structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0054] In one embodiment, Figure 1 As shown, Figure 1 A schematic diagram of a process for predicting pathogenicity of a genetic variation provided in an embodiment of the present application; the present application provides a method for predicting pathogenicity of a genetic variation, the method comprising:

[0055] S110: Obtain the gene variation site to be predicted.

[0056] In this step, when predicting the pathogenicity of genetic variation, it is necessary to obtain the gene variation site to be predicted. The source of the gene variation site in this application can be a variation site obtained by nucleic acid sequencing and sequence alignment, a variation site included in a public database or publicly released by others, or an artificially simulated variation site.

[0057] In this application, when using nucleic acid to obtain variant sites by sequencing and sequence alignment, multiple different samples containing nucleic acid can be obtained in advance, and the type of nucleic acid is not particularly limited, and can be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), preferably DNA. For RNA, it can also be reverse transcribed into DNA by experimental methods for subsequent detection and analysis.

[0058] Furthermore, after obtaining multiple samples, the present application can determine the gene variation site of each sample by sequence comparison. For example, the present application can detect the gene sequence of each sample by a gene sequencer and obtain sequencing data, and then map the sequencing data to the reference genome sequence and compare it with the reference genome sequence, so as to find the single base site in the sequencing data that is different from the reference genome sequence, that is, the gene variation site.

[0059] It is understood that the reference genome sequence here refers to a gene sequence fragment with a unique base arrangement, and the position on the chromosome can be accurately located by comparing with such a fragment. In the embodiments of the present application, a reference genome sequence with version number hg38, hg19 or hg18 can be used, without limitation.

[0060] Furthermore, since a person's whole genome sequencing data can mine four million SNVs (single base sites that are different from the reference genome) and half a million indels (insertions or deletions), there can be multiple gene variation sites in this application. The specific value can be determined according to the actual situation and is not limited here.

[0061] S120: Determine feature data corresponding to the gene variation site to be predicted.

[0062] In this step, after the gene variation site to be predicted is obtained through S110, feature data corresponding to the gene variation site to be predicted may be determined.

[0063] Among them, the feature data here can include functional conservation features corresponding to gene variation sites (GERP++, LRT, Phylop and PhastCons), functional impact features (SIFT, Polyphen2, PROVEAN, PrimateAI, MutationAssessor, ClinPred, M_CAP, MVP, F athmm_MKL, TraP, SilVa, usDSM, PrDSM and EnDSM), splicing impact features (SpliceAI), population frequency features (af_popmax) and other feature data.

[0064] The feature data in this application can be obtained by calling local annotation text or calling a web API interface. When calling local annotation text, it can be downloaded in advance from the official variant site annotation database or a third-party variant site annotation database. For example, the dbNSFP database provides functional prediction and annotation data for human gene variant sites, and the dbNSFP database provides a one-stop query and download function for single nucleotide variations (nsSNVs) of non-synonymous mutations.

[0065] Furthermore, after the characteristic data of each gene variation site is obtained, the characteristic data can be preprocessed so that the preprocessed characteristic data can be applied to the model training of the deep neural network.

[0066] When preprocessing the feature data, we can first determine the feature data of each gene variation site that needs to be obtained from multiple dimensions in this application, such as obtaining feature data of four categories: functional conservation features, functional impact features, splicing impact features, and population frequency features. Then, we can check the feature data of each gene variation site that has been obtained to see if there is any missing feature data. If there is any missing, we can fill in the missing feature data with a neutral threshold to enhance the model's ability to learn mutation types.

[0067] In addition, the present application can also perform normalization processing on the feature data, normalizing the feature values ​​of the feature data to the interval [-1, 1], so as to better train the model.

[0068] S130: Inputting the characteristic data into a pre-configured genetic variation pathogenicity prediction model to obtain a genetic variation pathogenicity prediction result output by the genetic variation pathogenicity prediction model.

[0069] In this step, after determining the characteristic data corresponding to the gene variation site through S120, the characteristic data can be input into a pre-configured genetic variation pathogenicity prediction model, so as to predict the characteristic data through the genetic variation pathogenicity prediction model, thereby obtaining a genetic variation pathogenicity prediction result, which can characterize the pathogenicity of the gene variation site to be predicted.

[0070] Among them, the genetic variation pathogenicity prediction model in this application is obtained by integrating multiple target prediction sub-models designed according to different pathogenicity level classifications. It can be understood that according to the ACMG's 5-level classification standard, the pathogenicity of the gene variation site can be divided into pathogenic (P), possibly pathogenic (LP), clinically unclear (VUS), possibly benign (LB) and benign (B). This application can be based on ACMG's 5-level classification standard (P / LP / VUS / LB / B), as well as the interpretation experience of the clinical interpretation team, and combined with the difficulty of distinguishing each pathogenicity level. Design multiple target prediction sub-models, after integrating multiple target prediction sub-models, form the final genetic variation pathogenicity prediction model, which comprehensively covers the multi-dimensional characteristics of gene variation and performs effectiveness screening, so it is closer to clinical application scenarios.

[0071] When integrating multiple target prediction sub-models, the integrated network architecture of the genetic variation pathogenicity prediction model can be determined based on the prediction direction and prediction range of each target prediction sub-model. For example, the target prediction sub-model with more prediction directions and a wider prediction range is used as the network layer close to the input layer in the genetic variation pathogenicity prediction model, and the target prediction sub-model with fewer prediction directions and a narrower prediction range is used as the network layer close to the output layer in the genetic variation pathogenicity prediction model, so that the overall network architecture of the genetic variation pathogenicity prediction model presents a division from coarse granularity to fine granularity, which is more conducive to improving the accuracy of model prediction.

[0072] Furthermore, before making predictions, the target prediction sub-model in the present application can be trained independently for different target prediction sub-models. During the training process, each target prediction sub-model can use the characteristic data of the training gene variant site of its corresponding pathogenicity level classification as a training sample, and use the pathogenicity level classification of the training gene variant site as a sample label to train the model. For example, some target prediction sub-models do not have VUS-rated variant sites in their training data, while some target prediction sub-models only have LB and VUS-rated variant sites in their training data; similarly, before entering the model training, the sample labels need to be reclassified, for example, the sample label of the LB-rated variant site in the target prediction sub-model can be BLB or LB, where BLB refers to a benign tendency.

[0073] In the above embodiment, when predicting the pathogenicity of the gene variation site to be predicted, the characteristic data corresponding to the gene variation site to be predicted is first determined, and then the characteristic data is predicted by a pre-configured genetic variation pathogenicity prediction model, and the genetic variation pathogenicity prediction result is obtained; in the present application, since the genetic variation pathogenicity prediction model is obtained by integrating multiple target prediction sub-models designed according to different pathogenicity level classifications, and each target prediction sub-model is trained using the characteristic data of the training gene variation site classified by its corresponding pathogenicity level as a training sample, and the pathogenicity level classification of the training gene variation site is used as a sample label for training, relative to the network model directly trained according to a single mutation type, the present application uses the characteristic data of the training gene variation site classified by different pathogenicity levels as a training sample, so that the training sample contains all mutation types. Therefore, the genetic variation pathogenicity prediction model is applicable to all mutation types, and can obtain accurate genetic variation pathogenicity prediction results based on the input characteristic data of the gene variation site, thereby effectively improving the prediction accuracy of genetic variation pathogenicity.

[0074] In one embodiment, determining the characteristic data corresponding to the gene variation site to be predicted in S120 may include:

[0075] S121: calling a local annotation file, wherein the local annotation file is an annotation file downloaded in advance from a variant site functional annotation library.

[0076] S122: Searching the local annotation file for feature data corresponding to the gene variation site to be predicted.

[0077] S123: If the feature data corresponding to the gene variation site to be predicted is not found in the local annotation file, calling a web API interface to obtain the feature data corresponding to the gene variation site through the web API interface.

[0078] In this embodiment, when acquiring feature data corresponding to a gene variant site, it can be acquired by calling a local annotation text, and when calling a local annotation text, it can be downloaded and stored locally in advance from an official variant site annotation database or a third-party variant site annotation database. For example, the dbNSFP database provides functional prediction and annotation data for human gene variant sites, and the dbNSFP database provides a one-stop query and download function for single nucleotide variations (nsSNVs) of non-synonymous mutations.

[0079] It is understandable that since the annotation file data in the local repository is large, the gene variation site and the feature data can be associated by creating an index. When acquiring the feature, the gene variation site can be used as a keyword to perform a local query to find the feature data corresponding to the gene variation site.

[0080] Furthermore, since some feature data do not provide download files, the feature data corresponding to the gene mutation site cannot be found through the local annotation file. At this time, the corresponding feature data can be obtained through web page query, such as by calling the web page API interface to obtain the feature data of this part.

[0081] In one embodiment, before inputting the feature data into a preconfigured genetic variation pathogenicity prediction model in S130, the following steps may also be included:

[0082] S210: Adding derived variables to the feature data according to the missing state of the feature data.

[0083] S220: The derived variables are displayed as missing feature data for completion.

[0084] S230: normalizing the completed feature data.

[0085] In this embodiment, after the characteristic data of the gene variation site is obtained, the characteristic data can also be preprocessed, so that the preprocessed characteristic data can be passed into the genetic variation pathogenicity prediction model to obtain the genetic variation pathogenicity prediction result.

[0086] When preprocessing the feature data, the dimensions of the feature data of each gene variation site that needs to be obtained in this application can be determined first, such as obtaining feature data of four categories: functional conservation features, functional impact features, splicing impact features, and population frequency features. Then, the feature data of each gene variation site that has been obtained is checked to see if there is any missing feature data. If so, the missing feature data is supplemented to enhance the model's ability to learn mutation types.

[0087] When checking whether there is any missing feature data, the present application can add a derivative variable to the feature data. The purpose of this variable is to mark whether the feature data is missing (i.e., the feature data does not exist or is not found). If it is marked as missing, the feature needs to be completed with a neutral threshold.

[0088] For example, 10 feature data correspond to 10 corresponding derived variables. If a feature data does not exist or is not found, that is, it is missing, the derived variable of the feature data is set to 1; conversely, if the feature data exists, the derived variable of the feature data is set to 0.

[0089] After adding derived variables to each feature data, since the missing feature data itself is still missing, in order to further improve the learning ability of the model, this application can complete the feature data displayed as missing in the derived variables, thereby enhancing the model's learning ability for mutation types. When completing the missing feature data, the official recommended threshold of the feature source corresponding to the missing feature data can be used to fill it. Since the official recommended threshold represents the most neutral indicator, filling it with the official recommendation will not affect the learning ability of the model, but can also improve the training effect of the model.

[0090] In addition, the present application can also perform normalization processing on the completed feature data. Specifically, the feature values ​​of the feature data can be normalized to the interval [-1, 1] to facilitate better training of the model.

[0091] In one embodiment, the training process of each target prediction sub-model may include:

[0092] S310: Determine an initial prediction sub-model, and obtain the pathogenicity grade classification corresponding to the initial prediction sub-model, and characteristic data of the training gene variation site of the pathogenicity grade classification.

[0093] In this step, when training each target prediction sub-model, the initial prediction sub-model may be determined first, and then the pathogenicity grade classification corresponding to the initial prediction sub-model and the characteristic data of the training gene variation sites of the pathogenicity grade classification are obtained.

[0094] Among them, when determining the initial prediction sub-model, a whole set of classification rules can be designed based on clinical interpretation experience, and different classification strategies can be adopted in different links. For example, some links need to distinguish three categories, then a three-classification model can be selected; if you only plan to distinguish two categories, you can choose a two-classification model.

[0095] When training the model, each target prediction sub-model in the present application mainly uses the supervised learning method, that is, the training data includes training samples and sample labels. The training samples in the present application can be the characteristic data of the training gene variant sites of the pathogenicity level classification corresponding to the initial prediction sub-model, and the sample labels can be the pathogenicity level classification corresponding to the initial prediction sub-model. For example, when the pathogenicity level classification corresponding to some target prediction sub-models is P, or LP, or B, or LB, there are no VUS-rated variant sites in the training data of this type of target prediction sub-model, and when the genetic variation types corresponding to some target prediction sub-models are LB and VUS, there are only LB and VUS-rated variant sites in the training data of this type of target prediction sub-model; similarly, the sample labels will also be reclassified, for example, the sample label of the LB-rated variant site in the target prediction sub-model can be BLB or LB.

[0096] Furthermore, when training the model, gene variant sites with a review status of 2 stars or above in the ClinVar database of the National Center for Biotechnology Information (NCBI) of the United States can be selected as training gene variant sites. Generally speaking, the Review status in the Clinvar database is divided into 0-4 stars. The higher the star rating, the more reliable the interpretation of the gene variant site. This application uses gene variant sites with 2 stars or above and the corresponding rating results for model training, which can effectively improve the reliability of the model.

[0097] S320: Using a k-fold cross-validation method, the feature data of the training gene variation site is divided into k data sets, and k rounds of model training are performed.

[0098] In this step, when each target prediction sub-model is trained, a k-fold cross validation method can be used for batch training, and the model with the highest prediction accuracy is finally retained as the final target prediction sub-model. In addition, each target prediction sub-model in this application is independently performed during k-fold cross validation without interfering with each other, which can improve the prediction accuracy of the model to a greater extent.

[0099] It can be understood that the k-fold cross validation method is to divide the training data into k parts, one of which is used as validation data and the other k-1 parts are used as training data, and the best model is obtained after k cross-repeating. In the present invention, k=5 can be set, and k rounds of model training can be implemented through the scikit-learn library.

[0100] S330: In each round of model training, one of the data sets is selected as a validation set and the remaining data sets are selected as training sets. The initial prediction sub-model is trained using the feature data in the training set, and the trained initial prediction sub-model is verified using the feature data in the validation set to obtain a verification result.

[0101] In this step, when performing k batches of model training, in each round of model training, one of the data sets can be selected as the validation set and the remaining data sets as the training set. The feature data in the training set is used to train the initial prediction sub-model, and the feature data in the validation set is used to verify the trained initial prediction sub-model, thereby obtaining a verification result.

[0102] For example, the present application can use Pytorch to construct a deep neural network, and the number of neurons in the input layer is the number of feature data used for training. For example, the input layer is {x i |x 1 ,x 2 ,...,x m}, where x i is the i-th feature data of the input, and m is the number of feature data of the input; the number of neurons in the output layer varies according to the number of classifications (e.g., 1 for binary classification and 3 for ternary classification). For example, the output of a deep neural network can be calculated as O = HW o +b o , H=R(XW h +b h ), where the output of the deep neural network is O, the output of the hidden layer is H, R is the activation function, X is the input feature matrix, and W h and b h are the weight and bias parameters of the hidden layer, W o and b o are the weight and bias parameters of the output layer respectively; the weight update formula can be Among them, w is the neural network weight, η is the learning rate, α is the weight parameter of L2 regularization, and Loss is the loss function.

[0103] It is understandable that learning rate, L2 regularization weight parameter, dropout ratio and batch size are all hyperparameters set during model training. Hyperparameters are provided before training, and constantly changing hyperparameters to train and evaluate models is also called the process of parameter adjustment. Regularization is often used in machine learning to control the complexity of the model and reduce overfitting. L2 regularization is one of the methods. The weight of L2 regularization and the dropout rate vary depending on the type of model.

[0104] Furthermore, for the binary classification model, the present application can use the BCELoss and sigmoid loss and activation function combination to form BCEWithLogits Loss. The specific formula is as follows:

[0105] l n =-w n [y n ·logσ(x n )+(1-y n )·log(1-σ(x n ))]

[0106] Where N is the batch size, n is the nth site in each batch, and x n Represents the neural network output value of the nth site in a batch, y n Indicates the sample label value of the nth site in a batch, l n represents the loss function value of the nth site in a batch, w is the weight value, σ(x n ) is an activation function, which can be specifically

[0107] For the three-classification model, this application can use the combination of NLLLoss and softmax loss and activation function to form CrossEntropy Loss. The specific formula is as follows:

[0108]

[0109] Among them, class is the classification label value, C is the total number of categories, j is the jth category, x[j] represents the neural network output value of the jth category, and x[class] specifies the neural network output value predicted as the target category (true category).

[0110] S340: According to the pathogenicity grade classification corresponding to the initial prediction sub-model, the model with the highest prediction accuracy in the verification results of k rounds of model training is selected as the target prediction sub-model.

[0111] In this step, after k rounds of model training for the initial prediction sub-model, the prediction accuracy after each round of model training can be verified according to the pathogenicity level classification corresponding to the initial prediction sub-model, and then the initial prediction sub-model with the highest prediction accuracy in the verification results of k rounds of model training is selected as the target prediction sub-model, which can improve the generalization ability of the model.

[0112] In one embodiment, the genetic variation pathogenicity prediction model may include a main direction prediction layer, a direction correction layer, a degree prediction layer, and a mapping layer.

[0113] Inputting the characteristic data into a pre-configured genetic variation pathogenicity prediction model in S130 to obtain a genetic variation pathogenicity prediction result output by the genetic variation pathogenicity prediction model may include:

[0114] S131: Input the feature data into the main direction prediction layer, predict the main direction of genetic variation of the feature data, and obtain a prediction result of the main direction of genetic variation.

[0115] S132: Correcting the main direction of the genetic variation of the main direction prediction result of the genetic variation by the direction correction layer to obtain a main direction correction result of the genetic variation.

[0116] S133: Using the degree prediction layer to perform probability value prediction on the correction result of the main direction of the genetic variation, to obtain a probability value prediction result.

[0117] S134: Mapping the probability value prediction result to the corresponding mapping interval through the mapping layer to obtain a final mapping score, and using the mapping score as the genetic variation pathogenicity prediction result.

[0118] In this embodiment, when predicting the pathogenicity of genetic abnormalities at gene mutation sites, the network architecture of the genetic variation pathogenicity prediction model can be determined based on the prediction direction and prediction range of each target prediction sub-model. For example, a target prediction sub-model with more prediction directions and a wider prediction range is used as a network layer close to the input layer in the genetic variation pathogenicity prediction model, and a target prediction sub-model with fewer prediction directions and a narrower prediction range is used as a network layer close to the output layer in the genetic variation pathogenicity prediction model, so that the overall network architecture of the genetic variation pathogenicity prediction model presents a division from coarse granularity to fine granularity, which is more conducive to improving the accuracy of model prediction.

[0119] Specifically, when predicting the pathogenicity of genetic abnormalities at gene mutation sites, the characteristic data of the gene mutation sites can be input into the main direction prediction layer, and the main direction prediction layer is used to predict the main direction of the genetic variation of the characteristic data to obtain the main direction prediction result of the genetic variation. Then, according to the main direction prediction result of the genetic variation, the respective direction correction layers are entered to perform direction correction. After obtaining the main direction correction result of the genetic variation, the degree prediction layer is entered to perform probability value prediction on the main direction correction result of the genetic variation to obtain the probability value prediction result. Finally, the mapping layer is entered to perform probability value prediction on the main direction correction result of the genetic variation to obtain the probability value prediction result.

[0120] In one embodiment, the target prediction sub-model may include a three-classification model and a two-classification model; the three-classification model is applied to the main direction prediction layer and the direction correction layer, and the two-classification model is applied to the direction correction layer and the degree prediction layer.

[0121] In this embodiment, the target prediction sub-model can be divided into a binary classification model and a three-classification model. Among them, since the three-classification model can predict three categories, the three-classification model can be applied to the main direction prediction layer and direction correction layer of the genetic variation pathogenicity prediction model, and the binary classification model can predict one of the two categories. Therefore, the binary classification model can be applied to the direction correction layer and degree prediction layer of the genetic variation pathogenicity prediction model.

[0122] For example, the present application can design eight binary or three-classification models based on clinical medical interpretation experience, and perform model training on these eight binary or three-classification models. Figure 2 , Figure 2 Schematic diagram of the training architecture of each target prediction sub-model provided in the embodiment of the present application; it can be understood that according to the ACMG's 5-level classification standard, the pathogenicity of the gene variant site can be divided into pathogenic (P), possibly pathogenic (LP), clinically unclear (VUS), possibly benign (LB) and benign (B). The present application can design multiple target prediction sub-models based on the ACMG's 5-level classification standard (P / LP / VUS / LB / B), the interpretation experience of the clinical interpretation team, and the difficulty of each rating classification, such as Figure 2 The PLP / VUS / BLB three-classification model, the P / LP / VUS three-classification model, the PLP / BLB two-classification model, the B / LB / VUS three-classification model, the P / LP two-classification model, the LP / VUS two-classification model, the B / LB two-classification model and the LB / VUS two-classification model.

[0123] Among them, the PLP / VUS / BLB three-classification model is a model formed by combining P, LP, VUS, LB, and B to predict the main direction of genetic variation of feature data, and the main direction of genetic variation may include pathogenic tendency (PLP), unclear clinical significance (VUS), and benign tendency (BLB); the P / LP / VUS three-classification model is a model formed by combining P, LP, and VUS to correct pathogenic, possible pathogenic, and unclear clinical significance in the main direction of genetic variation, and so on, thereby obtaining eight two-classification models or three-classification models.

[0124] Furthermore, when training eight binary or three-class models, the hyperparameters used in training are different due to differences in training samples and sample labels. For details, see Figure 3 , Figure 3 A schematic diagram of the structure of training parameters in the model training phase provided in an embodiment of the present application; Figure 3 The training parameters in are different according to the pathogenicity classification of the neural network structure, and the types of training parameters may include optimizer, activation function, loss function, regularization, batch size, dropout ratio, learning rate, etc., which are not limited here. In addition, Figure 3 The values ​​of each training parameter in are the final parameter values ​​determined after repeated parameter adjustments.

[0125] Furthermore, when training the model, sub-model training can be performed separately according to the pathogenicity level (P / LP / VUS / LB / B). Each gene mutation site can be predicted by all models to obtain multiple prediction scores, select the output value of the most reliable classification, and obtain a technical effect similar to that of the present application.

[0126] In one embodiment, the three-classification model applied to the main direction prediction layer may include a first classification model, which is used to predict the main direction of genetic variation of the feature data, wherein the main direction of genetic variation includes pathogenic tendency, unclear clinical significance and benign tendency.

[0127] The three-classification model applied to the direction correction layer may include a second classification model and a third classification model, wherein the second classification model is used to correct pathogenic, possibly pathogenic, and unclear clinical significance in the main direction of the genetic variation, and the third classification model is used to correct benign, possibly benign, and unclear clinical significance in the main direction of the genetic variation.

[0128] The binary classification model applied to the direction correction layer may include a fourth classification model, and the fourth classification model is used to correct the pathogenic tendency and benign tendency in the main direction of the genetic variation.

[0129] The binary classification model applied to the degree prediction layer may include a fifth classification model, a sixth classification model, a seventh classification model and an eighth classification model; wherein the fifth classification model is used to predict the probability value of pathogenic or possibly pathogenic in the main direction correction results of the genetic variation, the sixth classification model is used to predict the probability value of possibly pathogenic or clinically unclear significance in the main direction correction results of the genetic variation, the seventh classification model is used to predict the probability value of benign or possibly benign in the main direction correction results of the genetic variation, and the eighth classification model is used to predict the probability value of possibly benign or clinically unclear significance in the main direction correction results of the genetic variation.

[0130] In this embodiment, in the genetic variation pathogenicity prediction model, the eight target prediction sub-models are at different levels and play different classification effects. Figure 4 As shown, Figure 4 A network architecture diagram of a genetic variation pathogenicity prediction model provided in an embodiment of the present application; Figure 4 In the main direction prediction layer of the genetic variation pathogenicity prediction model, only the first classification model, such as the PLP / VUS / BLB model 10, is used to predict the main direction of the genetic variation, which are the PLP direction, the VUS direction and the BLB direction, namely pathogenic tendency, unclear clinical significance and benign tendency. The direction correction layer includes the second classification model, the third classification model and the fourth classification model, wherein the second classification model can be the P / LP / VUS model 20, the third classification model can be the B / LB / VUS model 30, and the fourth classification model can be the PLP / BLB model 40, which are respectively used for the correction of the three main directions of the main direction prediction layer. At the same time, the prediction results of this layer determine the mapping interval of the subsequent mapping layer, and the mapping interval is formulated in combination with bioinformatics and clinical medical interpretation experience.

[0131] Further, Figure 4 The degree prediction layer in the example may include the fifth classification model, the sixth classification model, the seventh classification model and the eighth classification model, such as the fifth classification model may be the P / LP model 50, the sixth classification model may be the LP / VUS model 60, the seventh classification model may be the B / LB model 70, and the eighth classification model may be the LB / VUS model 80. This layer needs to retain the probability value predicted by the model for the score mapping of the subsequent mapping layer. The mapping layer in the present application mainly combines the mapping interval of the direction correction layer and the probability value of the degree prediction layer to obtain the final mapping score, which is the prediction result of the pathogenicity of the genetic variation.

[0132] Combination Figure 4 For example, the feature matrix of a variant site can first be passed into the PLP / VUS / BLB model 10. If the classification result is VUS, it will be passed into the PLP / BLB model 40. If the classification result is PLP, it will be passed into the LP / VUS model 60. The final prediction score will be mapped in the range of [0.5, 0.85] as the prediction result of the pathogenicity of the genetic variation.

[0133] It is understandable that the output result of the genetic variation pathogenicity prediction model in this application is a score between [0, 1], and this application can determine the clear pathogenicity by dividing the threshold. The threshold can be customized by the user according to the trained model and application requirements, for example, [0.95, 1] ​​is defined as P, [0.85, 0.95) as LP, [0.5, 0.85) as VUS, [0.1, 0.5) as LB, and [0, 0.1) as B.

[0134] The genetic variation pathogenicity prediction device provided in the embodiments of the present application is described below. The genetic variation pathogenicity prediction device described below and the genetic variation pathogenicity prediction method described above can be referenced to each other.

[0135] In one embodiment, Figure 5 As shown, Figure 5 A schematic diagram of the structure of a genetic variation pathogenicity prediction device provided in an embodiment of the present application; the present application also provides a genetic variation pathogenicity prediction device, including a site acquisition module 210, a feature determination module 220, and a pathogenicity prediction module 230, which specifically include the following:

[0136] The site acquisition module 210 is used to acquire the gene variation site to be predicted.

[0137] The feature determination module 220 is used to determine feature data corresponding to the gene variation site to be predicted.

[0138] The pathogenicity prediction module 230 is used to input the characteristic data into a pre-configured genetic variation pathogenicity prediction model to obtain the genetic variation pathogenicity prediction result output by the genetic variation pathogenicity prediction model.

[0139] Among them, the genetic variation pathogenicity prediction model is obtained by integrating multiple target prediction sub-models designed according to different pathogenicity level classifications, and each target prediction sub-model uses the characteristic data of the training gene variation site of its corresponding pathogenicity level classification as a training sample, and is trained using the pathogenicity level classification of the training gene variation site as a sample label.

[0140] In the above embodiment, when predicting the pathogenicity of the gene variation site to be predicted, the characteristic data corresponding to the gene variation site of the sample to be tested is first determined, and then the characteristic data is predicted by a pre-configured genetic variation pathogenicity prediction model, and the genetic variation pathogenicity prediction result is obtained; in the present application, since the genetic variation pathogenicity prediction model is obtained by integrating multiple target prediction sub-models classified according to different pathogenicity levels, and each target prediction sub-model is trained using the characteristic data of the training gene variation site classified according to its corresponding pathogenicity level as a training sample, and the pathogenicity level classification of the training gene variation site is used as a sample label for training, relative to the network model directly trained according to a single mutation type, the present application uses the characteristic data of the training gene variation site classified according to different pathogenicity levels as a training sample, so that the training sample contains all mutation types. Therefore, the genetic variation pathogenicity prediction model is applicable to all mutation types, and can obtain accurate genetic variation pathogenicity prediction results according to the input characteristic data of the gene variation site, thereby effectively improving the prediction accuracy of genetic variation pathogenicity.

[0141] In one embodiment, the present application also provides a storage medium, characterized in that: the storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the genetic variation pathogenicity prediction method as described in any of the above embodiments.

[0142] In one embodiment, the present application further provides a computer device, characterized in that it includes: one or more processors, and a memory.

[0143] The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the method for predicting the pathogenicity of genetic variations as described in any one of the above embodiments is executed.

[0144] Indicatively, Figure 6 As shown, Figure 6 This is a schematic diagram of the internal structure of a computer device provided in an embodiment of the present application. The computer device 300 may be provided as a server. Figure 6 The computer device 300 includes a processing component 302, which further includes one or more processors, and a memory resource represented by a memory 301, for storing instructions that can be executed by the processing component 302, such as an application. The application stored in the memory 301 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 302 is configured to execute instructions to perform the genetic variation pathogenicity prediction method of any of the above embodiments.

[0145] The computer device 300 may further include a power supply component 303 configured to perform power management of the computer device 300, a wired or wireless network interface 304 configured to connect the computer device 300 to a network, and an input / output (I / O) interface 305. The computer device 300 may operate based on an operating system stored in the memory 301, such as Mac OSX™, Linux™, Free BSD™, or the like.

[0146] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0147] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0148] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can refer to each other.

[0149] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for predicting pathogenicity of genetic variation, characterized in that: The method comprises: Obtaining the gene variation site to be predicted; Determining feature data corresponding to the gene variation site to be predicted; Inputting the characteristic data into a pre-configured genetic variation pathogenicity prediction model to obtain a genetic variation pathogenicity prediction result output by the genetic variation pathogenicity prediction model; The genetic variation pathogenicity prediction model is obtained by integrating multiple target prediction sub-models designed according to different pathogenicity level classifications, and each target prediction sub-model is trained using the characteristic data of the training gene variation site of its corresponding pathogenicity level classification as a training sample, and the pathogenicity level classification of the training gene variation site as a sample label; The training process of each target prediction sub-model includes: Determining an initial prediction sub-model, and obtaining a pathogenicity grade classification corresponding to the initial prediction sub-model, and characteristic data of a training gene variation site of the pathogenicity grade classification; Using a k-fold cross-validation method, the characteristic data of the training gene mutation site is divided into k data sets, and k rounds of model training are performed; In each round of model training, one of the data sets is selected as a validation set, and the remaining data sets are selected as training sets. The initial prediction sub-model is trained using the feature data in the training set, and the trained initial prediction sub-model is verified using the feature data in the validation set to obtain a verification result. According to the pathogenicity grade classification corresponding to the initial prediction sub-model, the model with the highest prediction accuracy in the verification results of k rounds of model training is selected as the target prediction sub-model; The genetic variation pathogenicity prediction model includes a main direction prediction layer, a direction correction layer, a degree prediction layer and a mapping layer; Inputting the characteristic data into a preconfigured genetic variation pathogenicity prediction model to obtain a genetic variation pathogenicity prediction result output by the genetic variation pathogenicity prediction model, including: Inputting the characteristic data into the main direction prediction layer, predicting the main direction of genetic variation of the characteristic data, and obtaining a prediction result of the main direction of genetic variation; Correcting the main direction of the genetic variation of the main direction prediction result of the genetic variation by the direction correction layer to obtain a main direction correction result of the genetic variation; Using the degree prediction layer to predict the probability value of the correction result of the main direction of the genetic variation, to obtain a probability value prediction result; The probability value prediction result is mapped to the corresponding mapping interval through the mapping layer to obtain a final mapping score, and the mapping score is used as the genetic variation pathogenicity prediction result.

2. The method according to claim 1, characterized in that The determining of feature data corresponding to the gene variation site to be predicted includes: Calling a local annotation file, wherein the local annotation file is an annotation file downloaded in advance from a variable site functional annotation library; Searching the local annotation file for feature data corresponding to the gene variation site to be predicted; If the feature data corresponding to the gene variation site to be predicted is not found in the local annotation file, the webpage API interface is called to obtain the feature data corresponding to the gene variation site through the webpage API interface.

3. The method according to claim 1 or 2, characterized in that: Before inputting the feature data into a pre-configured genetic variation pathogenicity prediction model, the method further includes: Adding derived variables to the feature data according to the missing condition of the feature data; The derived variables are displayed as missing feature data for completion; Normalize the completed feature data.

4. The method according to claim 1, characterized in that: The target prediction sub-model includes a three-classification model and a two-classification model; The three-class model is applied to the main direction prediction layer and the direction correction layer, and the two-class model is applied to the direction correction layer and the degree prediction layer.

5. The method according to claim 1, characterized in that The three-classification model applied to the main direction prediction layer includes a first classification model, wherein the first classification model is used to predict the main direction of the genetic variation of the feature data, wherein the main direction of the genetic variation includes pathogenic tendency, unclear clinical significance, and benign tendency; The three-classification model applied to the direction correction layer includes a second classification model and a third classification model, wherein the second classification model is used to correct pathogenicity, possible pathogenicity, and unclear clinical significance in the main direction of the genetic variation, and the third classification model is used to correct benignity, possible benignity, and unclear clinical significance in the main direction of the genetic variation; The binary classification model applied to the direction correction layer includes a fourth classification model, and the fourth classification model is used to correct the pathogenic tendency and benign tendency in the main direction of the genetic variation; The binary classification models applied to the degree prediction layer include a fifth classification model, a sixth classification model, a seventh classification model and an eighth classification model; Among them, the fifth classification model is used to predict the probability value of pathogenic or possibly pathogenic in the main direction correction results of the genetic variation, the sixth classification model is used to predict the probability value of possibly pathogenic or clinically unclear in the main direction correction results of the genetic variation, the seventh classification model is used to predict the probability value of benign or possibly benign in the main direction correction results of the genetic variation, and the eighth classification model is used to predict the probability value of possibly benign or clinically unclear in the main direction correction results of the genetic variation.

6. A genetic variation pathogenicity prediction device, characterized in that: include: A site acquisition module is used to obtain the gene variation site to be predicted; A feature determination module, used to determine feature data corresponding to the gene variation site to be predicted; A pathogenicity prediction module, used to input the characteristic data into a pre-configured genetic variation pathogenicity prediction model to obtain a genetic variation pathogenicity prediction result output by the genetic variation pathogenicity prediction model; The genetic variation pathogenicity prediction model is obtained by integrating multiple target prediction sub-models designed according to different pathogenicity level classifications, and each target prediction sub-model is trained using the characteristic data of the training gene variation site of its corresponding pathogenicity level classification as a training sample, and the pathogenicity level classification of the training gene variation site as a sample label; The training process of each target prediction sub-model includes: Determining an initial prediction sub-model, and obtaining a pathogenicity grade classification corresponding to the initial prediction sub-model, and characteristic data of a training gene variation site of the pathogenicity grade classification; Using a k-fold cross-validation method, the characteristic data of the training gene mutation site is divided into k data sets, and k rounds of model training are performed; In each round of model training, one of the data sets is selected as a validation set, and the remaining data sets are selected as training sets. The initial prediction sub-model is trained using the feature data in the training set, and the trained initial prediction sub-model is verified using the feature data in the validation set to obtain a verification result. According to the pathogenicity grade classification corresponding to the initial prediction sub-model, the model with the highest prediction accuracy in the verification results of k rounds of model training is selected as the target prediction sub-model; The genetic variation pathogenicity prediction model includes a main direction prediction layer, a direction correction layer, a degree prediction layer and a mapping layer; The pathogenicity prediction module comprises: Inputting the characteristic data into the main direction prediction layer, predicting the main direction of genetic variation of the characteristic data, and obtaining a prediction result of the main direction of genetic variation; Correcting the main direction of the genetic variation of the main direction prediction result of the genetic variation by the direction correction layer to obtain a main direction correction result of the genetic variation; Using the degree prediction layer to predict the probability value of the correction result of the main direction of the genetic variation, to obtain a probability value prediction result; The probability value prediction result is mapped to the corresponding mapping interval through the mapping layer to obtain a final mapping score, and the mapping score is used as the genetic variation pathogenicity prediction result.

7. A storage medium, characterized in that: The storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the method for predicting the pathogenicity of genetic variations as described in any one of claims 1 to 5.

8. A computer device, characterized in that: include: one or more processors, and memory; The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the method for predicting pathogenicity of genetic variations as described in any one of claims 1 to 5 are performed.

Citation Information

Patent Citations

  • Genome unit point variation pathogenicity prediction method and system and storage medium

    CN110245685A

  • Gene variation site pathogenicity prediction method and system based on neural network

    CN113808662A