Intelligent analysis method and device for cancer-related genes based on weighted voting, medium

CN121789766BActive Publication Date: 2026-08-11XIAN ZHONGMEI HONGKANG BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

但是,由于每一种计算生物学工具的基本原理和适用场景不同,以及机器学习模型的性能依赖于训练数据的质量,因此依靠单一工具的癌症相关基因智能分析可靠性不足

Benefits of technology

本发明中,鉴于现有的癌症基因预测方法存在以下问题:由于每一种计算生物学工具的基本原理和适用场景不同,以及机器学习模型的性能依赖于训练数据的质量,因此依靠单一工具的癌症相关基因智能分析可靠性不足。为此,本发明首先使用不同的计算生物学工具进行初步的预测,由于不同的计算生物学工具的预测原理和适用场景不同,因此综合多种计算生物学工具能够更加全面可靠地对待测基因进行分析。然后再以构造非线性特征的方式融合多种计算生物学工具的预测分数,最后将待测变异点的预测分数、置信度和非线性扩展特征输入到预先训练好的多个一级模型中,获得多个一级模型的输出结果。此时多个模型的输出结果会因为不同一级模型的底层原理不同而存在不同,例如有些模型无法预测分析的变异结果但另一些模型就能够分析出来。因此,本发明采用加权融合的方法对不同的一级模型的输出结果进行融合,基于最终的融合结果,分析得到待测基因数据对应的功能预测结果。因此,本发明能够对待测基因数据进行更加全面综合的分析,综合不同分析工具和不同模型的分析结果,最终从整体上提高了分析结果的可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789766B_ABST
    Figure CN121789766B_ABST
Patent Text Reader

Abstract

This invention relates to the field of gene analysis technology and discloses a weighted voting-based intelligent analysis method for cancer-related genes. The invention first uses different computational biology tools to perform preliminary predictions on the target gene. Then, it fuses the prediction scores from multiple computational biology tools by constructing nonlinear features. Finally, the prediction scores, confidence levels, and nonlinear extension features of the target variant are input into multiple pre-trained first-level models to obtain the output results of these models. A weighted fusion method is used to fuse the output results of different first-level models. Based on the final fusion result, the functional prediction results corresponding to the target gene data are analyzed. Therefore, this invention enables a more comprehensive analysis of the target gene data, integrating the analysis results of different tools and models, ultimately improving the overall reliability of the analysis results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene analysis technology, and in particular to a method, device, and medium for intelligent analysis of cancer-related genes based on weighted voting. Background Technology

[0002] Single nucleotide variants (SNPs / SNVs) in cancer-related genes (such as driver genes, tumor suppressor genes, and key genes in signaling pathways) have significant clinical implications for tumor development, prognostic assessment, and personalized treatment. With the interdisciplinary development of artificial intelligence and bioinformatics, several relatively mature computational biology tools have emerged, including SIFT (Sorting Intolerant From Tolerant), CADD (Combined Annotation Dependent Depletion), PredictSNP, and REVEL (Rare Exome Variant Ensemble Learner). These tools can predict the potential functional impact of SNPs in cancer-related genes, providing prediction scores and confidence levels. Furthermore, well-trained machine learning models can also predict the potential functional impact of SNPs in cancer-related genes. However, due to the different fundamental principles and applicable scenarios of each computational biology tool, and the fact that the performance of machine learning models depends on the quality of training data, the reliability of intelligent analysis of cancer-related genes relying on a single tool is insufficient. Summary of the Invention

[0003] This invention provides a method, device, and medium for intelligent analysis of cancer-related genes based on weighted voting, which improves the reliability of intelligent analysis of cancer-related genes.

[0004] The first aspect of this invention discloses a method for intelligent analysis of cancer-related genes based on weighted voting, the method comprising: Acquire the gene data to be tested, and use a variety of computational biology tools to analyze the gene data to be tested, and obtain the prediction scores and confidence scores of all the variant points to be tested corresponding to the gene data to be tested, wherein the prediction scores are used to indicate the classification results of the corresponding variant points to be tested; Based on the predicted scores analyzed by various computational biology tools, a nonlinear extended feature of all test variants corresponding to the cancer-related gene data is constructed. The weighted voting prediction stage includes: inputting the prediction scores, confidence levels, and nonlinear extension features of all the variants to be tested into multiple pre-trained first-level models to obtain the output results of multiple first-level models; weighting and fusing the output results of all the first-level models according to the preset weight parameters corresponding to each first-level model to obtain a fused output result; and determining the functional prediction result corresponding to the gene data to be tested based on the fused output result.

[0005] A second aspect of this invention discloses a weighted voting-based intelligent analysis device for cancer-related genes, the device comprising: The data acquisition module is used to acquire the gene data to be tested, analyze the gene data to be tested using a variety of computational biology tools, and obtain the prediction scores and confidence levels of all the variant points to be tested corresponding to the gene data to be tested. The prediction scores are used to indicate the classification results of the corresponding variant points to be tested. The tool calculation module is used to construct nonlinear extended features of all test variants corresponding to the cancer-related gene data based on the predicted scores analyzed by various computational biology tools. The voting prediction module is used to perform the operations of the weighted voting prediction stage, including: inputting the predicted scores, confidence levels and nonlinear extension features of all the variants to be tested into multiple pre-trained first-level models to obtain the output results of multiple first-level models; weighting and fusing the output results of all the first-level models according to the preset weight parameters corresponding to each first-level model to obtain a fused output result; and determining the functional prediction result corresponding to the gene data to be tested based on the fused output result.

[0006] As an optional implementation, in a second aspect of the invention, the primary model is obtained through the following steps: Obtain a first training dataset and a validation dataset, wherein the first training dataset and the validation dataset include cancer-related gene data and variant classification labels corresponding to the cancer-related gene data; The first training dataset was analyzed using a variety of computational biology tools to obtain the prediction scores and confidence levels of all training variants corresponding to the cancer-related gene data. Based on the predicted scores analyzed by various computational biology tools, a nonlinear extended feature of all training variants corresponding to the cancer-related gene data is constructed. The model training steps include: The predicted scores, confidence levels, and nonlinear extension features of all training variants are used as the second training dataset to train multiple machine learning models, resulting in multiple well-trained first-level models. For each mutation classification in the validation dataset: validate multiple trained first-level models based on the validation data corresponding to the mutation classification, obtain the validation result for each first-level model, and determine the class weight Q for each first-level model based on the validation result; Obtain the class weight Q for each of the first-level models in each variant classification. ji ; Where j indicates the j-th first-level model and i indicates the i-th mutation classification.

[0007] As an optional implementation, in a second aspect of the invention, the output of the first-level model includes the predicted classification result for each variable point to be tested; and the preset weight parameters are obtained through the following steps: For each of the first-level models, based on the output of that first-level model, determine the proportion P of each predicted classification result in the output of that first-level model. i Based on the class weights Q corresponding to all variant categories in this first-level model. ji The weight parameter K corresponding to this first-level model is calculated using the following formula. j :

[0008] Where I refers to the number of categories predicted in the output of the first-level model.

[0009] As an optional implementation, in the second aspect of the invention, the various computational biology tools include a variety of tools such as SIFT, CADD, PredictSNP, REVEL, and MCAP. The various first-level models are of multiple types, including random forest model, decision tree model, logistic regression model, SVM model, and feedforward neural network model.

[0010] As an optional implementation, in a second aspect of the present invention, the first training dataset further includes pathological state labels corresponding to the cancer-related gene data; The device also includes: The weight pre-setting module is used to divide all training variants into multiple state groups according to the pathological state labels corresponding to the cancer-related genes; obtain the prediction scores, confidence levels, and nonlinear expansion features of all training variants in each state group; for each state group: re-trigger the model training step to obtain multiple trained first-level models corresponding to the state group, and the class weight Q of each first-level model corresponding to the state group for each variant classification. ji .

[0011] As an optional implementation, in a second aspect of the invention, the apparatus further includes: The pathology analysis module is used to input the predicted scores and confidence levels of all test variants analyzed by various computational biology tools into a pre-trained pathology state analysis model, obtain the output results of the pathology state analysis model, and determine the pathology state corresponding to the test gene based on the output results of the pathology state analysis model. The pathological state analysis model is obtained through the following steps: Determine the state label corresponding to each training variant point based on the pathological state label; The predicted scores, confidence levels, and state labels of all training variants are used as training data to train the preset initial pathological state analysis model, resulting in a trained pathological state analysis model.

[0012] As an optional implementation, in a second aspect of the invention, the voting prediction module performs the operations of the weighted voting prediction stage, specifically including: Based on the pathological state corresponding to the gene to be tested, the state group corresponding to the gene to be tested is determined, and then multiple trained first-level models corresponding to the state group are used as multiple target first-level models. The predicted scores, confidence levels, and nonlinear extension features of all the variants to be tested are input into multiple target primary models to obtain the output results of multiple target primary models; according to the preset weight parameters corresponding to each target primary model, the output results of all the target primary models are weighted and fused to obtain a fused output result; based on the fused output result, the functional prediction result corresponding to the gene data to be tested is determined.

[0013] As an optional implementation, in the second aspect of the invention, the predicted scores, confidence levels, and nonlinear extension features of all training variant points are used as a second training dataset to train multiple machine learning models, resulting in multiple trained first-level models, including: The predicted scores, confidence levels, and nonlinear extension features of all training variants are used as the second training dataset to train multiple machine learning models, resulting in multiple well-trained initial first-level models. Each initial level-one model is validated using a pre-defined validation dataset to obtain a validation result for each initial level-one model. An evaluation score for each initial level-one model is then determined based on the validation results. From the multiple initial first-level models, the initial first-level models whose evaluation scores exceed a preset score threshold are selected as first-level models.

[0014] As an optional implementation, in a second aspect of the invention, the step of analyzing the first training dataset using various computational biology tools to obtain the prediction scores and confidence levels of all training variants corresponding to the cancer-related gene data includes: For each of the aforementioned computational biology tools: The first training dataset is analyzed using the computational biology tool to obtain the initial training variants analyzed by the computational biology tool and the prediction score corresponding to each initial training variant. For the cancer-related gene data, a filter window of a preset size is used to move along the cancer-related gene data once every preset amount of data. During each move: initial training variants in the filter window whose distance is less than a preset distance threshold are merged, and then the initial training variant with the highest confidence in the filter window is extracted as the training variant selected by the filter window. Based on all training variants selected by the filtering window during all movements, the prediction scores and confidence levels of all training variants corresponding to the cancer-related gene data are obtained.

[0015] A third aspect of this invention discloses a weighted voting-based intelligent analysis system for cancer-related genes, the system comprising: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute some or all of the steps in the weighted voting-based intelligent analysis method for cancer-related genes according to any of the first aspects of the present invention.

[0016] The fourth aspect of the present invention discloses a computer storage medium storing computer instructions, which, when invoked, are used to execute some or all of the steps in the weighted voting-based intelligent analysis method for cancer-related genes described in any of the first aspects of the present invention.

[0017] Compared with the prior art, the present invention has the following beneficial effects: In this invention, existing cancer gene prediction methods suffer from the following problems: due to the different fundamental principles and applicable scenarios of each computational biology tool, and the fact that the performance of machine learning models depends on the quality of training data, the reliability of intelligent analysis of cancer-related genes relying on a single tool is insufficient. Therefore, this invention first uses different computational biology tools for preliminary prediction. Since different computational biology tools have different prediction principles and applicable scenarios, combining multiple computational biology tools can provide a more comprehensive and reliable analysis of the target gene. Then, the prediction scores of multiple computational biology tools are fused by constructing nonlinear features. Finally, the prediction scores, confidence levels, and nonlinear extended features of the target variant are input into multiple pre-trained first-level models to obtain the output results of multiple first-level models. At this point, the output results of multiple models will differ due to the different underlying principles of different first-level models; for example, some models may fail to predict the analyzed variant results, while others may be able to. Therefore, this invention employs a weighted fusion method to fuse the output results of different first-level models. Based on the final fusion result, the functional prediction results corresponding to the target gene data are analyzed. Therefore, this invention enables a more comprehensive analysis of the gene data to be tested, integrating the analysis results of different analytical tools and models, and ultimately improving the overall reliability of the analysis results. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating a weighted voting-based intelligent analysis method for cancer-related genes disclosed in an embodiment of the present invention. Figure 2 This is a flowchart illustrating a model training method disclosed in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a cancer-related gene intelligent analysis device based on weighted voting, as disclosed in an embodiment of the present invention. Figure 4 This is a schematic diagram of the structure of a cancer-related gene intelligent analysis system based on weighted voting, as disclosed in an embodiment of the present invention. Detailed Implementation

[0020] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or end that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or ends.

[0022] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0023] This invention discloses a method, device, and medium for intelligent analysis of cancer-related genes based on weighted voting, which improves the reliability of intelligent analysis of cancer-related genes. These are described in detail below.

[0024] Example 1 like Figure 1 As shown, Embodiment 1 of the present invention discloses a weighted voting-based intelligent analysis method for cancer-related genes. This method can be integrated into a weighted voting-based intelligent analysis device for cancer-related genes, which can in turn be integrated into a local server or a cloud server. Specifically, the weighted voting-based intelligent analysis method for cancer-related genes disclosed in Embodiment 1 of the present invention may include: Step 101: Obtain the gene data to be tested, and use a variety of computational biology tools to analyze the gene data to obtain the prediction scores and confidence levels of all the variants to be tested corresponding to the gene data.

[0025] In this embodiment of the invention, for the gene data to be tested, all single nucleotide variant (SNP / SNV) records can be extracted from clinical data, retaining key fields such as chromosome, locus, reference / substitute nucleotide, and clinical variant classification. For example, a data segment from the gene data to be tested is shown below:

[0026] In this embodiment of the invention, the predicted score is used to indicate the classification result of the corresponding variant to be tested; the classification result can be classified according to clinical significance (based on the ACMG / AMP guidelines) into the following five categories: pathogenic, possibly pathogenic, unclear significance, possibly benign, and benign. Alternatively, it can be classified according to the intensity of functional impact into high impact, moderate impact, low impact, and modifying factors, etc.

[0027] In this embodiment of the invention, optional computational biology tools may include multiple tools selected from SIFT, CADD, PredictSNP, REVEL, and MCAP. These tools are authoritative or commonly used in bioinformatics for predicting the functional impact of gene variations; they can score and classify variations based on different principles.

[0028] The principles and output results of the above computational biology tools are shown in the table below:

[0029] Step 102: Based on the predicted scores analyzed by various computational biology tools, construct nonlinear extended features for all test variants corresponding to cancer-related gene data.

[0030] Because different computational biology tools have different analytical principles and output results, the output results of different tools cannot be directly fused and used as input data for the next stage. In this embodiment of the invention, the outputs of different tools are used to construct nonlinear extended features. For example, nonlinear extended features such as square and cube are constructed on the predicted scores of five tools to enhance the fitting of complex decision boundaries and to better fuse the outputs of different tools.

[0031] For example, in an alternative embodiment, the nonlinear extension feature can be constructed as follows: Let S be the score predicted by the i-th computational biology tool. i The score predicted by the j-th computational biology tool is S. j Then, exemplary nonlinear extension features include: Single-tool nonlinear term: S i 2 S i 3 ; Inter-tool interaction project: Si ·S j ; Global statistics: max(S1...S5), std(S1...S5).

[0032] Step 103, the weighted voting prediction stage, may include: inputting the predicted scores, confidence levels, and nonlinear extension features of all the variants to be tested into multiple pre-trained first-level models to obtain the output results of multiple first-level models; weighting and fusing the output results of all first-level models according to the preset weight parameters corresponding to each first-level model to obtain the fused output result; and determining the functional prediction result corresponding to the gene data to be tested based on the fused output result.

[0033] In this embodiment of the invention, the types of multiple first-level models can be various of the following: random forest model, decision tree model, logistic regression model, SVM model, and feedforward neural network model.

[0034] In this embodiment of the invention, existing cancer gene prediction methods suffer from the following problems: due to the different fundamental principles and applicable scenarios of each computational biology tool, and the fact that the performance of machine learning models depends on the quality of training data, the reliability of intelligent analysis of cancer-related genes relying on a single tool is insufficient. Therefore, this invention first uses different computational biology tools for preliminary prediction. Since different computational biology tools have different prediction principles and applicable scenarios, combining multiple computational biology tools can provide a more comprehensive and reliable analysis of the target gene. Then, the prediction scores of multiple computational biology tools are fused by constructing nonlinear features. Finally, the prediction scores, confidence levels, and nonlinear extended features of the target variant are input into multiple pre-trained first-level models to obtain the output results of multiple first-level models. At this point, the output results of multiple models will differ due to the different underlying principles of different first-level models; for example, some models may be unable to predict variant results while others can. Therefore, this embodiment of the invention employs a weighted fusion method to fuse the output results of different first-level models. Based on the final fusion result, the functional prediction results corresponding to the target gene data are analyzed. Therefore, this invention enables a more comprehensive analysis of the gene data to be tested, integrating the analysis results of different analytical tools and models, and ultimately improving the overall reliability of the analysis results.

[0035] In an optional embodiment, the first-level model is obtained through the following steps: Obtain a first training dataset and a validation dataset. The first training dataset and the validation dataset may include cancer-related gene data and the variant classification labels corresponding to the cancer-related gene data. The first training dataset was analyzed using a variety of computational biology tools to obtain the prediction scores and confidence levels of all training variants corresponding to cancer-related gene data. Based on the predicted scores analyzed by various computational biology tools, a nonlinear extended feature of all training variants corresponding to cancer-related gene data is constructed. Model training steps may include: The predicted scores, confidence levels, and nonlinear extension features of all training variants are used as the second training dataset to train multiple machine learning models, resulting in multiple well-trained first-level models.

[0036] like Figure 2 As shown, the training process of the above model is illustrated with an example as follows: S1. Data Sources and Construction (Limited to Cancer-Related Genes) All SNP / SNV records were extracted from NCBI clinical data of 1196 cancer-related genes, retaining key fields such as chromosome, locus, reference / alternative nucleotides, and clinical variant classification.

[0037] The prediction scores and confidence levels of the corresponding variants were obtained from five tools: SIFT, CADD, PredictSNP, REVEL, and MCAP, and then the IDs and coordinates were aligned.

[0038] After cleaning, standardization, and consistency verification, a total of 300,228 valid data points were obtained that can be used for training and evaluation.

[0039] S2, Feature Engineering (Optimization for Tumor Scenarios) Nonlinear extended features such as square and cube are constructed for the predicted scores of the five tools to enhance the fitting of complex decision boundaries.

[0040] Perform scaling, score orientation consistency (reverse normalization if necessary), missing value imputation and robustness processing; optional tumor-related annotations (such as whether it is located in a known driver domain, protein conservation indicators, etc.) can be added.

[0041] The base score, confidence level, and nonlinear extension are retained as inputs to the model.

[0042] S3, Single Model Training and Parameter Optimization Random forest, decision tree, logistic regression, support vector machine, and feedforward neural network algorithms were used for modeling.

[0043] Key hyperparameters were optimized through cross-validation and grid / Bayesian search, and the optimal parameters and performance of each model were recorded.

[0044] The three models that perform best on the validation set are selected as candidate base models.

[0045] S4, Weighted Voting Integration Strategy Weighted voting: Weights can be determined based on comprehensive indicators such as accuracy, recall and AUC on the validation set, and the probability outputs of the three models can be weighted and fused.

[0046] In another alternative embodiment, the above training process assigns only a single weight to each first-level model. However, in reality, each model has different predictive performance for different functional categories of mutation points due to its different underlying principles. Therefore, a more reliable approach would be to assign a weight to each first-level model for each mutation category.

[0047] Specifically, in this optional embodiment, the model training step may further include: For each mutation classification in the validation dataset: validate multiple trained first-level models based on the validation data corresponding to that mutation classification, obtain the validation result for each first-level model, and determine the class weight Q for each first-level model based on the validation result; Through the above operations, the class weights of all first-level models for each mutation category are calculated, thus obtaining the class weight Q of each first-level model for each mutation category. ji ; Where j indicates the j-th first-level model, and i indicates the i-th mutation classification, for example, Q. 34 This represents the class weight of the third primary model in the fourth variant classification.

[0048] In another optional embodiment, the output of the first-level model may include the predicted classification result for each variant point to be tested; and preset weight parameters, obtained through the following steps: For each first-level model, based on the output of that first-level model, determine the proportion P of each predicted classification result in the output of that first-level model. i Based on the class weights Q corresponding to all variant categories in this first-level model. ji The weight parameter K corresponding to this first-level model is calculated using the following formula. j :

[0049] Where I refers to the number of categories predicted in the output of the first-level model.

[0050] Among them, P i = (Number of variants predicted to be of class i) / (Total number of variants), and, for the weight parameter K corresponding to the first-level model j; Before use, a normalization operation can also be included:

[0051] In this optional embodiment, the weight parameters are no longer directly adopted from pre-determined parameters. Instead, more reasonable weight parameters are recalculated based on the output results of the first-level models. During the calculation of the weight parameters, the focus is on the proportion of each predicted classification result in the output results of each first-level model. For example, for a support vector machine first-level model, which is more suitable for predicting benign variants, the weight parameters are obtained by accumulating the proportion of benign variants in the prediction results and their pre-defined weights in benign classification. Optionally, normalization can be performed after obtaining all weight parameters, or it can be performed during the weighted fusion process.

[0052] In this optional embodiment, the proportion Pi of each predicted classification result in the output of the first-level model can be determined by the following method: The number of each predicted classification result in the output of the first-level model is determined. When the proportion of a certain classification result exceeds a certain predefined threshold, the proportion of that classification result is considered to be 100%, and the proportion of other classification results is considered to be 0. In this way, once a certain output result in the output of the first-level model reaches a certain proportion, the first-level model is given a larger proportion of that result, so that the model has a greater influence on the prediction results of that category in the final prediction results, and further improves the reliability of the prediction process.

[0053] In another optional embodiment, the functional prediction of variants in cancer-related genes should not be limited to the gene itself, but should also consider the pathological state corresponding to the cancer gene. Unlike the classification of variants, the classification of variants in this embodiment refers to single nucleotide variants of cancer-related genes, while the pathological state refers to the entire cancer-related gene. All single nucleotide variants on the cancer gene collectively determine the pathological state of the entire cancer-related gene. For example, the pathological state can be divided into: GX: grade cannot be assessed; G1: highly differentiated - cancer cells are close to normal cells, grow slowly, and have low invasiveness; G2: moderately differentiated - between G1 and G3; G3: poorly differentiated - cancer cells are very different from normal cells, grow rapidly, and have high invasiveness; G4: undifferentiated - cancer cells are extremely primitive, their origin cannot be identified, and the malignancy is the highest. Therefore, different state groups can be set for different pathological states, and then a corresponding model can be trained for each state group to make the trained model more targeted, thereby improving the reliability of the prediction.

[0054] Therefore, in yet another optional embodiment, the first training dataset also includes pathological state labels corresponding to cancer-related gene data; The method may also include: Based on the pathological state labels corresponding to cancer-related genes, all training variants are divided into multiple state groups; the prediction scores, confidence scores and nonlinear expansion features of all training variants in each state group are obtained. For each state group: re-trigger the model training step to obtain multiple trained first-level models corresponding to that state group, and the class weight Q of each first-level model for each mutation classification. ji .

[0055] In this optional embodiment, the model training steps are divided according to different state groups. First, the predicted scores, confidence levels, and nonlinear expansion features of all training variants within each state group are obtained. Then, the corresponding model training steps are performed for each state group, resulting in multiple trained first-level models for each state group, and the class weight Q of each first-level model for each variant classification in each state group. ji This allows for a more detailed analysis of the primary model and weight parameters.

[0056] In yet another optional embodiment, the method may further include: The predicted scores and confidence levels of all the variants to be tested, analyzed by various computational biology tools, are input into a pre-trained pathological state analysis model to obtain the output results of the pathological state analysis model. Based on the output results of the pathological state analysis model, the pathological state corresponding to the gene to be tested is determined. The prediction of pathological status is not limited to the functional impact of a single nucleotide variation, but rather takes the entire cancer-related gene as the object of analysis and prediction. This enables the weighted voting-based intelligent analysis method for cancer-related genes in this embodiment of the invention to also predict and analyze pathological status information.

[0057] The pathological state analysis model is trained through the following steps: Determine the state label corresponding to each training variant point based on the pathological state label. The predicted scores, confidence levels, and state labels of all training variants are used as training data to train the preset initial pathological state analysis model, resulting in a trained pathological state analysis model.

[0058] In this optional embodiment, the pathological state analysis model may optionally rely on an XGBoost or Transformer classifier; the input features of the pathological state analysis model may include variable point statistics (such as pathogenicity ratio).

[0059] In yet another optional embodiment, the weighted voting prediction stage may specifically include: Based on the pathological state of the gene to be tested, the state group corresponding to the gene to be tested is determined, and then multiple trained first-level models corresponding to the state group are used as multiple target first-level models. The predicted scores, confidence levels, and nonlinear extension features of all the variants to be tested are input into multiple target-level models to obtain the output results of multiple target-level models. According to the preset weight parameters corresponding to each target-level model, the output results of all target-level models are weighted and fused to obtain the fused output result. Based on the fused output result, the functional prediction result corresponding to the gene data to be tested is determined.

[0060] In this optional embodiment, since the trained primary model is trained in multiple groups based on different pathological states, and each group of models is trained specifically for that state, when performing single nucleotide variant analysis, the current state group corresponding to the gene to be tested is first determined based on the pathological analysis results. Then, the primary model corresponding to that state group is used as the model for this analysis and prediction. Through the above operations, this optional embodiment can use a more targeted primary model to perform subsequent analysis and prediction operations, thereby further improving the reliability of the prediction.

[0061] In another optional embodiment, the predicted scores, confidence levels, and nonlinear extension features of all training variants are used as a second training dataset to train multiple machine learning models, resulting in multiple trained first-level models, which may include: The predicted scores, confidence levels, and nonlinear extension features of all training variants are used as the second training dataset to train multiple machine learning models, resulting in multiple well-trained initial first-level models. Each initial level 1 model is validated using a pre-defined validation dataset to obtain the validation results for each initial level 1 model. The evaluation score for each initial level 1 model is then determined based on the validation results. From multiple initial primary models, those with evaluation scores exceeding a preset score threshold are selected as primary models.

[0062] In this optional embodiment, accuracy, recall, and AUC may be the primary metrics, supplemented by precision, F1 score, PR-AUC, and confusion matrix; with emphasis placed on minority class recall as the basis for obtaining the evaluation score.

[0063] In this optional embodiment, the number of initial first-level models can be larger. After one round of training, the performance of the initial first-level models is evaluated by a validation dataset. Then, several first-level models with better performance are selected as the final first-level models used, thereby obtaining first-level models with better prediction performance.

[0064] In another alternative embodiment, the overall data quality of the first training dataset may be poor, and there may be a lot of invalid or inefficient data. When the first training dataset has a large amount of data, the low-quality data may lead to a decrease in training results, or even cause the trained model to give incorrect analysis results when there is a lot of inefficient data.

[0065] To address the aforementioned issues, in this optional embodiment, various computational biology tools are used to analyze the first training dataset to obtain prediction scores and confidence levels for all training variants corresponding to cancer-related gene data. This may include: For each computational biology tool: The computational biology tool was used to analyze the first training dataset to obtain the initial training variants and the prediction scores corresponding to each initial training variant. For cancer-related gene data, a filter window of a preset size is used to move along the cancer-related gene data once every preset amount of data. During each move: initial training variants in the filter window whose distance is less than a preset distance threshold are merged, and then the initial training variant with the highest confidence in the filter window is extracted as the training variant selected by the filter window; in this optional embodiment, the distance threshold can be set to 10bp.

[0066] Based on all training variants selected by the filtering window during all moves, the prediction scores and confidence levels of all training variants corresponding to cancer-related gene data are obtained.

[0067] In this optional embodiment, it was found that the occurrence of variant points tends to be concentrated in segments of gene data that are close to each other. That is, multiple locations identified as variant points may be clustered together in a segment of data. However, these variant points may be caused by a single nucleotide variation. Therefore, within a continuous segment of gene data, selecting the data with the highest training quality can improve the overall data quality. To address this, this optional embodiment employs a filtering window approach, moving along the data and performing a filtering operation each time it moves. This leverages the principle of concentrated variant point occurrences to filter out low-quality data.

[0068] In this optional embodiment, the parameters corresponding to the filtering window can be set as follows: filtering window size = 20bp, moving step size = 10bp; the filtering rule is the local maximum CADD score.

[0069] Optionally, the data filtering operations described above can be applied both during the model training phase and during the analysis of the gene to be tested. In this case, various computational biology tools can be used to analyze the gene data to obtain the prediction scores and confidence levels of all the variants corresponding to the gene data. This can include: The computational biology tool was used to analyze the gene data to be tested, and the initial variants to be tested and the prediction scores corresponding to each initial variant were obtained. For the gene data to be tested, a filter window of a preset size is used to move along the gene data to be tested once every preset amount of data. During each move: the initial variants to be tested that are less than a preset distance threshold in the filter window are merged, and then the initial variants to be tested with the highest confidence in the filter window are extracted as the variants to be tested selected by the filter window. Based on all the test variants selected by the filtering window during all the moves, the prediction scores and confidence levels of all test variants corresponding to the test gene data are obtained.

[0070] Example 2 Based on the same inventive concept, Embodiment 2 of this invention discloses a cancer-related gene intelligent analysis device based on weighted voting, such as... Figure 3 As shown, the device may include: The data acquisition module 201 is used to acquire the gene data to be tested, and to analyze the gene data to be tested using a variety of computational biology tools to obtain the prediction scores and confidence levels of all the variant points to be tested corresponding to the gene data to be tested. The prediction scores are used to indicate the classification results of the corresponding variant points to be tested. The tool calculation module 202 is used to construct nonlinear extended features of all test variants corresponding to cancer-related gene data based on the predicted scores analyzed by various computational biology tools. The voting prediction module 203 is used to perform the operations of the weighted voting prediction stage, which may include: inputting the prediction scores, confidence levels and nonlinear extension features of all the variants to be tested into multiple pre-trained first-level models to obtain the output results of multiple first-level models; weighting and fusing the output results of all first-level models according to the preset weight parameters corresponding to each first-level model to obtain the fused output result; and determining the functional prediction result corresponding to the gene data to be tested based on the fused output result.

[0071] In an optional embodiment, the first-level model is obtained through the following steps: Obtain a first training dataset and a validation dataset. The first training dataset and the validation dataset may include cancer-related gene data and the variant classification labels corresponding to the cancer-related gene data. The first training dataset was analyzed using a variety of computational biology tools to obtain the prediction scores and confidence levels of all training variants corresponding to cancer-related gene data. Based on the predicted scores analyzed by various computational biology tools, a nonlinear extended feature of all training variants corresponding to cancer-related gene data is constructed. Model training steps may include: The predicted scores, confidence levels, and nonlinear extension features of all training variants are used as the second training dataset to train multiple machine learning models, resulting in multiple well-trained first-level models. For each mutation classification in the validation dataset: validate multiple trained first-level models based on the validation data corresponding to that mutation classification, obtain the validation result for each first-level model, and determine the class weight Q for each first-level model based on the validation result; Obtain the class weight Q for each first-level model in each mutation classification. ji ; Where j indicates the j-th first-level model and i indicates the i-th mutation classification.

[0072] In another optional embodiment, the output of the first-level model may include the predicted classification result for each variant point to be tested; and preset weight parameters, obtained through the following steps: For each first-level model, based on the output of that first-level model, determine the proportion P of each predicted classification result in the output of that first-level model. i Based on the class weights Q corresponding to all variant categories in this first-level model. ji The weight parameter K corresponding to this first-level model is calculated using the following formula. j :

[0073] Where I refers to the number of categories predicted in the output of the first-level model.

[0074] In yet another alternative embodiment, the various computational biology tools may include multiple tools such as SIFT, CADD, PredictSNP, REVEL, and MCAP. The various types of first-level models include random forest models, decision tree models, logistic regression models, SVM models, and feedforward neural network models.

[0075] In yet another optional embodiment, the first training dataset may further include pathological state labels corresponding to cancer-related gene data; The device may also include: The weight pre-setting module is used to divide all training variants into multiple state groups based on the pathological state labels corresponding to cancer-related genes; it obtains the prediction scores, confidence levels, and nonlinear expansion features of all training variants within each state group; for each state group: it re-triggers the model training step to obtain multiple trained first-level models corresponding to that state group, and the class weight Q of each first-level model for each variant classification. ji .

[0076] In yet another alternative embodiment, the device may further include: The pathology analysis module is used to input the predicted scores and confidence levels of all the variants to be tested analyzed by various computational biology tools into a pre-trained pathology state analysis model, obtain the output results of the pathology state analysis model, and determine the pathology state corresponding to the gene to be tested based on the output results of the pathology state analysis model. The pathological state analysis model is trained through the following steps: Determine the state label corresponding to each training variant point based on the pathological state label. The predicted scores, confidence levels, and state labels of all training variants are used as training data to train the preset initial pathological state analysis model, resulting in a trained pathological state analysis model.

[0077] In yet another optional embodiment, the voting prediction module 203 performs the operations of the weighted voting prediction phase, which may specifically include: Based on the pathological state of the gene to be tested, the state group corresponding to the gene to be tested is determined, and then multiple trained first-level models corresponding to the state group are used as multiple target first-level models. The predicted scores, confidence levels, and nonlinear extension features of all the variants to be tested are input into multiple target-level models to obtain the output results of multiple target-level models. According to the preset weight parameters corresponding to each target-level model, the output results of all target-level models are weighted and fused to obtain the fused output result. Based on the fused output result, the functional prediction result corresponding to the gene data to be tested is determined.

[0078] In another optional embodiment, the predicted scores, confidence levels, and nonlinear extension features of all training variants are used as a second training dataset to train multiple machine learning models, resulting in multiple trained first-level models, which may include: The predicted scores, confidence levels, and nonlinear extension features of all training variants are used as the second training dataset to train multiple machine learning models, resulting in multiple well-trained initial first-level models. Each initial level 1 model is validated using a pre-defined validation dataset to obtain the validation results for each initial level 1 model. The evaluation score for each initial level 1 model is then determined based on the validation results. From multiple initial primary models, those with evaluation scores exceeding a preset score threshold are selected as primary models.

[0079] In another optional embodiment, analyzing the first training dataset using various computational biology tools to obtain prediction scores and confidence levels for all training variants corresponding to cancer-related gene data may include: For each computational biology tool: The computational biology tool was used to analyze the first training dataset to obtain the initial training variants and the prediction scores corresponding to each initial training variant. For cancer-related gene data, a filter window of a preset size is used to move along the cancer-related gene data once every preset amount of data. During each move: initial training variants in the filter window whose distance is less than a preset distance threshold are merged, and then the initial training variant with the highest confidence in the filter window is extracted as the training variant selected by the filter window. Based on all training variants selected by the filtering window during all moves, the prediction scores and confidence levels of all training variants corresponding to cancer-related gene data are obtained.

[0080] Example 3 Please see Figure 4 , Figure 4 This is a schematic diagram of a weighted voting-based intelligent analysis system for cancer-related genes, as disclosed in an embodiment of the present invention. The weighted voting-based intelligent analysis system for cancer-related genes may include: Memory 301 storing executable program code; Processor 302 coupled to memory 301; The processor 302 calls the executable program code stored in the memory 301 to execute some or all of the steps in any of the weighted voting-based intelligent analysis methods for cancer-related genes in Embodiment 1 of the present invention.

[0081] Example 4 This invention discloses a computer storage medium storing computer instructions. When executed by a processor, these computer instructions implement some or all of the steps in any of the weighted voting-based intelligent analysis methods for cancer-related genes in Embodiment 1 of this invention.

[0082] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each implementation method can be implemented using software plus necessary general-purpose hardware platforms, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or in terms of their contribution to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium.

Claims

1. A method for intelligent analysis of cancer-related genes based on weighted voting, characterized in that, The method includes: Acquire the gene data to be tested, and use a variety of computational biology tools to analyze the gene data to be tested, and obtain the prediction scores and confidence scores of all the variant points to be tested corresponding to the gene data to be tested, wherein the prediction scores are used to indicate the classification results of the corresponding variant points to be tested; Based on the predicted scores analyzed by various computational biology tools, a nonlinear extended feature of all test variants corresponding to the cancer-related gene data is constructed. The weighted voting prediction stage includes: inputting the prediction scores, confidence levels, and nonlinear extension features of all the variants to be tested into multiple pre-trained first-level models to obtain the output results of multiple first-level models; weighting and fusing the output results of all the first-level models according to the preset weight parameters corresponding to each first-level model to obtain a fused output result; and determining the functional prediction result corresponding to the gene data to be tested based on the fused output result.

2. The intelligent analysis method for cancer-related genes based on weighted voting according to claim 1, characterized in that, The first-level model is obtained through the following steps: Obtain a first training dataset and a validation dataset, wherein the first training dataset and the validation dataset include cancer-related gene data and variant classification labels corresponding to the cancer-related gene data; The first training dataset was analyzed using a variety of computational biology tools to obtain the prediction scores and confidence levels of all training variants corresponding to the cancer-related gene data. Based on the predicted scores analyzed by various computational biology tools, a nonlinear extended feature of all training variants corresponding to the cancer-related gene data is constructed. The model training steps include: The predicted scores, confidence levels, and nonlinear extension features of all training variants are used as the second training dataset to train multiple machine learning models, resulting in multiple well-trained first-level models. For each mutation classification in the validation dataset: validate multiple trained first-level models based on the validation data corresponding to the mutation classification, obtain the validation result for each first-level model, and determine the class weight Q for each first-level model based on the validation result; Obtain the class weight Q for each of the first-level models in each variant classification. ji ; Where j indicates the j-th first-level model and i indicates the i-th mutation classification.

3. The intelligent analysis method for cancer-related genes based on weighted voting according to claim 2, characterized in that, The output of the first-level model includes the predicted classification result for each variant point to be tested; and the preset weight parameters are obtained through the following steps: For each of the first-level models, based on the output of that first-level model, determine the proportion P of each predicted classification result in the output of that first-level model. i Based on the class weights Q corresponding to all variant categories in this first-level model. ji The weight parameter K corresponding to this first-level model is calculated using the following formula. j : Where I refers to the number of predicted classification categories in the output of the first-level model, and P... i This represents the proportion of category i in the output of the j-th first-level model.

4. The intelligent analysis method for cancer-related genes based on weighted voting according to any one of claims 1-3, characterized in that, The various computational biology tools include multiple tools such as SIFT, CADD, PredictSNP, REVEL, and MCAP. The various first-level models are of multiple types, including random forest model, decision tree model, logistic regression model, SVM model, and feedforward neural network model.

5. The intelligent analysis method for cancer-related genes based on weighted voting according to claim 2 or 3, characterized in that, The first training dataset also includes pathological state labels corresponding to the cancer-related gene data; Furthermore, the method further includes: Based on the pathological state labels corresponding to the cancer-related genes, all training variants are divided into multiple state groups; the prediction scores, confidence levels, and nonlinear expansion features of all training variants in each state group are obtained. For each state group: the model training step is retried to obtain multiple trained first-level models corresponding to that state group, and the class weight Q of each first-level model for each mutation classification. ji .

6. The intelligent analysis method for cancer-related genes based on weighted voting according to claim 5, characterized in that, The method further includes: The predicted scores and confidence levels of all the variants to be tested, analyzed by various computational biology tools, are input into a pre-trained pathological state analysis model to obtain the output results of the pathological state analysis model. Based on the output results of the pathological state analysis model, the pathological state corresponding to the gene to be tested is determined. The pathological state analysis model is obtained through the following steps: Determine the state label corresponding to each training variant point based on the pathological state label; The predicted scores, confidence levels, and state labels of all training variants are used as training data to train the preset initial pathological state analysis model, resulting in a trained pathological state analysis model.

7. The intelligent analysis method for cancer-related genes based on weighted voting according to claim 6, characterized in that, The weighted voting prediction stage specifically includes: Based on the pathological state corresponding to the gene to be tested, the state group corresponding to the gene to be tested is determined, and then multiple trained first-level models corresponding to the state group are used as multiple target first-level models. The predicted scores, confidence levels, and nonlinear extension features of all the variants to be tested are input into multiple target primary models to obtain the output results of multiple target primary models; according to the preset weight parameters corresponding to each target primary model, the output results of all the target primary models are weighted and fused to obtain a fused output result; based on the fused output result, the functional prediction result corresponding to the gene data to be tested is determined.

8. The intelligent analysis method for cancer-related genes based on weighted voting according to claim 2 or 3, characterized in that, The predicted scores, confidence levels, and nonlinear extension features of all training variants are used as the second training dataset to train multiple machine learning models, resulting in multiple well-trained first-level models, including: The predicted scores, confidence levels, and nonlinear extension features of all training variants are used as the second training dataset to train multiple machine learning models, resulting in multiple well-trained initial first-level models. Each initial level-one model is validated using a pre-defined validation dataset to obtain a validation result for each initial level-one model. An evaluation score for each initial level-one model is then determined based on the validation results. From the multiple initial first-level models, the initial first-level models whose evaluation scores exceed a preset score threshold are selected as the first-level models.

9. The intelligent analysis method for cancer-related genes based on weighted voting according to claim 2 or 3, characterized in that, The analysis of the first training dataset using multiple computational biology tools yields prediction scores and confidence levels for all training variants corresponding to the cancer-related gene data, including: For each of the aforementioned computational biology tools: The first training dataset is analyzed using the computational biology tool to obtain the initial training variants analyzed by the computational biology tool and the prediction score corresponding to each initial training variant. For the cancer-related gene data, a filter window of a preset size is used to move along the cancer-related gene data once every preset amount of data. During each move: initial training variants in the filter window whose distance is less than a preset distance threshold are merged, and then the initial training variant with the highest confidence in the filter window is extracted as the training variant selected by the filter window. Based on all training variants selected by the filtering window during all movements, the prediction scores and confidence levels of all training variants corresponding to the cancer-related gene data are obtained.

10. A smart analysis device for cancer-related genes based on weighted voting, characterized in that, The device includes: The data acquisition module is used to acquire the gene data to be tested, analyze the gene data to be tested using a variety of computational biology tools, and obtain the prediction scores and confidence levels of all the variant points to be tested corresponding to the gene data to be tested. The prediction scores are used to indicate the classification results of the corresponding variant points to be tested. The tool calculation module is used to construct nonlinear extended features of all test variants corresponding to the cancer-related gene data based on the predicted scores analyzed by various computational biology tools. The voting prediction module is used to perform the operations of the weighted voting prediction stage, including: inputting the predicted scores, confidence levels and nonlinear extension features of all the variants to be tested into multiple pre-trained first-level models to obtain the output results of multiple first-level models; weighting and fusing the output results of all the first-level models according to the preset weight parameters corresponding to each first-level model to obtain a fused output result; and determining the functional prediction result corresponding to the gene data to be tested based on the fused output result.

11. A cancer-related gene intelligent analysis system based on weighted voting, characterized in that, The system includes: a memory storing executable program code; a processor coupled to the memory; the processor calling the executable program code stored in the memory to execute the intelligent analysis method for cancer-related genes based on weighted voting as described in any one of claims 1-9.

12. A computer storage medium, characterized in that, The computer storage medium stores computer instructions, which, when executed by a processor, implement the intelligent analysis method for cancer-related genes based on weighted voting as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Cancer gene classification method and device based on two-stage depth feature selection and storage medium

    CN112926640A

  • Gene sequence prediction method based on deep learning model and related equipment

    CN118173166A