Auxiliary method for clinical early diagnosis of digestive tract tumors
By standardizing and encoding tumor-related data, combined with nonlinear relationships and interaction effect analysis, the problem of inability to effectively capture complex interaction effects and nonlinear relationships between genes in the prior art is solved, and more efficient and accurate tumor risk assessment and early diagnosis are achieved.
Patent Information
- Application Number
- CN202510300414.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-14
AI Technical Summary
Existing early tumor diagnosis methods cannot fully capture the complex interaction effects between genes, ignore the nonlinear relationship between gene mutations and expression, and are inefficient in processing multi-dimensional data, lacking effective data standardization, coding and missing value processing mechanisms.
Gene mutations, gene expression and clinical data from different sources are transformed into a unified format by standardizing, encoding and missing values of the input data. Combining mutation frequency, expression volume and nonlinear relationships reflect the contribution of gene mutations to tumor occurrence, the interaction effects between gene mutations are captured for component analysis and dimensionality reduction, the similarity between genes is measured and the interaction effect calculation is optimized, the patients' tumor risk is evaluated, and the risks are classified and adaptively optimized.
It significantly improves the accuracy, personalization and computing efficiency of tumor risk prediction, providing more reliable and efficient support for early clinical diagnosis.
Smart Images

Figure CN120164534A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tumor risk assessment and genomic data processing, and specifically to an auxiliary method for early clinical diagnosis of digestive tract tumors. Background Art
[0002] In recent years, with the rapid development of biomedicine and genomics, early diagnosis techniques for digestive tract tumors have received extensive attention. Traditional tumor screening methods, such as imaging examinations and tumor marker detections, although can help diagnose tumors to a certain extent, their sensitivity and specificity are relatively low, and it is often difficult to detect tumors at an early stage. With the progress of genomics and molecular biology, gene mutations, gene expressions, and their related biomarkers have become new directions for early tumor diagnosis. Using genomic data to evaluate tumor risk has become a research hotspot. In recent years, multi-level analysis methods based on gene mutation data and clinical data, especially multi-gene combination analysis, have achieved preliminary application effects in tumor screening. These methods provide potential for early tumor diagnosis through in-depth mining of genomic data, but still face challenges in data processing and analysis capabilities. Existing early tumor diagnosis methods, especially those based on gene mutation data analysis, mainly rely on a single-effect model of gene mutations and ignore the complex interaction effects between different genes. Although gene mutation frequencies and expression levels are widely used in tumor risk assessment, existing technologies often cannot handle multi-dimensional and multi-level data, especially the interaction effects of multiple genes are often not fully considered. Traditional tumor risk assessment methods often use simple linear models to describe the role of gene mutations, ignoring the non-linear relationship between gene mutations and gene expressions, which limits the accuracy of risk assessment. In addition, in terms of modeling the interaction effects of multiple genes, most methods only consider low-order interactions or simple binary interactions, without further exploring higher-order complex interaction effects between genes. Although principal component analysis (PCA) is applied to dimensionality reduction, in existing technologies, the use of PCA is often limited to data dimensionality reduction and fails to fully optimize the calculation of interaction effects, resulting in low processing efficiency and inaccurate results for high-dimensional gene data. Existing risk assessment methods also lack effective data standardization, coding, and missing value processing mechanisms when dealing with different types, sources, and scales of gene data, so they face high data complexity and inconsistency problems in practical applications. Therefore, existing technologies cannot provide effective solutions in multiple dimensions such as multi-gene interaction effects, non-linear relationships, data standardization, and optimized calculations like the present invention, which limits the accuracy and practicality of early tumor diagnosis. Summary of the Invention
[0003] In view of the above existing problems, the present invention is proposed.
[0004] Therefore, the technical problem solved by the present invention is that existing early tumor diagnosis methods cannot fully capture complex gene-gene interaction effects, ignore the non-linear relationship between gene mutations and expression levels, are inefficient in processing multi-dimensional data, and how to comprehensively improve diagnostic accuracy through standardization, coding, and optimized algorithms.
[0005] To solve the above technical problems, the present invention provides the following technical solution: An auxiliary method for clinical early diagnosis of digestive tract tumors, including standardizing, coding, and handling missing values of input data, and converting gene mutations, gene expressions, and clinical data from different sources into a unified format; reflecting the contribution of gene mutations to tumorigenesis by combining mutation frequency, expression level, and non-linear relationship; capturing the interaction effects between gene mutations for component analysis and dimensionality reduction; measuring the similarity between genes, optimizing the calculation of gene-gene interaction effects, and evaluating the tumor risk of patients; classifying the tumor risk of patients and an adaptive optimization and dynamic adjustment mechanism for dividing the tumor risk.
[0006] As a preferred embodiment of the auxiliary method for clinical early diagnosis of digestive tract tumors according to the present invention, wherein: the standardizing, coding, and handling missing values of the input data include collecting the input data; The input data includes gene mutation data, gene expression data, and clinical data.
[0007] As a preferred embodiment of the auxiliary method for clinical early diagnosis of digestive tract tumors according to the present invention, wherein: the standardizing, coding, and handling missing values of the input data include standardizing the gene mutation frequency, using min-max standardization to map the mutation frequency to the range of 0 to 1; Encoding the gene mutation data, using one-hot encoding to convert the mutation type into a numerical form; Standardizing the gene expression data, using Z-score standardization to normalize the data through the mean and standard deviation, so that the mean of the data set is 0 and the standard deviation is 1; Performing numerical standardization processing on the clinical data.
[0008] As a preferred embodiment of the auxiliary method for clinical early diagnosis of digestive tract tumors according to the present invention, wherein: the combining of mutation frequency, expression level, and non-linear relationship to reflect the contribution of gene mutations to tumorigenesis includes adopting a high-order non-linear regression model, considering the non-linear relationship between the mutation frequency and expression level of genes, and describing the risk contribution of gene mutations to digestive tract tumors; Introducing a non-linear amplification effect term to reflect the risk trend of tumorigenesis; Introducing a gene expression level amplification term to describe the non-linear effect of gene expression level on risk; Expressing the coupling effect of mutation frequency and gene expression level; Construct a risk contribution model of genes to reflect the contribution of gene mutations to tumorigenesis.
[0009] As a preferred embodiment of the auxiliary method for early clinical diagnosis of digestive tract tumors according to the present invention, wherein: capturing the interaction effect between gene mutations for component analysis and dimensionality reduction includes capturing the interaction effect between gene mutations by constructing a high-order interaction model; Characterize the interaction intensity between genes by adjusting the interaction effect coefficient; Introduce principal component analysis (PCA) for dimensionality reduction of the high-dimensional data of gene mutation interactions; Output the interaction effect value through the high-order interaction model to represent the contribution of gene joint mutations to tumorigenesis.
[0010] As a preferred embodiment of the auxiliary method for early clinical diagnosis of digestive tract tumors according to the present invention, wherein: measuring the similarity between genes, optimizing the calculation of the interaction effect between genes, and evaluating the tumor risk of patients includes introducing a Laplacian matrix for optimization during the evaluation of gene mutations and tumor risk, representing the similarity between nodes in graph theory; Quantify the similarity between genes through the Laplacian matrix, use the mutation risk values between genes as nodes, and the similarity between nodes as edge weights to optimize the relationship between gene mutations; Balance the weights of gene similarity and PCA; Represent the impact of a single gene on tumor risk by outputting and weighting the risk contribution of each gene.
[0011] As a preferred embodiment of the auxiliary method for early clinical diagnosis of digestive tract tumors according to the present invention, wherein: classifying the tumor risk of patients and an adaptive optimization and dynamic adjustment mechanism for dividing the tumor risk includes classifying the tumor risk of patients according to the total tumor risk value; The division of risk levels is based on a risk threshold, and the threshold is used to determine patients with low, medium, and high risks; Introduce an adaptive optimization mechanism into the optimization model to adjust the risk assessment according to the data; The risk threshold is dynamically adjusted according to patient data; Provide a diagnosis and treatment recommendation for the division of tumor risk through risk grading and early warning.
[0012] Another object of the present invention is to provide an auxiliary system for early clinical diagnosis of digestive tract tumors, which can reflect the contribution of gene mutations to tumorigenesis by combining mutation frequency, expression level, and non-linear relationships, and solves the problem that the current early tumor diagnosis methods cannot fully capture the complex interaction effects between genes.
[0013] As a preferred solution of the auxiliary system for the clinical early diagnosis of digestive tract tumors according to the present invention, it includes: a data input and preprocessing module, a gene mutation risk assessment and interaction effect modeling module, and a risk grading and early warning module; the data input and preprocessing module is used to uniformly standardize gene mutation data, gene expression data, and clinical data from different sources and convert them into a format for model calculation; the gene mutation risk assessment and interaction effect modeling module is used to evaluate the risk contribution of a single gene through a high-order non-linear regression model and capture the interaction effect between multiple genes; the risk grading and early warning module is used to classify the tumor risk of patients according to the calculated total risk value and provide early warning.
[0014] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the auxiliary method for the clinical early diagnosis of digestive tract tumors are implemented.
[0015] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of the auxiliary method for the clinical early diagnosis of digestive tract tumors are implemented.
[0016] The beneficial effects of the present invention: The auxiliary method for the clinical early diagnosis of digestive tract tumors provided by the present invention standardizes, encodes, and processes missing values for the input data, converts gene mutation, gene expression, and clinical data from different sources into a unified format, ensuring data consistency and operability. This processing not only improves the data quality but also provides a reliable basis for subsequent analysis. Secondly, the contribution of gene mutations to tumorigenesis is reflected by combining mutation frequency, expression level, and non-linear relationships. The non-linear relationship between gene mutations and expression levels is captured through a high-order non-linear regression model, accurately describing the contribution of genes to tumor risk. This step effectively improves the processing ability of mutation frequency, gene expression, and their mutual relationships, enhancing the sensitivity and accuracy of the model. Further, the interaction effect between gene mutations is captured and component analysis is performed for dimensionality reduction. The interaction effect calculation is optimized through a high-order interaction model and principal component analysis (PCA), enabling the effective modeling of the influence of multi-gene interactions. At the same time, the calculation efficiency is optimized through dimensionality reduction. Next, the similarity between genes is measured and the calculation of the interaction effect is optimized. The Laplacian matrix is introduced to optimize the measurement of the similarity between genes, and the calculation method of the interaction effect is adjusted to improve the evaluation accuracy of the complex interaction between multiple genes. Finally, the tumor risk of patients is classified and an adaptive optimization and dynamic adjustment mechanism is introduced. The model parameters and risk thresholds are dynamically adjusted according to the gene mutation data of patients, realizing personalized risk assessment. Through this multi-dimensional and multi-step collaborative work, the present invention significantly improves the accuracy, personalization, and calculation efficiency of tumor risk prediction, providing more reliable and efficient support for clinical early diagnosis. Brief Description of the Drawings
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0018] Figure 1 It is the overall flowchart of an auxiliary method for the clinical early diagnosis of digestive tract tumors provided by the first embodiment of the present invention. Detailed Embodiments
[0019] To make the above objects, features, and advantages of the present invention more obvious and understandable, the detailed embodiments of the present invention will be described below in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0020] Embodiment 1, referring to Figure 1 , is an embodiment of the present invention, providing an auxiliary method for the clinical early diagnosis of digestive tract tumors, including: S1: Standardize, encode, and process missing values for the input data, and convert gene mutations, gene expressions, and clinical data from different sources into a unified format.
[0021] Furthermore, the input data includes gene mutation data, gene expression data, and clinical data.
[0022] The gene mutation data includes gene selection, mutation frequency, mutation type, and mutation site.
[0023] The gene mutation data is the basis for model calculation. Common genes related to digestive tract tumors include, but are not limited to, p53, K-ras, APC, C-myc, etc. Mutations in each gene will affect the probability of tumor occurrence, so the mutation information of each gene must be accurately obtained.
[0024] The mutation frequency is the proportion of mutations of each gene in the sample, usually obtained through gene sequencing or chip technology. The mutation frequency is an indicator to measure the prevalence of gene mutations in the patient population.
[0025] The mutation type is the specific type of mutation of each gene (such as point mutation, deletion, insertion, frameshift mutation, etc.). The mutation type affects the function of the gene and its role in tumorigenesis, and needs to be marked in detail.
[0026] Mutation sites indicate the locations where gene mutations occur. Mutations at certain sites may be closely related to the occurrence of tumors. Analysis of specific sites can help subdivide the risks of gene mutations.
[0027] Gene expression data includes gene expression levels. Gene expression level refers to the abundance of gene transcripts (mRNA), usually obtained through RNA-seq or real-time quantitative PCR (qPCR). Gene expression level is an indicator of the activity level of a gene in a cell. Combining mutation frequency and expression level can more comprehensively evaluate the contribution of a gene to tumor risk.
[0028] Clinical data includes basic patient information, diagnosis results, and imaging examinations.
[0029] Basic patient information includes age, gender, lifestyle habits (such as smoking, drinking, etc.), family history, medical history, etc. This information helps improve the comprehensive evaluation ability of the model.
[0030] Diagnosis results are the patient's historical diagnosis data, including whether they are diagnosed with digestive tract tumors, the stage of the tumor, the type of tumor, etc.
[0031] Imaging examination results such as CT, MRI, ultrasound, etc. These data play an important role in early tumor screening.
[0032] It should be noted that the input data comes from different sources and has inconsistent formats. The input data is transformed into a unified format suitable for model calculation through standardization processing.
[0033] If the numerical range of the mutation frequency varies greatly (for example, a frequency difference from 0 to 1), then standardization processing is required.
[0034] Use min-max standardization: Map the mutation frequency to the range from 0 to 1, expressed as:
[0035] Where, is the original mutation frequency, and are the minimum and maximum values in the dataset respectively, is the value after standardization.
[0036] The types of gene mutations are usually discrete variables (such as point mutations, insertions, deletions, etc.). The model calculation requires converting them into numerical forms. Use one-hot encoding to convert the mutation types into numerical forms. Assume that gene has 3 mutation types (A, B, C), then it is converted into the following encoding:
[0037] Each type is represented by an independent binary vector and is applicable to multiple possible mutation types.
[0038] Normalization of gene expression levels. Gene expression levels are usually continuous variables, and there are significant differences in the expression levels of different genes. Therefore, normalization processing is required.
[0039] Use Z-score normalization to normalize the data through the mean and standard deviation, making the mean of the data set 0 and the standard deviation 1, expressed as:
[0040] where, is the original gene expression level, and are the mean and standard deviation of the gene expression levels respectively, is the normalized gene expression level.
[0041] Clinical data usually includes categorical variables and numerical variables. To adapt to the model input, categorical variables need to be numericalized, and numerical variables may need to be normalized.
[0042] For categorical variables (such as gender, whether smoking, etc.), use one-hot encoding or label encoding.
[0043] For numerical variables (such as age, tumor size, etc.), use min-max normalization or Z-score normalization.
[0044] Missing value handling. Missing values are common problems in data processing, and missing values must be handled to ensure data integrity.
[0045] For numerical data, use mean filling or median filling, or it can also be filled through machine learning algorithms (such as KNN filling).
[0046] For categorical data, use mode filling or the most frequent category filling.
[0047] After normalization and encoding, all input data will be converted into a data matrix in a unified format. The rows of the data matrix represent the samples of each patient, and the columns represent features such as gene mutations, gene expressions, and clinical information. For example, assuming there are n patient samples and m features (including the mutation frequencies, expression levels, etc. of multiple genes), the final input data matrix X is an n×m matrix, where each row represents all the features of a patient.
[0048] The final data input matrix will be used as the input of the calculation model and passed to the subsequent mathematical model for risk assessment and prediction.
[0049] S2: Reflect the contribution of gene mutations to tumorigenesis by combining mutation frequency, expression level, and non-linear relationships.
[0050] Furthermore, considering the complexity of gene mutations, a higher-order non-linear regression model was adopted to describe the impact of gene mutations on the risk of digestive tract tumors. The higher-order non-linear regression model is as follows:
[0051] where, is the risk contribution of the -th gene; is the mutation frequency of the -th gene; is the expression level of the -th gene; is the coefficient of this gene, controlling the mutation frequency, gene expression level, and their non-linear relationship.
[0052] It should be noted that the mutation frequency is not directly proportional to the tumor risk. As the mutation frequency increases, the risk of tumorigenesis also shows a non-linear increasing trend. Especially in the case of high-frequency mutations, the risk may increase sharply. Therefore, in the higher-order non-linear regression model, the term is introduced to reflect this non-linear amplification effect, where and are parameters controlling the non-linear effect of the mutation frequency, which can flexibly adjust the impact of the mutation frequency on the tumor risk.
[0053] The expression level of the gene is also closely related to the tumor risk. Generally speaking, genes with high expression may increase the risk of tumorigenesis, while the impact of genes with low expression is relatively small. By introducing the term, the non-linear impact of the gene expression level on the risk can be more accurately described, is the parameter controlling the intensity of the impact of the expression level. The exponential term makes the change in the gene expression level have a more prominent impact on the tumor risk.
[0054] The mutation frequency and the gene expression level do not independently affect tumorigenesis, and there may be an interaction between them. For example, the high-frequency occurrence of certain gene mutations may enhance the tumor risk in the context of high gene expression. Therefore, the term in the higher-order non-linear regression model combines the interaction effect of the mutation frequency and the gene expression level, reflecting their joint action in affecting the tumor risk.
[0055] To better reflect the complex effects of gene mutations, a high-order non-linear regression model uses high-order non-linear terms. Specifically, the mutation frequency term represents the non-linear effect of mutation frequency on tumor risk, and the gene expression level term represents the non-linear effect of gene expression level.
[0056] Furthermore, there may be complex interaction effects between mutations of multiple genes, and the interaction effects have an important impact on tumor risk. To capture the interaction effects, by designing a high-order interaction model and combining principal component analysis (PCA) to optimize the interaction effects, it is expressed as:
[0057] where is the risk value of gene gene , gene , gene ; is the interaction effect coefficient, representing the interaction influence between genes; PCA is used for principal component analysis to optimize the calculation of the interaction effects between genes.
[0058] It should be noted that in the high-order interaction model, represents the interaction effect coefficient between gene , gene and gene . By adjusting the coefficient, the high-order interaction model can depict different interaction intensities between genes.
[0059] The interaction effects are not limited to the combination of two-gene mutations, and the interaction effects of three genes are also considered.
[0060] Principal component analysis (PCA) is introduced. PCA is used to reduce the dimension of the high-dimensional data of the interaction of multiple gene mutations, extract the main components, and simplify the calculation process. PCA extracts the main interaction information between gene mutations through dimensionality reduction, compresses the high-dimensional interaction effects into fewer dimensions, thereby effectively reducing the calculation time and memory consumption. The interaction effects after dimensionality reduction can be directly used to calculate the total tumor risk.
[0061] S3: Capture the interaction effects between gene mutations for component analysis and dimensionality reduction.
[0062] Furthermore, the Laplacian matrix is used to measure the similarity between genes to optimize the calculation of the interaction effects between genes. Through the optimization of the Laplacian matrix, the calculation method of the similarity of mutations between genes can be improved, making the contribution of the interaction effects between genes more accurate and further improving the accuracy of tumor risk assessment.
[0063] The mutation similarity between genes can be represented by a risk value. A high similarity between two genes means that their contributions to tumors may be relatively close under the same conditions, or the mutation of one gene may amplify the risk of another gene mutation. Through the Laplacian matrix, this similarity between genes can be quantified, thus effectively measuring the impact of gene interaction effects.
[0064] The Laplacian matrix optimizes the relationship between gene mutations by taking the mutation risk values between genes as nodes and the similarity between nodes as edge weights, reducing the redundant part in the calculation of interaction effects.
[0065] The weight of the Laplacian matrix controls the weight balance between the similarity between genes and the result of PCA dimensionality reduction. The Laplacian matrix is introduced to measure the similarity between genes, and the coefficient is adjusted by optimizing the value of the Laplacian matrix, expressed as:
[0066] where, is the Laplacian matrix value between genes and ; is the balancing term, controlling the ratio of the similarity between different genes and the result of PCA dimensionality reduction; represents the result of the interaction between genes after PCA dimensionality reduction; is the Laplacian matrix value between genes and ; is the balancing term, controlling the ratio of the similarity between different genes and the result of PCA dimensionality reduction; represents the result of the interaction between genes after PCA dimensionality reduction.
[0067] S4: Measure the similarity between genes, optimize the calculation of gene interaction effects, and evaluate the tumor risk of patients.
[0068] Furthermore, the final total tumor risk needs to comprehensively consider the independent risk of each gene, the interaction effect between genes, and the similarity after Laplacian matrix optimization, so as to comprehensively evaluate the tumor risk of patients. By summing up the risk contributions of each part, the final risk value is output as the diagnostic basis. The risk contribution of each gene is calculated and weighted to reflect the impact of a single gene on tumor risk. The interaction effect between genes further enhances the prediction ability of the model by considering the complex synergistic effects between genes. The interaction effect can reflect the additive effect on the risk of tumor occurrence when multiple genes mutate jointly. The Laplacian matrix The measurement of gene - gene similarity is optimized, further improving the calculation accuracy of interaction effects and making the final tumor risk assessment more precise.
[0069] It should be noted that the final tumor risk combines the single - gene risk, the interaction effects between genes, and the optimized similarity matrix, expressed as:
[0070] By integrating the single - gene risk, interaction effects, and gene - gene similarity, the final total risk value can comprehensively reflect the patient's tumor risk and improve the accuracy of early diagnosis.
[0071] S5: Classify the patient's tumor risk and establish an adaptive optimization and dynamic adjustment mechanism for tumor risk classification.
[0072] Furthermore, according to the total tumor risk value , classify the patient's tumor risk to provide precise diagnostic decision - making support. The risk grading aims to divide the patient's tumor risk into different levels, so as to guide doctors to take different intervention measures. The division of risk levels is based on the risk thresholds obtained through prior model optimization and , and the thresholds are used to determine low, medium, and high - risk patients.
[0073] When the total risk value is lower than the threshold , the patient's tumor risk is low. At this time, the patient may not need immediate further examination or treatment, but still needs regular monitoring.
[0074] When the total risk value is between and , the patient's tumor risk is at a medium level, and further examination or early intervention is recommended to detect potential tumor risks in a timely manner.
[0075] When the total risk value exceeds the threshold , the patient's tumor risk is high, and urgent examination and treatment measures are required. More in - depth imaging examinations or histological analyses may be needed.
[0076] Based on the calculated total risk expressed as:
[0077] Among them, and are risk thresholds based on model optimization, which divide low, medium, and high - risk levels. When the patient's total risk value exceeds the high - risk threshold When this occurs, the system will automatically issue an early warning to prompt the doctor for further examination.
[0078] It should be noted that in order to improve the accuracy of the model and enable it to adapt to different datasets, an adaptive optimization mechanism is introduced. The adaptive optimization mechanism enables the model to self-adjust as new data is input by dynamically adjusting the interaction effect coefficient and the risk threshold, continuously optimizing its prediction accuracy.
[0079] Interaction effect coefficient and control the degree of influence of gene-gene interactions. The value of the coefficient will be adjusted according to new data, reflecting the actual impact of different gene mutation combinations on tumor risk. The risk threshold and can also be dynamically adjusted according to new patient data. Through optimization algorithms such as gradient descent, the model automatically adjusts these thresholds, thereby improving the accuracy of risk grading. Through gradient descent or other optimization algorithms, the model continuously minimizes the loss function .
[0080] To enhance the adaptability of the model, adaptive optimization and gradient descent are used to dynamically adjust the interaction effect coefficient and the risk threshold. It is expressed as:
[0081] where is the learning rate, controlling the magnitude of each update; is the loss function, representing the prediction error of the model.
[0082] Adaptive optimization enables the model to dynamically adjust parameters according to new data, ensuring that over time, the model can continuously optimize and adapt to changing clinical data. Gradient descent can adjust the interaction effect coefficient and the risk threshold through the backpropagation mechanism, thereby improving the accuracy of tumor risk assessment.
[0083] As new data is continuously input, the model can self-learn and adjust, adapt to the characteristics of different patient groups, and enhance its generalization ability. Dynamically adjusting the interaction effect coefficient and the risk threshold ensures that the risk assessment for each patient can be personalized according to their specific gene background. Dynamic adjustment and optimization can provide more accurate tumor risk prediction, helping doctors make more appropriate diagnosis and treatment decisions. Through the adaptive optimization mechanism, the model can gradually improve its disease prediction ability as clinical data increases, thus providing strong support in daily diagnosis and treatment.
[0084] The precise classification of tumor risk by the model provides clear diagnosis and treatment suggestions for doctors. Low-risk patients can be regularly monitored, medium-risk patients need further examinations, and high-risk patients require urgent intervention. Through the adaptive optimization and dynamic adjustment mechanism, the model can flexibly handle new data, improve the prediction accuracy, make the tumor risk assessment more personalized and precise, and ultimately provide strong support for clinical decision-making.
[0085] Example 2, an embodiment of the present invention, provides an auxiliary method for the early clinical diagnosis of digestive tract tumors. In order to verify the beneficial effects of the present invention, scientific demonstrations are carried out through economic benefit calculations and simulation experiments.
[0086] First, the experiment uses the data of 6 digestive tract tumor patients. The experimental data comes from the clinical records of 6 digestive tract tumor patients, and the data includes gene mutation frequencies, mutation types, gene expression levels, patient ages, imaging examination results, etc. The gene mutation frequencies are subjected to min-max standardization so that their values are between 0 and 1. One-hot encoding is used for the gene mutation types to convert different mutation types into numerical representations. The gene expression levels are subjected to Z-score standardization processing to convert the expression levels of all genes into standardized numerical values. The numerical variables (such as patient ages) in the clinical data are also subjected to min-max standardization, and the categorical variables (such as imaging examination results) use label encoding or one-hot encoding.
[0087] A high-order non-linear regression model is adopted to combine the mutation frequency and gene expression level to describe their non-linear relationship with tumor risk. A non-linear amplification effect term is introduced, and a high-order interaction model is used to capture the interaction effects between gene mutations, especially the contributions of two-gene and three-gene interactions to tumor risk.
[0088] To reduce the computational complexity and optimize the calculation of interaction effects, the experiment introduces principal component analysis (PCA) to perform dimensionality reduction on the gene mutation data. The Laplacian matrix is introduced to measure the similarity between genes, and by adjusting the weights, the gene similarity and the PCA dimensionality reduction results are balanced.
[0089] According to the total tumor risk value, the experiment divides the patients into three categories: low, medium, and high risk. An adaptive optimization mechanism is introduced, which can dynamically adjust the model parameters according to new data to adapt to the gene mutation data of different patients.
[0090] Table 1 Experimental data table As can be seen from the experimental data table, after standardizing the gene mutation frequency and expression level of each patient, all input data are converted into a unified format, ensuring data compatibility and operability. After processing by a high-order non-linear regression model, the non-linear relationship between the gene mutation frequency and gene expression level is effectively quantified, which can reflect the non-linear enhancement effect of the increase in mutation frequency on tumor risk. The total tumor risk values of patients P4 and P6 in the table are relatively high, 0.95 and 0.90 respectively, mainly due to their relatively high mutation frequencies (30% and 40%), and their gene expression levels are also relatively high (1.5 and 2.0 respectively). While the total risk value of patient P1 is the lowest (0.75), with a mutation frequency of 12% and a gene expression level of 1.2, indicating a relatively low tumor risk.
[0091] Further analyzing the situations of patients P2 and P3, it can be seen that although their mutation frequencies are relatively low (18% and 10%), the differences in gene expression levels (P2 is -0.5, P3 is 0.8) have a significant impact on the total risk value. Through the non-linear regression model, the interaction between the mutation frequency and gene expression level is accurately reflected, and finally the risk value of P2 is 0.85, while the risk value of P3 is 0.65.
[0092] These data analyses show that the method of the present invention can effectively improve the accuracy of tumor risk assessment by comprehensively considering the gene mutation frequency, expression level and their non-linear relationship. In addition, through interaction effect modeling and Laplacian matrix optimization, the joint action of multiple genes can be effectively captured, further enhancing the prediction ability of the model. Compared with the prior art, the method of the present invention can more accurately identify high-risk patients, especially in the case of large differences in multiple gene mutations and gene expressions, showing strong prediction ability and personalized assessment effect.
[0093] Generally speaking, the experimental results verify that the tumor risk assessment method provided by the present invention has significant innovation and advantages, and can provide accurate and reliable support for the early diagnosis of tumors.
[0094] Example 3, an embodiment of the present invention, provides an auxiliary system for the clinical early diagnosis of digestive tract tumors, including a data input and preprocessing module, a gene mutation risk assessment and interaction effect modeling module, and a risk grading and warning module.
[0095] Among them, the data input and preprocessing module is used to uniformly standardize the gene mutation data, gene expression data, and clinical data from different sources and convert them into the format for model calculation. The gene mutation risk assessment and interaction effect modeling module is used to evaluate the risk contribution of a single gene through a high-order nonlinear regression model and capture the interaction effects among multiple genes. The risk grading and early warning module is used to grade the tumor risk of patients according to the calculated total risk value and provide early warnings.
[0096] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0097] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0098] More specific examples (a non-exhaustive list) of computer-readable media include the following: electrical connections (electronic devices) having one or more wirings, portable computer disk cartridges (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0099] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc. It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
[0100] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. An auxiliary method for clinical early diagnosis of digestive tract tumors, characterized in that: include: Standardize, encode and process missing values of input data, and convert gene mutation, gene expression and clinical data from different sources into a unified format; Combining mutation frequency, expression level and nonlinear relationship reflects the contribution of gene mutation to tumorigenesis; Capture the interaction effects between gene mutations to perform component analysis and dimensionality reduction; Measure the similarity between genes, optimize the calculation of interaction effects between genes, and assess the patient's tumor risk; Classify patients' tumor risks and use adaptive optimization and dynamic adjustment mechanisms to divide tumor risks.
2. The auxiliary method for early clinical diagnosis of digestive tract tumors according to claim 1, characterized in that: The standardization, encoding and missing value processing of input data includes collecting input data; Input data include gene mutation data, gene expression data and clinical data.
3. The auxiliary method for clinical early diagnosis of digestive tract tumors according to claim 2, characterized in that: The standardization, encoding and missing value processing of the input data include standardizing the gene mutation frequency, using minimum-maximum standardization to map the mutation frequency to a range of 0 to 1; Encode the gene mutation data and use one-hot encoding to convert the mutation type into numerical form; Standardize the gene expression data using Z-score standardization, normalize the data by mean and standard deviation so that the mean of the data set is 0 and the standard deviation is 1; Clinical data were digitized and standardized.
4. The auxiliary method for clinical early diagnosis of digestive tract tumors according to claim 3, characterized in that: The combining of mutation frequency, expression level and nonlinear relationship to reflect the contribution of gene mutation to tumorigenesis includes adopting a high-order nonlinear regression model to consider the nonlinear relationship between gene mutation frequency and expression level, and describe the risk contribution of gene mutation to digestive tract tumors; The nonlinear amplification effect term is introduced to reflect the risk trend of tumor occurrence; Introduce gene expression amplification term to describe the nonlinear effect of gene expression on risk; The coupling effect between expression mutation frequency and gene expression level; Constructing a gene risk contribution model reflects the contribution of gene mutations to tumorigenesis.
5. The auxiliary method for early clinical diagnosis of digestive tract tumors according to claim 4, characterized in that: The capturing of the interaction effects between gene mutations to perform component analysis and dimensionality reduction includes capturing the interaction effects between gene mutations by constructing a high-order interaction model; By adjusting the interaction effect coefficient, the interaction strength between genes can be characterized; Component analysis PCA is introduced to reduce the dimensionality of high-dimensional data of gene mutation interactions; The interaction effect value is output through the high-order interaction model, indicating the contribution of gene joint mutation to tumor occurrence.
6. The auxiliary method for early clinical diagnosis of digestive tract tumors according to claim 5, characterized in that: The method of measuring the similarity between genes, optimizing the calculation of the interaction effect between genes, and evaluating the tumor risk of patients includes introducing Laplace matrix optimization in the process of evaluating gene mutation and tumor risk to represent the similarity between nodes in graph theory; Through the Laplace matrix, the similarity between genes is quantified, the mutation risk value between genes is used as the node, and the similarity between nodes is used as the edge weight to optimize the relationship between gene mutations; Balance the weight of gene similarity and PCA; The impact of individual genes on tumor risk is represented by outputting and weighting the risk contribution of each gene.
7. The auxiliary method for early clinical diagnosis of digestive tract tumors according to claim 6, characterized in that: The tumor risk classification of the patient and the adaptive optimization and dynamic adjustment mechanism include classifying the tumor risk of the patient according to the total tumor risk value; The risk classification is based on risk thresholds, which are used to identify low-, medium-, and high-risk patients; Introduce adaptive optimization mechanism into the optimization model to adjust risk assessment according to data; Risk thresholds are dynamically adjusted based on patient data; Provide diagnosis and treatment recommendations by dividing tumor risks through risk grading and early warning.
8. A system using the auxiliary method for early clinical diagnosis of digestive tract tumors as claimed in any one of claims 1 to 7, characterized in that: It includes data input and preprocessing module, gene mutation risk assessment and interaction effect modeling module, risk classification and early warning module; The data input and preprocessing module is used to standardize gene mutation data, gene expression data and clinical data from different sources and convert them into a format for model calculation; The gene mutation risk assessment and interaction effect modeling module is used to assess the risk contribution of a single gene through a high-order nonlinear regression model and capture the interaction effects between multiple genes; The risk grading and early warning module is used to grade the patient's tumor risk according to the calculated total risk value and provide early warning.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the processor implements the steps of the auxiliary method for early clinical diagnosis of gastrointestinal tumors described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the auxiliary method for early clinical diagnosis of digestive tract tumors described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method for systematically quantifying influence of mutation on expression profile based on neural network
CN118262794A
Kidney clear cell tumor prediction method, device and equipment and storage medium
CN118711678A
Data analysis system and method of gene regulation and control network based on deep regression algorithm
CN119252332A
Expert system for classification and prediction of generic diseases, and for association of molecular genetic parameters with clinical parameters
US20040076984A1
Methods and systems for determining personalized therapies
US20180107786A1
Cited By
Elderly NSC emergency treatment risk layering method, device and medium
CN120878244A