An auxiliary method for clinical early diagnosis of digestive tract tumors
By standardizing and encoding gene mutation, expression, and clinical data, combined with high-order nonlinear regression models and Laplace matrix optimization, the problem of insufficient capture of gene interaction effects in existing technologies is solved, achieving more accurate tumor risk assessment and early diagnosis.
Patent Information
- Application Number
- CN202510300414.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-03-14
AI Technical Summary
Existing early tumor diagnosis methods cannot fully capture the complex interaction effects between genes, ignore the nonlinear relationship between gene mutations and expression levels, and are inefficient in processing multi-dimensional data, resulting in limited diagnostic accuracy and practicality.
By standardizing, encoding and processing missing values of the input data, gene mutations, gene expression and clinical data from different sources are converted into a unified format. Combined with mutation frequency, expression level and nonlinear relationship, a high-order nonlinear regression model is used to capture the interaction effects between gene mutations. The Laplace matrix is used to optimize the similarity between genes, and an adaptive optimization mechanism is introduced for dynamic adjustment to assess patients' tumor risk.
It significantly improves the accuracy and personalization of tumor risk prediction, enhances the sensitivity and accuracy of the model, and provides more reliable and efficient support for early clinical diagnosis.
Smart Images

Figure CN120164534B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tumor risk assessment and genomic data processing, and specifically to an auxiliary method for clinical early diagnosis of digestive tract tumors. Background Art
[0002] In recent years, with the rapid development of biomedicine and genomics, early diagnosis technology for gastrointestinal tumors has received widespread attention. Traditional tumor screening methods, such as imaging examinations and tumor marker detection, can help diagnose tumors to a certain extent, but their sensitivity and specificity are relatively low, and it is often difficult to detect tumors in the early stages. With the advancement of genomics and molecular biology, gene mutations, gene expression and their related biomarkers have become new directions for early diagnosis of tumors. The use of genomic data to assess tumor risk has become a research hotspot. In recent years, multi-level analysis methods based on gene mutation data and clinical data, especially multi-gene combination analysis, have achieved initial application results in tumor screening. These methods provide potential for early diagnosis of tumors through in-depth mining of genomic data, but still face challenges in data processing and analysis capabilities.
[0003] Existing early cancer diagnosis methods, particularly those based on gene mutation data, primarily rely on single-effect models of gene mutations, neglecting the complex interactions between different genes. Although gene mutation frequency and expression levels are widely used in cancer risk assessment, existing technologies often fail to process multidimensional and multi-level data, particularly underconsidering the interactive effects of multiple genes. Traditional cancer risk assessment methods often use simple linear models to describe the effects of gene mutations, ignoring the nonlinear relationship between gene mutations and gene expression, which limits the accuracy of risk assessment. Furthermore, when modeling the interactive effects of multiple genes, most methods only consider low-order or simple binary interactions, failing to further explore higher-order, complex interactions between genes. Although principal component analysis (PCA) has been applied to dimensionality reduction, existing techniques often limit its use to data dimensionality reduction and fail to fully optimize the calculation of interaction effects, resulting in low processing efficiency and inaccurate results for high-dimensional genetic data. Existing risk assessment methods also lack effective data standardization, encoding, and missing value handling mechanisms when processing genetic data of varying types, sources, and scales. Consequently, they face significant data complexity and inconsistency in practical applications. Therefore, the existing technology cannot provide effective solutions in multiple dimensions such as multi-gene interaction effects, nonlinear relationships, data standardization and optimized calculations like the present invention, which limits the accuracy and practicality of early tumor diagnosis. Summary of the Invention
[0004] In view of the above-mentioned problems, the present invention is proposed.
[0005] Therefore, the technical problems solved by the present invention are: the existing early diagnosis methods for tumors are unable to fully capture the complex interaction effects between genes, ignore the nonlinear relationship between gene mutations and expression levels, are inefficient in processing multi-dimensional data, and how to comprehensively improve diagnostic accuracy through standardization, coding and optimization algorithms.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: an auxiliary method for the early clinical diagnosis of gastrointestinal tumors, comprising standardizing, encoding and processing missing values of input data, converting gene mutations, gene expression and clinical data from different sources into a unified format; combining mutation frequency, expression level and nonlinear relationship to reflect the contribution of gene mutations to tumor occurrence; capturing the interaction effects between gene mutations to perform component analysis and dimensionality reduction; measuring the similarity between genes, optimizing the calculation of interaction effects between genes, and evaluating the patient's tumor risk; classifying the patient's tumor risk and adaptively optimizing and dynamically adjusting the mechanism to divide the tumor risk.
[0007] As a preferred embodiment of the auxiliary method for early clinical diagnosis of digestive tract tumors of the present invention, wherein: the standardization, encoding and missing value processing of input data includes collecting input data;
[0008] Input data include gene mutation data, gene expression data and clinical data.
[0009] As a preferred embodiment of the auxiliary method for early clinical diagnosis of digestive tract tumors according to the present invention, the standardization, encoding and missing value processing of input data includes standardizing gene mutation frequencies, using minimum-maximum normalization to map mutation frequencies to a range of 0 to 1;
[0010] Encode the gene mutation data and use one-hot encoding to convert the mutation type into numerical form;
[0011] Gene expression data were normalized using Z-score normalization, which normalized the data by mean and standard deviation so that the mean of the data set was 0 and the standard deviation was 1;
[0012] Clinical data were digitized and standardized.
[0013] As a preferred embodiment of the auxiliary method for early clinical diagnosis of digestive tract tumors described in the present invention, the method of combining mutation frequency, expression level, and nonlinear relationship to reflect the contribution of gene mutation to tumorigenesis includes using a high-order nonlinear regression model to describe the risk contribution of gene mutation to digestive tract tumors by considering the nonlinear relationship between gene mutation frequency and expression level;
[0014] The nonlinear amplification effect term is introduced to reflect the risk trend of tumor occurrence;
[0015] Introducing a gene expression amplification term to describe the nonlinear effect of gene expression on risk;
[0016] The coupling effect between expression mutation frequency and gene expression level;
[0017] Constructing a gene risk contribution model reflects the contribution of gene mutations to tumorigenesis.
[0018] As a preferred embodiment of the auxiliary method for early clinical diagnosis of digestive tract tumors according to the present invention, the method of capturing the interaction effects between gene mutations and performing component analysis and dimensionality reduction includes capturing the interaction effects between gene mutations by constructing a high-order interaction model;
[0019] By adjusting the interaction effect coefficient, the interaction strength between genes can be characterized;
[0020] Component analysis PCA is introduced to reduce the dimensionality of high-dimensional data of gene mutation interactions;
[0021] The interaction effect value is output through the high-order interaction model to indicate the contribution of gene joint mutation to tumor occurrence.
[0022] As a preferred embodiment of the auxiliary method for early clinical diagnosis of digestive tract tumors described in the present invention, the method of measuring the similarity between genes, optimizing the calculation of the interaction effect between genes, and assessing the patient's tumor risk includes introducing Laplace matrix optimization to represent the similarity between nodes in graph theory during the assessment of gene mutation and tumor risk;
[0023] Through the Laplace matrix, the similarity between genes is quantified, the mutation risk value between genes is used as the node, and the similarity between nodes is used as the edge weight to optimize the relationship between gene mutations;
[0024] Balance the weight of gene similarity and PCA;
[0025] The impact of individual genes on tumor risk is expressed by outputting and weighting the risk contribution of each gene.
[0026] As a preferred embodiment of the auxiliary method for early clinical diagnosis of digestive tract tumors according to the present invention, wherein: the classification of the patient's tumor risk and the adaptive optimization and dynamic adjustment mechanism include classifying the patient's tumor risk according to the total tumor risk value;
[0027] Risk classification is based on risk thresholds, which are used to identify low-, medium-, and high-risk patients;
[0028] Introducing an adaptive optimization mechanism into the optimization model to adjust risk assessment based on data;
[0029] Risk thresholds are dynamically adjusted based on patient data;
[0030] Provide diagnosis and treatment recommendations by dividing tumor risks through risk grading and early warning.
[0031] Another object of the present invention is to provide an auxiliary system for the clinical early diagnosis of gastrointestinal tumors, which can reflect the contribution of gene mutations to tumor occurrence by combining mutation frequency, expression level and nonlinear relationship, thereby solving the problem that current early tumor diagnosis methods cannot fully capture the complex interaction effects between genes.
[0032] As a preferred embodiment of the auxiliary system for early clinical diagnosis of gastrointestinal tumors described in the present invention, it includes: a data input and preprocessing module, a gene mutation risk assessment and interaction effect modeling module, and a risk grading and early warning module; the data input and preprocessing module is used to standardize gene mutation data, gene expression data and clinical data from different sources, and convert them into a format for model calculation; the gene mutation risk assessment and interaction effect modeling module is used to evaluate the risk contribution of a single gene through a high-order nonlinear regression model, and capture the interaction effects between multiple genes; the risk grading and early warning module is used to grade the patient's tumor risk based on the calculated total risk value, and provide early warning.
[0033] A computer device includes a memory and a processor, wherein the memory stores a computer program and the processor executes the computer program to implement a step of an auxiliary method for clinical early diagnosis of digestive tract tumors.
[0034] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of an auxiliary method for clinical early diagnosis of digestive tract tumors.
[0035] Beneficial effects of the present invention: The auxiliary method for early clinical diagnosis of gastrointestinal tumors provided by the present invention standardizes, encodes, and handles missing values in input data, converting gene mutation, gene expression, and clinical data from various sources into a unified format to ensure data consistency and operability. This processing not only improves data quality but also provides a reliable foundation for subsequent analysis. Secondly, the contribution of gene mutations to tumorigenesis is reflected by combining mutation frequency, expression, and nonlinear relationships. A high-order nonlinear regression model is used to capture the nonlinear relationship between gene mutations and expression, accurately describing the contribution of genes to tumor risk. This step effectively improves the ability to process mutation frequency, gene expression, and their interrelationships, enhancing the model's sensitivity and accuracy. Furthermore, the method captures the interaction effects between gene mutations and performs component analysis and dimensionality reduction. High-order interaction models and principal component analysis (PCA) are used to optimize the calculation of interaction effects, enabling effective modeling of the effects of multi-gene interactions while also improving computational efficiency through dimensionality reduction. Next, the method measures inter-gene similarity and optimizes the calculation of interaction effects. The Laplacian matrix is introduced to optimize the inter-gene similarity metric and the calculation of interaction effects is adjusted, improving the accuracy of assessing complex interactions between multiple genes. Finally, the system categorizes patients' tumor risk and introduces an adaptive optimization and dynamic adjustment mechanism to dynamically adjust model parameters and risk thresholds based on the patient's genetic mutation data, achieving personalized risk assessment. Through this multi-dimensional, multi-step collaborative approach, the present invention significantly improves the accuracy, personalization, and computational efficiency of tumor risk prediction, providing more reliable and efficient support for early clinical diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0037] Figure 1 This is an overall flow chart of an auxiliary method for early clinical diagnosis of digestive tract tumors provided by the first embodiment of the present invention. DETAILED DESCRIPTION
[0038] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0039] Example 1, with reference to Figure 1 , which is an embodiment of the present invention, provides an auxiliary method for clinical early diagnosis of digestive tract tumors, comprising:
[0040] S1: Standardize, encode, and handle missing values of input data, and convert gene mutation, gene expression, and clinical data from different sources into a unified format.
[0041] Furthermore, the input data includes gene mutation data, gene expression data and clinical data.
[0042] Gene mutation data includes gene selection, mutation frequency, mutation type and mutation site.
[0043] Gene mutation data is the foundation of model calculations. Common genes associated with gastrointestinal tumors include, but are not limited to, p53, K-ras, APC, and C-myc. Mutations in each gene can affect the probability of tumor development, so accurate mutation information for each gene is essential.
[0044] Mutation frequency is the proportion of mutations in each gene in a sample, typically obtained through gene sequencing or microarray technology. Mutation frequency is a measure of the prevalence of gene mutations in a patient population.
[0045] The mutation type is the specific type of mutation for each gene (e.g., point mutation, deletion, insertion, frameshift mutation, etc.). The mutation type affects gene function and its role in tumorigenesis and must be clearly labeled.
[0046] Mutation sites indicate where a gene mutation occurs. Mutations at certain sites may be closely associated with tumor development. Analyzing specific sites can help quantify the risk of gene mutations.
[0047] Gene expression data includes gene expression levels. Gene expression refers to the abundance of gene transcripts (mRNA), typically obtained through RNA-seq or real-time quantitative PCR (qPCR). Gene expression is a measure of gene activity in cells. Combining mutation frequency and expression levels can more comprehensively assess a gene's contribution to cancer risk.
[0048] Clinical data include basic patient information, diagnosis results, and imaging examinations.
[0049] The patient's basic information includes age, gender, lifestyle habits (such as smoking, drinking, etc.), family history, medical history, etc. This information helps to improve the comprehensive evaluation ability of the model.
[0050] The diagnosis result is the patient's historical diagnostic data, including whether it is diagnosed as a gastrointestinal tumor, the tumor stage, the tumor type, etc.
[0051] The results of imaging examinations such as CT, MRI, and ultrasound play an important role in early tumor screening.
[0052] It should be noted that the input data come from different sources and are not in uniform formats. Through standardization, the input data are converted into a unified format suitable for model calculation.
[0053] If the mutation frequency values vary widely (for example, the frequency difference ranges from 0 to 1), normalization is required.
[0054] Use min-max normalization: map the mutation frequency to the range of 0 to 1, expressed as:
[0055]
[0056] in, is the original mutation frequency, and are the minimum and maximum values in the data set, respectively. is the standardized value.
[0057] Gene mutation types are usually discrete variables (such as point mutations, insertions, deletions, etc.), and model calculations need to convert them into numerical form. One-hot encoding is used to convert mutation types into numerical form. Assume that the gene There are three mutation types (A, B, C), which are converted into the following codes:
[0058]
[0059] Each type is represented by a separate binary vector, which applies to multiple possible mutation types.
[0060] Standardization of gene expression. Gene expression is usually a continuous variable, and the expression levels of different genes vary greatly, so standardization is required.
[0061] Using Z-score standardization, the data is normalized by the mean and standard deviation so that the mean of the data set is 0 and the standard deviation is 1, which is expressed as:
[0062]
[0063] in, is the original gene expression level, and are the mean and standard deviation of gene expression, is the normalized gene expression level.
[0064] Clinical data usually include categorical variables and numerical variables. In order to adapt to model input, categorical variables need to be quantified and numerical variables may need to be standardized.
[0065] For categorical variables (such as gender, smoking status, etc.), use one-hot encoding or label encoding.
[0066] For numerical variables (such as age and tumor size), minimum-maximum standardization or Z-score standardization was used.
[0067] Missing value processing,Missing values are a common problem in data processing, and missing values must be processed to ensure data integrity.
[0068] For numerical data, mean filling or median filling is used, or machine learning algorithms (such as KNN filling) can be used for filling.
[0069] For categorical data, mode filling or most frequent category filling is used.
[0070] After normalization and encoding, all input data is converted into a unified data matrix. The rows of the data matrix represent each patient's sample, and the columns represent features such as gene mutations, gene expression, and clinical information. For example, assuming there are n patient samples and m features (including mutation frequencies and expression levels of multiple genes), the final input data matrix X is an n×m matrix, where each row represents all the features of a single patient.
[0071] The final data input matrix will serve as the input of the computational model and be passed to the subsequent mathematical model for risk assessment and prediction.
[0072] S2: Combining mutation frequency, expression level and nonlinear relationship reflects the contribution of gene mutation to tumorigenesis.
[0073] Furthermore, considering the complexity of gene mutations, a high-order nonlinear regression model was used to describe the impact of gene mutations on the risk of gastrointestinal tumors. The high-order nonlinear regression model is as follows:
[0074]
[0075] in, For the The risk contribution of each gene; For the The mutation frequency of each gene; For the The expression level of each gene; is the coefficient of the gene, controlling the mutation frequency, gene expression and the nonlinear relationship between them.
[0076] It should be noted that the mutation frequency It is not directly proportional to the risk of tumor. As the mutation frequency increases, the risk of tumor occurrence also shows a nonlinear increasing trend. Especially in the case of high-frequency mutations, the risk may rise sharply. Therefore, by introducing The term is used to reflect this nonlinear amplification effect, where and It is a parameter that controls the nonlinear effect of mutation frequency and can flexibly adjust the impact of mutation frequency on tumor risk.
[0077] Gene expression It is also closely related to the risk of tumors. Generally speaking, highly expressed genes may increase the risk of tumors, while low-expressed genes have less impact. item, can more accurately describe the nonlinear effect of gene expression on risk. A parameter that controls the intensity of the effect of expression. The exponential term makes the impact of changes in gene expression on tumor risk more prominent.
[0078] Mutation frequency and gene expression They do not affect tumorigenesis independently; there may be interactions between them. For example, the high frequency of certain gene mutations may enhance tumor risk in the context of high gene expression. Therefore, the terms in the high-order nonlinear regression model are The combination of the interactive effects of mutation frequency and gene expression level reflects the joint effect of the two in affecting tumor risk.
[0079] To better reflect the complex effects of gene mutations, high-order nonlinear regression models use high-order nonlinear terms. Specifically, the mutation frequency term Indicates the nonlinear effect of mutation frequency on tumor risk, gene expression Represents the nonlinear effect of gene expression.
[0080] Furthermore, there may be complex interactions between mutations in multiple genes, and the interaction effects have a significant impact on tumor risk. To capture the interaction effects, a high-order interaction model was designed and combined with principal component analysis (PCA) to optimize the interaction effects, which can be expressed as:
[0081]
[0082] in, gene , gene , gene Value at risk; is the interaction effect coefficient, which indicates the interaction effect between genes; PCA is used for principal component analysis to optimize the calculation of interaction effects between genes.
[0083] It should be noted that in the high-order interaction model, Representative genes ,Gene and genes By adjusting the coefficients, the high-order interaction model can describe the different interaction strengths between genes.
[0084] The interaction effect is not limited to the combination of two-gene mutations, but also considers the interaction of three genes.
[0085] Principal component analysis (PCA) was introduced. PCA is used to reduce the dimensionality of high-dimensional data on the interactions of multiple gene mutations, extracting the principal components and simplifying the computational process. PCA extracts the key interactions between gene mutations through dimensionality reduction, compressing the high-dimensional interactions into a smaller number of dimensions, thereby effectively reducing computational time and memory consumption. The reduced interactions can then be directly used to calculate the overall cancer risk.
[0086] S3: Capture the interaction effects between gene mutations to perform component analysis and dimensionality reduction.
[0087] Furthermore, the Laplace matrix is used to measure the similarity between genes and optimize the calculation of gene interaction effects. By optimizing the Laplace matrix, the calculation method of mutation similarity between genes can be improved, making the contribution of gene interaction effects more precise and further improving the accuracy of tumor risk assessment.
[0088] The similarity of mutations between genes can be expressed using a risk value. High similarity between two genes indicates that they may contribute similarly to a tumor under the same conditions, or that mutations in one gene may amplify the risk of mutations in another. The Laplacian matrix can be used to quantify this similarity between genes, effectively measuring the impact of gene interactions.
[0089] The Laplace matrix optimizes the relationship between gene mutations and reduces the redundancy in the calculation of interaction effects by using the mutation risk values between genes as nodes and the similarity between nodes as edge weights.
[0090] The weight of the Laplace matrix controls the weight balance between the similarity between genes and the PCA dimensionality reduction results. The Laplace matrix is introduced to measure the similarity between genes, and the coefficients are adjusted by optimizing the value of the Laplace matrix, which is expressed as:
[0091]
[0092] in, It's genes and The Laplacian matrix values between ; It is a balance term that controls the ratio of similarity between different genes to the PCA dimensionality reduction results; It represents the result of PCA dimensionality reduction of the interaction between genes; It's genes and The Laplacian matrix values between ; It is a balance term that controls the ratio of similarity between different genes to the PCA dimensionality reduction results; It represents the result of PCA dimensionality reduction of the interaction between genes.
[0093] S4: Measure the similarity between genes, optimize the calculation of interaction effects between genes, and assess patients' tumor risk.
[0094] Furthermore, the final total cancer risk It is necessary to comprehensively consider the independent risk of each gene, the interaction effect between genes, and the similarity after Laplace matrix optimization to comprehensively assess the patient's tumor risk. By summing up the risk contribution of each part, the final risk value is output as the basis for diagnosis. The risk contribution of each gene The interaction effects between genes are calculated and weighted to reflect the effect of individual genes on cancer risk. By considering the complex synergistic effects between genes, the predictive ability of the model is further enhanced. The interaction effect can reflect the additive effect of multiple gene mutations on the risk of tumor development. The measurement of similarity between genes has been optimized, further improving the calculation accuracy of interaction effects, making the final tumor risk assessment more accurate.
[0095] It should be noted that the final tumor risk The risk of a single gene, the interaction effect between genes and the optimized similarity matrix are integrated and expressed as:
[0096]
[0097] By integrating single-gene risk, interaction effects, and similarity between genes, the final total risk value can comprehensively reflect the patient's tumor risk and improve the accuracy of early diagnosis.
[0098] S5: Classify patients’ tumor risks and use adaptive optimization and dynamic adjustment mechanisms to divide tumor risks.
[0099] Furthermore, according to the total tumor risk value , classify the patient's tumor risk to provide accurate diagnostic decision support. Risk grading aims to divide the patient's tumor risk into different levels, thereby guiding doctors to take different intervention measures. The risk level division is based on the risk threshold obtained in advance through model optimization. and , thresholds are used to identify low-, medium-, and high-risk patients.
[0100] When the total risk Below threshold At this point, the patient may not need further testing or treatment immediately, but will still need regular monitoring.
[0101] When the total risk lie in and When the patient's tumor risk is between 2 and 3 months, it is considered to be at a moderate level, and further examination or early intervention is recommended to detect potential tumor risks in a timely manner.
[0102] When the total risk Exceeding the threshold When the patient's tumor risk is high, urgent investigation and treatment measures are required, and more in-depth imaging or histological analysis may be required.
[0103] Based on the calculated total risk Expressed as:
[0104]
[0105] in, and The risk threshold is optimized based on the model, and is divided into low, medium and high risk levels. When the patient's total risk value exceeds the high risk threshold When the patient is ill, the system will automatically issue an early warning, prompting the doctor to conduct further examination.
[0106] It should be noted that to improve the accuracy of the model and adapt it to different data sets, an adaptive optimization mechanism was introduced. This mechanism dynamically adjusts the interaction effect coefficient and risk threshold, enabling the model to self-adjust as new data is input, continuously optimizing its prediction accuracy.
[0107] Interaction effect coefficient and Controls the degree of influence of gene interactions. The coefficient value will be adjusted according to new data to reflect the actual impact of different gene mutation combinations on tumor risk. Risk threshold and It can also be dynamically adjusted based on new patient data. Through optimization algorithms such as gradient descent, the model automatically adjusts these thresholds to improve the accuracy of risk classification. Through gradient descent or other optimization algorithms, the model continuously minimizes the loss function .
[0108] To enhance the adaptability of the model, adaptive optimization and gradient descent method are used to dynamically adjust the interaction effect coefficient and risk threshold. It can be expressed as:
[0109]
[0110] in, is the learning rate, which controls the magnitude of each update; is the loss function, which represents the prediction error of the model.
[0111] Adaptive optimization allows the model to dynamically adjust parameters based on new data, ensuring continuous optimization over time and adapting to changing clinical data. Gradient descent can improve the accuracy of tumor risk assessment by adjusting interaction coefficients and risk thresholds through backpropagation.
[0112] As new data is continuously input, the model can self-learn and adapt, adapting to the characteristics of different patient groups and enhancing its generalization capabilities. Dynamic adjustments to interaction effect coefficients and risk thresholds ensure that each patient's risk assessment is personalized based on their specific genetic background. Dynamic adjustment and optimization can provide more accurate tumor risk predictions, helping doctors make more appropriate diagnosis and treatment decisions. Through this adaptive optimization mechanism, the model's predictive capabilities can gradually improve as clinical data increases, providing strong support for daily diagnosis and treatment.
[0113] The model accurately categorizes tumor risk, providing physicians with clear diagnostic and treatment recommendations. Low-risk patients can be monitored regularly, medium-risk patients require further examination, and high-risk patients require urgent intervention. Through adaptive optimization and dynamic adjustment mechanisms, the model can flexibly respond to new data, improving prediction accuracy and making tumor risk assessment more personalized and accurate, ultimately providing strong support for clinical decision-making.
[0114] Example 2, an embodiment of the present invention, provides an auxiliary method for clinical early diagnosis of gastrointestinal tumors. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.
[0115] First, the experiment used data from six patients with gastrointestinal cancers. The experimental data were derived from the clinical records of these patients. The data included gene mutation frequency, mutation type, gene expression, patient age, and imaging findings. Gene mutation frequency was min-max normalized to a value between 0 and 1. Gene mutation types were encoded using one-hot encoding to convert different mutation types into numerical representations. Gene expression levels were normalized using Z-scores to convert the expression levels of all genes into standardized numerical values. Numerical variables in the clinical data (such as patient age) were also min-max normalized, and categorical variables (such as imaging findings) were encoded using label encoding or one-hot encoding.
[0116] A high-order nonlinear regression model was used to combine mutation frequency and gene expression to describe the nonlinear relationship between mutation frequency and tumor risk. A nonlinear amplification effect term was introduced, and a high-order interaction model was used to capture the interactive effects between gene mutations, particularly the contribution of two-gene and three-gene interactions to tumor risk.
[0117] To reduce computational complexity and optimize the calculation of interaction effects, the experiment introduced principal component analysis (PCA) to reduce the dimensionality of gene mutation data. The Laplace matrix was introduced to measure the similarity between genes, and weights were adjusted to balance gene similarity and PCA dimensionality reduction results.
[0118] Based on the total tumor risk score, the experiment categorized patients into three risk groups: low, medium, and high. An adaptive optimization mechanism was introduced to dynamically adjust model parameters based on new data, adapting to the genetic mutation data of different patients.
[0119] Table 1 Experimental data table
[0120]
[0121] As shown in the experimental data table, after standardizing the gene mutation frequency and expression level for each patient, all input data were converted to a unified format, ensuring data compatibility and operability. Using a high-order nonlinear regression model, the nonlinear relationship between gene mutation frequency and gene expression level was effectively quantified, reflecting the nonlinear enhancing effect of increased mutation frequency on cancer risk. Patients P4 and P6 in the table have higher total cancer risk values of 0.95 and 0.90, respectively, primarily due to their higher mutation frequencies (30% and 40%) and higher gene expression levels (1.5 and 2.0, respectively). Patient P1, on the other hand, has the lowest total risk value (0.75), with a mutation frequency of 12% and a gene expression level of 1.2, indicating a lower cancer risk.
[0122] Further analysis of patients P2 and P3 revealed that despite relatively low mutation frequencies (18% and 10%), differences in gene expression (-0.5 for P2 and 0.8 for P3) significantly impacted the overall risk. The nonlinear regression model accurately captured the interaction between mutation frequency and gene expression, ultimately resulting in a risk of 0.85 for P2 and 0.65 for P3.
[0123] These data analyses demonstrate that the method of the present invention can effectively improve the accuracy of tumor risk assessment by comprehensively considering gene mutation frequency, expression levels, and their nonlinear relationships. Furthermore, through interaction effect modeling and Laplace matrix optimization, it can effectively capture the combined effects of multiple genes, further enhancing the model's predictive power. Compared to existing technologies, the method of the present invention can more accurately identify high-risk patients, especially in cases of multiple gene mutations and large differences in gene expression, demonstrating strong predictive power and personalized assessment results.
[0124] In general, the experimental results verify that the tumor risk assessment method provided by the present invention has significant innovation and advantages, and can provide accurate and reliable support for the early diagnosis of tumors.
[0125] Example 3, an embodiment of the present invention, provides an auxiliary system for clinical early diagnosis of gastrointestinal tumors, including a data input and preprocessing module, a gene mutation risk assessment and interaction effect modeling module, and a risk grading and early warning module.
[0126] The data input and preprocessing module is used to standardize gene mutation data, gene expression data and clinical data from different sources and convert them into a format for model calculation. The gene mutation risk assessment and interaction effect modeling module is used to evaluate the risk contribution of a single gene through a high-order nonlinear regression model and capture the interaction effects between multiple genes. The risk grading and warning module is used to grade the patient's tumor risk based on the calculated total risk value and provide early warning.
[0127] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0128] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0129] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting, or processing it in another suitable manner as necessary, and then storing it in a computer memory.
[0130] It should be understood that various aspects of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following technologies known in the art may be used: a discrete logic circuit having logic gates for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gates, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc. It should be noted that the above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to be limiting. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art will understand that modifications or equivalent substitutions may be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and such modifications are intended to be encompassed by the claims of the present invention.
[0131] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. An auxiliary method for clinical early diagnosis of digestive tract tumors, characterized in that: include: Standardize, encode, and handle missing values of input data, converting gene mutation, gene expression, and clinical data from different sources into a unified format; Combining mutation frequency, expression level and nonlinear relationship to reflect the contribution of gene mutation to tumorigenesis; The combination of mutation frequency, expression level and nonlinear relationship to reflect the contribution of gene mutation to tumorigenesis includes using a high-order nonlinear regression model to consider the nonlinear relationship between gene mutation frequency and expression level to describe the risk contribution of gene mutation to digestive tract tumors; The nonlinear amplification effect term is introduced to reflect the risk trend of tumor occurrence; Introducing a gene expression amplification term to describe the nonlinear effect of gene expression on risk; The coupling effect between expression mutation frequency and gene expression level; Constructing a gene risk contribution model to reflect the contribution of gene mutations to tumorigenesis; The high-order nonlinear regression model is as follows: ; in, For the The risk contribution of each gene; For the The mutation frequency of each gene; For the The expression level of each gene; is the coefficient of the gene, controlling the mutation frequency, gene expression and the nonlinear relationship between them; Mutation frequency term Indicates the nonlinear effect of mutation frequency on tumor risk, gene expression Indicates the nonlinear effect of gene expression; The interaction effect is optimized by combining principal component analysis, which is expressed as: ; in, For genes Value at risk; is the interaction effect coefficient, which indicates the interaction effect between genes; PCA is used for principal component analysis to optimize the calculation of the interaction effect between genes; Capture the interaction effects between gene mutations to perform component analysis and dimensionality reduction; Measure the similarity between genes, optimize the calculation of interaction effects between genes, and assess patients' tumor risk; Final cancer risk The risk of a single gene, the interaction effect between genes and the optimized similarity matrix are integrated and expressed as: ; Classify patients' tumor risks and use adaptive optimization and dynamic adjustment mechanisms to divide tumor risks.
2. The auxiliary method for early clinical diagnosis of digestive tract tumors according to claim 1, characterized in that: The standardization, encoding and missing value processing of input data includes collecting input data; Input data include gene mutation data, gene expression data and clinical data.
3. The auxiliary method for early clinical diagnosis of digestive tract tumors according to claim 2, characterized in that: The standardization, encoding and missing value processing of the input data includes standardizing the gene mutation frequency, using minimum-maximum standardization to map the mutation frequency to a range of 0 to 1; Encode the gene mutation data and use one-hot encoding to convert the mutation type into numerical form; Gene expression data were normalized using Z-score normalization, which normalized the data by mean and standard deviation so that the mean of the data set was 0 and the standard deviation was 1; Clinical data were digitized and standardized.
4. The auxiliary method for early clinical diagnosis of digestive tract tumors according to claim 3, characterized in that: The capturing of the interaction effects between gene mutations to perform component analysis and dimensionality reduction includes capturing the interaction effects between gene mutations by constructing a high-order interaction model; By adjusting the interaction effect coefficient, the interaction strength between genes can be characterized; Component analysis PCA is introduced to reduce the dimensionality of high-dimensional data of gene mutation interactions; The interaction effect value is output through the high-order interaction model to indicate the contribution of gene joint mutation to tumor occurrence.
5. The auxiliary method for early clinical diagnosis of digestive tract tumors according to claim 4, characterized in that: The method of measuring the similarity between genes, optimizing the calculation of the interaction effect between genes, and evaluating the patient's tumor risk includes introducing Laplace matrix optimization to represent the similarity between nodes in graph theory during the evaluation process of gene mutation and tumor risk; Through the Laplace matrix, the similarity between genes is quantified, the mutation risk value between genes is used as the node, and the similarity between nodes is used as the edge weight to optimize the relationship between gene mutations; Balance the weight of gene similarity and PCA; The impact of individual genes on tumor risk is expressed by outputting and weighting the risk contribution of each gene.
6. The auxiliary method for early clinical diagnosis of digestive tract tumors according to claim 5, characterized in that: The tumor risk classification of patients and the adaptive optimization and dynamic adjustment mechanism include classifying the tumor risk of patients according to the total tumor risk value; Risk classification is based on risk thresholds, which are used to identify low-, medium-, and high-risk patients; Introducing an adaptive optimization mechanism into the optimization model to adjust risk assessment based on data; Risk thresholds are dynamically adjusted based on patient data; Provide diagnosis and treatment recommendations by dividing tumor risks through risk grading and early warning.
7. A system using the auxiliary method for early clinical diagnosis of digestive tract tumors according to any one of claims 1 to 6, characterized in that: It includes data input and preprocessing module, gene mutation risk assessment and interaction effect modeling module, and risk grading and early warning module; The data input and preprocessing module is used to standardize gene mutation data, gene expression data and clinical data from different sources and convert them into a format suitable for model calculation; The gene mutation risk assessment and interaction effect modeling module is used to assess the risk contribution of a single gene through a high-order nonlinear regression model and capture the interaction effects between multiple genes; The risk grading and early warning module is used to grade the patient's tumor risk according to the calculated total risk value and provide early warning.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the processor implements the steps of the auxiliary method for early clinical diagnosis of digestive tract tumors according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the auxiliary method for early clinical diagnosis of digestive tract tumors according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method for systematically quantifying influence of mutation on expression profile based on neural network
CN118262794A
Methods and systems for determining personalized therapies
US20180107786A1