A Digestive Tumor Risk Assessment System Based on Multi-omics Data
By constructing a multi-omics data-driven risk assessment system for digestive tumors, integrating circulating tumor DNA methylation, serum biochemistry, and clinical phenotypic data, this system addresses the limitations of existing technologies, such as reliance on single data sources and insufficient organ-specific analysis, enabling early and accurate screening and clear medical guidance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BAOTOU MEDICAL COLLEGE OF INNER MONGOLIA UNIV OF SCI & TECH
- Filing Date
- 2026-04-14
- Publication Date
- 2026-06-02
AI Technical Summary
Existing digestive tumor risk assessment technologies suffer from problems such as reliance on a single data source, insufficient integration of multi-dimensional information, inability to effectively distinguish risks at specific anatomical sites, and lack of organ-specific analysis engines, resulting in low assessment accuracy and poor guidance.
A digestive tumor risk assessment system based on multi-omics data was constructed, including a multi-source heterogeneous data acquisition terminal, a data standardization and preprocessing module, a broad-spectrum digestive system risk assessment engine, an organ-specific feature decoupling module, and a multimodal cross-omics fusion reasoning module. By integrating circulating tumor DNA methylation, serum biochemical indicators, and clinical phenotypic data, multi-dimensional information fusion and organ-specific risk assessment were achieved.
It significantly improves the sensitivity and specificity of risk assessment, enabling early detection of digestive system lesions, providing clear directions for medical treatment, saving medical resources, reducing the invasiveness of examinations, and achieving a seamless connection between cutting-edge life science technologies and practical clinical applications.
Smart Images

Figure CN122135997A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics and medical big data processing, specifically relating to a digestive tumor risk assessment system based on multi-omics data. Background Technology
[0002] Cancer risk assessment is a crucial component of modern precision medicine. It involves the comprehensive analysis of individual physiological indicators, lifestyle habits, and genetic information to pre-determine the probability of developing malignant tumors. In the field of digestive system diseases, due to the high physiological correlation and pathological differences among organs such as the esophagus, stomach, and colorectal region, early and accurate screening has become a key means of reducing cancer mortality. With the development of life science technologies, the in-depth analysis of the human internal microenvironment and molecular markers using multi-omics data integration has become a core technological approach for digestive health monitoring and disease prevention.
[0003] Among them, the multi-omics-based digestive tumor risk assessment system focuses on predicting the risk of disease in specific organs of the digestive system. It aims to collect and jointly model multi-source heterogeneous data, including circulating tumor DNA methylation levels, biochemical indicators, and clinical history, to output clinically significant quantitative risk results. This type of system strives to identify subtle biomolecular changes before the onset of clinical symptoms, providing examinees with precise medical advice and intervention guidance, thereby establishing an effective connection from general population screening to precision screening.
[0004] Current technologies for assessing the risk of digestive tumors still face significant challenges. Many assessment models rely on single data sources, such as clinical questionnaires or routine blood indicators, leading to insufficient integration of multi-dimensional information and limited assessment accuracy. This single-dimensional analysis model often only provides a broad conclusion on gastrointestinal risk, failing to effectively distinguish the specific anatomical source of the risk, resulting in poor guidance for clinical action. Furthermore, although liquid biopsy technology has been applied in some single-cancer screenings, due to significant differences in molecular markers between the upper and lower gastrointestinal tracts, current technologies generally lack organ-specific analysis engines, making it difficult to achieve precise site-specific analysis in complex data streams. These shortcomings not only lead to redundant consumption of testing resources but also prevent the system from outputting targeted screening recommendations, becoming a key bottleneck restricting the effectiveness of early diagnosis and treatment of digestive system tumors. Summary of the Invention
[0005] The purpose of this invention is to provide a digestive tumor risk assessment system and method based on multi-omics data, in order to solve the technical problems existing in digestive tumor risk assessment technology, such as reliance on a single data source, insufficient fusion of multi-dimensional information, inability to effectively distinguish the risk of specific anatomical sites, and lack of organ-specific analysis engine.
[0006] A gastrointestinal tumor risk assessment system based on multi-omics data includes: A multi-source heterogeneous data acquisition terminal is used to simultaneously acquire circulating tumor DNA methylation data, serum biochemical index data, and clinical phenotypic characteristic data of the subjects; A data standardization preprocessing module is connected to the multi-source heterogeneous data acquisition terminal and is used to normalize, denoise, and construct features for the multi-source heterogeneous data to generate a standardized tensor. A broad-spectrum digestive system risk assessment engine is connected to the data standardization preprocessing module. It is used to assess the overall risk score of the digestive system based on the standardized tensor and activate the targeted parsing program when the risk score exceeds the first safety threshold. An organ-specific feature decoupling module is connected to the broad-spectrum digestive system risk assessment engine, used to extract differential feature vectors of the esophagus, stomach and colon and rectum by comparing tissue-specific methylation maps; The multimodal cross-omics fusion inference module is connected to the organ-specific feature decoupling module. It is used to fuse the differential feature vector, serum biochemical index data and clinical phenotypic feature data based on the dynamic weight allocation model, and output a targeted risk probability matrix for a specific organ. The clinical decision support and early warning terminal is connected to the multimodal cross-omics fusion reasoning module and is used to generate multi-level risk warning signals and clinical action plans based on the directional risk probability matrix.
[0007] Preferably, the data standardization preprocessing performed by the data standardization preprocessing module includes: Sequencing quality control, methylation frequency calculation, and piecewise linear mapping normalization algorithm were performed on circulating tumor DNA methylation data to map the methylation frequency to the interval of 0 to 1. An outlier removal procedure based on the absolute deviation of the median was performed on serum biochemical data; Natural language processing techniques were used to extract structured features from clinical phenotypic data. The preprocessing module also introduces a generative adversarial network to complete missing data and constructs a multidimensional normalized tensor containing feature interaction terms.
[0008] Preferably, the broad-spectrum digestive system risk assessment engine integrates a deep residual network and a multilayer perceptron, and is trained by minimizing the cross-entropy loss function to extract common signals of malignant proliferation in the digestive system; the engine also introduces an attention mechanism to focus on key methylation sites, and uses an ensemble learning algorithm and a gradient boosting decision tree model to perform weighted fusion of risk scores.
[0009] Preferably, the organ-specific feature decoupling module utilizes pre-stored organ-specific methylation reference atlases of the esophagus, stomach, and colorectal region to perform tissue-specific methylation site mining through a tissue origination classifier; the classifier is based on a random forest architecture and uses a composite similarity evaluation index to calculate the similarity between the subject's feature vector and the reference vector, using the following formula:
[0010] in, For composite similarity scores, For cosine similarity, For Euclidean distance This is the adjustment coefficient.
[0011] Preferably, the multimodal cross-omics fusion reasoning module adopts a dynamic reasoning model based on evidence weight allocation, and the weight calculation formula is as follows:
[0012] in, Representing the Decision weights for each omics dimension This represents the mutual information value between this dimension's feature and the target risk variable. The confidence factor represents the data in this dimension; the multimodal cross-omics fusion reasoning module also constructs a Bayesian reasoning network, using clinical phenotypes as prior probabilities and molecular biological data and biochemical indicators as likelihood information to calculate the posterior organ-specific risk probability.
[0013] Preferably, the system further includes a dynamic risk monitoring module, which acquires time-series data through the historical data retrieval unit of the multi-source heterogeneous data acquisition terminal, and the data standardization preprocessing module performs smoothing processing using a weighted moving average method and introduces a background noise subtraction algorithm to remove non-tumor-specific methylation signals.
[0014] Preferably, the system further includes a large-scale concurrent processing module, a data standardization preprocessing module that uses a containerized elastic scaling architecture and a field-programmable gate array accelerator for sequence alignment, and a multimodal cross-omics fusion inference module that introduces Bayesian posterior probability solution details and is equipped with homomorphic encryption and blockchain traceability auditing to ensure data security.
[0015] Preferably, the system further includes in-depth analysis of clinical phenotypes, a multi-source heterogeneous data acquisition terminal equipped with a medical entity recognition engine to extract weighted risk factors from electronic medical records, and a multimodal cross-omics fusion reasoning module that adopts a multi-task learning framework to simultaneously predict tumor pathological subtyping and evolution speed.
[0016] Preferably, the system further includes a robust optimization module, a data standardization preprocessing module that uses unique molecular identifiers and adaptive smoothing algorithms to process low sequencing depth data, and a broad-spectrum digestive system risk assessment engine that integrates multi-scale convolutional neural networks to extract local and global methylation features.
[0017] Preferably, the system also includes a standardized workflow module, which includes an automated process for data acquisition, preprocessing, risk assessment, organ decoupling, fusion reasoning, and early warning; the clinical decision support and early warning terminal provides visualized risk profiles and closed-loop follow-up management functions, and is deeply integrated with the internal systems of medical institutions.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention fundamentally solves the problem of traditional assessment methods' strong dependence on a single data source by constructing a multi-source heterogeneous data acquisition terminal and a multimodal cross-omics fusion inference module. The system integrates multi-dimensional information such as circulating tumor DNA methylation, serum biochemistry, and clinical phenotypes, achieving in-depth cross-validation at the molecular, biochemical, and clinical levels. This multi-dimensional fusion mechanism significantly improves the sensitivity and specificity of risk assessment, enabling the earlier detection of occult digestive system lesion signals compared to traditional methods, and effectively reducing the false negative or false positive rate that may be caused by single tests.
[0019] 2. This invention introduces an organ-specific feature decoupling module, enabling precise analysis from broad-spectrum digestive system risk to risk at specific anatomical sites; utilizing tissue-specific methylation mapping technology, the system can effectively distinguish tumor signals from different organs such as the esophagus, stomach, and colorectal region; this technology overcomes the bottleneck of existing technologies that cannot locate the source of risk, providing examinees with a clear direction for medical treatment, avoiding blind screening of the entire digestive tract, greatly saving medical resources and reducing the burden of invasive examinations for examinees.
[0020] 3. This invention employs a dynamic reasoning model based on evidence weight allocation and a hierarchical assessment architecture, enabling the evaluation process to possess high logical rigor and adaptability. The system can automatically adjust the decision contribution of each module according to the quality and type of input data, ensuring that a stable and reliable risk probability matrix can be output under different data completeness levels. At the same time, the hierarchically designed decision support terminal can transform complex omics conclusions into intuitive clinical guidance suggestions, achieving a seamless connection from cutting-edge life science technologies to actual clinical applications. This has significant application value for building a closed-loop management system for the early diagnosis and treatment of digestive system tumors.
[0021] 4. This invention ensures the scalability and continuous evolution of the evaluation system through a digital and standardized data processing workflow and an incrementally learnable model design. As the sample size and pathological feedback data accumulate, the system can automatically optimize the feature extraction algorithm and decision threshold, so that its analysis accuracy for complex cases steadily improves with the expansion of the scale of use. This engineered architecture design provides efficient and practical technical support for the accurate screening of large-scale populations, and has significant social and economic benefits. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of the overall technical architecture of the digestive tumor risk assessment system based on multi-omics data proposed in this invention; Figure 2 This is a schematic diagram of the core principle framework of the multimodal cross-omics fusion inference module in this invention; Figure 3 This is a logical flowchart of the data standardization preprocessing and feature construction in this invention; Figure 4 This is a logical flowchart of the organ-specific feature decoupling module in this invention; Figure 5 This is a schematic diagram of the multi-level interaction relationship and data flow of the clinical decision support and early warning terminal in this invention. Detailed Implementation
[0023] Example 1 Please refer to the attached document. Figure 1 This embodiment provides a digestive tumor risk assessment system based on multi-omics data. Built on a high-performance computing cluster architecture, the system aims to achieve early and accurate screening of digestive system malignancies through deep integration of multidimensional biological information. The entire system consists of multiple interconnected hardware terminals and software logic modules, ensuring automation and standardization throughout the entire process from obtaining raw biological samples to generating final clinical decision recommendations.
[0024] The multi-source heterogeneous data acquisition terminal, acting as the system's sensing layer, is responsible for simultaneously acquiring circulating tumor DNA methylation data, serum biochemical indicators, and clinical phenotypic characteristics from the subjects. When acquiring circulating tumor DNA methylation data, the terminal first receives a plasma sample obtained through a peripheral blood collection procedure. This process involves collecting 10 ml of peripheral blood from the subject using an EDTA anticoagulant tube and separating the plasma via double centrifugation within 2 hours to minimize genomic DNA contamination from leukocyte lysis. Subsequently, the system uses an automated nucleic acid extractor to extract free DNA from the plasma using magnetic beads. For the extracted free DNA, the terminal performs bisulfite conversion, a chemical process that deaminoides unmethylated cytosine in the sequence, converting it to uracil, while leaving methylated cytosine unchanged, thus converting the chemical modification signal into detectable base sequence differences. After conversion, the system constructs a high-performance sequencing library and performs capture sequencing targeting specific genomic regions. The specific genomic region contains at least 128 methylation-sensitive sites that are highly associated with the occurrence of gastrointestinal tumors. These sites are distributed in key promoter and enhancer regions of organs such as the esophagus, stomach, and colorectal region.
[0025] In the serum biochemical data collection phase, a multi-source heterogeneous data acquisition terminal was connected to an automated biochemical analyzer to perform quantitative analysis on the collected serum samples from the subjects. The tests covered the concentrations of several tumor markers associated with the digestive system, specifically carcinoembryonic antigen (CEA), carbohydrate antigen 19-9, and pepsinogen. The pepsinogen assay was further refined to the concentrations of pepsinogen 1 and pepsinogen 2, and the ratio between the two was calculated. In addition, clinical phenotypic data were collected through a structured electronic questionnaire. The terminal automatically extracted the subjects' age, gender, dietary habits (including salt intake, smoking and alcohol consumption), family medical history (including cancer history in first-degree relatives), and past history of gastrointestinal diseases (such as atrophic gastritis, intestinal metaplasia, polyp removal records, etc.). All collected raw data were transmitted to the data center via an encrypted protocol.
[0026] Please refer to the attached document. Figure 3The data standardization and preprocessing module receives various raw information from multi-source heterogeneous data acquisition terminals and performs rigorous normalization and noise reduction. For circulating tumor DNA methylation data, this module first performs sequencing quality control, removing adapter sequences and low-quality reads. Subsequently, it eliminates signal bias caused by differences in library concentration between different samples by calibrating the sequencing depth. For each target methylation-sensitive site, the module calculates its methylation frequency, i.e., the proportion of methylated cytosine in the total coverage. To ensure comparability between different batches of data, the module uses a normalization algorithm based on piecewise linear mapping to map the frequency value of each methylation site to a continuous interval between 0 and 1.
[0027] For serum biochemical indicator data, the data standardization preprocessing module performs an outlier removal procedure based on the absolute deviation of the median. This procedure identifies and excludes extreme outliers caused by non-specific inflammation or laboratory random errors by calculating the median of the sample population and measuring the absolute distance of each observation from the median. The processed biochemical indicator data also undergoes dimensional transformation, converting it into a specific dimension in the standardized tensor. For clinical phenotypic feature data, this module applies natural language processing techniques to extract features from unstructured text, transforming qualitative descriptions into quantitative or Boolean-valued numerical features. Finally, all preprocessed features are concatenated and encapsulated into a multidimensional standardized tensor, which serves as the input to subsequent models.
[0028] Please refer to the attached document. Figure 1 The broad-spectrum digestive system risk assessment engine initiates the first phase of overall risk assessment based on the generated standardized tensor. This engine integrates a deep residual network structure, which, by introducing a skip connection mechanism, solves the gradient vanishing problem that easily occurs during the training of deep neural networks, enabling the model to effectively extract high-order nonlinear features from the data. In the pre-training phase, the global assessment model uses a large-scale multi-omics dataset containing tens of thousands of patients with digestive tract tumors and healthy controls. By minimizing the cross-entropy loss function, it learns the essential differences between healthy and diseased groups in methylation profiles, biochemical profiles, and clinical characteristics.
[0029] The broad-spectrum digestive system risk assessment engine, upon receiving a standardized tensor, extracts common malignant proliferation signals through collaborative computation between a multilayer perceptron and convolutional layers. These signals reflect the presence of early tumorigenesis trends in the digestive system as a whole. The engine ultimately calculates an overall digestive system cancer risk score for the examinee, set within a range of 0 to 100. The engine incorporates a first safety threshold, a risk threshold determined based on large-scale population epidemiological data. When the overall risk score exceeds this first safety threshold, the system determines that the examinee has a potential risk of digestive system lesions and activates subsequent targeted analysis procedures. If the score is below the threshold, the system outputs a low-risk conclusion and recommends regular follow-up.
[0030] Please refer to the attached document. Figure 4 The organ-specific feature decoupling module is activated when the overall risk score is abnormal. Its core task is to deeply mine tissue-specific methylation sites in circulating tumor DNA. This module uses pre-stored organ-specific methylation reference atlases of the esophagus, stomach, and colorectal region to perform a refined comparison of the subject's methylation map. These reference atlases are methylation feature libraries extracted from thousands of tissue samples diagnosed by gold standard pathology, covering the characteristic methylation patterns of each organ in both normal and cancerous states.
[0031] The tissue-specific methylation site mining process performed by the organ-specific feature decoupling module involves the following steps. First, the methylation frequency located in known organ-specific enhancer regions is retrieved from the subject data. Enhancers, as key cis-acting elements regulating gene expression, exhibit extremely high specificity across different tissues, and their methylation levels directly reflect the tissue identity of the source cells. Second, the module calculates the deviation between the frequency values of these specific regions in the subjects and a healthy tissue reference baseline. Finally, these deviation values are input into a tissue origination classifier. This tissue origination classifier employs a random forest-based architecture, capable of identifying the most likely source organ of abnormal methylation signals by integrating the voting results of multiple decision trees. The classifier outputs a tissue contribution index for each subject organ and transforms it into a differential feature vector characterizing damage to that specific organ.
[0032] Please refer to the attached document. Figure 2 The multimodal cross-omics fusion inference module receives differential feature vectors output by the organ-specific feature decoupling module and, in conjunction with serum biochemical index data and clinical phenotypic feature data, constructs an inference model based on evidence weight allocation. The core logic of this module lies in identifying the complementarity and consistency between data from different dimensions. To achieve this goal, the module first calculates the mutual information between different omics evidence, that is, measures the information gain of one set of data given the knowledge of another set of data.
[0033] The multimodal cross-omics fusion inference module employs an adaptive weight adjustment mechanism. When circulating tumor DNA signals exhibit high-confidence organ-specificity—that is, when the probability of methylation signals being traced back to the source is much higher than the random background—the system automatically increases the weight of genomic features to above 0.7. This processing logic ensures that highly specific molecular evidence dominates the final decision. Conversely, when molecular biological signals are weak due to low circulating tumor DNA concentrations, and serum markers or clinical phenotypic features show strong correlations, the system increases the weight proportions of these dimensions.
[0034] During evidence fusion, the system constructs a Bayesian inference network, using clinical phenotypes as prior probabilities and molecular biological data and biochemical indicators as likelihood information, thereby calculating the posterior organ-specific risk probability. The specific weight allocation calculation follows these principles:
[0035] in, Representing the Decision weights for each omics dimension This represents the mutual information value between this dimension's feature and the target risk variable. This represents the confidence factor for the data in this dimension. Through this dynamic adjustment, the module can output a targeted risk probability matrix for specific anatomical sites such as the esophagus, stomach, and colon / rectum. This matrix not only contains the cancer risk probability value for each organ but also the corresponding confidence interval.
[0036] Please refer to the attached document. Figure 5 The clinical decision support and early warning terminal generates multi-level risk warning signals based on the score distribution in the targeted risk probability matrix. This terminal is configured to output corresponding endoscopic examination recommendations or pathological biopsy guidance when the risk probability of a specific organ exceeds a second safety threshold. The second safety threshold is a differentiated threshold set for different organs. For example, for gastric risk, if the probability exceeds a preset 65%, the terminal will automatically display a list of abnormal indicators for the examinee, such as abnormal pepsinogen ratios and changes in specific methylation sites, and recommend performing gastroscopy and necessary biopsy sampling.
[0037] The clinical decision support and early warning terminal has a dynamic report generation function, classifying assessment results into three levels: low risk, medium risk, and high risk. For the low-risk level, the terminal generates annual follow-up recommendations and provides guidelines for gastrointestinal health management. For the medium-risk level, the terminal generates targeted non-invasive examination recommendations based on the specific location of the affected organs, such as suggesting a carbon-13 breath test or repeat serum tumor marker testing. For the high-risk level, the terminal immediately triggers an emergency alert, displaying a prominent reminder icon on the doctor's workstation interface and automatically generating a detailed risk assessment report, including trend charts of various omics indicators and suggested clinical reference pathways.
[0038] The system in this embodiment operates on a cloud server architecture with high-performance graphics processing units and large-scale random access memory. Modules communicate via a representational state transfer application programming interface (API) to ensure efficient data transmission. To protect privacy, the system is equipped with encrypted storage units, utilizing advanced encryption standards to desensitize and hierarchically store the examinee's identity information and biological data. Regarding model evolution, the system features real-time feedback. When an examinee obtains a pathological diagnosis in a subsequent clinical pathway, this result serves as a ground truth label and is fed back to the cloud. The broad-spectrum digestive system risk assessment engine and tissue origination classifier utilize this new data to perform online incremental learning, continuously optimizing feature extraction accuracy and risk prediction weights, thereby enabling the system to steadily improve accuracy as its usage scales.
[0039] Through the aforementioned complete technical architecture, this embodiment achieves a step-by-step assessment model from broad-spectrum digestive system screening to precise organ localization. The system can not only capture extremely weak circulating tumor DNA methylation signals, but also eliminate false positive results caused by benign inflammation or physiological fluctuations through cross-validation of multi-omics data. The introduction of the organ-specific feature decoupling module directly addresses the clinical pain point of "knowing the danger but not knowing where the danger lies" in existing technologies, providing scientific and technical support for the early detection, early diagnosis, and early treatment of digestive tumors.
[0040] Example 2 Building upon Example 1, Example 2 further optimizes the ability of the multi-source heterogeneous data acquisition terminal to capture dynamic trends and refines the degradation processing logic of the multimodal cross-omics fusion inference module in cases of incomplete data. This example is particularly suitable for long-term dynamic risk monitoring of high-risk populations, improving the predictive foresight by introducing time series analysis.
[0041] Please refer to the attached document. Figure 1In this embodiment, the multi-source heterogeneous data acquisition terminal adds a historical data retrieval unit. This unit supports the input of the examinee's physical examination indicator trends over the past 5 years, with a focus on the fluctuation trajectory of serum biochemical indicators. For example, the system not only collects the current carcinoembryonic antigen concentration but also automatically retrieves and compares the test records from the previous two years. By calculating the dynamic evolution rate of the indicators, the system can identify those weak abnormalities that, although not yet exceeding the clinical reference range, show a significant upward trend.
[0042] The data standardization preprocessing module employs a special processing logic design for time series data. When processing serum biochemical indicators from consecutive years, the module uses a weighted moving average method to smooth the data, eliminating random fluctuations. The smoothed values, along with their first derivatives (i.e., the rate of change), are incorporated into the standardization tensor. For circulating tumor DNA methylation data, if multiple detection records exist, the module focuses on comparing the cumulative increase in methylation frequency to assess the growth burden trend of tumor signals.
[0043] Please refer to the attached document. Figure 3 In this embodiment, the data standardization preprocessing module also introduces a stripping algorithm for non-specific background noise. In circulating tumor DNA sequencing data, background noise often originates from normal DNA fragments produced by leukocyte apoptosis or methylation changes caused by certain non-cancerous inflammations. The module establishes a large-scale methylation background database of non-cancerous digestive diseases (such as common gastritis, Crohn's disease, ulcerative colitis, etc.) and subtracts these common but non-tumor-specific methylation signals from the subject samples. This process greatly improves the recognition accuracy of the organ-specific feature decoupling module in subsequent stages.
[0044] In this embodiment, the broad-spectrum digestive system risk assessment engine employs a combination of ensemble learning algorithms and multilayer perceptrons. In addition to deep residual networks, the engine also incorporates a gradient-boosting decision tree model. These two models, with their distinct properties, independently predict the standardized tensor, and their outputs are weighted and fused by a meta-learner. This ensemble strategy significantly enhances the model's generalization ability to small-sample, high-dimensional omics data, ensuring stable predictive performance across subjects of different ages and ethnicities.
[0045] Please refer to the attached document. Figure 4 The organ-specific feature decoupling module further refines the methylation reference atlas during tissue tracing analysis. In this embodiment, the reference atlas not only distinguishes organs but also specific locations within organs, such as the upper, middle, and lower segments of the esophagus, and the antrum and body of the stomach. During the comparison process, the module employs a composite similarity evaluation index, the calculation logic of which is as follows:
[0046] in, Represents the composite similarity score. The cosine similarity between the subject's feature vector and the reference vector. This represents the Euclidean distance between the two. This is the adjustment coefficient. By combining angular similarity and distance similarity in the calculation, the system can more robustly identify the anatomical source of anomalous methylation signals.
[0047] Please refer to the attached document. Figure 2 In this embodiment, the multimodal cross-omics fusion inference module enhances the in-depth analysis of clinical phenotypic features. By applying advanced natural language processing models, the system can identify complex logical relationships from medical record descriptions. For example, if a medical record mentions "atrophic gastritis with moderate intestinal metaplasia," the system automatically parses it into two independent risk modulators and assigns corresponding prior probability increments. When the system detects missing data in a certain omics (e.g., the subject has not undergone methylation sequencing and only has biochemical indicators and clinical characteristics), the multimodal cross-omics fusion inference module automatically switches to a degraded inference mode. In this mode, the system derives the risk assessment conclusion with the highest probability by reallocating the weights of the remaining dimensions and combining them with the conditional probability distribution in a Bayesian network, and clearly marks the confidence level of the data in the output report.
[0048] Please refer to the attached document. Figure 5 In this embodiment, the clinical decision support and early warning terminal achieves deep integration with the internal systems of medical institutions. When a high-risk warning signal is output, the terminal not only displays endoscopic examination recommendations but also automatically queries the appointment schedules of endoscopy departments in tertiary hospitals near the examinee's geographical location. For the generated risk level report, the terminal supports digital signatures and encrypted sharing, facilitating examinees to directly present detailed omics assessment evidence to doctors during medical visits. In addition, the terminal also has a built-in deep interpretation engine for the abnormal indicator list, which can transform abstract methylation frequencies and biochemical concentration values into understandable health risk descriptions, greatly improving examinee compliance with examinations.
[0049] Example 2 further improves the system's performance in complex clinical scenarios by introducing time series analysis and background noise subtraction techniques. The system not only focuses on static risk scores but also on the evolutionary logic of risk over time, which has significant clinical value in identifying slow-growing but highly invasive early-stage gastrointestinal tumors. Through digital data flow and a closed-loop incremental learning mechanism, the assessment system described in this example can continuously absorb the latest clinical knowledge, constantly improving the accuracy and coverage of its risk warnings.
[0050] Example 3 Example 3 focuses on describing the concurrent processing mechanism and data security protection system of this system in a large-scale population screening scenario, and details the Bayesian posterior probability calculation details when processing multimodal cross-omics data. This example aims to demonstrate the technical support capabilities of this evaluation system as a city-level or national-level public health service platform.
[0051] Please refer to the attached document. Figure 1 In this embodiment, the multi-source heterogeneous data acquisition terminal adopts a distributed deployment mode. The hardware front-end of the terminal is widely distributed in physical examination centers and community health centers at all levels, while the core computing logic is concentrated on the cloud server. In order to support large-scale concurrent sample data processing, the data standardization preprocessing module adopts an elastic scaling architecture based on containerization technology. When the instantaneous peak of the methylation sequencing data stream to be processed exceeds the preset load, the system automatically triggers the horizontal scaling of computing nodes to ensure the real-time performance of data processing.
[0052] When processing circulating tumor DNA methylation data, the data standardization preprocessing module employs a highly efficient base recognition algorithm and a sequence alignment accelerator. This accelerator utilizes field-programmable gate array (FPGA) hardware to perform parallel alignment of sequencing reads, significantly shortening the conversion cycle from raw data to the methylation frequency matrix. Simultaneously, for serum biochemical index data generated by different laboratories, this module performs rigorous inter-laboratory standard normalization, eliminating systematic errors caused by differences in test kit brands and instrument models by introducing calibration parameters from international standard reference materials.
[0053] Please refer to the attached document. Figure 2 The multimodal cross-omics fusion inference module introduces a complex Bayesian inference network when performing targeted risk probability calculations. This network treats the subject's digestive system health status as a latent variable and the observed multi-omics data as manifest variables. The inference process first sets an initial prior probability distribution using large-scale population epidemiological data. For example, for men over 50 years of age with a history of smoking, the prior probability of esophageal and gastric cancer is adjusted accordingly.
[0054] With the integration of multi-source data, the system iteratively updates the probability distribution using Bayesian formulas. When processing differentiated feature vectors, the system utilizes a likelihood function to measure the probability of the current methylation pattern occurring in a specific organ cancer state. The parameters of the likelihood function are obtained through statistical modeling of tens of thousands of confirmed cases, reflecting the sensitivity and specificity of specific methylation sites in different cancer types. Finally, the system calculates the posterior organ-specific risk probability matrix and performs a confidence test on it. Only when the posterior probability distribution exhibits sufficient determinism will the system classify it as high-risk.
[0055] Please refer to the attached document. Figure 4In this embodiment, the organ-specific feature decoupling module adds a non-specific signal interception layer. This layer is specifically designed to identify non-tumor-specific methylation changes caused by the body's overall immune response, systemic metabolic abnormalities, or aging. By comparing the methylation characteristics of circulating immune cells in the blood, the system can accurately determine the source of background signals mixed in the blood. If the abnormal signal is found to originate primarily from immune cells rather than epithelial cells, the system will reduce the weight of the tissue origination classifier, thereby effectively avoiding misjudgments caused by systemic non-tumor inflammation.
[0056] Please refer to the attached document. Figure 5 In this embodiment, the clinical decision support and early warning terminal provides full-lifecycle health record management functions. In addition to real-time risk assessment, the terminal can also develop personalized prevention and intervention strategies for examinees based on the evolution of multi-omics data. For example, for examinees identified as having medium-risk gastric conditions, the terminal not only recommends endoscopic screening but also automatically generates a low-salt, vitamin-rich dietary intervention plan based on their dietary preferences in their clinical phenotype, and regularly pushes reminders via mobile devices.
[0057] The system and methods operate on a cloud server that complies with the medical information security level protection standards. To ensure absolute data security, the system employs homomorphic encryption technology, enabling the multimodal cross-omics fusion inference module to perform matrix operations without decrypting the original data. The subjects' personal privacy information and bioomics data are physically isolated at the storage layer. All cross-module data transmissions undergo traceability auditing based on blockchain technology, ensuring that the process of arriving at each risk assessment conclusion is traceable and verifiable.
[0058] Example 3 demonstrates the stability and reliability of a digestive tumor risk assessment system based on multi-omics data in a large-scale application environment through engineered architecture optimization and rigorous probabilistic statistical modeling. By deeply integrating molecular and clinical approaches, the system not only improves the accuracy of individual diagnostics but also provides a low-cost, high-efficiency technical tool for precise prevention at the population level.
[0059] Example 4 This embodiment further elaborates on the system's structured extraction logic when processing complex clinical phenotypic feature data, as well as the specific details of the multi-source heterogeneous data acquisition terminal performing stratified scanning of various digestive tumor markers. This embodiment demonstrates the system's processing capabilities when dealing with unstructured medical big data.
[0060] Please refer to the attached document. Figure 1The multi-source heterogeneous data acquisition terminal is equipped with a deep learning-based medical entity recognition engine when collecting clinical phenotypic feature data. This engine can automatically scan the chief complaint, present illness history, and physical examination records in electronic medical records to identify entity words highly related to digestive system risks. For example, when descriptions such as "change in stool characteristics," "dull pain in the upper abdomen," or "foreign body sensation in the esophagus" appear in the medical record, the engine not only identifies the symptoms themselves but also combines them with modifiers such as duration and frequency to transform them into weighted risk factors. This structured extraction process transforms linguistic information that is originally difficult for computers to process directly into effective features in a multidimensional tensor.
[0061] During the collection of serum biochemical indicators, the terminal performed a more detailed stratified scan. For gastric tumor risk, the system focused on detecting the concentrations of pepsinogen 1 and pepsinogen 2, combined with Helicobacter pylori antibody detection results. For hepatobiliary and pancreatic digestive system risks, the system simultaneously acquired total bilirubin, direct bilirubin, and specific carbohydrate antigen indicators. For colorectal risks, the system integrated carcinoembryonic antigen and fecal occult blood immunoassay data. This stratified scanning strategy allows the multi-source heterogeneous data acquisition terminal to dynamically adjust the detection focus based on the subject's initial presentation.
[0062] Please refer to the attached document. Figure 3 In this embodiment, the data standardization preprocessing module introduces a missing data completion algorithm based on generative adversarial networks. In actual clinical applications, subjects may have missing indicators for various reasons. Traditional elimination processes lead to a significant waste of useful information. This module utilizes a pre-trained generative model to perform probabilistic prediction and completion of missing dimensions based on the known data distribution patterns. The completed data is then labeled using a specific confidence mask to ensure differentiated treatment in subsequent inference stages.
[0063] Please refer to the attached document. Figure 1 In this embodiment, the broad-spectrum digestive system risk assessment engine incorporates an attention mechanism. When extracting features from the standardized tensor, the attention mechanism automatically learns and focuses on key feature sites that have the greatest impact on overall risk. For example, in thousands of methylation frequency data points, small changes at certain sites may have higher diagnostic value than large fluctuations at other sites. By assigning higher attention weights to these key sites, the engine significantly improves the sensitivity of capturing very early-stage tumor signals.
[0064] Please refer to the attached document. Figure 2In this embodiment, the multimodal cross-omics fusion inference module employs a multi-task learning framework. This framework not only outputs a directional risk probability matrix but also simultaneously predicts the potential pathological subtype and evolution rate of the tumor. By sharing and learning common features across different tasks, the system enhances its understanding of complex biological laws. When calculating the directional risk probability of each digestive organ, the system constructs a complex nonlinear mapping function to transform multi-dimensional mutual information into a final probability distribution. This calculation process involves manifold expansion of a high-dimensional feature space, ensuring the logical consistency of evidence from different sources during the fusion process.
[0065] Please refer to the attached document. Figure 4 In this embodiment, the organ-specific feature decoupling module adds logic for excluding signals from non-digestive system organs. In actual screening, certain lung or urinary system tumors may also release circulating tumor DNA into the bloodstream. To avoid false positives for digestive system tumors, this module performs secondary filtering on detected signals by comparing them to specific methylation reference atlases of lung and kidney tissues. Only when the probability of a signal originating in a digestive system-related organ is significantly higher than in other unrelated organs will it proceed to the subsequent fusion inference stage.
[0066] Please refer to the attached document. Figure 5 In this embodiment, the clinical decision support and early warning terminal provides closed-loop follow-up management functionality. The system can automatically trigger personalized follow-up notifications based on the patient's risk trends. For medium- and high-risk patients, the terminal records whether they have undergone the recommended endoscopic examination and feeds the results back to the cloud in real time. This closed-loop management based on real-world data not only directly benefits the patients but also provides valuable training material for the continuous iteration of the system's underlying model.
[0067] The technical solution described in this embodiment further expands the application boundaries of the system by deeply mining clinical phenotypic data and introducing a multi-task learning framework. The system is not only a risk assessment tool, but also an intelligent decision-making center connecting the front end of molecular biology with the back end of clinical diagnosis and treatment, building a solid digital foundation for the precision prevention and control system of digestive system tumors.
[0068] Example 5 This embodiment details the robustness optimization measures of the system when dealing with different sequencing depths and sample qualities, and elaborates on the key technical details of the data standardization preprocessing module in multi-omics feature extraction.
[0069] Please refer to the attached document. Figure 1The multi-source heterogeneous data acquisition terminal employs an enhanced library construction scheme to address the trace characteristics of circulating tumor DNA. At the sequencing front end, each raw DNA molecule is labeled with a unique molecular identifier. In subsequent data processing, the data normalization preprocessing module accurately distinguishes between genuine biological mutations or methylation modifications and random errors introduced by the sequencer by comparing sequence reads with the same molecular identifier. This technique allows the system to maintain an extremely high signal-to-noise ratio even at relatively low sequencing depths.
[0070] The data normalization preprocessing module employs an adaptive smoothing algorithm when handling methylation frequencies. This algorithm dynamically adjusts the size of the smoothing window based on the sequencing coverage depth of each site. For regions with higher coverage depth, the algorithm retains more original fluctuation details; for regions with lower coverage depth, it reduces random variance by fusing signals from adjacent sites. The processed methylation features exhibit extremely high stability, providing high-quality input for the broad-spectrum digestive system risk assessment engine.
[0071] Please refer to the attached document. Figure 3 In this embodiment, the standardized tensor construction process not only includes basic feature values but also introduces interaction terms between features. For example, the module automatically calculates the product terms between certain specific methylation sites and serum tumor marker concentrations. These interaction features capture the synergistic effects between different omics dimensions, reflecting deep biological associations that are difficult to characterize with a single indicator.
[0072] Please refer to the attached document. Figure 1 In this embodiment, the broad-spectrum digestive system risk assessment engine integrates a multi-scale convolutional neural network structure. This structure can simultaneously extract methylation correlation features of local neighboring sites and global methylation distribution pattern features. The global evaluation model achieves high-dimensional extraction of common tumor features of the digestive system through joint optimization of millions of parameters. During training, the system employs sample resampling technology to address the common imbalance between healthy and diseased samples in medical data, ensuring that the model has sufficient sensitivity to low-frequency, high-risk events.
[0073] Please refer to the attached document. Figure 4 In this embodiment, the organ-specific feature decoupling module introduces a similarity-based tracing logic. After the system acquires the subject's differentiated feature vector, it not only compares it to the overall reference atlas but also calculates its similarity to specific organ subtype atlases (such as early-stage gastric cancer and advanced-stage gastric cancer atlases). Through this multi-level comparison, the module can output more detailed diagnostic clues, providing clinicians with a reference for determining the stage of the disease.
[0074] Please refer to the attached document. Figure 2In this embodiment, the multimodal cross-omics fusion inference module enhances the removal of clinical confounding factors. For example, when assessing colorectal risk, the system automatically identifies "recent bowel preparation records" or "acute enteritis attack records" among clinical phenotypic features. These confounding factors temporarily reduce the weight of serum biochemical indicators and increase the decision weight of methylation features. Through this intelligent context-aware mechanism, the system demonstrates strong adaptability in complex real-world clinical environments.
[0075] Please refer to the attached document. Figure 5 In this embodiment, the clinical decision support and early warning terminal provides a highly visualized risk profile. The terminal distributes the examinee's various indicators on a multi-dimensional radar chart, intuitively displaying the degree of risk deviation in molecular, biochemical, and clinical dimensions. For examinees identified as high-risk, the terminal will also automatically retrieve and push relevant pathophysiological mechanism explanations to help examinees understand the scientific basis behind the risk conclusion.
[0076] This embodiment significantly improves the system's tolerance to fluctuations in real-world sample quality through meticulous optimization of the signal processing layer. Whether in high-throughput laboratory environments or initial sampling scenarios in primary healthcare institutions, the system can output high-confidence risk assessment results, demonstrating its extremely high engineering practical value.
[0077] Example 6 This embodiment focuses on describing the specific steps of the evaluation method proposed in this invention during actual operation, as well as the data interaction protocols between each step. This embodiment demonstrates the rigorous logic of the system as a standardized process.
[0078] Step 1: The system activates the multi-source heterogeneous data acquisition terminal to begin the data integration process for the subjects. At this time, the peripheral blood samples have completed bisulfite conversion and library construction, and the raw sequences generated by the sequencing process are transmitted to the preprocessing module via a high-speed fiber optic channel. Simultaneously, the serum biomarker concentrations and clinical questionnaire information of the subjects are digitized and entered into the system.
[0079] Step 2: The data standardization and preprocessing module is activated. The system first performs base quality scoring and filtering on the sequencing data to ensure that each methylation site reading used in the calculation has an accuracy of over 99.9%. Next, the module performs linear normalization, mapping the original concentration and frequency values to a standard computational tensor space. During this process, the module automatically detects data integrity; if key dimensions are missing, an alarm is triggered, and the terminal is guided to perform data supplementation or initiate predictive completion.
[0080] Step 3: The broad-spectrum digestive system risk assessment engine loads a global evaluation model. Standardized tensors are fed into a pre-defined deep residual network, undergoing dozens of nonlinear transformations to extract common signals characterizing malignant proliferation in the digestive system. The risk score calculated by the engine is written into the examinee's temporary file in real time.
[0081] Step 4: The system performs logical judgment. If the risk score is lower than the first safety threshold, the system directly jumps to the result output stage. If the score exceeds the threshold, the system automatically triggers the next stage of organ-specific analysis.
[0082] Step 5 activates the organ-specific feature decoupling module. The system extracts tissue-specific features from a massive database of methylation sites and performs high-dimensional comparisons with a pre-stored organ atlas library. This step generates differential feature vectors that clearly point to the potential damaged organ.
[0083] Step 6: The multimodal cross-omics fusion inference module receives feature information from all dimensions. The system uses a Bayesian inference model for evidence fusion, dynamically adjusting the contribution weights of each omics. The system calculates the final targeted risk probability according to the following formula:
[0084] in, Represents the final probability of targeted risk. The dynamic weights for the three types of omics data (methylation, biochemistry, and clinical). This serves as the output function for the corresponding sub-model. Through this combination of linear and nonlinear reasoning, the system generates the final directional risk probability matrix.
[0085] Step 7: The clinical decision support and early warning terminal matches risk response rules. Based on the maximum value in the probability matrix and its corresponding organ category, the system retrieves the optimal treatment recommendation from a pre-set clinical pathway database. Finally, the system generates a comprehensive report including quantitative risk level, indicator anomaly analysis, and clinical action plan, and pushes it to the examinee and attending physician.
[0086] The method described in this embodiment ensures that every step of the evaluation process is traceable and highly automated. By solidifying complex bioinformatics processing logic into standardized computational steps, this invention significantly lowers the technical threshold for early screening of digestive tumors and provides a standardized implementation paradigm for accurate screening of large-scale populations.
[0087] In summary, the digestive tumor risk assessment system and method based on multi-omics data proposed in this invention achieve high sensitivity and specificity in assessing digestive system malignancies through deep integration of multi-source data, precise decoupling of organ-specific data, and adaptive adjustment of multimodal weights. This system not only overcomes the limitations of single data sources at the technical level but also demonstrates strong robustness and scalability at the engineering implementation level. With the continuous accumulation of data and iterative evolution of the model, this invention will undoubtedly play a crucial role in the early diagnosis and treatment of digestive tumors, generating profound social and economic benefits.
Claims
1. A digestive tumor risk assessment system based on multi-omics data, characterized in that, include: A multi-source heterogeneous data acquisition terminal is used to simultaneously acquire circulating tumor DNA methylation data, serum biochemical index data, and clinical phenotypic characteristic data of the subjects; A data standardization preprocessing module is connected to the multi-source heterogeneous data acquisition terminal and is used to normalize, denoise, and construct features for the multi-source heterogeneous data to generate a standardized tensor. A broad-spectrum digestive system risk assessment engine is connected to the data standardization preprocessing module. It is used to assess the overall risk score of the digestive system based on the standardized tensor and activate the targeted parsing program when the risk score exceeds the first safety threshold. An organ-specific feature decoupling module is connected to the broad-spectrum digestive system risk assessment engine, used to extract differential feature vectors of the esophagus, stomach and colon and rectum by comparing tissue-specific methylation maps; The multimodal cross-omics fusion inference module is connected to the organ-specific feature decoupling module. It is used to fuse the differential feature vector, serum biochemical index data and clinical phenotypic feature data based on the dynamic weight allocation model, and output a targeted risk probability matrix for a specific organ. The clinical decision support and early warning terminal is connected to the multimodal cross-omics fusion reasoning module and is used to generate multi-level risk warning signals and clinical action plans based on the directional risk probability matrix.
2. The digestive tumor risk assessment system based on multi-omics data according to claim 1, characterized in that, The data standardization preprocessing module performs the following data standardization preprocessing: Sequencing quality control, methylation frequency calculation, and piecewise linear mapping normalization algorithm were performed on circulating tumor DNA methylation data to map the methylation frequency to the interval of 0 to 1. An outlier removal procedure based on the absolute deviation of the median was performed on serum biochemical data; Natural language processing techniques were used to extract structured features from clinical phenotypic data. The preprocessing module also introduces a generative adversarial network to complete missing data and constructs a multidimensional normalized tensor containing feature interaction terms.
3. The digestive tumor risk assessment system based on multi-omics data according to claim 1, characterized in that, The broad-spectrum digestive system risk assessment engine integrates a deep residual network and a multilayer perceptron. It is trained by minimizing the cross-entropy loss function to extract common signals of malignant proliferation in the digestive system. The engine also introduces an attention mechanism to focus on key methylation sites and uses an ensemble learning algorithm and a gradient boosting decision tree model to perform weighted fusion of risk scores.
4. The digestive tumor risk assessment system based on multi-omics data according to claim 1, characterized in that, The organ-specific feature decoupling module utilizes pre-stored organ-specific methylation reference atlases for the esophagus, stomach, and colorectal region to perform tissue-specific methylation site mining through a tissue origination classifier. The classifier is based on a random forest architecture and uses a composite similarity evaluation index to calculate the similarity between the subject's feature vector and the reference vector, using the following formula: , in, For composite similarity scores, For cosine similarity, For Euclidean distance This is the adjustment coefficient.
5. A digestive tumor risk assessment system based on multi-omics data according to claim 1, characterized in that, The multimodal cross-omics fusion reasoning module adopts a dynamic reasoning model based on evidence weight allocation, and the weight calculation formula is as follows: , in, Representing the Decision weights for each omics dimension This represents the mutual information value between this dimension's feature and the target risk variable. The confidence factor represents the data in this dimension; the multimodal cross-omics fusion reasoning module also constructs a Bayesian reasoning network, using clinical phenotypes as prior probabilities and molecular biological data and biochemical indicators as likelihood information to calculate the posterior organ-specific risk probability.
6. The digestive tumor risk assessment system based on multi-omics data according to claim 1, characterized in that, The system also includes a dynamic risk monitoring module, which acquires time-series data through the historical data retrieval unit of the multi-source heterogeneous data acquisition terminal. The data standardization preprocessing module uses a weighted moving average method for smoothing and introduces a background noise subtraction algorithm to remove non-tumor-specific methylation signals.
7. A digestive tumor risk assessment system based on multi-omics data according to claim 1, characterized in that, The system also includes a large-scale concurrent processing module, a data standardization preprocessing module that uses a containerized elastic scaling architecture and a field-programmable gate array accelerator for sequence alignment, and a multimodal cross-omics fusion inference module that introduces Bayesian posterior probability calculation details and is equipped with homomorphic encryption and blockchain traceability auditing to ensure data security.
8. A digestive tumor risk assessment system based on multi-omics data according to claim 1, characterized in that, The system also includes in-depth analysis of clinical phenotypes, a multi-source heterogeneous data acquisition terminal equipped with a medical entity recognition engine to extract weighted risk factors from electronic medical records, and a multimodal cross-omics fusion reasoning module that adopts a multi-task learning framework to simultaneously predict tumor pathological subtyping and evolution speed.
9. A digestive tumor risk assessment system based on multi-omics data according to claim 1, characterized in that, The system also includes a robust optimization module, a data standardization preprocessing module that uses unique molecular identifiers and adaptive smoothing algorithms to process low sequencing depth data, and a broad-spectrum digestive system risk assessment engine that integrates multi-scale convolutional neural networks to extract local and global methylation features.
10. A digestive tumor risk assessment system based on multi-omics data according to claim 1, characterized in that, The system also includes a standardized workflow module, which includes an automated process for data acquisition, preprocessing, risk assessment, organ decoupling, fusion reasoning, and early warning. The clinical decision support and early warning terminal provides visualized risk profiles and closed-loop follow-up management functions, and is deeply integrated with the internal systems of medical institutions.