Gastric cancer risk assessment method and system based on peripheral blood multi-omics dynamic analysis
By using peripheral blood multi-omics dynamic analysis, we can detect the inflection point of gastric cancer risk evolution and recommend personalized follow-up intervals. This solves the problem that existing models cannot identify the nonlinear evolution process of gastric cancer and improves the accuracy of early gastric cancer identification and risk stratification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG PROVINCIAL HOSPITAL AFFILIATED TO SHANDONG FIRST MEDICAL UNIVERSITY (SHANDONG PROVINCIAL HOSPITAL)
- Filing Date
- 2026-06-10
- Publication Date
- 2026-07-10
AI Technical Summary
Existing gastric cancer risk assessment models are mostly cross-sectional static models, which fail to effectively capture the dynamic changes of biomarkers, cannot identify key signals in the nonlinear evolution process from precancerous lesions to early gastric cancer, and lack personalized follow-up recommendations, leading to missed diagnoses and over-testing.
By acquiring peripheral blood multi-omics data from the same subject at at least four consecutive time points, numerical difference method was used to detect the rate of change and acceleration inflection point of biomarkers. Combined with gradient boosting tree model and deep Q network reinforcement learning, personalized follow-up intervals and dynamic risk scores were output.
It improves the sensitivity and specificity of early gastric cancer identification and risk stratification, reduces the risk of missed diagnosis, optimizes the utilization of medical resources, and enables personalized follow-up recommendations.
Smart Images

Figure CN122369958A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical and health information technology, specifically to a method and system for assessing gastric cancer risk based on dynamic analysis of peripheral blood multi-omics. Background Technology
[0002] Gastric cancer is a leading cause of cancer-related morbidity and mortality worldwide. Early diagnosis and accurate risk stratification are crucial for improving patient prognosis. In recent years, peripheral blood-based multi-omics biomarker detection technologies have rapidly developed. By simultaneously detecting protein biomarkers, methylated DNA biomarkers, and metabolic biomarkers, and combining them with machine learning algorithms to construct gastric cancer risk assessment models, these technologies have shown initial clinical value surpassing traditional single-biomarker screening. For example, studies have used exosomal miRNA combined with clinical factors to construct diagnostic models with an AUC exceeding 0.95; machine learning models based on plasma cfDNA multi-omics characteristics are also being used for early gastric cancer detection.
[0003] However, existing technologies generally have the following limitations: First, most existing gastric cancer risk assessment models are cross-sectional static models, which only use multi-omics test values from a single blood sample to output a risk score, ignoring the dynamic changes in biomarker levels during the evolution of gastric cancer from precancerous lesions (such as atrophic gastritis and intestinal metaplasia) to early gastric cancer. In fact, the temporal characteristics of biomarkers, such as the rate and acceleration of change over time, are often more sensitive predictors of malignant transformation than single absolute values.
[0004] Second, although a few studies have attempted to use longitudinal follow-up data for dynamic analysis, their methods are still limited to simple linear fitting of the rate of change, and cannot identify the acceleration abrupt change point (inflection point) in the process of biomarker change. In clinical practice, the transformation of precancerous lesions into gastric cancer often manifests as a non-linear process of "slow change followed by accelerated deterioration." The appearance of the inflection point is a key signal of a qualitative change in the nature of risk, and current technology lacks automatic detection and quantitative assessment methods for this inflection point.
[0005] Third, existing models typically use fixed thresholds to define follow-up recommendations after outputting risk scores (e.g., 3 months for high risk, 6 months for medium risk, and 12 months for low risk), failing to adaptively calculate the optimal follow-up interval based on individual dynamic risk evolution characteristics. This leads to under-monitoring of some rapidly deteriorating patients and over-testing of some long-term stable patients, reducing the accuracy of clinical decision-making and the efficiency of resource utilization.
[0006] Therefore, there is an urgent need for a comprehensive assessment method that can automatically detect the inflection point of gastric cancer risk evolution based on peripheral blood multi-omics longitudinal data and recommend personalized follow-up intervals accordingly. Summary of the Invention
[0007] To address the aforementioned technical problems, this invention provides a method and system for assessing gastric cancer risk based on dynamic multi-omics analysis of peripheral blood.
[0008] The technical solution adopted in this invention is as follows: A gastric cancer risk assessment method based on peripheral blood multi-omics dynamic analysis includes the following steps: S1. Obtain peripheral blood multi-omics detection datasets from the same subject at at least four consecutive time points. The multi-omics detection datasets include protein biomarkers, methylated DNA biomarkers, and metabolic biomarkers at each time point. The sampling intervals between the time points may be unequal. S2. For each biomarker in the multi-omics detection dataset, extract the detection value of the biomarker at each time point and the corresponding timestamp to construct time series data. Use the numerical difference method to calculate the first derivative and the second derivative of the time series data in sequence to obtain the rate of change data and acceleration data of each biomarker at each time point. The time point when the absolute value of the second derivative changes from below the first threshold to above the second threshold is determined as the acceleration change inflection point position of the biomarker. Inflection point flags are generated for each biomarker. The inflection point flag of the biomarker that is detected is assigned a value of 1, and the inflection point flag of the biomarker that is not detected is assigned a value of 0. S3. For each marker with an inflection point value of 1, extract all detection values after the inflection point from the time series data of the marker, use linear regression to fit the rate of change and acceleration after the inflection point, and subtract the average rate of change before the inflection point from the rate of change after the inflection point to obtain the difference in the rate of change before and after the inflection point. S4. Input the current detection value, inflection point marker, rate of change after the inflection point, and acceleration after the inflection point of each marker into the pre-trained risk assessment model, and output a dynamic risk score. S5. Input the dynamic risk score, the inflection point flag of at least one marker, and the maximum value of the rate of change after the inflection point among all markers with inflection point flags set to 1 into the reinforcement learning agent, and output the recommended next follow-up interval. S6. Generate and output a comprehensive evaluation report that includes the dynamic risk score, the inflection point detection result list, and the recommended next follow-up interval.
[0009] The average rate of change before the inflection point in S3 is calculated as follows: calculate the rate of change data at each time point from the starting time point to the inflection point of the time series data, sum the rate of change data at each time point, and divide by the number of time points to obtain the average rate of change before the inflection point.
[0010] The risk assessment model in S4 is an ensemble learning model based on gradient boosting trees. The training samples of this model include longitudinal multi-omics data of each subject in the historical cohort and the corresponding gastric cancer outcome labels.
[0011] The reinforcement learning agent in S5 adopts a deep Q-network structure. Its state space includes dynamic risk score, inflection point marker and maximum rate of change after the inflection point, action space is a preset discrete follow-up interval value, and reward function is set as a weighted sum of negative missed diagnosis risk loss and detection cost.
[0012] The decision rule for the next follow-up interval recommended in S5 is as follows: When an inflection point exists and the maximum rate of change after the inflection point is greater than 0.1, output an interval of 30 to 90 days; when an inflection point exists and the maximum rate of change after the inflection point is not greater than 0.1, output an interval of 90 to 180 days; when no inflection point exists and the dynamic risk score is lower than the preset risk threshold, output an interval of 180 to 365 days; when no inflection point exists and the dynamic risk score is not lower than the preset risk threshold, output an interval of 90 to 180 days.
[0013] The protein biomarker group includes at least one selected from OLFM4, ECM1, GSN, IGFBP2, and IQGAP1; the methylated DNA biomarker group includes at least one selected from ZNF154, CDO1, and ELMO1; and the metabolic biomarker group includes at least one selected from adenosine, fumarate, tryptophan, and acylcarnitine.
[0014] The inflection point detection result list in S6 includes the name of each inflection point marker with a value of 1, the location of the inflection point in acceleration change, the rate of change after the inflection point, and the acceleration after the inflection point.
[0015] It also includes S7: when the recommended next follow-up interval is less than 90 days, an emergency warning sign is added to the comprehensive assessment report.
[0016] A gastric cancer risk assessment system based on peripheral blood multi-omics dynamic analysis includes: The longitudinal multi-omics data acquisition module is used to acquire peripheral blood multi-omics detection datasets of the same subject at at least four consecutive time points. The multi-omics detection datasets include protein biomarkers, methylated DNA biomarkers, and metabolic biomarkers at each time point, and the sampling intervals between the time points may be unequal. The acceleration inflection point detection and marker generation module is used to extract the detection value and corresponding timestamp of each marker in the multi-omics detection dataset to construct time series data. The first and second derivatives of the time series data are calculated sequentially using the numerical difference method to obtain the rate of change data and acceleration data of each marker at each time point. The time point when the absolute value of the second derivative changes from below the first threshold to above the second threshold is determined as the acceleration change inflection point position of the marker. An inflection point marker is generated for each marker, wherein the inflection point marker of the marker that is detected is assigned a value of 1, and the inflection point marker of the marker that is not detected is assigned a value of 0. The inflection point post-trend feature extraction module is used to extract all detection values after the inflection point time point from the time series data of each marker for which the inflection point marker is assigned a value of 1. Linear regression fitting is used to obtain the rate of change and acceleration after the inflection point, and the difference in the rate of change before and after the inflection point is obtained by subtracting the average rate of change before the inflection point from the rate of change after the inflection point. The dynamic risk score calculation module is used to input the current detection value, inflection point marker, rate of change after the inflection point, and acceleration after the inflection point of each marker into the pre-trained risk assessment model and output a dynamic risk score. The personalized follow-up interval recommendation module is used to input the dynamic risk score, the existence of an inflection point flag of at least one marker with a value of 1, and the maximum value of the rate of change after the inflection point among all markers with an inflection point flag of 1 into the reinforcement learning agent, and output the recommended next follow-up interval. The comprehensive assessment report generation module is used to generate and output a comprehensive assessment report that includes the dynamic risk score, the inflection point detection result list, and the recommended next follow-up interval.
[0017] The beneficial effects of this invention are: This invention collects peripheral blood protein, methylated DNA, and metabolic multi-omics data from at least four consecutive time points from the same subject. It incorporates the inflection point of the biomarker's change acceleration as a key signal of qualitative change in gastric cancer risk into the assessment system. A gradient boosting tree model is used to integrate current detection values, inflection point markers, post-inflection rate of change, and acceleration, improving the sensitivity and specificity of early gastric cancer identification and risk stratification. This addresses the challenge of traditional cross-sectional static models failing to capture the nonlinear evolution of precancerous lesions into gastric cancer. Furthermore, a deep Q-network reinforcement learning agent is introduced to adaptively output personalized follow-up intervals based on dynamic risk scores, the existence of global inflection points, and the maximum post-inflection rate of change. Four-category clinical decision rules ensure that the recommended results meet clinical requirements, effectively reducing the risk of missed diagnoses in high-risk patients while minimizing over-testing in low-risk populations and optimizing the utilization of medical resources. Attached Figure Description
[0018] Figure 1 This is a flowchart of a gastric cancer risk assessment method based on peripheral blood multi-omics dynamic analysis, according to an embodiment of the present invention. Figure 2 This is an overall flowchart of a gastric cancer risk assessment method based on peripheral blood multi-omics dynamic analysis according to an embodiment of the present invention; Figure 3 This is a block diagram of a gastric cancer risk assessment system based on peripheral blood multi-omics dynamic analysis, according to an embodiment of the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] like Figures 1-2 As shown, the gastric cancer risk assessment method based on peripheral blood multi-omics dynamic analysis according to an embodiment of the present invention includes the following steps: S1. Obtain peripheral blood multi-omics detection datasets from the same subject at at least four consecutive time points. The multi-omics detection datasets include protein biomarkers, methylated DNA biomarkers, and metabolic biomarkers at each time point. The sampling intervals between time points may be unequal.
[0021] The protein biomarker group includes at least one selected from OLFM4, ECM1, GSN, IGFBP2, and IQGAP1; the methylated DNA biomarker group includes at least one selected from ZNF154, CDO1, and ELMO1; and the metabolic biomarker group includes at least one selected from adenosine, fumarate, tryptophan, and acylcarnitine.
[0022] In this embodiment of the invention, the sampling time points are at least four consecutive time points. This quantity and attribute requirement are designed based on the necessity of mathematical calculation and the clinical evolution characteristics of gastric cancer. From a mathematical calculation perspective, this invention subsequently needs to calculate the first derivative (rate of change) and second derivative (acceleration of change) of the biomarker using the non-equidistant finite forward difference method. The calculation of the first derivative requires the detection values of two adjacent time points, and the calculation of the second derivative requires two adjacent first derivative values. If the number of sampling time points is less than 4, the number of second derivative values obtained will be ≤1, which is insufficient to complete the abrupt trend judgment of the "absolute value of the second derivative changing from below the first threshold to above the second threshold" required for the inflection point of acceleration change. Therefore, the number of sampling time points must satisfy n≥4, which is the minimum sample size requirement for subsequent inflection point detection. From the perspective of clinical testing attributes, it is required to sample at continuous time points for the same subject. The same subject ensures the individual uniqueness of the test data and eliminates the bias in biomarker detection values caused by individual differences such as age, gender, underlying diseases, and genetic background between different subjects. This ensures that the extracted time-series change features only reflect the evolution of the subject's own gastric cancer risk. Continuous time points refer to the continuous follow-up sampling of the subject, rather than unrelated scattered sampling. This can capture the continuous dynamic process of biomarker changes from steady changes to accelerated changes, which is consistent with the non-linear clinical characteristics of gastric cancer evolution from precancerous lesions to early gastric cancer.
[0023] This invention requires that the multi-omics detection dataset must simultaneously include proteomic biomarkers, methylated DNA biomarkers, and metabolic biomarkers, and at least one biomarker from each group must be selected from the candidate set defined in this invention. This design is based on the multi-molecular evolutionary patterns of gastric cancer development and is also an advantage of multi-omics detection compared to single-biomarker detection. It should be noted that the occurrence of gastric cancer is the result of multi-level regulation by the genome, epigenome, proteome, and metabolome. A single biomarker can only reflect changes at a certain molecular level, while the multi-omics fusion of these three types of biomarkers can achieve a comprehensive and multi-dimensional characterization of gastric cancer risk, significantly improving the sensitivity and specificity of risk assessment.
[0024] The protein biomarker group can be selected from OLFM4, ECM1, GSN, IGFBP2, and IQGAP1. These are core molecules whose protein expression levels change abnormally during the development and progression of gastric cancer. Their expression changes are closely related to abnormal proliferation of gastric mucosal epithelial cells and precancerous lesions. For example, OLFM4 is a biomarker for gastric mucosal epithelial cell proliferation, and its high expression is closely related to gastric mucosal intestinal metaplasia and dysplasia. IGFBP2 participates in the proliferation and invasion of gastric cancer cells and is an important protein signal for precancerous lesions of gastric cancer. The methylated DNA biomarker group can be selected from ZNF154, CDO1, and ELMO1. The occurrence of gastric cancer is often accompanied by the methylation silencing of tumor suppressor genes. This type of biomarker is a classic epigenetic biomarker in the field of gastric cancer. Its elevated methylation level can appear before changes in protein levels, achieving earlier risk warning. The metabolic biomarker group can be selected from adenosine, fumarate, tryptophan, and acylcarnitine. Tumor cells exhibit metabolic reprogramming characteristics, and gastric cancer cells show abnormalities in glycolysis, the tricarboxylic acid cycle, and amino acid metabolism. These biomarkers are core characteristic markers of gastric cancer metabolic reprogramming, and changes in their peripheral blood concentrations can directly reflect the abnormal metabolic state of the gastric mucosa. In this invention, at least one biomarker is selected from each group, ensuring coverage at the molecular level across multiple omics dimensions and avoiding the one-sidedness of risk assessment due to the absence of a biomarker at a certain molecular level. In clinical practice, the number of biomarkers in each group can be flexibly selected according to the testing equipment, clinical needs, and patient conditions, balancing comprehensiveness and economy of testing. For example, two protein biomarkers, one methylated DNA biomarker, and two metabolic biomarkers can be selected. In this invention, the sampling intervals between different time points can be unequal. This design is key to adapting the invention to real-world clinical scenarios, enhancing the clinical operability and scalability of the technical solution. Traditional time-series data analysis methods often require equidistant sampling. However, in clinical follow-up of gastric cancer patients, the follow-up time cannot be strictly fixed due to factors such as individual physical condition, access to medical care, and clinical treatment needs. For example, patients may be followed up at 1 month, 3 months, 6 months, and 12 months post-surgery, with sampling intervals of 2 months, 3 months, and 6 months, respectively. Forcing equidistant sampling would significantly increase the difficulty of clinical implementation and reduce clinical operability. This invention achieves the calculation of temporal characteristics of data at arbitrary sampling intervals through a non-equidistant finite forward difference method. Mathematically, the detection values can be normalized using the actual difference in sampling timestamps, ensuring the accuracy of the calculation of the rate of change and acceleration. Therefore, the sampling intervals can be unequal, which does not affect the subsequent mathematical model calculations and significantly reduces the difficulty of clinical sampling, thus improving the clinical applicability and scalability of the invention.
[0025] To ensure the accuracy of subsequent mathematical model calculations, this step, when acquiring the peripheral blood multi-omics detection dataset, must adhere to standardized procedures for peripheral blood sample collection, processing, and detection. Furthermore, the detection of various biomarkers should utilize clinically-specific quantitative detection kits, which are the mainstream and preferred method for achieving peripheral blood biomarker detection in this invention. During sample collection, all peripheral blood samples from the same subject must be collected using the same method (e.g., 5 mL of fasting venous blood) and the same anticoagulant (e.g., EDTA anticoagulant) to avoid deviations in detection values due to differences in collection methods. During sample processing, all samples must undergo centrifugation to separate plasma / serum within the same time window, such as centrifuging plasma / serum within one hour of blood collection. Centrifugation speed, time, temperature, and other processing conditions must be kept consistent to prevent biomarker degradation or denaturation. In the testing process, the use of reagent kits must follow strict standardization requirements, and different biomarker groups are compatible with corresponding commercially available mature quantitative reagent kits: protein biomarker groups are compatible with enzyme-linked immunosorbent assay (ELISA) kits and chemiluminescent immunoassay kits; methylated DNA biomarker groups are compatible with methylation-specific PCR (MSP) quantitative kits and pyrosequencing methylation detection kits; metabolic biomarker groups are compatible with liquid chromatography-mass spectrometry (LC-MS) quantitative kits and enzymatic quantitative detection kits. Alternatively, a combination of reagent kits and matching small detection devices can be used to improve the portability of the test and meet the rapid sample testing needs of primary healthcare institutions and gastric cancer screening sites. Meanwhile, for the detection of the same biomarker at different time points, the same brand and batch of reagent kits and supporting auxiliary detection equipment must be used. The operating environment, personnel, and procedures for sample testing must be consistent, and the selected reagent kits must be quantitative detection type. Qualitative reagent kits that can only output "positive / negative" are prohibited to ensure that specific quantitative detection values of biomarkers can be obtained to meet the needs of subsequent mathematical calculations. In addition, the selected reagent kits must be validated for peripheral blood samples to ensure good sensitivity and specificity for gastric cancer-related target biomarkers, avoid false positive and false negative results, and ensure the reliability of the original test data.
[0026] In this step, the raw data obtained from clinical testing needs to be structured. Through standardized structured processing, the scattered quantitative values are transformed into standardized data that can be directly calculated. Specifically, firstly, all biomarker detection data of the same subject are sorted according to the sampling timestamp to obtain a timestamp set. ( The system matches corresponding protein biomarker, methylated DNA biomarker, and metabolic biomarker values for each time point, ensuring a one-to-one mapping between the biomarkers and timestamps. Based on this, a three-dimensional tensor of the multi-omics detection dataset is constructed. The first dimension The total number of multi-omics biomarkers, satisfying , , , The number of biomarkers for proteins, methylated DNA, and metabolic biomarkers, respectively, with each of the three being ≥1; the second dimension. The number of sampling time points, and The third dimension (3) corresponds to the protein biomarker dimension (dimension 1), the methylated DNA biomarker dimension (dimension 2), and the metabolic biomarker dimension (dimension 3), respectively; tensor elements Indicates the first The first marker, the first The time point, the first Quantitative detection values (real numbers) of each marker group. Finally, from the three-dimensional tensor. For each marker Extracting one-dimensional time series ,in as a marker In the Time points The detection values, i=1,2,...,n, are the basic units for calculating time series features. The structured extraction ensures that subsequent steps can perform independent inflection point detection for each marker and then perform multi-marker fusion analysis.
[0027] For example, if a subject selects the protein biomarker OLFM4, the methylated DNA biomarker CDO1, and the metabolic biomarker adenosine (… ), sampling at 4 time points, then the tensor It is a 3×4×3 three-dimensional matrix, where For OLFM4 in The detected value, For CDO1 in methylation level, For adenosine in The concentration.
[0028] S2. For each biomarker in the multi-omics detection dataset, extract the detection value of the biomarker at each time point and the corresponding timestamp to construct time series data. Use the numerical difference method to calculate the first and second derivatives of the time series data in sequence to obtain the rate of change data and acceleration data of each biomarker at each time point. The time point when the absolute value of the second derivative changes from below the first threshold to above the second threshold is determined as the inflection point position of the acceleration change of the biomarker. Inflection point flags are generated for each biomarker. The inflection point flag of the biomarker that is detected is assigned a value of 1, and the inflection point flag of the biomarker that is not detected is assigned a value of 0.
[0029] This step extracts the temporal dynamic features of biomarkers using the numerical difference method, determines the inflection point of acceleration change based on the second derivative threshold mutation, and generates standardized inflection point markers. This overcomes the technical limitations of existing technologies, which can only perform simple linear fitting of biomarker changes and cannot identify key nodes of acceleration mutations in the evolution of gastric cancer.
[0030] Before performing acceleration inflection point detection and marker generation, it is necessary to complete the standardization construction of single-marker time series data. This process is a fundamental prerequisite for subsequent numerical calculations and inflection point detection. This step uses the multi-omics detection tensor constructed in step S1. Extracting individual markers Detection value sequence and match the corresponding set of sampling timestamps. To form the original time pair Subsequently, the raw time-series data were standardized to convert all timestamps to the same time unit (clinically preferred to be days, but months or years can also be selected depending on the follow-up period), and the initial sampling time could be included. Set the baseline value to 0, and subsequent timestamps represent the time difference from the first sampling. Simultaneously, maintain consistency between the dimensions of the biomarker detection values and those in step S1 (e.g., protein ng / mL, methylation rate %, metabolite μmol / L) to avoid modifying dimensions across time points. Finally, perform mild smoothing and denoising on the detection value sequence (clinically preferred moving average method, such as a 3-point moving average). Remove extreme outliers caused by sample detection errors, such as values exceeding the normal reference range by more than 10 times. The denoising process must preserve the overall trend of the original data to prevent the smoothing of acceleration mutation features due to excessive denoising. The final result is the value for each biomarker. Standardized time series ,satisfy Sampling interval (Can be any positive real number), and the detected value has no extreme anomalies.
[0031] Based on the standardized single-marker time series, this invention employs a non-equidistant finite forward difference method that adapts to features where sampling intervals may be unequal. It sequentially calculates the first derivative (rate of change) and second derivative (acceleration of change) of the marker detection values, achieving accurate extraction of the marker's temporal dynamic features. Among these features, the marker... In the Time points rate of change The calculation formula is: ; In the formula For the first The time point and the first The change in the detected value at each time point The actual sampling interval between the two time points. The unit is measured in units of detection value / time, such as ng / mL / day or % / month. The positive and negative values, and the absolute values, respectively characterize the direction and rate of change of the biomarker's detection value. A positive value indicates that the detection value increases over time; a negative value indicates that the detection value decreases over time. The larger the absolute value, the more significant the change. Based on the rate of change, the biomarker is further calculated. In the Time points Change in acceleration The calculation formula is: ; In the formula For the first The time point and the first The rate of change at each time point, with the actual sampling interval in the denominator. The unit is the unit of detection value / unit of time. 2 , such as ng / mL / day 2 ,% / moon 2 Its positive and negative values and absolute values reflect the trend and magnitude of the rate of change. Positive values indicate that the rate of change is accelerating, while negative values indicate that the rate of change is slowing down. The larger the absolute value, the more drastic the rate of change. The abrupt change in the absolute value of the acceleration is the core mathematical characteristic of the transformation of gastric cancer risk from quantitative to qualitative change.
[0032] This invention uses the abrupt change in the absolute value of the acceleration of a biomarker to achieve precise and automatic determination of the inflection point of acceleration change. Clinically, this inflection point is a critical time node for the transformation of precancerous lesions into early-stage gastric cancer, and also represents the optimal window for early intervention in gastric cancer. Two core positive real-value hyperparameters are defined in the inflection point determination process, namely the first threshold... With the second threshold And satisfy Both thresholds were determined from clinical cohort data of gastric cancer; if biomarkers The acceleration sequence contains an index This makes its acceleration satisfy Then determine the index. corresponding time point This indicates the location of the inflection point in the acceleration change of this marker. Inflection point determination follows the core rules of trend abrupt change and unique inflection point, and must satisfy the condition that "the absolute value of the previous acceleration is lower than..." The current absolute value of acceleration is higher than The trend of adjacent values, rather than a single acceleration value higher than ", is observed. To avoid mistaking accidental acceleration fluctuations for inflection points; if multiple indexes satisfy the conditions... Only the first one that meets the condition is selected. As a turning point, the first accelerating mutation in the evolution of gastric cancer is the most critical risk signal, while subsequent mutations only represent a continuation of malignant development. This leads to the determination of the turning point index. Then, it is mapped to the normalized actual timestamp from step S2. This timestamp is a clinically interpretable inflection point and will be included in subsequent comprehensive assessment reports to provide clinicians with specific nodes of risk transformation.
[0033] After detecting inflection points in acceleration, this step generates standardized inflection point markers for each marker, transforming continuous inflection point location features into discrete 0 / 1 binary features. This enables seamless integration with subsequent machine learning risk assessment models while ensuring consistent feature dimensions across all markers, simplifying the model input structure. (Marker) The inflection point is marked as Its assignment formula is: ; in This indicates that the marker has detected an inflection point in the acceleration change, suggesting a core signal that the risk of gastric cancer has changed from quantitative to qualitative. This marker will be used as a high-weight feature in the subsequent dynamic risk score calculation. This indicates that no inflection point in acceleration change was detected for this marker, and the rate of change of its detected value remained stable, with no obvious signs of a qualitative change in risk. This step generates a set of inflection point markers for all multi-omics markers. Together with the rate of change sequence, acceleration sequence, and inflection point location set, this will serve as the final output of this step.
[0034] The first threshold used for inflection point determination in this step With the second threshold Its value directly determines the sensitivity and specificity of inflection point detection. Too small a value can easily lead to false positives inflection point detection. Excessive thresholds can easily lead to false negative inflection point detections. Therefore, both thresholds need to be scientifically calibrated based on large-sample clinical cohort data of gastric cancer, rather than being subjectively set. The specific process for threshold calibration is as follows: First, longitudinal multi-omics follow-up data and corresponding clinical outcomes of precancerous lesions of gastric cancer, early gastric cancer, and healthy individuals are collected. The time series data of each biomarker are labeled with the true inflection point determined by clinical experts. Subsequently, a 5-fold cross-validation method can be used to further validate the inflection point. , A grid search was performed, searching the 0th to 95th quantiles of absolute acceleration values in the clinical cohort. Next, the parameter combination that maximizes the Youden index of the inflection point detection model was selected, balancing the sensitivity and specificity of the detection. Finally, based on clinical guidelines for gastric cancer screening, the optimal parameter combination was slightly fine-tuned to adapt to the screening needs of different clinical scenarios. (The calibrated parameter combination is shown in the image.) , The hyperparameters of the model are fixed and can be adjusted individually according to the actual needs of the clinical center. For example, the hyperparameters can be appropriately reduced for high-risk population screening. To improve detection sensitivity, screening of the general population can be appropriately increased. To reduce the false positive rate.
[0035] S3. For each marker with an inflection point value of 1, extract all detection values after the inflection point from the time series data of that marker. Use linear regression to fit and obtain the rate of change and acceleration after the inflection point. Subtract the average rate of change before the inflection point from the rate of change after the inflection point to obtain the difference in the rate of change before and after the inflection point. The average rate of change before the inflection point is calculated as follows: calculate the rate of change data at each time point from the start time point of the time series data to the inflection point, sum the rate of change data at each time point, and divide by the number of time points to obtain the average rate of change before the inflection point.
[0036] This step is used to perform targeted trend feature quantification extraction on the biomarkers with an inflection point value of 1 in step S2. It transforms the nonlinear temporal changes of these biomarkers during the evolution of gastric cancer risk into standardized, quantitative features that can be directly input into machine learning models. This provides input evidence reflecting the evolutionary trend after qualitative changes in risk for subsequent dynamic risk assessment of gastric cancer. The feature extraction operation in this step is only performed on biomarkers that have detected inflection points in acceleration changes. For biomarkers with an inflection point value of 0 in step S2, no feature extraction is performed; only zeros are added to ensure the consistency of the dimensionality of the input features for subsequent risk assessment models and avoid model calculation errors caused by inconsistent feature dimensions.
[0037] Specifically, when extracting features from markers exhibiting acceleration inflection points, the first step is to calculate the average rate of change before the inflection point. This indicator serves as a benchmark feature for quantifying the trend of change during the stable evolution phase before the inflection point, providing a standardized reference value for subsequent comparisons of trends after the inflection point. The formula for calculating the average rate of change before the inflection point is: ; The summation term To arithmetically sum all rates of change of the marker m from the first time point to the inflection point, reflecting the overall trend of all rates of change before the inflection point; denominator This represents the rate of change before the inflection point, consistent with the inflection point index, ensuring the unbiasedness of the average value. It is the arithmetic mean of all rates of change before the inflection point, and is a quantitative summary of the trend of change during the steady phase before the inflection point. This indicates that the marker m shows an average upward trend before the inflection point, with an upward rate of... This suggests that molecular signals associated with gastric cancer increase slowly during a stable phase. This indicates that the marker m shows an average decreasing trend before the inflection point, with a decreasing rate of... This suggests that molecular signals associated with gastric cancer gradually weaken during a stable phase. This indicates that the biomarker m showed no significant change trend before the inflection point, and the molecular signals related to gastric cancer were in a stable state.
[0038] It should be noted that the calculation of this indicator must be strictly limited to the summation range from the start point of the time series to the inflection point, and must not include any rate of change value after the inflection point, so as to prevent the accelerating trend after the inflection point from interfering with the quantitative results of the benchmark characteristics.
[0039] After calculating the average rate of change before the inflection point, it is necessary to extract the detected values after the inflection point and perform linear regression fitting to quantify the accelerated change trend of the biomarker after the inflection point. First, from the total time series of the biomarker, all detected values and corresponding timestamps from the inflection point to the last time point are extracted to form a specific data subset after the inflection point. During the extraction process, the detection values and timestamps (i=r) at the inflection point must be retained. This time point serves as the boundary between the steady-state and accelerated change phases of the biomarker. Including it in the fitting ensures the continuity of the trend. Simultaneously, at least two data subsets after the inflection point are required to provide a valid data foundation for linear regression fitting. A univariate linear regression is performed on the data subsets after the inflection point, and the fitting equation is: ; in as a marker After the inflection point The fitted value of the detection value at each time point is the optimal estimate of the actual detection value through a linear model; The slope term of the linear regression, which is also the rate of change of this biomarker after the inflection point, is a quantitative value of the accelerating trend of gastric cancer risk. The larger the value, the more dramatic the change in molecular signals after the inflection point, and the faster the risk of gastric cancer worsens. The first after the inflection point Standardized timestamps for each point in time; The intercept term in linear regression is used only to ensure the mathematical accuracy of the fitted equation. It has no independent clinical significance and is not included in the input of subsequent risk assessment models.
[0040] Based on the mathematical properties of linear functions, where the first derivative is a constant and the second derivative is zero, and considering the definition of acceleration as the rate of change of the rate of change in this invention, and aligning with the clinical progression of gastric cancer, after precancerous lesions transform into early-stage cancer, the rate of change of relevant biomarkers remains constant and accelerated, without secondary mutations. Therefore, the inflection point acceleration of biomarkers is defined. This value indicates that after the inflection point, the marker enters a linear evolution stage with a constant rate of change, and the acceleration tends to stabilize.
[0041] After obtaining the average rate of change before and after the inflection point, it is necessary to calculate the difference in rates of change before and after the inflection point. This indicator is the core quantitative result integrating the first two features, directly reflecting the relative magnitude of the change in the marker's rate of change after the inflection point, and is also a high-weight input feature for subsequent risk assessment models. The formula for calculating the difference in rates of change before and after the inflection point is as follows: ; in as a marker The difference in the rate of change before and after the inflection point, The rate of change of this marker after the inflection point. This represents the average rate of change of this biomarker before the inflection point. Its positive or negative value directly corresponds to the evolutionary trend of gastric cancer risk. When the value is high, it indicates that the rate of change of the biomarker after the inflection point is faster than the average rate before the inflection point. The larger the value, the more obvious the acceleration. This suggests that the molecular signals related to gastric cancer show a significant acceleration and enhancement after the inflection point, and the risk of gastric cancer enters an accelerated deterioration stage from a stable stage. It is a core high-risk signal in gastric cancer screening. When the rate of change after the inflection point is consistent with the average rate before the inflection point, it indicates that the appearance of the inflection point has not changed the trend of the marker's change, and there is no obvious signal of a qualitative change in risk. False positive inflection points need to be ruled out in conjunction with the patient's clinical baseline data. When the inflection point is reached, it indicates that the rate of change after the inflection point is slower than the average rate before the inflection point, suggesting that the appearance of the inflection point causes the change trend of gastric cancer-related molecular signals to become more gradual, and the risk of gastric cancer does not worsen significantly or even shows a decreasing trend.
[0042] It should be noted that if the number of data subsets after the inflection point is less than two, making it impossible to complete linear regression fitting, then the biomarker should be treated as having no inflection point, and its inflection point flag should be reset to 0. The goodness of fit of the fitting results needs to be verified, and it is recommended that the goodness of fit be no less than 0.8. If it is lower than this value, it indicates that the trend after the inflection point has significant non-linear characteristics, and it is necessary to combine clinical data to exclude false positive inflection points. The data subsets after the inflection point need to be checked for extreme outliers. If outliers exist, the nearest neighbor interpolation method can be used for correction to avoid the deviation of fitting results caused by outliers.
[0043] S4. Input the current detection value, inflection point marker, rate of change after the inflection point, and acceleration after the inflection point for each biomarker into the pre-trained risk assessment model, and output a dynamic risk score. The risk assessment model is an ensemble learning model based on gradient boosting trees. The training samples of this model include longitudinal multi-omics data of each subject in the historical cohort and the corresponding gastric cancer outcome labels.
[0044] Before inputting the model, standardized single-marker feature vectors need to be independently constructed for each multi-omics marker. This integrates the multi-dimensional risk features of individual markers into a unified feature unit, which is the foundation for achieving global fusion of multi-marker features. For the [specific example]... The formula for constructing the feature vector of a single marker is as follows: ; in Representing the A 4-dimensional single-marker feature vector of a marker, with superscript... This is the transpose symbol for a vector, used to convert a row vector into a column vector, which facilitates the subsequent concatenation of the total feature vectors. To reflect the current expression / methylation / concentration level of the biomarker, it is a static baseline signal of gastric cancer risk; To reflect whether the biomarker has an acceleration mutation inflection point, it is a binary variable that takes only 0 or 1, which belongs to the mutation characteristics that reflect the qualitative change in the risk of gastric cancer; To reflect the accelerated rate of change of biomarkers after the inflection point, it is a quantitative signal of the speed at which the risk of gastric cancer worsens. To reflect the stability of acceleration after the inflection point of the biomarker, a constant value of 0 indicates that the risk has entered a linear evolution stage of continuous acceleration. The feature vector has a fixed dimension of 4, which is independent of the type of biomarker (protein, methylated DNA, metabolism), realizing feature standardization for different types of biomarkers. Furthermore, the vector integrates continuous numerical features and binary categorical features, allowing the gradient boosting tree model to directly process this mixed-type feature without additional feature encoding operations.
[0045] After constructing the feature vectors of individual markers, the feature vectors of all markers need to be vertically concatenated according to their numerical order to form a total input feature vector that the model can directly recognize. This achieves global fusion of multi-marker, multi-dimensional risk features, fully leveraging the advantages of multi-omics detection. The formula for concatenating the total input feature vector is: ; in This is the total input feature vector of the model. The total number of multi-omics biomarkers. For the first The row vector obtained by transposing the feature vector of each marker. The dimension of the total input feature vector is... The total dimension is 12 if three biomarkers are selected (1 protein + 1 methylated DNA + 1 metabolism); and 20 if five biomarkers are selected. The dimension increases linearly with the number of biomarkers, which is suitable for gradient boosting tree models to handle high-dimensional features. Furthermore, this vector includes core risk features such as static detection, inflection point mutations, and temporal trends for all biomarkers, achieving comprehensive fusion of multi-omics features from protein, methylated DNA, and metabolism biomarkers. During the assembly process, the biomarker numbering order must remain fixed, and the assembly order must be completely consistent between the model training phase and the actual clinical application phase to avoid distortion of model prediction results due to feature misalignment.
[0046] After inputting the standardized total input feature vector into the pre-trained gradient boosting tree ensemble learning model, the model can output the subject's dynamic gastric cancer risk score through forward propagation, realizing the transformation from multi-dimensional features to a single risk quantification value. The calculation formula is as follows: And the score satisfies the range constraint. In the formula The model outputs a dynamic risk score for gastric cancer. This is an ensemble learning model of gradient boosting trees pre-trained on clinical historical cohort data. This serves as the total input feature vector for the completed model. This gradient boosting tree model is essentially a... 3D feature space to The nonlinear mapping function of the risk space can accurately capture the complex nonlinear interaction between multiple biomarker features, breaking through the feature fusion limitations of traditional linear models. At the same time, the model maps the prediction results to a fixed interval of 0 to 1 through the logistic regression output layer, effectively avoiding meaningless extreme values, making the score clinically interpretable and able to intuitively reflect the subject's current overall risk level of gastric cancer.
[0047] The gradient boosting tree ensemble learning model used in this step needs to be pre-trained on a large sample of historical clinical cohort data of gastric cancer before it can be applied to actual clinical risk assessment. The model training process is a supervised machine learning process, which relies heavily on a fully labeled training dataset. The model's training dataset is denoted as [dataset name missing]. ,in For the first in the historical queue The total input feature vector of each training sample is constructed according to the same rules as the feature vector construction rules of subjects in actual clinical applications, ensuring the consistency between model training and application. For the first The gastric cancer outcome labels corresponding to each training sample are binary variables that take only 0 or 1. This indicates that the subject was diagnosed with stomach cancer during the follow-up period. This indicates that the subject was not diagnosed with gastric cancer during the follow-up period. During model training, a logarithmic loss function is used to measure the error between the predicted risk score and the actual outcome label. Minimizing this logarithmic loss function is the core optimization objective. Training is completed through a serial ensemble of multiple decision trees. This model has advantages such as strong resistance to overfitting, insensitivity to missing values, and the ability to directly handle mixed-type features, making it suitable for the fusion requirements of multi-omics time-series features in this invention.
[0048] Meanwhile, to ensure the accuracy and stability of the model's prediction results, the feature vectors input to the model need to be standardized, especially the feature complements for markers without inflection points. This ensures that the feature dimensions of all markers are consistent and the values are valid, avoiding model prediction failures due to missing features or inconsistent formats. For inflection point markers... The absence of inflection point markers indicates a lack of actual post-inflection point change rates because step S3 (inflection point trend feature extraction) was not performed. acceleration after the inflection point Value, must be calculated according to , The rules are used to perform complementation, and after complementation, its single-marker feature vector is... This ensures that the feature vectors of all biomarkers are complete 4-dimensional vectors, allowing the model to identify risk signals of accelerated change after the biomarker has no inflection point through a 0 value. Simultaneously, due to the difference in the dimensions of detection values for different types of biomarkers, Z-score standardization is required for all features during model training. This transforms the feature values into dimensionless standard normal distribution values, avoiding excessive focus on large numerical features due to dimensional differences. Furthermore, in actual clinical applications, the mean and standard deviation of the model training dataset should be used for feature standardization, and the use of statistics from new subject data is prohibited. Additionally, extreme outliers in the feature vectors can be truncated to prevent distortion of model predictions.
[0049] Dynamic risk score output by the model The score ranges from 0 to 1, and its magnitude is positively correlated with the subject's risk level of gastric cancer. A higher score indicates a higher current risk of gastric cancer and a greater probability of developing gastric cancer. In clinical practice, the risk stratification threshold can be determined based on the ROC curve Youden index of a gastric cancer clinical cohort, dividing the dynamic risk score into three risk levels: low, medium, and high. This provides clinicians with an intuitive and practical basis for interpreting risk. A low-risk level indicates that the subject has an extremely low risk of gastric cancer, and there are no obvious abnormal risk signals in the relevant markers. If the subject has precancerous lesions, the lesions are in a stable state. The risk level is medium, indicating that the subject has a certain risk of gastric cancer and some biomarkers show abnormal values, but there is no clear acceleration inflection point mutation signal. A high-risk level indicates that the subject has an extremely high risk of gastric cancer, with the presence of one or more biomarkers showing a significant acceleration inflection point mutation, and a marked acceleration trend in the biomarkers after the inflection point. The risk stratification thresholds mentioned above are not fixed values and can be appropriately adjusted according to the screening needs of different clinical centers, such as targeting high-risk gastric cancer cohorts or general population cohorts, to ensure an optimal balance between the sensitivity and specificity of risk stratification.
[0050] S5. Input the dynamic risk score, the existence of an inflection point flag of at least one marker with a value of 1, and the maximum value of the rate of change after the inflection point among all markers with an inflection point flag of 1 into the reinforcement learning agent, and output the recommended next follow-up interval.
[0051] The reinforcement learning agent adopts a deep Q-network structure. Its state space includes dynamic risk score, inflection point marker and maximum rate of change after the inflection point, and the action space is a preset discrete follow-up interval. The reward function is set as a weighted sum of negative missed diagnosis risk loss and detection cost.
[0052] In one embodiment of the present invention, the decision rule for the recommended next follow-up interval is as follows: When an inflection point exists and the maximum rate of change after the inflection point is greater than 0.1, output an interval of 30 to 90 days; when an inflection point exists and the maximum rate of change after the inflection point is not greater than 0.1, output an interval of 90 to 180 days; when no inflection point exists and the dynamic risk score is lower than the preset risk threshold, output an interval of 180 to 365 days; when no inflection point exists and the dynamic risk score is not lower than the preset risk threshold, output an interval of 90 to 180 days.
[0053] This step integrates the optimality of reinforcement learning algorithms with the practical needs of gastric cancer clinical follow-up. It enables adaptive recommendation of follow-up intervals through the algorithm and constrains the algorithm's output through clinically customized rules, ultimately achieving precise, personalized, and efficient gastric cancer follow-up. This reduces the risk of missed diagnoses of gastric cancer while minimizing over-testing in low-risk individuals, improving the utilization efficiency of medical resources. This step inputs three risk features—dynamic risk score, global inflection point marker (whether an inflection point exists), and the maximum rate of change after the inflection point—into a pre-trained deep Q-network reinforcement learning agent. Combined with pre-defined clinical decision rules, it outputs a discrete next follow-up interval that meets clinical requirements. The state space, action space, and reward function of the reinforcement learning agent are all customized based on the needs of gastric cancer clinical follow-up, ensuring a high degree of adaptability between the algorithm's output and clinical applications.
[0054] Before inputting features into the reinforcement learning agent, the single-marker risk features obtained in steps S2 and S3 must be aggregated into global risk features that reflect the overall risk status of the subjects, specifically including global inflection point markers. Maximum rate of change after the inflection point Aggregation calculation of two core features. Let the set of multi-omics biomarkers for the subjects be... The single-marker inflection point marker is ( Then the global inflection point flag The mathematical formula is: ; in For the first The inflection point marker of each indicator takes a value of 0 or 1. The formula represents the maximum value of the inflection point values of all single markers, where the total number of multi-omics biomarkers is represented by the formula. If there is an inflection point value for at least one marker, the formula is used to determine the maximum value. ,but This characterizes the accelerating mutation signal indicating a shift in the risk of gastric cancer from quantitative to qualitative change in the subject's body; if the inflection point of all biomarkers is 0, then... The rate of change in gastric cancer-related biomarkers in the subjects was stable, with no risk of qualitative change.
[0055] Let the subset of markers that have inflection points be . Its rate of change after the inflection point is Then the maximum speed The mathematical formula is: ; like Then uniform regulations ,in For the first The rate of change after the inflection point of an inflection point marker, for Taking the absolute value and then finding the maximum value aims to eliminate the influence of positive or negative rates, quantifying only the drastic change after the inflection point of the marker. The higher the value, the more significant the accelerated change of the biomarker after the inflection point, and the faster the risk of gastric cancer in the subject deteriorates.
[0056] After aggregating the global risk features, they need to be fused with the dynamic risk score output in step S4 to construct a standardized reinforcement learning state vector, which serves as the sole input to the deep Q-network reinforcement learning agent. The reinforcement learning state vector uses... It means that its expression is: ; superscript This represents the transpose of a vector, that is, converting a row vector into a column vector that fits the input format of a deep Q-network; The dynamic risk score output in step S4 ranges from 0 to 1 and is used to quantify the overall gastric cancer risk level of the subject. This state vector is a 3-dimensional real number vector that integrates continuous numerical features and binary categorical features. It can be directly processed by a deep Q-network without additional feature encoding and comprehensively covers the overall risk quantification of gastric cancer. ), risk qualitative change signal ( ), speed of risk deterioration ( The three clinical dimensions provide a comprehensive quantitative summary of the subject's current gastric cancer risk status, providing standardized and structured feature inputs for the optimal action reasoning of the subsequent reinforcement learning agent. Furthermore, the splicing order of features remains completely consistent during model training and practical application, avoiding distortion of reasoning results due to feature misalignment.
[0057] In this embodiment of the invention, the reinforcement learning agent used in this step is a Deep Q-Network (DQN) structure, which is an intelligent algorithm model combining deep learning and Q-learning. It can achieve accurate mapping from a continuous state space to a discrete action space. Its core consists of three main elements: state space, action space, and reward function, all of which are customized based on the needs of clinical follow-up of gastric cancer. The agent's state space is the aforementioned 3-dimensional real space, and it corresponds to the reinforcement learning state vector. The dimensions are consistent; the action space is used This represents a pre-defined set of discrete follow-up interval values, with the unit uniformly set to days. A commonly used clinical design scheme is... It can be fine-tuned based on the characteristics of the screening population (e.g., high-risk groups, general population) and follow-up practice habits of clinical centers, without setting follow-up interval values that are not routinely used in clinical practice. Reward function It is the core learning basis of deep Q-networks, using The design is a weighted sum of negative missed diagnosis risk and testing costs, calculated using the following formula: ; in This represents a single discrete follow-up interval value in the action space. , For the weighting coefficients, satisfying Since the core goal of gastric cancer screening is early detection and early diagnosis, priority is given to reducing the risk of missed diagnoses, and the design is as follows: ; The risk loss for missed diagnosis is represented by a non-negative real number. The higher the risk status of the subject and the longer the selected follow-up interval, the greater the risk loss. The larger the value; Let $\mathbf{ ... Input pre-trained deep Q-network reinforcement learning agent The agent calculates and outputs the optimal action through forward propagation. The calculation formula is: ; in For a deep Q-network, the Q-value function represents the state... Next action Subsequently, the cumulative expected reward during the long-term follow-up of the subjects, Indicates from action space Among all discrete values, the action that maximizes the Q-value function is selected, and this optimal action is the initially recommended next follow-up interval.
[0058] To ensure that the algorithm's output fully aligns with clinical screening guidelines and practical practices for gastric cancer, the optimal action output by the deep Q-network needs to be determined using pre-defined four-category clinical decision rules. After applying interval constraints, the recommended next follow-up interval is finally determined and output. This decision-making rule follows the clinical logic of "the higher the risk, the shorter the follow-up interval." Recommended follow-up interval... The interval determination formula is: ; in (Fixed critical value), when At that time, it indicated that the malignant transformation of the gastric mucosa in the subject had entered a period of rapid progression; The preset risk threshold for dynamic risk scoring, with a value range of 0 to 1, is determined by using ROC curves and the Youden index from gastric cancer clinical cohort data. This threshold is also relevant for the general gastric cancer screening population. The value is usually 0.5, and it is also found in high-risk groups for stomach cancer. The typical value is 0.3. This rule categorizes the risk status of subjects into four classes. and This is a high-risk acute transformation type, requiring a short-term follow-up of 30-90 days to closely monitor changes in the condition; and For medium-risk, progressively changing cases, a mid-term follow-up of 90-180 days is required to balance the costs of disease monitoring and testing. and For low-risk, stable cases, long-term follow-up of 180-365 days is required to reduce excessive testing. and For patients with a moderate to stable risk profile, a mid-term follow-up of 90-180 days is required to promptly identify signals of risk changes. The final recommended follow-up interval is as follows. For the optimal discrete action within the corresponding interval This ensures the clinical feasibility of the results.
[0059] The deep Q-network reinforcement learning agent requires pre-training on large-sample clinical follow-up data of gastric cancer before it can be deployed in actual clinical applications. Its training process is a model-free reinforcement learning process, solving the problems of sample correlation and training instability in traditional Q-learning. To achieve this goal, the agent uses a dual-network structure of an experience replay pool and a target network + main network for training. The experience replay pool stores "state-action-reward-next state" experience sets generated during training. During training, batches of experience are randomly drawn from the pool for learning, effectively eliminating the temporal correlation of samples. In the dual-network structure, the main network updates parameters in real time to calculate the Q-value in the current state, while the target network has fixed parameters and periodically copies parameters from the main network to calculate the target Q-value, avoiding bootstrapping bias in Q-values during training and improving training stability. The optimization objective of this agent is to minimize the mean square error between the Q-value of the main network and the Q-value of the target network. The target Q-value is calculated according to the Bellman equation, taking into account both current rewards and future cumulative rewards. Through training with this optimization objective, the follow-up decisions learned by the agent can achieve the minimum weighted sum of missed diagnosis risk and detection cost during long-term follow-up, ensuring that the follow-up interval output by the algorithm is both in line with clinical benefits and has long-term optimality.
[0060] S6. Generate and output a comprehensive assessment report containing a dynamic risk score, a list of inflection point detection results, and a recommended next follow-up interval. The list of inflection point detection results includes the name of each inflection point marker with a value of 1, the location of the inflection point in acceleration, the rate of change after the inflection point, and the acceleration after the inflection point.
[0061] This step is used to structure, standardize, and clinically integrate scattered quantitative features (dynamic risk scores), time-series features (inflection point detection details), and clinical decision results (personalized follow-up intervals) to generate a comprehensive assessment report that can be quickly interpreted by clinicians. The report retains the accuracy of quantitative calculations while also taking into account the readability of clinical applications, solving the problem that the preceding technical features are difficult to use directly for clinical decision-making.
[0062] First, the basic results are accurately extracted from the preliminary steps of this invention, namely the dynamic risk score output in step S4. The recommended next follow-up interval output in step S5 The inflection point-related features of all markers in steps S2 to S4, among which are consecutive real numbers and satisfy It is a core indicator for quantifying the overall gastric cancer risk of subjects. A discrete positive real number in days is a core indicator for clinical follow-up decisions, and inflection point-related features include inflection point markers. Logo Name Location of the inflection point of acceleration change Rate of change after the inflection point acceleration after the inflection point This provides complete data support for constructing the subsequent inflection point detection result list. Secondly, based on the above filtering rules, a subset of inflection point markers is selected. ,like If it is an empty set, then directly determine the inflection point detection result list. Empty; if If not empty, then a 4-element composite information item is independently constructed for each inflection point marker. ,in The clinically standardized name of the biomarker. A standardized timestamp for the inflection point. The rate of change after the inflection point. The acceleration after the inflection point is constant at 0. This information item is the basic building block of the inflection point detection result list, realizing the comprehensive integration of the risk mutation characteristics of a single biomarker. Subsequently, all four-element composite information items are... The results of inflection point detection are compiled in the order of routine clinical testing for protein biomarkers, methylated DNA biomarkers, and metabolic biomarkers. This ensures the list is standardized and readable. Finally, dynamic risk scoring will be implemented. List of inflection point detection results Recommended interval for next follow-up visit The report integrates three main parts: overall risk, details of local risk, and clinical priority of clinical decisions, to form a standardized comprehensive risk assessment report for gastric cancer. Its overall set expression is: ; In this invention, the three core components of the comprehensive assessment report complement each other and progress step by step, forming a complete gastric cancer risk assessment system that quantifies overall risk, traces local risk sources, and facilitates clinical decision-making. Each component has clear clinical implications and standardized interpretation guidelines, enabling clinicians to quickly grasp the gastric cancer risk status of subjects. Dynamic risk scoring. As an overall risk quantification indicator, the value ranges from 0 to 1. The higher the value, the higher the risk of gastric cancer in the subject. In clinical practice, it can be combined with the risk threshold preset in step S4. The report directly provides It labels the risk levels as low, medium, and high, directly outputting the overall risk conclusion without requiring additional calculations by clinicians; inflection point detection result list. It can serve as a local source tracing module for risky mutations and can be presented in a clinically common three-line table format, with a table header and four-element composite information items. The components correspond one-to-one, and only relevant information with inflection point markers is included. High-risk indicators, such as those with absolute values greater than 0.1, can be highlighted in bold red. If no inflection point is detected, this module can clearly indicate "No inflection point in the acceleration change of any multi-omics biomarker was detected; all biomarkers showed stable temporal changes." Clinicians can use this module to trace the specific time points of core risk biomarkers and risk mutations, providing a basis for targeted monitoring; the recommended next follow-up interval... It can serve as a direct decision indicator for clinical follow-up. In the report, it is presented as specific discrete numerical values combined with units, without ambiguous interval values. At the same time, the basis for determining the follow-up interval will be marked, clearly explaining the technical logic behind the decision, so that clinicians can understand the basis for the decision and realize the direct implementation from risk quantification to clinical decision-making.
[0063] In specific embodiments of this invention, to ensure the applicability and consistency of the technical solution across different clinical centers, the report adopts the clinically common A4 document format, divided into six modules: title bar, basic subject information bar, core assessment results bar, detailed analysis bar, clinical decision-making recommendations bar, and remarks bar. This aligns with the clinical logic of "first look at the core conclusions, then the detailed analysis, and finally the decision-making recommendations," improving the efficiency of clinical decision-making. Regarding information presentation, numerical indicators are uniformly retained to two decimal places with the original units indicated; high-risk indicators are highlighted in bold red; character indicators use standardized clinical nomenclature; protein and methylated DNA biomarkers use their official English names; and metabolic biomarkers use their standard Chinese names to avoid ambiguity. In terms of content labeling, all core results clearly indicate the data source, all quantitative indicators are labeled with their corresponding clinical reference standards, and all special cases, such as data supplementation, outlier handling, and absence of inflection points, are explained in detail in the remarks bar. The remarks bar also includes technical information such as the detection method, biomarker selection, and model version, enhancing the report's scientific rigor and traceability.
[0064] In one embodiment of the present invention, S7 is further included: when the recommended next follow-up interval is less than 90 days, an emergency warning sign is added to the comprehensive evaluation report.
[0065] This step, through emergency warning indicators, allows clinicians to quickly identify high-risk acute-transformation cases when reviewing reports, prioritizing follow-up and intervention, thus solving the clinical problems of high-risk signals being easily overlooked and interventions being delayed in traditional reports.
[0066] Specifically, the early warning determination in this step is based solely on the recommended next follow-up interval output in step S5, while a fixed clinical threshold is set. The determination logic is as follows: ; in, The recommended next follow-up interval is in days and is calculated by a deep Q-network reinforcement learning agent by combining the subject's dynamic risk score, global inflection point features, and the maximum rate after the inflection point. The threshold for emergency warning is fixed at 90 days, which can be determined based on current clinical treatment guidelines for gastric cancer and multicenter, large-sample longitudinal follow-up studies. When the recommended follow-up interval is less than 90 days, the subject is considered to be at high risk of blast crisis in gastric cancer, with a possibility of rapid deterioration, requiring an emergency warning. When the follow-up interval is greater than or equal to 90 days, the subject is considered to be in the low-to-medium risk group, and follow-up can be carried out at the regular pace without additional warning prompts.
[0067] Based on the above judgment results, the emergency warning icon is standardized and binarized, and the assignment rule is expressed in the form of a piecewise function: ; In the formula This is a quantitative representation of the emergency warning indicator, and can only take the value 0 or 1; Representatives are required to attach visual emergency warning signs to the comprehensive assessment report. This indicates that no warning indicators are attached. This binary assignment method can achieve structured and quantitative storage of warning status, adapt to the needs of hospital information systems for electronic classification, screening and management of reports, and directly correspond to the visual presentation of paper and electronic reports, achieving a unified correspondence between numerical features and clinical visualization.
[0068] After assigning values to the warning labels, the emergency warning labels and their corresponding standardized text descriptions can be simultaneously attached to the generated basic comprehensive assessment report to form the final comprehensive assessment report with warnings. Its structured set can be represented as follows: ; In the formula The dynamic risk score output in step S4. The list of inflection point detection results generated in step S6 is integrated. To recommend the follow-up interval, This serves as an emergency warning label. When attaching warning labels and explanations, all core content of the basic comprehensive assessment report must be retained in its entirety. No modifications or deletions may be made to the risk scores, inflection point information, or follow-up recommendations obtained from previous steps, ensuring the completeness, consistency, and traceability of the report data.
[0069] The addition and presentation of emergency warning labels must adhere to unified clinical standardization requirements. Visual warning labels should be prominently displayed in a red background with white text, permanently attached to the core visual area of the clinical decision recommendation section of the report, and presented adjacent to the recommended follow-up interval results, ensuring that clinicians can immediately identify high-risk signals. Accompanying textual descriptions must clearly indicate the reason for the warning, the risk level, and specific clinical intervention requirements, providing direct guidance for subsequent follow-up examinations, detailed gastroscopy, and other diagnostic and treatment actions. For electronic comprehensive assessment reports, electronic warning mechanisms such as pop-up reminders and in-hospital message pushes can be added simultaneously, linking with the hospital's follow-up management system to achieve precise tracking and closed-loop management of high-risk cases.
[0070] The gastric cancer risk assessment method based on peripheral blood multi-omics dynamic analysis according to embodiments of the present invention collects peripheral blood protein, methylated DNA, and metabolic multi-omics data from at least four consecutive time points of the same subject. It incorporates the inflection point of the acceleration of biomarker changes as a key signal of qualitative change in gastric cancer risk into the assessment system. Furthermore, it utilizes a gradient boosting tree model to integrate current detection values, inflection point markers, post-inflection rate of change, and acceleration, thereby improving the sensitivity and specificity of early gastric cancer identification and risk stratification. This addresses the challenge of traditional cross-sectional static models failing to capture the nonlinear evolution of precancerous lesions into gastric cancer. Building upon this, a deep Q-network reinforcement learning agent is further introduced to assess risk based on dynamic risk scores, the existence of global inflection points, and the post-inflection rate of change. The system maximizes the rate and adaptively outputs personalized follow-up intervals. It also uses four-category clinical decision rules to ensure the recommended results meet clinical practice requirements. This effectively reduces the risk of missed diagnoses in high-risk patients while minimizing over-testing in low-risk populations, thus optimizing the utilization of medical resources. Finally, it generates a structured comprehensive assessment report containing dynamic risk scores, a list of inflection point detection results, and specific follow-up intervals. An emergency warning label is added for high-risk blast crisis cases with follow-up intervals less than 90 days. This achieves a closed-loop assessment from multi-omics longitudinal data collection, automatic identification of acceleration inflection points, dynamic risk quantification, to personalized follow-up decision-making and early warning. It provides a practical and scalable solution for accurate screening, dynamic monitoring, and clinical management of precancerous and early-stage gastric cancer patients.
[0071] Corresponding to the above embodiments, the present invention also proposes a gastric cancer risk assessment system based on peripheral blood multi-omics dynamic analysis.
[0072] like Figure 3 As shown, the gastric cancer risk assessment system based on peripheral blood multi-omics dynamic analysis according to an embodiment of the present invention includes a longitudinal multi-omics data acquisition module, an acceleration inflection point detection and marker generation module, a post-inflection point trend feature extraction module, a dynamic risk score calculation module, a personalized follow-up interval recommendation module, and a comprehensive assessment report generation module.
[0073] The longitudinal multi-omics data acquisition module is used to acquire peripheral blood multi-omics test datasets from the same subject at at least four consecutive time points. The multi-omics test datasets include protein biomarkers, methylated DNA biomarkers, and metabolic biomarkers at each time point, and the sampling intervals between time points may be unequal.
[0074] The acceleration inflection point detection and marker generation module is used to extract the detection value and corresponding timestamp of each marker in the multi-omics detection dataset to construct time series data. The first and second derivatives of the time series data are calculated sequentially using the numerical difference method to obtain the rate of change and acceleration data of each marker at each time point. The time point when the absolute value of the second derivative changes from below the first threshold to above the second threshold is determined as the acceleration change inflection point position of the marker. An inflection point marker is generated for each marker. The inflection point marker is assigned a value of 1 for the marker that is detected and a value of 0 for the marker that is not detected.
[0075] The post-inflection point trend feature extraction module is used to extract all detection values after the inflection point time point from the time series data of each marker for which the inflection point marker is assigned a value of 1. Linear regression fitting is used to obtain the post-inflection point change rate and post-inflection point acceleration. The post-inflection point change rate of the marker is subtracted from the average change rate before the inflection point to obtain the difference in change rate before and after the inflection point.
[0076] The dynamic risk score calculation module is used to input the current detection value, inflection point marker, rate of change after the inflection point, and acceleration after the inflection point of each marker into the pre-trained risk assessment model, and output a dynamic risk score.
[0077] The personalized follow-up interval recommendation module is used to input the dynamic risk score, the existence of an inflection point flag of at least one marker (assigned to 1), and the maximum rate of change after the inflection point among all markers with an inflection point flag of 1 into the reinforcement learning agent, and output the recommended next follow-up interval.
[0078] The comprehensive assessment report generation module is used to generate and output a comprehensive assessment report that includes a dynamic risk score, a list of inflection point detection results, and a recommended next follow-up interval.
[0079] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for assessing gastric cancer risk based on dynamic multi-omics analysis of peripheral blood, characterized in that, Includes the following steps: S1. Obtain peripheral blood multi-omics detection datasets from the same subject at at least four consecutive time points. The multi-omics detection datasets include protein biomarkers, methylated DNA biomarkers, and metabolic biomarkers at each time point. The sampling intervals between the time points may be unequal. S2. For each biomarker in the multi-omics detection dataset, extract the detection value of the biomarker at each time point and the corresponding timestamp to construct time series data. Use the numerical difference method to calculate the first derivative and the second derivative of the time series data in sequence to obtain the rate of change data and acceleration data of each biomarker at each time point. The time point when the absolute value of the second derivative changes from below the first threshold to above the second threshold is determined as the acceleration change inflection point position of the biomarker. Inflection point flags are generated for each biomarker. The inflection point flag of the biomarker that is detected is assigned a value of 1, and the inflection point flag of the biomarker that is not detected is assigned a value of 0. S3. For each marker with an inflection point value of 1, extract all detection values after the inflection point from the time series data of the marker, use linear regression to fit the rate of change and acceleration after the inflection point, and subtract the average rate of change before the inflection point from the rate of change after the inflection point to obtain the difference in the rate of change before and after the inflection point. S4. Input the current detection value, inflection point marker, rate of change after the inflection point, and acceleration after the inflection point of each marker into the pre-trained risk assessment model, and output a dynamic risk score. S5. Input the dynamic risk score, the inflection point flag of at least one marker, and the maximum value of the rate of change after the inflection point among all markers with inflection point flags set to 1 into the reinforcement learning agent, and output the recommended next follow-up interval. S6. Generate and output a comprehensive evaluation report that includes the dynamic risk score, the inflection point detection result list, and the recommended next follow-up interval.
2. The gastric cancer risk assessment method based on peripheral blood multi-omics dynamic analysis according to claim 1, characterized in that, The average rate of change before the inflection point in S3 is calculated as follows: Calculate the rate of change data for each time point from the start time point to the inflection point of the time series data. Sum the rate of change data for each time point and divide by the number of time points to obtain the average rate of change before the inflection point.
3. The gastric cancer risk assessment method based on peripheral blood multi-omics dynamic analysis according to claim 1, characterized in that, The risk assessment model in S4 is an ensemble learning model based on gradient boosting trees. The training samples of this model include longitudinal multi-omics data of each subject in the historical cohort and the corresponding gastric cancer outcome labels.
4. The gastric cancer risk assessment method based on peripheral blood multi-omics dynamic analysis according to claim 1, characterized in that, The reinforcement learning agent in S5 adopts a deep Q-network structure. Its state space includes dynamic risk score, inflection point marker and maximum rate of change after the inflection point, action space is a preset discrete follow-up interval value, and reward function is set as a weighted sum of negative missed diagnosis risk loss and detection cost.
5. The gastric cancer risk assessment method based on peripheral blood multi-omics dynamic analysis according to claim 4, characterized in that, The decision rule for the next follow-up interval recommended in S5 is as follows: When an inflection point exists and the maximum rate of change after the inflection point is greater than 0.1, output an interval of 30 to 90 days; When an inflection point exists and the maximum rate of change after the inflection point is no greater than 0.1, output an interval of 90 to 180 days. When there is no inflection point and the dynamic risk score is lower than the preset risk threshold, output an interval of 180 to 365 days. When there is no inflection point and the dynamic risk score is not lower than the preset risk threshold, output an interval of 90 to 180 days.
6. The gastric cancer risk assessment method based on peripheral blood multi-omics dynamic analysis according to claim 1, characterized in that, The protein biomarker group includes at least one selected from OLFM4, ECM1, GSN, IGFBP2, and IQGAP1; The set of methylated DNA markers includes at least one selected from ZNF154, CDO1, and ELMO1; The metabolic biomarker group includes at least one of adenosine, fumarate, tryptophan, and acylcarnitine.
7. The gastric cancer risk assessment method based on peripheral blood multi-omics dynamic analysis according to claim 1, characterized in that, The inflection point detection result list in S6 includes the name of each inflection point marker with a value of 1, the location of the inflection point in acceleration change, the rate of change after the inflection point, and the acceleration after the inflection point.
8. The gastric cancer risk assessment method based on peripheral blood multi-omics dynamic analysis according to claim 1, characterized in that, It also includes S7: when the recommended next follow-up interval is less than 90 days, an emergency warning sign is added to the comprehensive assessment report.
9. A gastric cancer risk assessment system based on peripheral blood multi-omics dynamic analysis, used to implement the gastric cancer risk assessment method based on peripheral blood multi-omics dynamic analysis as described in any one of claims 1 to 8, characterized in that, include: The longitudinal multi-omics data acquisition module is used to acquire peripheral blood multi-omics detection datasets of the same subject at at least four consecutive time points. The multi-omics detection datasets include protein biomarkers, methylated DNA biomarkers, and metabolic biomarkers at each time point, and the sampling intervals between the time points may be unequal. The acceleration inflection point detection and marker generation module is used to extract the detection value and corresponding timestamp of each marker in the multi-omics detection dataset to construct time series data. The first and second derivatives of the time series data are calculated sequentially using the numerical difference method to obtain the rate of change data and acceleration data of each marker at each time point. The time point when the absolute value of the second derivative changes from below the first threshold to above the second threshold is determined as the acceleration change inflection point position of the marker. An inflection point marker is generated for each marker, wherein the inflection point marker of the marker that is detected is assigned a value of 1, and the inflection point marker of the marker that is not detected is assigned a value of 0. The inflection point post-trend feature extraction module is used to extract all detection values after the inflection point time point from the time series data of each marker for which the inflection point marker is assigned a value of 1. Linear regression fitting is used to obtain the rate of change and acceleration after the inflection point, and the difference in the rate of change before and after the inflection point is obtained by subtracting the average rate of change before the inflection point from the rate of change after the inflection point. The dynamic risk score calculation module is used to input the current detection value, inflection point marker, rate of change after the inflection point, and acceleration after the inflection point of each marker into the pre-trained risk assessment model and output a dynamic risk score. The personalized follow-up interval recommendation module is used to input the dynamic risk score, the existence of an inflection point flag of at least one marker with a value of 1, and the maximum value of the rate of change after the inflection point among all markers with an inflection point flag of 1 into the reinforcement learning agent, and output the recommended next follow-up interval. The comprehensive assessment report generation module is used to generate and output a comprehensive assessment report that includes the dynamic risk score, the inflection point detection result list, and the recommended next follow-up interval.