A data processing model and its construction method for gastric cancer risk assessment based on dynamic changes in circulating tumor DNA ac4C.

CN122575490APending Publication Date: 2026-08-14NORTHERN JIANGSU PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0006]本发明目的在于针对现有胃癌风险评估的数据处理模型生物标志物或临床指标单一,难以精准捕捉早期胃癌的生物学特性的不足,提供一种基于循环肿瘤DNA ac4C动态变化的胃癌风险评估的数据处理模型及其构建方法,解决现有早期胃癌评估技术有创性、依从性低、早期判断准确性差以及缺乏动态监测能力等问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575490A_ABST
    Figure CN122575490A_ABST
Patent Text Reader

Abstract

This invention relates to a data processing model and its construction method for gastric cancer risk assessment based on dynamic changes in circulating tumor DNA (ct4C), belonging to the field of medical testing technology. The method includes: obtaining blood samples from study subjects at at least two different time points and extracting plasma ctDNA; detecting ct4C modification levels using ct4C RNA immunoprecipitation sequencing technology; processing the data to obtain relative ct4C levels and calculating their dynamic change rate; integrating and screening this change rate with clinical characteristics (such as tumor stage) to obtain a core feature set; and training and validating the gastric cancer risk assessment model based on this feature set using a machine learning algorithm (such as XGBoost). The model constructed in this invention integrates dynamic characteristics of ctDNA ct4C with clinical indicators, achieving non-invasive, highly sensitive, and highly specific early risk screening for gastric cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical testing technology, specifically relating to a data processing model for gastric cancer risk assessment based on dynamic changes in circulating tumor DNA ac4C and its construction method. Background Technology

[0002] Gastric cancer is a highly prevalent malignant tumor worldwide, and early screening is crucial for improving patient prognosis. Currently used clinical methods for assessing gastric cancer have significant limitations: while gastroscopy is the gold standard for screening, it is an invasive procedure with poor patient compliance and makes it difficult to dynamically monitor disease progression; traditional tumor markers such as carcinoembryonic antigen (CEA) and carbohydrate antigen 19-9 (CA19-9) have low sensitivity and specificity in early-stage gastric cancer, failing to meet the needs of early clinical screening.

[0003] Circulating tumor DNA (ctDNA), as an emerging liquid biopsy biopsy marker, has shown potential in early tumor screening and prognostic assessment. Existing studies have confirmed the value of ctDNA in predicting postoperative recurrence risk and monitoring treatment response in gastric cancer patients, but most studies focus on gene mutations or epigenetic modifications of ctDNA at single time points, without exploring in depth the dynamic changes of specific modifications (such as ac4C) in the early development of gastric cancer.

[0004] N4-Acetylcytidine (ac4C), a key RNA modification, plays an important role in tumorigenesis and development. Previous studies have shown that ac4C modification levels are significantly elevated in gastric cancer tissues and are associated with high NAT10 expression. However, current research on ac4C modification in ctDNA is still in its early stages. Its dynamic changes in the early development of gastric cancer are not yet clear, and there is a lack of early assessment models based on this modification, resulting in its clinical application value not being fully explored.

[0005] Furthermore, existing data processing models for gastric cancer risk assessment mostly rely on single biomarkers or clinical indicators, failing to achieve multi-dimensional integration of ctDNA modification characteristics and clinicopathological parameters. This makes it difficult to accurately capture the biological characteristics of early gastric cancer, thus hindering the improvement of screening accuracy. Therefore, developing early risk assessment technology based on dynamic changes in ctDNA ac4C has become a key direction for solving the current dilemma in early gastric cancer screening. Summary of the Invention

[0006] The purpose of this invention is to address the shortcomings of existing data processing models for gastric cancer risk assessment, which rely on single biomarkers or clinical indicators and are difficult to accurately capture the biological characteristics of early gastric cancer. This invention provides a data processing model for gastric cancer risk assessment based on the dynamic changes of circulating tumor DNA ac4C and its construction method, thereby solving the problems of existing early gastric cancer assessment technologies, such as invasiveness, low compliance, poor accuracy in early judgment, and lack of dynamic monitoring capabilities.

[0007] This invention provides a method for constructing a data processing model for gastric cancer risk assessment based on dynamic changes in circulating tumor DNA (ctDNA). The gastric cancer risk assessment model is constructed by using machine learning methods based on the detection of ac4C modification levels and dynamic changes in circulating tumor DNA (ctDNA) in blood samples.

[0008] Furthermore, the detection of the ac4C modification level employs ac4C RNA immunoprecipitation sequencing technology. Specifically, the ac4C RNA immunoprecipitation sequencing technology involves: incubating an ac4C-specific antibody with a ctDNA fragment overnight at 4°C; capturing the antibody-ctDNA complex with magnetic beads and eluting to obtain ac4C-modified ctDNA; constructing a sequencing library with an insert fragment of 300-500 bp; and performing high-throughput sequencing at a sequencing depth of ≥50×. This specific step clarifies the key parameters for antibody incubation, magnetic bead capture, library construction, and sequencing depth. The 300-500 bp insert fragment is suitable for the short fragment characteristics of cfDNA, and the sequencing depth of ≥50× ensures the reliable detection of low-frequency modification signals. These specific process conditions collectively guarantee the specificity and sensitivity of ac4C modification detection, which is the key technology for reproducibly achieving the effects of this invention.

[0009] Furthermore, the dynamic changes in ac4C are obtained by calculating the rate of change of the relative levels of ac4C at different time points. By transforming the ac4C modification level information at multiple time points into a derived feature reflecting its changing trend—the "ac4C dynamic change rate"—it is possible to capture temporal biological information related to the occurrence, development, or treatment response of gastric cancer. Compared to detection relying on static levels at a single time point, this dynamic feature can more sensitively indicate disease state transitions and is key to improving the model's ability to identify early gastric cancer and disease progression.

[0010] Furthermore, the dynamic rate of change of ac4C is calculated using the formula (relative level of ac4C at the next time point - relative level of ac4C at the previous time point) / relative level of ac4C at the previous time point. Using this standardized formula to calculate the rate of change eliminates differences in baseline ac4C levels among individuals, quantitatively and normally representing the magnitude of ac4C modification changes over time. This calculation method makes dynamic features comparable across different samples, providing standardized and quantifiable input for subsequent effective integration with machine learning models, and is a crucial technical step in ensuring model stability and repeatability.

[0011] Furthermore, this includes the following steps: (1) Blood samples were obtained from the study subjects at at least two different time points. Plasma ctDNA was extracted and its ac4C modification level was detected. A prospective cohort design was adopted, selecting patients with pathologically confirmed gastric cancer in stages T1-T2 as the case group and healthy individuals undergoing physical examinations during the same period as the control group. Fasting venous blood was collected from both groups of study subjects at at least two different time points, including before the first sampling and at the time of sampling. Plasma was separated and stored at low temperature for later use. These multiple time points are generally taken before sampling, at the time of sampling (i.e., at the time of clinical diagnosis), and after intervention, including but not limited to 1, 3, and 6 months after intervention. Other time points (such as 2 months, 4 months, etc.) can also be selected according to clinical monitoring needs. By setting multiple time points such as "before the first sampling and at the time of sampling", continuous dynamic data reflecting the occurrence and development of the disease can be obtained, which provides a data basis for the subsequent calculation of the key feature of "ac4C dynamic change rate". This is a prerequisite for realizing dynamic risk assessment and can capture richer disease time-series information compared to single-time-point detection.

[0012] ctDNA was extracted from plasma using magnetic beads and quantified to a concentration ≥10 ng / μL. The ac4C modification level of ctDNA was detected using ac4C RNA immunoprecipitation sequencing. Maintaining a ctDNA concentration ≥10 ng / μL ensured sufficient starting material for subsequent sequencing and guaranteed data reliability. This specific technique of ac4C RNA immunoprecipitation sequencing can directly and accurately capture and quantify the specific epigenetic modification ac4C on ctDNA, providing a novel and specific biomarker input for the model.

[0013] (2) Process the sequencing data to calculate the relative level of ac4C and the dynamic change rate of ac4C for each sample; process the raw sequencing data obtained in step (1), including quality control, sequence alignment, deduplication and standardization, calculate the relative level of ac4C for each sample, and calculate the dynamic change rate of ac4C based on the relative level of ac4C at different time points; process the raw data through standardized bioinformatics processes (such as FastQC, Bowtie2, Samtools), which can effectively remove technical noise and obtain clean, comparable ac4C modified data. Calculating the "dynamic change rate of ac4C" transforms the absolute level information at multiple time points into a derived feature that reflects the trend of change. This feature can more sensitively indicate the biological dynamics related to the disease process and is the key to achieving high sensitivity of the model.

[0014] The data processing workflow is as follows: FastQC quality control → Trimmomatic removal of low-quality reads with Q < 20 → Bowtie2 alignment to the human genome hg38 → Samtools deduplication → calculation of ac4C relative levels. A standardized and reproducible bioinformatics analysis chain from raw data to final information is disclosed. Using FastQC and other methods for quality control and low-quality read removal ensures data quality; alignment to the latest reference genome hg38 improves localization accuracy; and Samtools deduplication eliminates bias introduced by PCR amplification. This disclosed and specific workflow ensures the accuracy and comparability of the calculated "ac4C relative levels".

[0015] (3) Integrate and screen ac4C-related features with clinical features to obtain a core feature set; screen the dynamic change rate of ac4C obtained in step (2), as well as clinical variables such as age, gender, and tumor stage; use t-test / Mann-Whitney U test for dynamic change rate of ac4C and age, and chi-square test / Fisher exact probability test for gender and tumor stage for univariate screening, and then use LASSO regression and XGBoost algorithm for multivariate feature selection, eliminate collinear variables, and retain features that are significantly related to early gastric cancer screening (screening criterion P<0.05). The two-step method of "univariate statistical test + multivariate machine learning screening" is adopted. First, irrelevant variables are quickly eliminated, and then LASSO, XGBoost and other methods are used to automatically select the most predictive and low collinear feature combination from high-dimensional data. This method takes into account both screening efficiency and scientificity, can effectively prevent overfitting, ensure that the features that finally enter the model are concise and strongly correlated, and improve the generalization ability and interpretability of the model. Among them, continuous features and numerical features are the same concept, including the ac4C dynamic rate of change and age.

[0016] By integrating ctDNA ac4C levels and dynamic change rates with clinical indicators such as age, sex, and tumor stage, a composite assessment feature was constructed. Furthermore, by integrating novel molecular dynamic features (ac4C levels and change rates) with traditional clinicopathological indicators (age, sex, and stage), a "multi-dimensional composite feature" was created. This integration fully leverages the complementary nature of information from different types of data, enabling a comprehensive characterization of disease status from both molecular biological dynamics and clinical macroscopic phenotypes, thereby significantly improving the comprehensiveness and accuracy of risk assessment.

[0017] (4) Based on the core feature set, a gastric cancer risk assessment model is trained and constructed using machine learning algorithms; continuous features are standardized or normalized, nonlinear features are transformed using logarithmic transformation or Box-Cox transformation, and categorical variables are encoded using one-hot encoding or label encoding; by standardizing, transforming, or encoding features of different scales and types, all features are unified into a numerical space suitable for machine learning algorithms. This step eliminates the influence of dimensions, improves the distribution of data, and in particular makes nonlinear relationships easier for the model to learn, thus laying a technical foundation for the efficient training and stable convergence of the subsequent model.

[0018] The importance of features is evaluated using XGBoost and Random Forest algorithms to identify the features that contribute most to the model's predictions. By utilizing the feature importance evaluation function built into the tree model, the core features that contribute most to the prediction results can be quantitatively identified. This not only helps to understand the model's internal decision-making logic and increase the model's credibility (interpretability), but also further verifies the effectiveness of feature selection in step (4) and may provide data support for clinical focus areas.

[0019] Based on the screening and construction of multi-dimensional composite features, namely ctDNA ac4C level and ac4C dynamic change rate, and molecular dynamic features such as age, gender, and tumor stage, plus clinical indicators, models were constructed using logistic regression, support vector machine, random forest, and XGBoost machine learning algorithms, respectively. The model with the best performance was selected as the final data processing model for gastric cancer risk assessment through comparative evaluation.

[0020] The model with the area under the curve (AUC) is primarily evaluated and selected using 10-fold cross-validation. Multiple machine learning algorithms with different principles (linear, nonlinear, and ensemble learning) are constructed and compared in parallel. 10-fold cross-validation robustly evaluates model performance within the training set, avoiding biases caused by limitations of a single algorithm or randomness in data partitioning. Using AUC as the core evaluation metric comprehensively measures the model's overall classification performance under different decision thresholds, thus scientifically and objectively selecting the model with the best generalization ability.

[0021] (5) Model Validation. The optimal model was externally validated on an independent validation set. The predictive performance of the model was evaluated using ROC curves, calibration curves, and the Hosmer-Lemeshow test. Regularization and ensemble learning methods were used to validate the model, resulting in a stable and reproducible data processing model for gastric cancer risk assessment. External validation on data independent of the training set is the "gold standard" for evaluating the real-world performance of the model. Through multi-dimensional evaluation using ROC curves, calibration curves, and the HL test, not only was the model's discriminative ability (AUC) verified, but also the accuracy of its predicted probabilities (calibration) and overall goodness of fit were verified. Regularization and other optimization methods were used to prevent overfitting, ultimately resulting in a stable, reliable, and reproducible data processing tool applicable to new data.

[0022] Furthermore, in step (3), the screening includes univariate screening using t-test or Mann-Whitney U test, and multivariate feature selection using LASSO regression.

[0023] Furthermore, in step (4), the machine learning algorithm includes at least one of XGBoost, random forest, support vector machine, or logistic regression.

[0024] Furthermore, in step (4), the machine learning algorithm is XGBoost. As a highly efficient gradient boosting ensemble learning algorithm, XGBoost excels at handling mixed-type features, capturing complex nonlinear relationships, and effectively preventing overfitting. Its selection as the optimal model proves that the feature set constructed in this invention, combined with advanced algorithms, can produce outstanding technical results (high AUC, sensitivity, and specificity).

[0025] Furthermore, the core feature set obtained in step (3) includes: the relative level of ctDNA ac4C at sampling time, the dynamic change rate of ac4C, and the tumor stage. The core feature combination obtained after screening and construction is thus clarified. This combination integrates three key information categories: "current state" (level at sampling time), "dynamic change" (change rate), and "clinical macro-indicators" (stage), representing the preferred solution for achieving optimal prediction results with the fewest features. This set of features is one of the core technical contributions of this invention, and its effectiveness has been confirmed by the high performance of the model in the embodiments.

[0026] Secondly, this invention also provides a data processing model for gastric cancer risk assessment based on dynamic changes in circulating tumor DNA ac4C, constructed by any of the aforementioned construction methods. This claim protects the product directly obtained by the aforementioned methods—namely, the specific, operable data processing model itself. This model is the final output and value carrier of all the aforementioned technical steps, the direct and unique product of the aforementioned methods, and a software entity containing specific algorithms, parameters, and feature weights. Protecting this model means protecting its direct deployment and application value without having to perform a complex construction process again.

[0027] Beneficial effects: (1) The gastric cancer risk assessment model constructed by the present invention has high-precision screening capability and integrates dynamic characteristics of ctDNA ac4C and clinical indicators in multiple dimensions. According to literature reports, the sensitivity of CEA for assessing early gastric cancer is about 30%, and that of CA19-9 is about 35%. The validation set of this model has a sensitivity of 88.38% and a specificity of 94.23%, which is significantly better than traditional tumor markers and meets the needs of high sensitivity and high specificity for early screening.

[0028] (2) This invention is the first to systematically reveal the dynamic changes of ctDNA ac4C in the early development of gastric cancer. Preliminary experiments show that the ac4C level in early gastric cancer patients is significantly higher than that in healthy controls (P<0.05), providing a new biomarker for early screening.

[0029] (3) The data processing model for gastric cancer risk assessment constructed using the method of this invention has significant clinical translational value. The non-invasive detection method improves patient compliance, and dynamic monitoring can reflect the treatment effect and recurrence risk in real time, providing a basis for decision-making for individualized treatment and is expected to reduce the rate of missed diagnosis and mortality of gastric cancer. Attached Figure Description

[0030] Figure 1 A graph showing the dynamic trend of ctDNA ac4C in gastric cancer patients at different time points; Figure 2 The ROC curve is shown for the data processing model of gastric cancer risk assessment based on ctDNA ac4C, where: training set (249 cases, accounting for 70% of the total sample size of 356 cases)6; validation set (107 cases, accounting for 30%), and the reference line is the randomized disease test (AUC=0.5). Detailed Implementation

[0031] The technical solution of the present invention will be described in detail below through specific embodiments, but the scope of protection of the present invention is not limited to the embodiments described. Experiment 1: Detection and Data Acquisition of Dynamic Changes in ctDNA ac4C

[0032] 1. Research Subjects A prospective cohort design was adopted, selecting patients with early-stage gastric cancer (T1-T2 stage) diagnosed at Subei People's Hospital in Jiangsu Province from July 2025 to June 2026 and healthy controls as the study subjects.

[0033] Inclusion criteria: age ≥18 years; diagnosed with gastric cancer by endoscopic biopsy or surgery; patient voluntarily participates and signs informed consent form; able to conduct regular follow-up and sample collection.

[0034] Exclusion criteria: history of other malignant tumors; presence of serious comorbidities (such as severe heart disease, liver or kidney dysfunction, etc.); having received chemotherapy or radiotherapy within the past three months; participation in other clinical trials; inability to cooperate with follow-up.

[0035] The study subjects were divided into two groups: Early gastric cancer group (T1-T2 stage): 70 cases, all diagnosed by endoscopic biopsy or surgical pathology, and meeting the inclusion criteria (age ≥18 years, voluntary participation, etc.), excluding patients with other malignant tumors, severe heart, liver and kidney dysfunction, etc.

[0036] Healthy control group: 70 people, who were undergoing physical examinations during the same period, excluded from those with a history of gastrointestinal diseases and cancer, and matched for gender and age with the gastric cancer group.

[0037] All participants signed informed consent forms, and this study was approved by the hospital's ethics committee.

[0038] 2. Sample Collection and Processing Collection time points: Fasting venous blood was collected at multiple time points, including before the first sampling, during sampling, and after the intervention. The plasma was separated and stored at low temperature.

[0039] Procedure: Collect 10 mL of fasting venous blood, inject it into an EDTA anticoagulant tube, centrifuge at 3000 rpm for 10 minutes within 2 hours to separate the plasma, aliquot it into 2 mL cryovials, and store at -80℃ in an ultra-low temperature freezer to avoid repeated freeze-thaw cycles.

[0040] 3. Detection of ctDNA ac4C level ctDNA extraction: The MagUltra Circulating Cell-free DNA Isolation Kit (catalog number N913-01) from Nanjing Novizan Biotechnology Co., Ltd. was used to extract ctDNA from plasma according to the instructions. The ctDNA concentration in the plasma was typically between 10 ng / μL and 50 ng / μL.

[0041] ac4C modification detection: ac4C RNA immunoprecipitation sequencing technology (ac4C-RIP-seq) was used, and the specific steps included: ① The N4-acetylglucosinolate (ac4C) monoclonal antibody (Abcam product number ab230544) was used to incubate the ctDNA obtained in the previous steps at 4°C overnight.

[0042] ② The magnetic beads capture the antibody-ctDNA complex, and the elution yields ac4C-modified ctDNA; ③ Construct sequencing libraries (insertion fragment length 300-500bp), and perform high-throughput sequencing using the Illumina NovaSeq 6000 platform with a sequencing depth ≥50×.

[0043] 4. Data Preprocessing Quality control was performed using FastQC software. Low-quality reads (Q<20) were removed using Trimmomatic. Bowtie2 was aligned to the human genome hg38, and duplicates were removed using Samtools. Finally, the relative ac4C level for each sample was calculated. Results are shown below. Figure 1 ,Depend on Figure 1 It is evident that the relative levels of ctDNA ac4C in patients with T1-T2 stage gastric cancer were significantly higher than those in the healthy control group before and during sampling. After treatment, the levels gradually decreased over time, indicating that ctDNA ac4C levels can dynamically reflect the disease status of early gastric cancer and can serve as an effective indicator for early screening and efficacy monitoring of gastric cancer.

[0044] The formula for calculating the ac4C rate of change is: ac4C rate of change = (ac4C relative level at sampling time - ac4C relative level at the previous time point) / ac4C relative level at the previous time point. For example, in one embodiment, the previous time point is one week before sampling.

[0045] Experiment 2: Construction of a data processing model for gastric cancer risk assessment 1. Feature Engineering 1.1 Dataset and Grouping To expand the sample size for model building and validation, subjects meeting the same criteria were further included in the initial cohort, resulting in a total of 356 samples, including 178 cases in the early gastric cancer group and 178 cases in the healthy control group. Using a computer-generated random seed, the samples were divided in a 7:3 ratio, with 249 cases in the training set (125 gastric cancer cases and 124 healthy cases) and 107 cases in the validation set (53 gastric cancer cases and 54 healthy cases). These groups were used for model training and independent external validation.

[0046] 1.2 Feature Inclusion Five features were included: relative ctDNA ac4C level at sampling time, ac4C change rate from before the first sampling (e.g., 1 week before sampling) to the sampling time, age, sex, and tumor stage (T1 / T2).

[0047] 1.3 Feature Filtering ① Univariate analysis: Based on the relative level and rate of change of ac4C obtained in Example 1, t-tests were used to screen continuous variables and chi-square tests were used to screen categorical variables. The screening criterion was P<0.05 as statistically significant. The results showed that both the ac4C level (at sampling time) and the rate of change of ac4C (1 week before sampling to sampling time) were significantly associated with gastric cancer (P<0.05). ② Multivariate LASSO regression: The analysis was performed using the glmnet package in R, with gastric cancer status as the dependent variable and the above five features as independent variables. A 10-fold cross-validation was used to select the optimal λ value. Clinical indicators of age, gender, and tumor stage were included, and three core features were selected: ac4C level (coefficient 0.62), ac4C change rate (coefficient 0.35), and tumor stage (coefficient 0.21). The above selection results are consistent with the core composite features defined in claim 4.

[0048] 2. Model Training 2.1 Operating Environment System: Windows 10 64-bit; Programming language: Python 3.8; Dependencies: Scikit-learn 1.0.2, XGBoost 1.5.0.

[0049] 2.2 Data Input Format The training set is organized into a standard machine learning format: X_train is the sample feature matrix, with each column corresponding to 3 core features; y_train is the sample label, with gastric cancer marked as 1 and healthy as 0.

[0050] 2.3 Feature Preprocessing The features are standardized using StandardScaler with the formula (x − mean) / standard deviation, which unifies the feature dimensions to improve model stability.

[0051] 2.4 Construction and Parameter Setting of Four Models ① Logistic Regression: Parameters C=0.1, solver='liblinear', random_state=42; After 10-fold cross-validation, the training set AUC=0.76, sensitivity 72.15%, and specificity 80.33%; ② Random Forest: Parameters n_estimators=500, max_depth=8, random_state=42; After 10-fold cross-validation, the training set AUC=0.81, sensitivity 79.42%, and specificity 85.71%; ③ Support Vector Machine: Parameters were set as kernel='rbf', C=1.0, probability=True, random_state=42; after 10-fold cross-validation, the training set AUC=0.78, sensitivity 75.60%, and specificity 82.14%. ④XGBoost: Set parameters learning_rate=0.1, n_estimators=200, max_depth=5, random_state=42; After 10-fold cross-validation, the training set AUC=0.85, sensitivity 85.21%, and specificity 90.17%.

[0052] 2.5 Methods for obtaining AUC and ROC curves The gastric cancer prediction probabilities output by each algorithm and the true labels of the samples are input into the Scikit-learn library. The true positive rate (TPR) and false positive rate (FPR) are calculated using the roc_curve function, and the area under the curve, i.e., the AUC value, is calculated using the auc function. The ROC curve is plotted with FPR as the horizontal axis and TPR as the vertical axis.

[0053] 2.6 Optimal Model Selection Using AUC as the primary evaluation metric, while also considering sensitivity and specificity, the XGBoost gradient boosting tree algorithm was found to have the highest AUC value and the best overall performance in the training set. Therefore, the XGBoost ensemble learning model was selected as the final data processing model for gastric cancer risk assessment. Experiment 3: Validation of the data processing model for gastric cancer risk assessment

[0054] A general validation method for clinical machine learning models was adopted, using the standard toolkit of Python 3.8 and Scikit-learn 1.0.2 to evaluate the model performance on an independent validation set (107 cases). The specific steps are as follows: 1. Data Input The three core features of the 107 samples in the validation set (relative level of ac4C at sampling time, rate of change of ac4C from before sampling to sampling time, and tumor stage) were input into the trained XGBoost model.

[0055] 2. Predicted probability output Call the model's built-in function predict_proba() to output the predicted probability value for each sample as gastric cancer.

[0056] 3. Performance evaluation methods (conventional methods in this field) Using pathological screening results as the gold standard, the following indicators were calculated using conventional methods commonly used in this field: Sensitivity (True Positive Rate, TPR) Specificity (True Negative Rate, TNR) ROC curve AUC value Calibration curve The Hosmer-Lemeshow goodness-of-fit test (can be performed using the statsmodels.stats.diagnostic module or corresponding statistical functions) 4. Software and Functions (Completely Open Source) All evaluation metrics were implemented using the Python third-party library Scikit-learn: roc_curve(): Calculates and plots the ROC curve, such as... Figure 2 As shown: auc(): Calculates the AUC value confusion_matrix(): Calculates sensitivity and specificity. calibration_curve(): Plots the calibration curve hosmer_lemeshow_test(): Model fit test 5. Verification Results In the independent validation set (107 cases), the performance evaluation results of the XGBoost model are as follows: Sensitivity: 88.38% Specificity: 94.23% AUC: 0.82 (95%CI: 0.75-0.89) The calibration curves show good agreement between the predicted probability and the actual incidence. The Hosmer-Lemeshow test showed a p-value of 0.32, indicating that the model fit was satisfactory. Experiment 4: A System Integration Example of a Data Processing Model for Gastric Cancer Risk Assessment

[0057] To demonstrate the usability of the model, the trained XGBoost model is packaged and integrated. Specifically, it can be deployed as a standalone software module or integrated into an existing clinical information analysis system.

[0058] In this integrated example, the system receives input data containing three feature values: "relative level of ctDNA ac4C at sampling time," "dynamic change rate of ac4C," and "tumor stage." The system calls the model to perform calculations and finally outputs a gastric cancer risk probability value between 0 and 1.

[0059] This probability value can serve as an objective quantitative reference indicator to assist clinicians in making comprehensive assessments. The model's output results show a significant correlation with pathological examination results, validating its technical capability in identifying high-risk individuals.

[0060] As described above, although the invention has been shown and described with reference to specific preferred embodiments, it should not be construed as limiting the invention itself. Various changes in form and detail may be made without departing from the spirit and scope of the invention as defined in the appended claims.

Claims

1. A method for constructing a data processing model for gastric cancer risk assessment based on dynamic changes in circulating tumor DNA ac4C, characterized in that, Based on the ac4C modification level and dynamic changes of circulating tumor DNA (ctDNA) in blood samples, a gastric cancer risk assessment model was constructed using machine learning methods.

2. The construction method according to claim 1, characterized in that: The level of ac4C modification was detected using ac4C RNA immunoprecipitation sequencing technology.

3. The construction method according to claim 1 or 2, characterized in that: The dynamic changes of ac4C are obtained by calculating the rate of change of the relative level of ac4C at different time points.

4. The construction method according to claim 3, characterized in that: The dynamic rate of change of ac4C is calculated using the formula (relative level of ac4C at the next time point - relative level of ac4C at the previous time point) / relative level of ac4C at the previous time point.

5. The construction method according to claim 1, characterized in that, Includes the following steps: (1) Obtain blood samples from the subjects at at least two different time points, extract plasma ctDNA and detect its ac4C modification level; (2) Process the sequencing data and calculate the relative level of ac4C and the dynamic change rate of ac4C for each sample; (3) Integrate and screen ac4C-related features with clinical features to obtain a core feature set; (4) Based on the core feature set, a gastric cancer risk assessment model is trained and constructed using machine learning algorithms; (5) Validate the model.

6. The construction method according to claim 5, characterized in that, In step (3), the screening includes univariate screening using t-test or Mann-Whitney U test, and multivariate feature selection using LASSO regression.

7. The construction method according to claim 5 or 6, characterized in that, In step (4), the machine learning algorithm includes at least one of XGBoost, random forest, support vector machine or logistic regression.

8. The construction method according to claim 7, characterized in that, In step (4), the machine learning algorithm is XGBoost.

9. The construction method according to claim 5, characterized in that, The core feature set obtained in step (3) includes: the relative level of ctDNA ac4C at the time of sampling, the dynamic change rate of ac4C, and the tumor stage.

10. A data processing model for gastric cancer risk assessment based on dynamic changes in circulating tumor DNA ac4C, characterized in that, It is constructed by the construction method described in any one of claims 1-9.