Early gastric cancer ESD (Electro-Static Discharge) incomplete cure risk early warning method based on big data technology

By processing and evaluating multi-source heterogeneous data after ESD surgery in early gastric cancer based on big data technology, the problems of inconsistent data processing and high complexity of risk assessment models in the existing technology are solved, and more accurate and timely risk warning support is achieved.

CN119964831APending Publication Date: 2025-05-09ZHEJIANG CANCER HOSPITAL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510051931.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art has problems such as inconsistent multi-source heterogeneous data processing, difficulty in capturing the dynamic characteristics of disease progression, high computational complexity and insufficient real-time performance in postoperative risk warning of ESD in early gastric cancer, resulting in insufficient warning timeliness and low risk quantification accuracy.

Method used

Using a method based on big data technology, structured feature tensors are constructed and feature dimensionality reduction is performed by acquiring and preprocessing CT image data, blood index data and ESD postoperative pathological data. Then, preliminary risk rating is performed based on the preset feature template, and a secondary evaluation is performed using historical case similarity analysis to calculate the probability of non-complete cure positive and generate a risk warning report.

Benefits of technology

It realizes unified processing and efficient feature extraction of multi-source heterogeneous medical data, which can accurately capture the dynamic characteristics of disease progression, reduces the computational complexity and improves real-timeness, and provides more accurate and timely risk warning support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964831A_ABST
    Figure CN119964831A_ABST
Patent Text Reader

Abstract

The invention discloses an early gastric cancer ESD (Electro-Static Discharge) incomplete cure risk early warning method based on a big data technology, and relates to the technical field of medical big data analysis, and the method comprises the steps: obtaining first related data of a test sample, and carrying out the preprocessing; inputting the preprocessed first related data into a risk prediction unit; the risk prediction unit comprises a pre-screening module and an evaluation module, the pre-screening module performs preliminary risk grading based on a preset feature template, and the evaluation module performs secondary evaluation on the samples with the preliminary risk grading as high risk by adopting historical case similarity analysis; and according to the evaluation result of the evaluation module, calculating the non-complete cure positive probability of the test sample, and generating a risk early warning report based on the non-complete cure positive probability. According to the method, the double-module design of the pre-screening module and the evaluation module is adopted, and the limitation of a traditional single-index evaluation method is overcome through calculation of a morphological risk index and a dynamic change index and a risk level dynamic adjustment mechanism based on historical case similarity analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical big data analysis, and specifically to a risk warning method for early-stage gastric cancer ESD incomplete cure based on big data technology. Background Art

[0002] Endoscopic submucosal dissection (ESD), as a minimally invasive treatment for early gastric cancer, has made significant technological progress in the past two decades, and its cure rate and postoperative survival rate have been greatly improved. With the in-depth application of artificial intelligence and big data technologies in the medical field, surgical prognosis evaluation based on multimodal medical data has gradually become a research hotspot. At present, risk assessment is mainly carried out by preoperative CT imaging examination, intraoperative endoscopic real-time observation, and postoperative pathological evaluation in clinical practice. Traditional evaluation methods are usually based on imaging feature extraction technology, such as lesion segmentation and feature recognition based on deep learning, or statistical methods are used to analyze the dynamic changes of blood indicators. However, these methods are often limited to the analysis of a single data source, lack the comprehensive utilization of multi-source heterogeneous medical data, and have problems with computational efficiency and accuracy when processing high-dimensional and multi-time series medical data.

[0003] Existing technologies still face multiple challenges in early gastric cancer post-ESD risk warning: First, there is a lack of a unified framework for standardized processing and feature extraction of multi-source heterogeneous data such as medical images, blood indicators and pathological data; second, existing risk assessment models mostly use static thresholds or simple linear combination methods, which are difficult to accurately capture the dynamic characteristics of disease progression; third, traditional methods have high computational complexity when processing high-dimensional medical features, lack of real-time performance, and lack of effective use of historical case data. Especially when dealing with borderline cases, due to the lack of a flexible risk level adjustment mechanism, misjudgment or missed judgment is prone to occur. In addition, existing methods generally have problems such as insufficient warning timeliness and low risk quantification accuracy, making it difficult to provide timely and reliable decision support for clinicians. Summary of the invention

[0004] In view of the above-mentioned problems, the present invention is proposed.

[0005] Therefore, the present invention provides a risk warning method for incomplete cure of ESD in early gastric cancer based on big data technology, which can solve the problems mentioned in the background technology.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: an early gastric cancer ESD incomplete cure risk warning method based on big data technology, comprising: obtaining first relevant data of a test sample and preprocessing it; the first relevant data comprises CT image data, blood index data and ESD postoperative pathological data; the preprocessed first relevant data is input into a risk prediction unit; the risk prediction unit comprises a pre-screening module and an evaluation module, the pre-screening module performs preliminary risk grading based on a preset feature template, and the evaluation module performs a secondary evaluation on the samples with the preliminary risk grading as high risk by using a historical case similarity analysis; according to the evaluation results of the evaluation module, the incomplete cure positive probability of the test sample is calculated, and a risk warning report is generated based on the incomplete cure positive probability.

[0007] As a preferred solution of the early gastric cancer ESD incomplete cure risk warning method based on big data technology described in the present invention, the first relevant data of the test sample is obtained, including the following steps: obtaining a medical image sequence output by a CT image acquisition device as the CT image data; obtaining the values ​​of various blood indicators in a blood test report as the blood indicator data; obtaining the scanning imaging results of post-ESD pathological sections as the post-ESD pathological data.

[0008] As a preferred scheme of the early gastric cancer ESD non-complete cure risk warning method based on big data technology described in the present invention, wherein: the inputting of the first relevant data after preprocessing into the risk prediction unit includes the following steps: constructing the first relevant data into a structured feature tensor; performing feature dimension reduction processing on the structured feature tensor; weighted integration of the reduced-dimensional features according to a preset weight coefficient to generate a unified feature representation, and inputting the unified feature representation into the pre-screening module and the evaluation module respectively; wherein the structured feature tensor includes CT image feature dimensions, blood index feature dimensions and pathological feature dimensions, the CT image feature dimensions correspond to space-time series information, the blood index feature dimensions correspond to multi-time point detection values, and the pathological feature dimensions correspond to multi-scale pathological image features.

[0009] As a preferred solution of the early gastric cancer ESD non-complete cure risk warning method based on big data technology described in the present invention, wherein: the preset feature template includes a first preset template and a second preset template; the first preset template performs morphological feature modeling for the CT image information and pathological information in the unified feature representation; the pre-screening module performs feature standardization processing on the CT image information and pathological information in the unified feature representation based on the first preset template to obtain a standardized value, and determines the morphological risk index in combination with the anatomical structure weight and the time correlation weight; the second preset template performs time series change feature modeling for the blood index information in the unified feature representation; the pre-screening module performs time series feature extraction on the blood index information in the unified feature representation based on the second preset template, and determines the dynamic change index in combination with the change characteristics of the detection values ​​at multiple time points.

[0010] As a preferred solution of the early gastric cancer ESD non-complete cure risk warning method based on big data technology described in the present invention, wherein: the morphological risk index and the dynamic change index are subjected to feature fusion and time evolution analysis to obtain a preliminary risk score, specifically, the morphological risk index and the dynamic change index are subjected to feature fusion and time evolution analysis to obtain the preliminary risk score; if the preliminary risk score is greater than the risk warning threshold, the corresponding test sample is marked as a high-risk sample, otherwise it is marked as a low-risk sample; the confidence score is calculated for the high-risk sample, and the confidence score is calculated by the following method: The steps are as follows: calculating the quality score of the high-risk sample; the quality score is determined based on the variance weight of the feature dimension and the weighted combination of the feature variance; calculating the data consistency value of the high-risk sample; the data consistency value is determined based on the degree of deviation between each feature value and its weighted average value; calculating the spatiotemporal uncertainty value of the high-risk sample; the spatiotemporal uncertainty value is determined based on the standard deviation of the time dimension and the space dimension and their corresponding weights; multiplying the quality score, the data consistency value and the spatiotemporal uncertainty value to obtain a confidence score, and if the confidence score is greater than a first preset threshold, a risk warning is triggered.

[0011] As a preferred solution of the early gastric cancer ESD non-complete cure risk warning method based on big data technology described in the present invention, the evaluation module uses historical case similarity analysis to perform a secondary evaluation on the samples that are initially risk-classified as high-risk, including the following steps: determining a similarity score based on the cosine similarity between the feature vector of the high-risk sample and the feature vector of each sample in the historical case database.

[0012] The samples in the historical case database are sorted in descending order according to the similarity score, and K historical cases with the highest similarity scores in the sorting results are selected as the reference sample set, where K is a second preset value set based on the total number of samples.

[0013] Based on the actual treatment results of each historical case in the reference sample set, the proportion of incompletely cured cases in the reference sample set is calculated to obtain the pathogenicity probability of similar historical cases.

[0014] The confidence level of the pathogenicity probability of the historical similar cases is calibrated; when the pathogenicity probability of the historical similar cases is greater than the first risk threshold, and the average morphological risk index of the non-completely cured cases in the reference sample set is greater than the morphological risk index of the high-risk samples, if the average dynamic change index of the non-completely cured cases in the reference sample set is greater than the dynamic change index of the high-risk samples, the warning level of the high-risk samples is maintained; if the average dynamic change index of the non-completely cured cases in the reference sample set is less than or equal to the dynamic change index of the high-risk samples, the warning level of the high-risk samples is reduced to medium risk.

[0015] When the probability of pathogenicity of the historically similar cases is less than or equal to the first risk threshold or the average morphological risk index of the non-completely cured cases in the reference sample set is less than or equal to the morphological risk index of the high-risk samples, if the probability of pathogenicity of the historically similar cases is greater than the second risk threshold and the average similarity score of the reference sample set is greater than the similarity threshold, the warning level of the high-risk samples is reduced to medium risk; if the probability of pathogenicity of the historically similar cases is less than or equal to the second risk threshold or the average similarity score of the reference sample set is less than or equal to the similarity threshold, the warning level of the high-risk samples is reduced to low risk.

[0016] As a preferred scheme of the early gastric cancer ESD incomplete cure risk warning method based on big data technology described in the present invention, the calculation of the positive probability of incomplete cure includes the following steps: obtaining the basic probability value of the test sample according to the preset probability mapping relationship corresponding to the warning level of the evaluation module; the preset probability mapping relationship is that the high-risk warning level corresponds to the first basic probability value, the medium-risk warning level corresponds to the second basic probability value, and the low-risk warning level corresponds to the third basic probability value; the basic probability value and the pathogenic probability of similar historical cases are weighted averaged to obtain the positive probability of incomplete cure of the test sample.

[0017] To further solve the above technical problems, the present invention provides the following technical solutions: an early gastric cancer ESD incomplete cure risk warning system based on big data technology, comprising: an acquisition and processing unit, used to obtain the first relevant data of the test sample and perform preprocessing; a risk prediction unit, including a pre-screening module and an evaluation module, the pre-screening module performs preliminary risk classification based on a preset feature template, and the evaluation module uses historical case similarity analysis to perform a secondary evaluation of the samples with the preliminary risk classification as high risk; a risk warning unit, used to calculate the incomplete cure positive probability of the test sample according to the evaluation result of the evaluation module, and generate a risk warning report based on the incomplete cure positive probability.

[0018] A computer device includes a memory and a processor, wherein the memory stores a computer program, and is characterized in that when the processor executes the computer program, the steps of the early gastric cancer ESD incomplete cure risk warning method based on big data technology as described above are implemented.

[0019] A computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of the early gastric cancer ESD incomplete cure risk warning method based on big data technology as described above are implemented.

[0020] Beneficial effects of the present invention: The present invention performs multi-dimensional preprocessing of CT image data, blood index data and ESD postoperative pathological data through a data processing unit, effectively solving the quality and standardization problems of multi-source heterogeneous medical data; in the risk prediction link, a dual-module design is adopted, in which the pre-screening module realizes the preliminary screening of high-risk samples through the calculation of morphological risk index and dynamic change index, and the evaluation module establishes a complete risk level dynamic adjustment mechanism based on historical case similarity analysis, overcoming the limitations of the traditional single indicator evaluation method; in the final risk assessment stage, through the preset probability mapping relationship and weighted average strategy, the accurate quantification of the risk of incomplete cure is achieved, providing clinicians with a more valuable decision support tool. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0022] Figure 1 This is a schematic diagram of the overall process of the early gastric cancer ESD incomplete cure risk warning method based on big data technology proposed by the present invention;

[0023] Figure 2 This is a workflow diagram of the risk prediction unit in the early gastric cancer ESD incomplete cure risk warning method based on big data technology proposed in the present invention.

[0024] Figure 3 This is the overall structure diagram of the early gastric cancer ESD incomplete cure risk warning system based on big data technology proposed by the present invention.

[0025] Figure 4 This is a diagram of the computer equipment in the early gastric cancer ESD incomplete cure risk warning method based on big data technology proposed by the present invention. DETAILED DESCRIPTION

[0026] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.

[0027] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0028] Example 1, reference Figure 1 and Figure 2 , which is an embodiment of the present invention, provides a risk warning method for incomplete cure of ESD in early gastric cancer based on big data technology.

[0029] In the relevant technologies, firstly, there is a lack of a unified framework for the standardized processing and feature extraction of multi-source heterogeneous data such as medical images, blood indicators and pathological data; secondly, the existing risk assessment models mostly use static thresholds or simple linear combination methods, which are difficult to accurately capture the dynamic characteristics of disease progression; thirdly, traditional methods have high computational complexity and lack of real-time performance when processing high-dimensional medical features, and lack effective use of historical case data. Especially when dealing with borderline cases, due to the lack of a flexible risk level adjustment mechanism, misjudgment or missed judgment is prone to occur. In addition, existing methods generally have problems such as insufficient warning timeliness and low risk quantification accuracy, making it difficult to provide timely and reliable decision support for clinicians.

[0030] The present application provides an effective solution to the above-mentioned problems. Next, multiple embodiments will be combined to explain in detail how to implement the early gastric cancer ESD incomplete cure risk warning method based on big data technology.

[0031] Figure 1 The overall flow chart of the early gastric cancer ESD incomplete cure risk warning method based on big data technology is shown, including:

[0032] S1: Acquire first relevant data of a test sample, and pre-process the relevant data through a data processing unit.

[0033] Specifically, the data processing unit includes a data cleaning module and a feature standardization module. The first related data includes CT image data, blood index data and ESD postoperative pathology data.

[0034] In a specific embodiment of the present invention, when obtaining the first relevant data of the test sample, the patient's CT image data is first obtained through the hospital's PACS (Picture Archiving and Communication Systems) system. The obtained CT image sequence includes scanned images in three directions: cross-section, sagittal and coronal planes. The scanning layer thickness is 1.25 mm, the layer spacing is 1.0 mm, and the matrix size is 512×512 pixels. For the acquisition of blood index data, the patient's blood test report within 7 days before surgery is exported through the hospital's LIS (Laboratory Information System) system, including but not limited to: tumor markers such as carcinoembryonic antigen (CEA), carbohydrate antigen 199 (CA199), alpha-fetoprotein (AFP), and inflammatory factor indicators such as white blood cell count (WBC), neutrophil percentage (NEU%), and C-reactive protein (CRP). The acquisition of postoperative pathological data after ESD is performed by scanning paraffin-embedded tissue slices through a digital slice scanner, and the scanning resolution is set to 0.25 μm / pixel, and the standardized pathological diagnosis report issued by the pathology department is collected at the same time.

[0035] In the data preprocessing stage, the data processing unit first evaluates and cleans the three types of data obtained. For CT image data, an image quality assessment method based on histogram analysis is used to eliminate low-quality images caused by patient movement, metal artifacts, etc., and an anisotropic diffusion filter algorithm is used to perform image denoising and enhancement. For blood index data, an outlier detection method based on the interquartile range (IQR) is used to identify and process outliers, and missing values ​​are reasonably filled according to the normal value range recommended by clinical guidelines. For pathological data, an adaptive threshold segmentation algorithm is used to remove background and extract regions of interest from digital slice images, and image registration technology is used to spatially align pathological images obtained at different magnifications. Subsequently, the feature standardization module of the data processing unit unifies the preprocessed data: the grayscale values ​​of CT images are uniformly normalized to the [0,1] interval, the blood index data are standardized according to the clinical reference range of each index, and the pathological image features are standardized by the Z-score method to achieve comparability between different samples. In addition, a strict data quality control process has been established, including data integrity check, format consistency verification and processing log recording, to ensure the reliability and traceability of preprocessing results.

[0036] S2: Input the preprocessed first relevant data into the risk prediction unit.

[0037] S2.1: Construct the first related data into a structured feature tensor.

[0038] Specifically, the structured feature tensor includes CT image feature dimension, blood index feature dimension and pathological feature dimension. Among them, the CT image feature dimension corresponds to space-time sequence information, the blood index feature dimension corresponds to multi-time point detection values, and the pathological feature dimension corresponds to multi-scale pathological image features.

[0039] It should be noted that in the present invention, the design of the structured feature tensor is optimized for the characteristics of postoperative prediction of early gastric cancer ESD. The CT image feature dimension contains spatial information and temporal information. By comparing the CT scan results at different time points, the changing characteristics of the lesion area can be effectively captured. The blood index feature dimension records the detection values ​​at multiple time points within 7 days before surgery to construct a characteristic curve reflecting the changes in the body state. The pathological feature dimension adopts a multi-scale analysis method, combined with the tissue structure characteristics under a 40x microscope and the cytological characteristics under a 400x microscope, to achieve feature extraction at the tissue and cell levels. This feature tensor design method can more completely reflect the patient's pathophysiological state.

[0040] S2.2: Perform feature dimensionality reduction on the structured feature tensor, including using principal component analysis to reduce feature redundancy, using autoencoders to extract feature representations, and using the t-SNE algorithm for nonlinear dimensionality reduction.

[0041] Specifically, the present invention adopts a three-step dimensionality reduction strategy: first, principal component analysis is used to reduce the linear correlation between features; then feature compression is performed through an autoencoder network, in which the encoder part adopts a residual connection structure to maintain important feature information; finally, the t-SNE algorithm is used for nonlinear dimensionality reduction, and the Barnes-Hut algorithm is used to improve computational efficiency, so that the present invention can meet the needs of real-time processing of new samples. This dimensionality reduction strategy not only maintains the discriminability of features, but also improves the processing efficiency of risk prediction.

[0042] S2.3: Perform weighted integration on the reduced-dimensional features according to preset weight coefficients to generate a unified feature representation, and input the unified feature representation into the pre-screening module and the evaluation module respectively.

[0043] It should be noted that in the present invention, the preset weight coefficients are determined by an adaptive weighting mechanism. The weight coefficients include γ1 (CT image feature weight), γ2 (blood index feature weight) and γ3 (pathological feature weight), satisfying γ1+γ2+γ3=1. The weight determination process includes two steps: first, the initial weights are determined on the validation set by the Bayesian optimization method; then, dynamic adjustments are made according to the prediction performance of each feature dimension. Practice shows that when γ1∈[0.4,0.5], γ2∈[0.2,0.3], and γ3∈[0.3,0.4], the prediction results are relatively stable. When the feature quality of a certain dimension is low (such as the presence of artifacts in CT images), this method reduces the weight of the corresponding dimension and appropriately increases the weights of other dimensions to maintain the reliability of the prediction results. This weight adjustment mechanism can adapt to prediction needs under different feature quality conditions.

[0044] Furthermore, Figure 2 The risk prediction unit includes a pre-screening module and an evaluation module. The pre-screening module performs preliminary risk classification based on a preset feature template.

[0045] Specifically, the preset feature template includes a first preset template and a second preset template. The first preset template performs morphological feature modeling for CT image information and pathological information in the unified feature representation, and the pre-screening module performs feature normalization processing on the CT image information and pathological information in the unified feature representation based on the first preset template to obtain a standardized value, and combines the anatomical structure weight and the time correlation weight to determine the morphological risk index M:

[0046]

[0047] Among them, α i represents the adaptive weight of the importance of the i-th feature, f i is the measured value of the ith medical feature, μ iis the historical statistical mean of the i-th feature, σ i is the standard deviation of the ith feature, S(d i ) is the anatomical structure space weight function, T i (t) is the time correlation function.

[0048] Preferably, the calculation formula of the morphological risk index M is obtained by introducing the adaptive weight α i With standardized items The square terms are combined and the anatomical structure space weight S(d i ) and the time correlation weight T i (t) is used as a modulation factor to realize the comprehensive evaluation of CT images and pathological characteristics in spatial and temporal dimensions. Compared with the traditional method of only considering a single eigenvalue, this formula can more accurately reflect the spatial distribution characteristics and temporal evolution characteristics of the lesions.

[0049] Among them, the anatomical structure spatial weight is calculated based on the lesion location information and anatomical structure characteristics:

[0050]

[0051] Among them, d i represents the Euclidean distance from the feature to the lesion center, D m is the medical space influence range constant, L(p i ) is the anatomical position characteristic function, p i is the anatomical space coordinate.

[0052] Preferably, the calculation of the anatomical structure spatial weight adopts the product of the distance-based exponential decay function and the anatomical position characteristic function, through The term realizes the continuous and smooth attenuation of weight with distance, avoids the mutation effect caused by the discrete threshold in the traditional method, and improves the rationality of spatial weight allocation.

[0053] According to the anatomical structure weight and position correlation, the anatomical position feature is calculated:

[0054]

[0055] Among them, ω j represents the weight coefficient of the jth anatomical structure, R ij is the correlation between position i and anatomical structure j, and k is the number of key anatomical structures.

[0056] Preferably, the anatomical position characteristic function L(p i ) The Sigmoid function is used to nonlinearly combine the effects of multiple anatomical structures. This method realizes the adaptive weighting of the importance of different anatomical structures, overcomes the limitation of the traditional method that the weights of anatomical structures are fixed, and makes the calculation of position features more in line with medical practice.

[0057] Perform time-effectiveness analysis on the standardized values ​​and calculate the time correlation function T i (t):

[0058]

[0059] Among them, η i Represents the data type timeliness coefficient, θ i is the time sensitivity parameter, τ i is the time decay constant, Δt i is the time interval.

[0060] Preferably, the time correlation function T i (t) By introducing the data type timeliness coefficient η i , time sensitivity parameter θ i and the time decay constant τ i , constructing (1+θ i ·Δt i )and The composite form of time decay model can accurately model the timeliness characteristics of different types of medical data and has better adaptability than the traditional linear time decay model.

[0061] The second preset template models the time series variation characteristics of the blood index information in the unified feature representation.

[0062] The pre-screening module extracts the time series features of the blood index information in the unified feature representation based on the second preset template, and determines the dynamic change index D by combining the change characteristics of the detection values ​​at multiple time points:

[0063]

[0064] Where N represents the total number of tests, b is the index number of the blood index, m is the total number of blood indexes, ψ b represents the importance weight of the bth blood index, v b,c is the detection value of the bth indicator at the cth time point, s c is the cth detection time point, P b is the indicator correlation coefficient.

[0065] Preferably, the calculation formula of the dynamic change index introduces a normalization term of 1 / N and combines ψ b The root mean square of the time series change rate is calculated to achieve robust extraction of the change characteristics of blood indicators at multiple time points, avoiding the shortcomings of traditional methods that only consider the changes in the starting and ending points while ignoring the intermediate processes.

[0066] Analyze the indicator change trends in multiple time windows and calculate the indicator correlation coefficient P b :

[0067]

[0068] Among them, w is the index number of the time window, W is the total number of time windows, μ w represents the weight coefficient of the w-th time window, U b,w is the changing trend coefficient of the bth indicator in the wth time window, and sign is the sign function.

[0069] Preferably, the correlation coefficient of the indicator is calculated using a signed power function |U b,w | 1 / 2 , by introducing the time window weight μ w It achieves adaptive fusion of changing trends at different time scales, overcoming the problem that traditional methods are insensitive to sudden changes.

[0070] The morphological risk index and dynamic change index were subjected to feature fusion and time evolution analysis, and the preliminary risk score was calculated according to the following formula:

[0071]

[0072] Among them, λ k Represents the weight coefficient of the k-th feature, F k is the comprehensive index value of the kth type of feature (when k = 1, that is, F1, corresponding to the morphological risk index M; when k = 2, that is, F2, corresponding to the dynamic change index D), H(F) is the feature synergy function, and V(t) is the time evolution function.

[0073] Preferably, the calculation of the preliminary risk score adopts the geometric mean form of the sum of squares of features, and through the modulation of feature synergy function and time evolution function, it realizes the nonlinear fusion of features from different sources, avoiding the feature masking effect in the traditional linear weighted method.

[0074] Analyze the correlation and synergy between features and calculate the feature synergy function:

[0075]

[0076] Among them, ∈ pq represents the correlation coefficient between features p and q, δ pq is the synergy coefficient of the feature pair.

[0077] Preferably, by introducing the correlation coefficient ∈ pq and the synergy coefficient δ pq, a quantitative description model of the interaction between features was established, which overcame the limitation of traditional methods that ignore the interaction of features and improved the accuracy of risk assessment.

[0078] According to the time series changes of characteristic indicators, the time evolution function is calculated:

[0079]

[0080] in, represents the change rate weight of the rth indicator, Δu r (z) is the time series change rate of the rth indicator, n r is the number of evolution indicators.

[0081] Furthermore, the confidence score is calculated for the high-risk samples, and the confidence score is obtained through the following calculation steps:

[0082] Calculate the quality score of high-risk samples, which is determined based on the weighted combination of the variance weights of the feature dimensions and the feature variance;

[0083] Calculate the data consistency value of high-risk samples, which is determined based on the degree of deviation of each eigenvalue from its weighted average;

[0084] Calculate the spatiotemporal uncertainty value of high-risk samples, which is determined based on the standard deviation of the time dimension and the space dimension and their corresponding weights;

[0085] The quality score, data consistency value and spatiotemporal uncertainty value are multiplied together to obtain a confidence score. If the confidence score is greater than a first preset threshold, a risk warning is triggered.

[0086] It should be noted that the present invention realizes accurate early warning of the risk of incomplete cure after ESD surgery for early gastric cancer by constructing a dual-index evaluation system of morphological risk index and dynamic change index. Among them, the morphological risk index realizes dynamic evaluation of CT images and pathological characteristics by introducing adaptive weights and space-time modulation factors; the dynamic change index uses a multi-time point sequence analysis method to extract the change characteristics of blood indicators; the two indicators are integrated based on a nonlinear feature fusion mechanism, and the modulation of the feature synergy function and the time evolution function is coordinated, which not only improves the accuracy and reliability of risk assessment, but also realizes early identification and early warning of high-risk cases, providing important decision-making support for clinicians to formulate personalized treatment plans and follow-up strategies.

[0087] Furthermore, the evaluation module uses historical case similarity analysis to conduct a secondary evaluation of samples that are initially classified as high risk.

[0088] S2.4: Calculate the cosine similarity between the feature vector of the high-risk sample and the feature vector of each sample in the historical case database to obtain a similarity score.

[0089] S2.5: Sort the samples in the historical case database in descending order according to the similarity scores, and select the K historical cases with the highest similarity scores in the sorting results as the reference sample set, where K is a second preset value set based on the total number of samples.

[0090] S2.6: Based on the actual treatment results of each historical case in the reference sample set, calculate the proportion of incompletely cured cases in the reference sample set and obtain the pathogenicity probability of similar historical cases.

[0091] S2.7: Confidence calibration is performed on the pathogenicity probability of similar historical cases, where when the pathogenicity probability of similar historical cases is greater than the first risk threshold, and the average morphological risk index of non-completely cured cases in the reference sample set is greater than the morphological risk index of high-risk samples, if the average dynamic change index of non-completely cured cases in the reference sample set is greater than the dynamic change index of high-risk samples, the warning level of high-risk samples is maintained; if the average dynamic change index of non-completely cured cases in the reference sample set is less than or equal to the dynamic change index of high-risk samples, the warning level of high-risk samples is reduced to medium risk.

[0092] When the probability of pathogenicity of historically similar cases is less than or equal to the first risk threshold or the average morphological risk index of non-completely cured cases in the reference sample set is less than or equal to the morphological risk index of high-risk samples, if the probability of pathogenicity of historically similar cases is greater than the second risk threshold and the average similarity score of the reference sample set is greater than the similarity threshold, the warning level of the high-risk samples is reduced to medium risk; if the probability of pathogenicity of historically similar cases is less than or equal to the second risk threshold or the average similarity score of the reference sample set is less than or equal to the similarity threshold, the warning level of the high-risk samples is reduced to low risk.

[0093] Preferably, the secondary evaluation mechanism based on the similarity of historical cases proposed in the present invention realizes the accurate screening of high-risk samples through cosine similarity calculation and the construction of dynamic reference sample set; especially in the confidence calibration link, by performing multi-dimensional comparison of morphological risk index, dynamic change index and pathogenicity probability of similar historical cases, a complete set of dynamic adjustment mechanism of risk level is established, which not only improves the reliability of evaluation results, but also realizes the refined grading of risk warning, provides clinicians with more valuable decision-making basis, effectively reduces the misjudgment rate, and improves the accuracy of prognosis evaluation after ESD surgery for early gastric cancer.

[0094] S3: According to the evaluation results of the evaluation module, the non-complete cure positive probability of the test sample is calculated, and a risk warning report is generated based on the non-complete cure positive probability.

[0095] Specifically, the risk warning report includes specific screening recommendations to avoid secondary surgery and analysis of key influencing factors.

[0096] S3.1: According to the preset probability mapping relationship corresponding to the warning level of the assessment module, the basic probability value of the test sample is obtained. The preset probability mapping relationship is that the high risk warning level corresponds to the first basic probability value, the medium risk warning level corresponds to the second basic probability value, and the low risk warning level corresponds to the third basic probability value.

[0097] S3.2: Take the weighted average of the basic probability value and the pathogenicity probability of similar historical cases to obtain the non-complete cure positive probability of the test sample.

[0098] In summary, the present invention performs multi-dimensional preprocessing of CT image data, blood index data and ESD postoperative pathological data through a data processing unit, effectively solving the quality and standardization problems of multi-source heterogeneous medical data; in the risk prediction link, a dual-module design is adopted, in which the pre-screening module realizes the preliminary screening of high-risk samples through the calculation of morphological risk index and dynamic change index, and the evaluation module establishes a complete risk level dynamic adjustment mechanism based on the historical case similarity analysis, overcoming the limitations of the traditional single indicator evaluation method; in the final risk assessment stage, the preset probability mapping relationship and weighted average strategy are used to achieve accurate quantification of the risk of incomplete cure, providing clinicians with a more valuable decision support tool.

[0099] Example 2, reference Figure 3 , which is an embodiment of the present invention, provides an early gastric cancer ESD incomplete cure risk warning system based on big data technology, including:

[0100] An acquisition and processing unit, used for acquiring first relevant data of the test sample and performing preprocessing;

[0101] The risk prediction unit includes a pre-screening module and an evaluation module. The pre-screening module performs preliminary risk classification based on a preset feature template, and the evaluation module performs a secondary evaluation on samples that are initially classified as high risk using a historical case similarity analysis.

[0102] The risk warning unit is used to calculate the non-complete cure positive probability of the test sample according to the evaluation results of the evaluation module, and generate a risk warning report based on the non-complete cure positive probability.

[0103] Example 3, reference Figure 4, is an embodiment of the present invention, which is different from the previous embodiment in that: if the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0104] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0105] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0106] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0107] Example 4 is an embodiment of the present invention, which provides an early gastric cancer ESD incomplete cure risk warning method based on big data technology. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0108] To verify the effectiveness of the present invention, this embodiment collected a total of 300 cases of early gastric cancer ESD surgery from 2020 to 2023 in the Department of Gastroenterology of a tertiary hospital as an experimental data set. The experiment adopted a five-fold cross-validation method to randomly divide the data set into a training set (240 cases) and a test set (60 cases). The comparative experiment selected a risk assessment method based on a single CT image feature (Scheme A) and a dynamic monitoring method that only considers blood indicators (Scheme B) as the control group. In the feature extraction stage, this embodiment performed standardized preprocessing on the CT image data, including image denoising, spatial alignment and feature normalization; performed time series analysis and outlier processing on the blood indicator data; and used a multi-scale feature extraction method for pathological data. In the risk prediction stage, the screening accuracy of the pre-screening module was optimized by adjusting the feature weight coefficient and the risk threshold; at the same time, a dynamic reference sample set construction strategy was adopted in the evaluation module to improve the reliability of risk level determination.

[0109] During the experiment, this embodiment focused on the following evaluation indicators: risk warning accuracy, high-risk sample identification rate, warning timeliness and clinical interpretability. For each test sample, its actual treatment results and predicted results were recorded, and various performance indicators were calculated. At the same time, this embodiment also performed a time series analysis of the warning results to evaluate the warning lead time and stability of the method. All experiments were conducted in the same hardware environment (Intel Core i7 processor, 32GB memory, NVIDIA RTX 3080 graphics card), implemented in the Python 3.7 programming environment, and model training and evaluation were performed using machine learning frameworks such as sklearn and pytorch.

[0110] Table 1 Comparison of performance indicators of different methods

[0111]

[0112] As shown in Table 1, the present invention is superior to the control group in all key performance indicators. Specifically, the risk warning accuracy rate reaches 92.5%, which is 14.2 and 16.7 percentage points higher than Scheme A and Scheme B respectively; in terms of the high-risk sample recognition rate, the present invention reaches 88.7%, which is significantly better than the 70% to 72% level of the control group; at the same time, the false positive rate of the present invention is only 7.3%, which is more than half lower than that of the control group; in terms of warning timeliness, the present invention can issue a risk warning 12.5 days in advance, which is 4 to 5 days earlier than the control group; the missed reporting rate is reduced to 4.8%, which is only about one-third of that of the control group. These data fully demonstrate the superiority of the present invention in risk warning after ESD surgery for early gastric cancer.

[0113] Table 2 Analysis of prediction accuracy at different risk levels

[0114] Risk Level Sample size Correct predictions Prediction accuracy (%) Recall rate (%) Confidence score High risk 85 79 92.9 94.2 0.89 Medium risk 125 114 91.2 90.8 0.85 Low risk 90 85 94.4 93.5 0.87

[0115] As shown in Table 2, the present invention shows high prediction accuracy for samples of different risk levels. 79 out of 85 samples in the high-risk group were correctly predicted, with a prediction accuracy of 92.9% and a recall rate of 94.2%; 114 out of 125 samples in the medium-risk group were correctly predicted, with a prediction accuracy of 91.2% and a recall rate of 90.8%; 85 out of 90 samples in the low-risk group were correctly predicted, with a prediction accuracy of 94.4% and a recall rate of 93.5%. The confidence scores of the three risk levels are all above 0.85, indicating that the prediction results have high credibility. It is particularly noteworthy that for the high-risk group, which is of greatest clinical concern, the present invention shows the highest recall rate (94.2%), which is of great significance for early detection of patients with potential risks. The data also show that the difference in prediction performance between different risk levels is small (standard deviation <2%), indicating that the present invention has good stability and universality.

[0116] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A risk warning method for incomplete cure of ESD in early gastric cancer based on big data technology, characterized in that: include: Acquire first relevant data of the test sample and perform preprocessing; the first relevant data includes CT image data, blood index data and ESD postoperative pathological data; Inputting the preprocessed first relevant data into a risk prediction unit; the risk prediction unit includes a pre-screening module and an evaluation module, the pre-screening module performs preliminary risk classification based on a preset feature template, and the evaluation module performs a secondary evaluation on the samples that are initially classified as high risk by using a historical case similarity analysis; According to the evaluation result of the evaluation module, the incomplete cure positive probability of the test sample is calculated, and a risk warning report is generated based on the incomplete cure positive probability.

2. The early gastric cancer ESD incomplete cure risk early warning method based on big data technology as claimed in claim 1, characterized in that: The step of obtaining the first relevant data of the test sample comprises the following steps: Acquire a medical image sequence output by a CT image acquisition device as the CT image data; Obtaining the values ​​of various blood indexes in the blood test report as the blood index data; The scanning imaging results of the post-ESD pathological sections are obtained as the post-ESD pathological data.

3. The early gastric cancer ESD incomplete cure risk early warning method based on big data technology as claimed in claim 2, characterized in that: The step of inputting the preprocessed first relevant data into a risk prediction unit comprises the following steps: constructing the first relevant data into a structured feature tensor; Performing feature dimensionality reduction processing on the structured feature tensor; The features after dimensionality reduction are weighted and integrated according to a preset weight coefficient to generate a unified feature representation, and the unified feature representation is input into the pre-screening module and the evaluation module respectively; Among them, the structured feature tensor includes CT image feature dimension, blood index feature dimension and pathological feature dimension, the CT image feature dimension corresponds to space-time series information, the blood index feature dimension corresponds to multi-time point detection value, and the pathological feature dimension corresponds to multi-scale pathological image features.

4. The early gastric cancer ESD incomplete cure risk early warning method based on big data technology as claimed in claim 3, characterized in that: The preset feature template includes a first preset template and a second preset template; The first preset template performs morphological feature modeling for the CT image information and pathological information in the unified feature representation; The pre-screening module performs feature normalization processing on the CT image information and pathological information in the unified feature representation based on the first preset template to obtain a normalized value, and determines a morphological risk index in combination with an anatomical structure weight and a time correlation weight; The second preset template performs temporal variation feature modeling for the blood index information in the unified feature representation; The pre-screening module extracts time series features from the blood index information in the unified feature representation based on the second preset template, and determines a dynamic change index in combination with change features of detection values ​​at multiple time points.

5. The early gastric cancer ESD incomplete cure risk early warning method based on big data technology as claimed in claim 4, characterized in that: Performing feature fusion and time evolution analysis on the morphological risk index and the dynamic change index to obtain a preliminary risk score, specifically performing feature fusion and time evolution analysis on the morphological risk index and the dynamic change index to obtain the preliminary risk score; If the preliminary risk score is greater than the risk warning threshold, the corresponding test sample is marked as a high-risk sample, otherwise it is marked as a low-risk sample; A confidence score is calculated for the high-risk sample, and the confidence score is obtained by the following calculation steps: Calculating a quality score of the high-risk sample; the quality score is determined based on a weighted combination of the variance weight of the feature dimension and the feature variance; Calculating the data consistency value of the high-risk sample; the data consistency value is determined based on the degree of deviation between each characteristic value and its weighted average value; Calculating the spatiotemporal uncertainty value of the high-risk sample; the spatiotemporal uncertainty value is determined based on the standard deviation of the time dimension and the space dimension and their corresponding weights; The quality score, the data consistency value and the spatiotemporal uncertainty value are multiplied to obtain a confidence score. If the confidence score is greater than a first preset threshold, a risk warning is triggered.

6. The early gastric cancer ESD incomplete cure risk early warning method based on big data technology as claimed in claim 5, characterized in that: The evaluation module uses historical case similarity analysis to perform a secondary evaluation on the samples that are initially classified as high risk, including the following steps: Determining a similarity score based on the cosine similarity between the feature vector of the high-risk sample and the feature vector of each sample in the historical case database; Sorting the samples in the historical case database in descending order according to the similarity score, and selecting K historical cases with the highest similarity scores in the sorting results as the reference sample set, where K is a second preset value set based on the total number of samples; Based on the actual treatment results of each historical case in the reference sample set, the proportion of incompletely cured cases in the reference sample set is calculated to obtain the disease probability of similar historical cases; Calibrate the confidence level of the pathogenicity probability of similar historical cases; When the pathogenicity probability of the historical similar cases is greater than the first risk threshold, and the average morphological risk index of the non-completely cured cases in the reference sample set is greater than the morphological risk index of the high-risk samples, if the average dynamic change index of the non-completely cured cases in the reference sample set is greater than the dynamic change index of the high-risk samples, the warning level of the high-risk samples is maintained; If the average dynamic change index of the non-completely cured cases in the reference sample set is less than or equal to the dynamic change index of the high-risk sample, the warning level of the high-risk sample is reduced to medium risk; When the pathogenicity probability of the historical similar cases is less than or equal to the first risk threshold or the average morphological risk index of the non-completely cured cases in the reference sample set is less than or equal to the morphological risk index of the high-risk samples, if the pathogenicity probability of the historical similar cases is greater than the second risk threshold and the average similarity score of the reference sample set is greater than the similarity threshold, the warning level of the high-risk samples is reduced to medium risk; If the pathogenic probability of the historical similar cases is less than or equal to the second risk threshold or the average similarity score of the reference sample set is less than or equal to the similarity threshold, the warning level of the high-risk sample is reduced to low risk.

7. The early gastric cancer ESD incomplete cure risk early warning method based on big data technology as claimed in claim 6, characterized in that: The calculation of the non-complete cure positive probability comprises the following steps: According to the preset probability mapping relationship between the warning level of the evaluation module and the warning level of the evaluation module, the basic probability value of the test sample is obtained; the preset probability mapping relationship is that the high risk warning level corresponds to the first basic probability value, the medium risk warning level corresponds to the second basic probability value, and the low risk warning level corresponds to the third basic probability value; The basic probability value and the pathogenic probability of similar historical cases are weighted averaged to obtain the incomplete cure positive probability of the test sample.

8. An early gastric cancer ESD incomplete cure risk early warning system based on big data technology, based on the early gastric cancer ESD incomplete cure risk early warning method based on big data technology according to any one of claims 1 to 7, characterized in that: include, An acquisition and processing unit, used for acquiring first relevant data of the test sample and performing preprocessing; The risk prediction unit includes a pre-screening module and an evaluation module. The pre-screening module performs preliminary risk classification based on a preset feature template, and the evaluation module performs a secondary evaluation on samples that are initially classified as high risk using a historical case similarity analysis. A risk warning unit is used to calculate the incomplete cure positive probability of the test sample according to the evaluation result of the evaluation module, and generate a risk warning report based on the incomplete cure positive probability.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the early gastric cancer ESD incomplete cure risk warning method based on big data technology are implemented as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the early gastric cancer ESD incomplete cure risk warning method based on big data technology are implemented as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Alarm precision improving method and device, equipment, storage medium and product

    CN120473065A