Nonlinear survival risk modeling method based on pathology after neoadjuvant chemotherapy for gastric cancer
By enhancing the Cox model with Gaussian processes and applying local monotonic constraints, a nonlinear survival risk modeling method for pathological changes after neoadjuvant chemotherapy in gastric cancer is constructed. This method solves the problem of large prognostic differences in traditional stratification methods and achieves more detailed and stable risk stratification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG CANCER HOSPITAL
- Filing Date
- 2026-04-23
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies are insufficient to fully reflect the continuous risk gradient of pathological burden and treatment response after neoadjuvant chemotherapy for gastric cancer. Traditional discrete stratification methods result in large prognostic differences, and nonlinear relationships are not fully utilized, affecting the consistency and practicality of stratification results.
A Gaussian process-enhanced Cox model was used to construct locally comparable sample pairs using pathological data such as the number of lymph nodes dissected, the number of positive lymph nodes, the post-treatment pathological T stage, and the tumor regression grade. A continuous risk function was generated, and a stable contour-based ypTNM risk stratification label was generated through local monotonic constraints and boundary stability screening.
It achieves continuous risk expression and stable stratification of pathological information after neoadjuvant chemotherapy for gastric cancer, reduces information loss caused by discrete grouping, and improves the detail and stability of prognostic stratification.
Smart Images

Figure CN122091218A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of prognostic prediction model construction technology, and more specifically, to a nonlinear survival risk modeling method based on pathology after neoadjuvant chemotherapy for gastric cancer. Background Technology
[0002] After receiving neoadjuvant chemotherapy, gastric cancer patients experience changes in tumor burden, lymph node status, and treatment response. Postoperative pathological data becomes a crucial basis for prognostic assessment. Current clinical practice often employs methods such as post-treatment pathological staging, positive lymph node grouping, and tumor regression grading for risk assessment. However, these methods mostly use discrete stratification, primarily relying on fixed groupings or simple combinations of a few indicators for classification.
[0003] For gastric cancer patients undergoing neoadjuvant chemotherapy, there are often complex relationships between the number of lymph nodes dissected, the number of positive lymph nodes, the post-treatment pathological T stage, and the tumor regression grade. The impact of different variables on survival outcomes is not a simple linear change. When using traditional discrete stratification methods, it is easy to find that patients within the same stratum still have large differences in prognosis, making it difficult to fully reflect the continuous risk gradient corresponding to pathological burden and treatment response.
[0004] At the same time, if conventional survival analysis models are directly applied to this type of data, the nonlinear relationships between local samples and the stability of contour boundaries are easily overlooked, which in turn affects the consistency and practicality of the stratification results.
[0005] Therefore, there is still a need for a survival risk modeling method based on pathological information after neoadjuvant chemotherapy for gastric cancer, capable of jointly processing multiple pathological variables and generating continuous risk expression and stable stratification results. Summary of the Invention
[0006] This application provides a nonlinear survival risk modeling method based on pathology after neoadjuvant chemotherapy for gastric cancer, in order to at least solve some of the technical problems existing in the related technologies described above.
[0007] According to a first aspect of the embodiments of this application, a nonlinear survival risk modeling method based on pathology after neoadjuvant chemotherapy for gastric cancer is provided, comprising: Pathological and follow-up data of gastric cancer patients after neoadjuvant chemotherapy were obtained. The pathological data included at least the number of lymph nodes dissected, the number of positive lymph nodes, the pathological T stage after treatment, and the tumor regression grade. The pathological data were subjected to legality checks, missing data handling, scaling transformation, and ordered coding to generate case feature vectors. Based on the case feature vectors, locally comparable sample pairs are constructed, and a set of locally monotonic constraints is generated. The case feature vector, the follow-up data, and the set of local monotonic constraints are input into the Gaussian process-enhanced Cox model, and the model is trained based on the Cox partial likelihood term and the local monotonic violation term to obtain a continuous risk function. The subspace is divided according to the post-treatment pathological T stage and tumor regression grade. The continuous risk function is then sampled in a grid within each subspace to generate a continuous risk surface. Candidate contour boundaries are extracted based on the continuous risk surface, and stable contour boundaries are selected based on the boundary stability index obtained by resampling. The case to be evaluated is input into the continuous risk function to obtain a continuous risk score, and the profile-based ypTNM risk stratification label of the case to be evaluated is determined by combining the stable profile boundary.
[0008] As an optional approach, the construction of locally comparable sample pairs and the generation of a set of locally monotonic constraints includes: A local comparison window based on the transformation result of the number of lymph nodes dissected is set in the case feature vector; within the local comparison window where the post-treatment pathological T stage and tumor regression grade are consistent, a first type of sample pair is constructed according to the number of positive lymph nodes; within the local comparison window where the number of positive lymph nodes and tumor regression grade are consistent, a second type of sample pair is constructed according to the size of the post-treatment pathological T stage; within the local comparison window where the number of positive lymph nodes and post-treatment pathological T stage are consistent, a third type of sample pair is constructed according to the pathological risk corresponding to the tumor regression grade in ascending order; the first type of sample pair, the second type of sample pair, and the third type of sample pair are summarized into the local monotonic constraint set.
[0009] As an optional approach, the Gaussian process-enhanced Cox model includes an input module, a kernel mapping module, a risk function module, a survival objective module, a constraint regularization module, and a parameter update module. The input module receives case feature vectors; the kernel mapping module generates latent risk function values based on the similarity between case feature vectors; the risk function module outputs a continuous risk function; the survival objective module constructs a Cox partial likelihood term based on follow-up data; the constraint regularization module constructs a locally monotonically violating term based on a set of locally monotonically bounded constraints; and the parameter update module updates the model parameters based on the total loss function.
[0010] As an optional approach, the kernel mapping module uses a radial basis function kernel function and sets length scale parameters for each feature dimension in the case feature vector. Each feature dimension includes the transformation results of lymph node dissection count, the transformation results of positive lymph node count, the pathological T-phase coding results after treatment, and the coding results of tumor regression grading.
[0011] As an optional approach, the training based on the Cox partial likelihood term and the local monotonic violation term to obtain the continuous risk function includes: calculating the risk set corresponding to cases where the endpoint event occurs, and generating the Cox partial likelihood term; calculating the difference in risk function values for each sample pair in the local monotonic constraint set, and generating the local monotonic violation term based on the difference in risk function values; performing a weighted summation of the Cox partial likelihood term and the local monotonic violation term to generate the total loss function; and updating the parameters of the Gaussian process-enhanced Cox model based on the total loss function to obtain the continuous risk function.
[0012] As an optional approach, the local monotonic violation term is generated based on the continuous risk function values of the two corresponding cases and a preset safety interval for each sample; when the continuous risk function values of the two cases satisfy the local monotonic constraint relationship and the difference between the corresponding risk function values is not less than the preset safety interval, the local monotonic violation term for the sample is set to zero.
[0013] As an optional approach, the method involves dividing the space into subspaces according to the post-treatment pathological T phase and tumor regression grade, and performing grid sampling on the continuous risk function within each subspace to generate a continuous risk surface, including: The pathological T-phase coding results and tumor regression grade coding results after fixation treatment are used to establish a two-dimensional regular grid based on the transformation results of lymph node dissection number and positive lymph node number. The continuous risk function is called for each grid point to generate the risk matrix of the corresponding subspace. The risk matrices of each subspace are summarized into the continuous risk surface.
[0014] As an optional approach, the extraction of candidate contour boundaries based on the continuous risk surface includes: The number of training cases in the neighborhood of each grid point is counted. For grid points where the number of training cases in the neighborhood is lower than a preset lower limit, risk value smoothing is performed. A set of candidate risk thresholds is generated based on the distribution of continuous risk scores of training cases. For each candidate risk threshold, isorisk lines are extracted from the continuous risk surface to generate candidate contour boundaries.
[0015] As an optional approach, the selection of stable contour boundaries based on the boundary stability index obtained from resampling includes: The training cases are resampled multiple times. For each resampled case, the following processes are repeated: local monotonic constraint set construction, Gaussian process enhanced Cox model training, continuous risk surface generation, and candidate contour boundary extraction. The positional difference of each candidate contour boundary in multiple resamplings is calculated to generate a boundary stability index. Candidate contour boundaries whose boundary stability index meets preset conditions are determined as stable contour boundaries. The boundary stability index is generated based on the average symmetrical distance between the candidate contour boundary and the corresponding resampled contour boundary. The smaller the positional difference, the larger the corresponding boundary stability index.
[0016] According to a second aspect of the embodiments of this application, a nonlinear survival risk modeling system based on pathology after neoadjuvant chemotherapy for gastric cancer is also provided, comprising: The data acquisition and preprocessing module is used to acquire pathological data and follow-up data of gastric cancer patients after neoadjuvant chemotherapy. The pathological data includes at least the number of lymph nodes dissected, the number of positive lymph nodes, the pathological T stage after treatment, and the tumor regression grade. The pathological data is subjected to legality checks, missing data handling, scaling transformation, and ordered coding to generate case feature vectors. The constraint construction module is used to construct locally comparable sample pairs based on the case feature vector and generate a set of locally monotonic constraints. The model training module is used to input the case feature vector, the follow-up data, and the set of local monotonic constraints into the Gaussian process-enhanced Cox model, and train it based on the Cox partial likelihood term and the local monotonic violation term to obtain a continuous risk function; The surface generation module is used to divide the subspace according to the post-treatment pathological T phase and tumor regression grade, and to perform grid sampling on the continuous risk function in each subspace to generate a continuous risk surface; The boundary filtering module is used to extract candidate contour boundaries based on the continuous risk surface and to filter stable contour boundaries based on the boundary stability index obtained by resampling. The risk stratification module is used to input the case to be evaluated into the continuous risk function to obtain a continuous risk score, and to determine the contour-based ypTNM risk stratification label of the case to be evaluated in combination with the stable contour boundary.
[0017] According to a third aspect of the embodiments of this application, an electronic device is provided, including: a processor; a memory for storing a computer program executable by the processor; wherein the processor is configured to execute the computer program in the memory to implement the method described in the first aspect.
[0018] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided, which, when an executable computer program in the storage medium is executed by a processor, enables the implementation of the method described in the first aspect.
[0019] Compared with existing technologies, this invention uses the number of lymph nodes dissected, the number of positive lymph nodes, the post-treatment pathological T stage, and the tumor regression grade as joint inputs. It introduces a Gaussian process into the proportional hazards regression framework to construct a continuous risk function, and further extracts contour-based stratification results based on the continuous risk surface. This approach can more completely preserve the continuous information in the pathological variables, reduce the information loss caused by discrete grouping, and at the same time, reduce the impact of local fluctuations in the risk surface on the stratification results through local monotonic constraints and boundary stability screening, making the prognostic stratification of gastric cancer patients after neoadjuvant chemotherapy more detailed and stable.
[0020] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Furthermore, no embodiment in this disclosure is required to achieve all the effects described above. Attached Figure Description
[0021] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0022] Figure 1 This is a schematic diagram of a nonlinear survival risk modeling method based on pathology after neoadjuvant chemotherapy for gastric cancer, provided in an embodiment of this disclosure.
[0023] Figure 2 This is a schematic diagram illustrating the process of generating a set of locally monotonic constraints provided in an embodiment of this disclosure.
[0024] Figure 3 A schematic diagram of the structure of the Gaussian process-enhanced Cox model provided in an embodiment of this disclosure.
[0025] Figure 4 This is a schematic diagram of the process for screening stable contour boundaries provided in an embodiment of this disclosure.
[0026] Figure 5 This is a schematic diagram of the structure of a nonlinear survival risk modeling system based on pathology after neoadjuvant chemotherapy for gastric cancer, provided in an embodiment of this disclosure.
[0027] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0029] It should be noted that the information collected in this application (including but not limited to user device information, user personal information, collected data, used data, generated data, processed data, etc.) and the data (including but not limited to data used for analysis, stored data, displayed data, collected information, used information, generated information, processed information, etc.) are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.
[0030] In one implementation, the server can be deployed in a hospital information system, a research data platform, or a standalone electronic device. The server receives structured pathological data and follow-up data generated after neoadjuvant chemotherapy in gastric cancer patients. The pathological data includes at least the number of lymph nodes dissected, the number of positive lymph nodes, the post-treatment pathological T stage, and the tumor regression grade. The server sequentially performs data verification, feature processing, local monotonic constraint construction, Gaussian process enhanced proportional hazards regression model training, continuous risk surface generation, contour boundary screening, and individual case reasoning.
[0031] To facilitate program implementation, the server can pre-create a case data table, a feature data table, a constraint set table, a model parameter table, a risk surface table, a boundary table, and a results table. The case data table stores the original pathology fields and follow-up fields; the feature data table stores the features after scaling and ordered encoding; the constraint set table stores locally comparable sample pairs; the model parameter table stores kernel function parameters, constraint weight parameters, local comparison window parameters, sparse region determination parameters, and boundary screening parameters; the risk surface table stores continuous risk values obtained from grid sampling; the boundary table stores candidate contour boundaries and stable contour boundaries; and the results table stores continuous risk scores, contour-based ypTNM risk stratification labels, and boundary proximity markers.
[0032] The implementation process of the method described in this application will be described in detail below with reference to specific embodiments. It should be noted that this embodiment is only used to explain this application and is not intended to limit the scope of protection of this application. Conventional adjustments or substitutions of each step by those skilled in the art without departing from the concept of this application should be included in the scope of protection of this application.
[0033] Please see Figure 1 , Figure 1 This is a flowchart of a nonlinear survival risk modeling method based on pathology after neoadjuvant chemotherapy for gastric cancer according to an embodiment of the present invention, such as... Figure 1 As shown, the method includes steps S1-S6: In step S1, pathological data and follow-up data of gastric cancer patients after neoadjuvant chemotherapy are obtained. The pathological data includes at least the number of lymph nodes dissected, the number of positive lymph nodes, the pathological T stage after treatment, and the tumor regression grade. The pathological data is subjected to legality checks, missing data processing, scaling transformation, and ordered coding to generate a case feature vector.
[0034] Current assessment systems most commonly used after neoadjuvant therapy primarily rely on pathological staging and related discrete stratification methods. These methods depend heavily on discrete combinations of a few indicators, failing to adequately represent the risk gradient differences resulting from continuous changes in lymph node burden, and also struggling to describe the nonlinear relationship between treatment response variables and staging variables. Addressing these technical issues, this implementation uses information such as the number of lymph nodes dissected, the number of positive lymph nodes, the post-treatment pathological T stage, and tumor regression grade as inputs. It introduces a Gaussian process into a proportional hazards regression framework to output an individualized continuous risk score, further generating risk stratification results that are closer to clinical applications.
[0035] Specifically, after a case is received, the server first performs a raw field check, reading each case record one by one and verifying whether the number of lymph nodes dissected is a positive integer, the number of positive lymph nodes is a non-negative integer, and the number of positive lymph nodes is no greater than the number of lymph nodes dissected. Simultaneously, it checks whether the post-treatment pathological T-phase and tumor regression grade belong to the preset classification set. For records with obvious abnormalities such as field conflicts or missing items, the server writes them to an exception table and does not send them to the subsequent modeling stage. The preset classification set comes directly from the grading system used in the pathology report; in one example, the server can save the allowed post-treatment pathological T-phase set and tumor regression grade set in the configuration file, and match them item by item after receiving the data.
[0036] After completing field checks, the server performs missing data handling. The number of lymph nodes dissected and the number of positive lymph nodes together represent the lymph node burden; if either is missing, the case is immediately excluded from the training set. Post-treatment pathological T-phase and tumor regression grade are both ordered discrete variables. If a small number of missing values exist, the server can fill in the missing values using the median grade or the highest frequency grade from the same data source, the same statistical period, or the same subset of pathological templates. The server records the completion flag during completion for subsequent tracking. If the missing percentage of a discrete field exceeds a preset limit, the server also removes the corresponding case to avoid systematically disturbing the sorting relationship and kernel function training due to sparse fields.
[0037] For the number of lymph nodes dissected and the number of positive lymph nodes, the server uses logarithmic transformation to compress the long-tailed distribution and reduce the amplifying effect of extreme values on the distance between samples. Specifically, the server generates two continuous features:
[0038] in, Indicates the number of lymph nodes removed. Indicates the number of positive lymph nodes; and This represents the continuous features after logarithmic transformation of two terms. For the post-treatment pathological T phase and tumor regression grade, the server encodes them in a pre-defined ordered order as follows: and For example, post-treatment pathological T-phase is encoded in ascending order of invasion depth; tumor regression grades are uniformly encoded in ascending order of pathological risk; in one example, the degree of regression is encoded in ascending order of good to poor. The same encoding direction is used in the training and inference phases. After encoding, the server generates a feature vector for each case. All cases were compiled into a training dataset, with each sample including follow-up time and outcome status.
[0039] In step S2, locally comparable sample pairs are constructed based on the case feature vectors to generate a set of locally monotonic constraints.
[0040] Specifically, according to embodiments of this disclosure, a novel contour-based hierarchical model is constructed using a Gaussian process-enhanced proportional hazards regression (PPR) process. First, variable quality control, missing data handling, and scaling are completed. Then, Gaussian process fitting is introduced into the PRR framework to fit the nonlinearity and interaction relationships between variables. Simultaneously, local monotonic constraints are constructed before training to address the issue of risk reversal in local regions of the continuous risk surface. Local risk reversal refers to the situation where, under otherwise approximately identical conditions, cases with a heavier pathological burden are assigned a lower risk by the model. This directly affects the shape of the continuous risk surface, resulting in jagged edges, islands, and local folds during boundary extraction, thus impacting the stability of the contour-based hierarchical model. By constructing locally comparable sample pairs and incorporating the comparison relationships into the loss function, these phenomena can be controlled.
[0041] Please see Figure 2 , Figure 2 A schematic diagram illustrating the process for generating a set of locally monotonic constraints provided in an embodiment of this disclosure is shown. Figure 2 As shown, in step S201, the local comparison window is set.
[0042] The server first sets a local comparison window based on the results of the lymph node dissection count. This local comparison window is denoted as... For two cases in The dimensions can still be considered as similar within the allowable range of differences. The server can use the training set... Determining the degree of dispersion For example, it can be configured to select based on standard deviation, interquartile range, or cross-validation results. If the training data source is fixed over a long period and the case structure is relatively stable, the server can maintain... No update; if the training set is continuously expanded incrementally, the server will recalculate this parameter during each retraining session, since all subsequent local comparisons depend on it. The server writes it into the model parameter table and calls it uniformly at each stage.
[0043] In step S202, the first type of sample pairs is constructed. The server enumerates any two cases in the training set. and When two cases meet , and At this time, both are locally similar in terms of post-treatment pathological T phase, tumor regression grade, and degree of tumor clearance. Case The positive lymph node burden was higher than that of the case. The server generates a pair of sequential samples. And it is stipulated that the sample is required to meet the criteria for the case. The continuous risk is no less than that of cases. Establish the constraint relationship and record the sample pair as a first-class sample pair.
[0044] In step S203, the second type of sample pair is constructed. When two cases meet the following conditions... , and At that time, the two are locally similar in terms of positive lymph node burden, tumor regression grade, and degree of dissection. Case The post-treatment pathological T phase was higher than that of the case. The server generates a second type of sample pair. And write the same local monotonic relation.
[0045] In step S204, the third type of sample pair is constructed. When two cases meet... , and At that time, the server compares the tumor regression grading coding results of the two. If the case For cases with a worse regression grade, i.e., if the tumor regression grade coding value of case i is greater than that of case j, it indicates that case i corresponds to a worse regression grade, and a third type of sample pair is generated. It also stipulates that the continuous risk of case i should not be lower than that of case j.
[0046] In step S205, the three types of sample pairs are merged into a set of locally monotonic constraints. This set only stores the sample index and order relationship, thus allowing the server to provide local pathological partial order constraints for continuous risk functions while maintaining nonlinear modeling capabilities.
[0047] Existing discrete stratification methods are still not precise enough in characterizing the heterogeneity of real risk, and the prognostic differences among patients in the same stratum are still large. By adding the above-mentioned local constraints to the server before training, the local instability problem during the transformation of continuous risk curves to contour-based stratification can be solved, and the pathologically acceptable local ordering relationship can be incorporated into the training objective, thereby making the subsequent boundary smoother and the stratification more stable.
[0048] In step S3, the case feature vector, the follow-up data, and the set of local monotonic constraints are input into the Gaussian process enhanced Cox model, and the model is trained based on the Cox partial likelihood term and the local monotonic violation term to obtain a continuous risk function.
[0049] In some embodiments, please refer to Figure 3 , Figure 3 A schematic diagram of the Gaussian process-enhanced Cox model structure provided in an embodiment of this disclosure is shown. Figure 3 As shown, the model is divided into an input module, a kernel mapping module, a risk function module, a survival objective module, a constraint regularization module, and a parameter update module.
[0050] The system comprises the following modules: an input module reads feature vectors, follow-up time, and outcome status; a kernel mapping module calculates the similarity between samples; a risk function module outputs the continuous risk value for each case; a survival objective module constructs a partial likelihood term for proportional hazards regression based on follow-up information; a constraint regularization module constructs violation terms based on a set of locally monotonic constraints; and a parameter update module summarizes the two losses and updates the kernel function parameters and correlation coefficients. These modules are sequentially connected: the input module feeds features into the kernel mapping module, the kernel mapping module feeds the similarity structure into the risk function module, the output of the risk function module simultaneously enters both the survival objective module and the constraint regularization module, and the parameter update module updates each parameter inversely based on the total loss function.
[0051] The kernel mapping module uses the radial basis function (RBF) kernel, also known as the Gaussian kernel, for any two case feature vectors. and The kernel function value is calculated according to the following formula:
[0052] in, This represents the kernel function amplitude parameter, which the server uses to control the overall fluctuation range of the potential risk function. Indicates the first The server uses length scale parameters for each feature dimension to control the smoothness of the risk surface along that dimension. The four length scale parameters correspond to... , , and This includes the transformation results of lymph node dissection count, the transformation results of positive lymph node count, the post-treatment pathological T-phase coding results, and the tumor regression grade coding results. The server initializes with candidate ranges for each length scale; subsequently, during training, these parameters are updated based on the total loss function. If the length scale of a certain dimension is small, it indicates that even small changes in that dimension can significantly alter the risk function; if the length scale is large, it indicates that the risk surface generated by the server in that dimension is smoother.
[0053] After the kernel mapping module outputs the similarity matrix, the risk function module calculates the latent risk value for each case according to the inference rules of the Gaussian process. The server directly records this latent risk value as a continuous risk function. The output of for the in the training set One case, the server received For cases to be evaluated, the server receives... .
[0054] The server records the follow-up time as... Record the final state as ,when When, it indicates that the endpoint event has occurred in this case; when At that time, it indicates deletion. The server constructs a risk set indexed by cases where the endpoint event occurred. ,in Includes all follow-up periods of not less than For the cases, the server constructs a partial likelihood term for proportional hazards regression:
[0055] in, This represents the loss of the survival goal. The server uses this formula to measure the consistency between the continuous risk function output by the model and the actual survival outcome. If a case experiences an event earlier and the model assigns it a higher risk, then the case's contribution to the loss of the survival goal is more reasonable; if the model assigns risks in the opposite order, the loss will be greater.
[0056] The server reads samples pair by pair from the set of locally monotonic constraints and calculates the locally monotonic violation term based on the risk difference of each sample pair. The server defines the safety interval parameter. , which is the minimum risk interval allowed to be retained in local comparisons. This parameter can be selected over a small non-negative range using a validation set for a pair of constrained samples. The server calculates the number of violations:
[0057] If the case The risk is no less than that of cases If the risk difference between the two samples satisfies the safety margin, then this pair of samples will not incur a penalty; if the case The risk is lower than the number of cases The server then generates a positive violation. Summing all violations yields a locally monotonic violation term:
[0058] in, Represents the set of locally monotonic constraints. This represents the total amount of local risk reversal. The smaller this term is, the fewer local risk directions that contradict the pathological order appear in the model.
[0059] Constraint weight parameters The total loss function is used to balance the survival objective loss and the local monotonic violation term. Its numerical source can be cross-validation results or the results of a manually pre-defined candidate set. The server constructs the total loss function:
[0060] Then, the server iteratively updates the training set. Various length scales Other Gaussian process-related parameters are monitored on the validation set to assess total loss, survival discrimination ability, and local violation level. If the validation metrics no longer improve within several consecutive iterations, or if the preset number of iterations is reached, the server terminates training and saves the final continuous risk function. The server also records these parameters in the model parameter table. , , , and If model retraining is subsequently performed, the server will re-estimate these parameters according to the same rules.
[0061] This implementation adds local monotonic constraints to the Gaussian process-enhanced Cox model, so that the trained continuous risk function not only maintains its discriminative ability, but also conforms more to pathological logic in local regions. The resulting risk surface is more coherent when extracting the contour boundary in the subsequent process, and is more suitable for forming stable stratification.
[0062] In step S4, the subspace is divided according to the post-treatment pathological T stage and tumor regression grade. The continuous risk function is then sampled in a grid within each subspace to generate a continuous risk surface.
[0063] The server divides the space into subspaces based on the post-treatment pathological T-phase coding results and the tumor regression grading coding results. Each subspace corresponds to a fixed set of subspaces. and In a fixed group , Subsequently, the server only samples on the two-dimensional plane formed by the transformation results of lymph node scan count and positive lymph node count.
[0064] The server, within each subspace, is based on the training set. and A two-dimensional regular grid is established based on the range of values. The server records the sampling step size of the two dimensions as follows: and These two parameters represent the grid precision, and their range can be set according to the distribution density of the training samples. If the data volume is large, a finer step size is used; if the data volume is small, the step size is appropriately increased to maintain local smoothness. Each grid point is denoted as . ,in and With the current subspace fixed, the server sequentially writes all grid points into the risk surface table.
[0065] For each grid point, the trained continuous risk function is called to obtain the risk value of that point. The risk matrices are formed in row and column order. Each risk matrix corresponds to a fixed set of post-treatment pathological T phase and tumor regression grade. The risk matrices of all subspaces are summed to form a set of continuous risk surfaces.
[0066] In some embodiments, a sparse region identification step is also included. Because the training samples are unevenly distributed across different subspaces, the continuous risk surface may exhibit less-than-smooth numerical variations in local regions; therefore, a sparse region identification step is added.
[0067] Specifically, define the neighborhood radius parameter. It also counts the number of training cases in the neighborhood of each grid point, and determines the number of cases when the number of cases in the neighborhood falls below a preset lower limit. When this happens, the grid point is marked as a sparse point. Risk value smoothing is performed on sparse points; in one example, the risk value of the current point is updated using a weighted average method, with weights set according to the distance of neighboring grid points from the current point. and All of these are obtained from the statistical distribution of training data. Through the above processing, the spikes and isolated depressions caused by insufficient local samples in the risk surface will be significantly reduced.
[0068] In step S5, candidate contour boundaries are extracted based on the continuous risk surface, and stable contour boundaries are selected based on the boundary stability index obtained by resampling.
[0069] Please see Figure 4 , Figure 4 A schematic diagram of the process for screening stable contour boundaries provided in an embodiment of this disclosure is shown. Figure 4 As shown, in step S401, candidate contour boundaries are extracted based on the continuous risk surface.
[0070] For each subspace and each candidate risk threshold, the server extracts isorisk lines on the corresponding risk matrix to generate candidate contour boundaries. The server stores the boundaries of the same threshold in different subspaces in a unified boundary table. If the area enclosed by a candidate contour boundary is smaller than the minimum connected component threshold, the server considers it a small region caused by numerical noise and deletes it. This minimum connected component threshold is determined by both the grid step size and the subspace scale. The server can specify its source rules in the modeling configuration, thereby obtaining a set of geometrically more continuous candidate contour boundaries.
[0071] In some embodiments, the candidate boundaries obtained from a single training iteration may still be sensitive to sample sampling. The server continues to perform resampling training and boundary stability screening to ensure that the final boundaries remain close to each other even when the data changes slightly.
[0072] Specifically, in step S402, the original training set is sampled with replacement at the case granularity for resampling training. The server defines the number of resampling operations. This parameter represents the number of repetitions for boundary stability assessment. The original training set is sampled with replacement at the case granularity, forming a resampled training set each time. For each resampled training set, the following processes are repeated: construction of a set of local monotonic constraints, Gaussian process-enhanced proportional hazards regression training, continuous risk surface generation, and candidate contour boundary extraction, to obtain a set of boundaries corresponding to the original training set.
[0073] In step S403, the candidate contour boundaries obtained from the original training set are used as reference boundaries, and the corresponding boundaries obtained from each resampling are paired with the reference boundaries. The average symmetric distance between the two boundaries is calculated and denoted as... Indicates the first The positional difference between the resampling boundary and the reference boundary, if the boundary position change is small, then Smaller; if the boundary position fluctuates significantly, then Relatively large.
[0074] In step S404, a boundary stability index is generated based on all average symmetric distances. This can be expressed by the formula:
[0075] in, Indicates the boundary stability index; Indicates the number of resampling attempts; Indicates the first The smaller the positional difference of the average symmetric distance corresponding to each resampling, the greater the boundary stability index. The stability index of each candidate contour boundary is written into the boundary table. Through the above processing, this implementation method further applies stability to specific boundary selection rules, so that contour layering depends not only on the model output, but also on the positional consistency of the boundary during repeated training.
[0076] In step S405, stable contour boundaries are selected based on a boundary stability threshold. The boundary stability threshold is defined. If the stability index corresponding to a candidate contour boundary is not lower than Retain the boundary; if it is lower than The boundary is then eliminated. The server can statistically analyze the distribution of stability indicators for all candidate boundaries. For example, the settings can be set according to the median level or the upper quantile level. After the filtering is completed, the final set of stable contour boundaries is obtained and written into the boundary table according to the risk threshold from low to high.
[0077] In step S6, the case to be evaluated is input into the continuous risk function to obtain a continuous risk score, and the contour-based ypTNM risk stratification label of the case to be evaluated is determined in combination with the stable contour boundary.
[0078] The same data verification, missing data handling, scaling, and ordered coding as in the training phase are performed on the cases to be evaluated, resulting in a feature vector for each case. This feature vector is then input into a continuous risk function to obtain a continuous risk score. The post-treatment pathological T-stage coding result and tumor regression grade coding result for each case are read and located in the corresponding subspace. Based on the case's position relative to the stable contour boundary within this subspace, its corresponding contour-based ypTNM risk stratification label is determined. The stratification label can be categorized into several levels according to risk from low to high, maintaining a fixed naming convention in the configuration. It should be noted that the contour-based ypTNM risk stratification label is a model stratification label generated based on the aforementioned pathological variables and continuous risk surface, and is not equivalent to the standard pathological TNM staging results.
[0079] To further characterize the relationship between cases and boundaries, the server also calculates the shortest distance from the case to be evaluated to the nearest stable contour boundary and generates boundary proximity markers accordingly. A boundary proximity threshold is defined. If the distance from the case to be evaluated to the nearest boundary is no greater than Record this case as a case near the boundary; if the distance is greater than This case was recorded as a non-border-adjacent case. It can be obtained from the average spacing of adjacent stable boundaries, or a reasonable value can be given based on the grid step size during the initial deployment.
[0080] In one implementation, the final output is written to a results table, which includes at least the case number, continuous risk score, profile-based ypTNM risk stratification label, and boundary proximity marker. If the server receives a batch of cases, the results are sorted from highest to lowest continuous risk score. This implementation further transforms the continuous risk score into a location-stable profile-based stratification result through local monotonic constraints and boundary stability filtering.
[0081] In some embodiments, the server can also perform periodic updates. In practice, after a new case is added to the case data table, the server completes field verification and feature organization according to the original rules. When the number of new cases reaches a set scale or the preset update cycle is reached, the server recalculates the local comparison window parameters, kernel function parameters, constraint weight parameters, sparse region determination parameters, candidate risk threshold, and boundary stability threshold, and retrains the model and updates the boundary table according to predetermined steps. If the hospital's data sources, pathology templates, and case structure are stable in the long term, the server can reduce the update frequency and only update when the sample size increases significantly.
[0082] Therefore, by unifying conventional pathological variables into the same feature space, Gaussian process is used to enhance proportional hazard regression to output continuous risk scores, and local monotonic constraints are used to suppress local risk reversal. Contour-type ypTNM risk stratification labels are generated through continuous risk surfaces and resampling stable boundaries. A continuous data transfer relationship is formed between risk scores, risk surfaces, and stratification boundaries. The stratification results retain continuous risk information and maintain the stability of boundary positions.
[0083] Please see Figure 5 , Figure 5 This is a schematic diagram of a nonlinear survival risk modeling system based on pathology after neoadjuvant chemotherapy for gastric cancer, provided in an embodiment of this application. Figure 5 As shown, the system includes: The data acquisition and preprocessing module 501 is used to acquire pathological data and follow-up data of gastric cancer patients after neoadjuvant chemotherapy. The pathological data includes at least the number of lymph nodes dissected, the number of positive lymph nodes, the pathological T stage after treatment, and the tumor regression grade. The pathological data is subjected to legality checks, missing data handling, scaling transformation, and ordered coding to generate case feature vectors. The constraint construction module 502 is used to construct locally comparable sample pairs based on the case feature vector and generate a set of locally monotonic constraints. Model training module 503 is used to input the case feature vector, the follow-up data and the set of local monotonic constraints into the Gaussian process enhanced Cox model, and train it based on the Cox partial likelihood term and the local monotonic violation term to obtain a continuous risk function; The surface generation module 504 is used to divide the subspace according to the post-treatment pathological T stage and tumor regression grade, and to perform grid sampling on the continuous risk function in each subspace to generate a continuous risk surface; The boundary filtering module 505 is used to extract candidate contour boundaries based on the continuous risk surface and to filter stable contour boundaries based on the boundary stability index obtained by resampling. The risk stratification module 506 is used to input the case to be evaluated into the continuous risk function to obtain a continuous risk score, and to determine the contour-based ypTNM risk stratification label of the case to be evaluated in combination with the stable contour boundary.
[0084] Each processing unit and / or module in the embodiments of this application can be implemented by an analog circuit that implements the functions described in the embodiments of this application, or by software that executes the functions described in the embodiments of this application.
[0085] Based on the same inventive concept, this application also provides an electronic device, the method corresponding to which can be the method in the foregoing embodiments, and its problem-solving principle is similar to that method. For example... Figure 6 As shown, Figure 6This is a schematic diagram of an electronic device structure provided in an embodiment of the present disclosure. The device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the methods and / or technical solutions of the foregoing embodiments of the present application.
[0086] In particular, the methods and / or embodiments in this application can be implemented as computer software programs. For example, the embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. When the computer program is executed by a processor, it performs the functions defined in the methods of this application.
[0087] Another embodiment of this application provides a storage medium storing computer program instructions thereon, which can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of this application described above.
[0088] In the above embodiments, the descriptions of each embodiment have different focuses. Parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. The above descriptions are merely preferred embodiments of this application and explanations of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by specific combinations of the above technical features, but should also cover other technical solutions formed by arbitrary combinations of the above technical features or their equivalent features without departing from the inventive concept.
Claims
1. A nonlinear survival risk modeling method based on pathology after neoadjuvant chemotherapy for gastric cancer, characterized in that, include: Pathological and follow-up data of gastric cancer patients after neoadjuvant chemotherapy were obtained. The pathological data included at least the number of lymph nodes dissected, the number of positive lymph nodes, the pathological T stage after treatment, and the tumor regression grade. The pathological data were subjected to legality checks, missing data handling, scaling transformation, and ordered coding to generate case feature vectors. Based on the case feature vectors, locally comparable sample pairs are constructed, and a set of locally monotonic constraints is generated. The case feature vector, the follow-up data, and the set of local monotonic constraints are input into the Gaussian process-enhanced Cox model, and the model is trained based on the Cox partial likelihood term and the local monotonic violation term to obtain a continuous risk function. The subspace is divided according to the post-treatment pathological T stage and tumor regression grade. The continuous risk function is then sampled in a grid within each subspace to generate a continuous risk surface. Candidate contour boundaries are extracted based on the continuous risk surface, and stable contour boundaries are selected based on the boundary stability index obtained by resampling. The case to be evaluated is input into the continuous risk function to obtain a continuous risk score, and the profile-based ypTNM risk stratification label of the case to be evaluated is determined by combining the stable profile boundary.
2. The method according to claim 1, characterized in that, The construction of locally comparable sample pairs and the generation of a set of locally monotonic constraints include: A local comparison window based on the transformation result of the number of lymph nodes dissected is set in the case feature vector; within the local comparison window where the post-treatment pathological T stage and tumor regression grade are consistent, a first type of sample pair is constructed according to the number of positive lymph nodes; within the local comparison window where the number of positive lymph nodes and tumor regression grade are consistent, a second type of sample pair is constructed according to the size of the post-treatment pathological T stage; within the local comparison window where the number of positive lymph nodes and post-treatment pathological T stage are consistent, a third type of sample pair is constructed according to the pathological risk corresponding to the tumor regression grade in ascending order; the first type of sample pair, the second type of sample pair, and the third type of sample pair are summarized into the local monotonic constraint set.
3. The method according to claim 2, characterized in that, The Gaussian process-enhanced Cox model includes an input module, a kernel mapping module, a risk function module, a survival objective module, a constraint regularization module, and a parameter update module. The input module receives case feature vectors; the kernel mapping module generates latent risk function values based on the similarity between case feature vectors; the risk function module outputs a continuous risk function; the survival objective module constructs a Cox partial likelihood term based on follow-up data; the constraint regularization module constructs a locally monotonically violating term based on a set of locally monotonically bound constraints; and the parameter update module updates the model parameters based on the total loss function.
4. The method according to claim 3, characterized in that, The kernel mapping module uses a radial basis function kernel and sets length scale parameters for each feature dimension in the case feature vector. Each feature dimension includes the transformation results of lymph node dissection count, the transformation results of positive lymph node count, the pathological T-phase coding results after treatment, and the coding results of tumor regression grading.
5. The method according to claim 4, characterized in that, The method of training based on the Cox partial likelihood term and the local monotonic violation term to obtain the continuous risk function includes: calculating the risk set corresponding to cases where the endpoint event occurs, and generating the Cox partial likelihood term; calculating the difference in risk function values for each sample pair in the local monotonic constraint set, and generating the local monotonic violation term based on the difference in risk function values; performing a weighted summation of the Cox partial likelihood term and the local monotonic violation term to generate the total loss function; and updating the parameters of the Gaussian process enhanced Cox model based on the total loss function to obtain the continuous risk function.
6. The method according to claim 5, characterized in that, The local monotonic violation term is generated based on the continuous risk function values of the two corresponding cases for each sample and a preset safety interval. When the continuous risk function values of the two cases satisfy the local monotonic constraint relationship and the difference between the corresponding risk function values is not less than the preset safety interval, the local monotonic violation term for the sample is set to zero.
7. The method according to claim 1, characterized in that, The process involves dividing the space into subspaces based on the post-treatment pathological T phase and tumor regression grade, and performing grid sampling on the continuous risk function within each subspace to generate a continuous risk surface, including: The pathological T-phase coding results and tumor regression grade coding results after fixation treatment are used to establish a two-dimensional regular grid based on the transformation results of lymph node dissection number and positive lymph node number. The continuous risk function is called for each grid point to generate the risk matrix of the corresponding subspace. The risk matrices of each subspace are summarized into the continuous risk surface.
8. The method according to claim 7, characterized in that, The extraction of candidate contour boundaries based on the continuous risk surface includes: The number of training cases in the neighborhood of each grid point is counted. For grid points where the number of training cases in the neighborhood is lower than a preset lower limit, risk value smoothing is performed. A set of candidate risk thresholds is generated based on the distribution of continuous risk scores of training cases. For each candidate risk threshold, isorisk lines are extracted from the continuous risk surface to generate candidate contour boundaries.
9. The method according to claim 8, characterized in that, The method of selecting stable contour boundaries based on the boundary stability index obtained from resampling includes: The training cases are resampled multiple times. For each resampled case, the following processes are repeated: local monotonic constraint set construction, Gaussian process enhanced Cox model training, continuous risk surface generation, and candidate contour boundary extraction. The positional difference of each candidate contour boundary in multiple resamplings is calculated to generate a boundary stability index. Candidate contour boundaries whose boundary stability index meets preset conditions are determined as stable contour boundaries. The boundary stability index is generated based on the average symmetrical distance between the candidate contour boundary and the corresponding resampled contour boundary. The smaller the positional difference, the larger the corresponding boundary stability index.
10. A nonlinear survival risk modeling system based on pathology after neoadjuvant chemotherapy for gastric cancer, characterized in that, include: The data acquisition and preprocessing module is used to acquire pathological data and follow-up data of gastric cancer patients after neoadjuvant chemotherapy. The pathological data includes at least the number of lymph nodes dissected, the number of positive lymph nodes, the pathological T stage after treatment, and the tumor regression grade. The pathological data is subjected to legality checks, missing data handling, scaling transformation, and ordered coding to generate case feature vectors. The constraint construction module is used to construct locally comparable sample pairs based on the case feature vector and generate a set of locally monotonic constraints. The model training module is used to input the case feature vector, the follow-up data, and the set of local monotonic constraints into the Gaussian process-enhanced Cox model, and train it based on the Cox partial likelihood term and the local monotonic violation term to obtain a continuous risk function; The surface generation module is used to divide the subspace according to the post-treatment pathological T phase and tumor regression grade, and to perform grid sampling on the continuous risk function in each subspace to generate a continuous risk surface; The boundary filtering module is used to extract candidate contour boundaries based on the continuous risk surface and to filter stable contour boundaries based on the boundary stability index obtained by resampling. The risk stratification module is used to input the case to be evaluated into the continuous risk function to obtain a continuous risk score, and to determine the contour-based ypTNM risk stratification label of the case to be evaluated in combination with the stable contour boundary.