Method, system, device and medium for predicting recurrence of colorectal cancer based on pathological features
By using a pathological feature-based approach to screen colorectal cancer subjects and construct a nomogram prediction model, the problem of inaccurate colorectal cancer recurrence risk assessment in existing technologies is solved, enabling early identification and personalized treatment support for patients with high recurrence risk.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 核工业四一六医院
- Filing Date
- 2026-04-08
- Publication Date
- 2026-07-03
AI Technical Summary
Current technologies lack reliable methods for predicting colorectal cancer recurrence, making it impossible to accurately assess patients' recurrence risk. This results in a lack of personalized strategies and the ability to identify patients at high risk of recurrence in clinical treatment.
By using a pathological feature-based approach, colorectal cancer subjects were screened, and paraffin-embedded tumor tissue, clinicopathological data, and long-term follow-up data were obtained. Tertiary lymphoid structures and traditional pathological features were identified, and a nomogram prediction model was constructed to assess the risk of colorectal cancer recurrence.
It improves the accuracy and reliability of predicting the risk of colorectal cancer recurrence, enhances the ability to identify patients at high risk of recurrence early, and supports personalized treatment decisions.
Smart Images

Figure CN122337555A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer technology, and in particular relates to a method, system, device and medium for predicting recurrence of colorectal cancer based on pathological characteristics. Background Technology
[0002] Colorectal cancer, the most common malignant tumor of the digestive tract, ranks third in incidence and second in mortality among cancers worldwide. In 2022, there were over 1.88 million new cases and 910,000 deaths globally. In my country, there were approximately 517,100 new cases, accounting for 10.7% of all new cancer cases, and approximately 240,000 deaths, accounting for 9.3% of cancer deaths. Furthermore, the 5-year survival rate is only 30%-40%, significantly lower than in developed countries in Europe and America. The prognosis of colorectal cancer patients is closely related to tumor stage. Stage I patients have a low recurrence rate after surgery and a 5-year survival rate close to 100%, while stage IV patients have a very poor prognosis due to distant metastasis at diagnosis. The recurrence rate for stage II / III patients is as high as 30%-40%, and recurrence and metastasis have become the main causes of death for patients at this stage. Currently, the treatment of colorectal cancer is mainly radical surgery, and postoperative adjuvant chemotherapy, radiotherapy and other comprehensive treatments are gradually becoming more common. However, the lack of reliable prognostic markers and efficient recurrence prediction strategies remains a major bottleneck in clinical treatment. Existing prediction methods mostly rely on TNM (Tumor-Node-Metastasis Staging) staging and some pathological features, which makes it difficult to accurately assess the risk of recurrence in patients and cannot provide sufficient support for the development of individualized treatment and follow-up plans.
[0003] With in-depth research into the tumor microenvironment, its crucial role in the development of malignant tumors has been widely confirmed. The tumor microenvironment is composed of immune cells, stromal cells, extracellular matrix, and tumor blood vessels. Among these, tertiary lymphoid structures, as an important feature of the tumor microenvironment, are ectopic lymphoid structures enriched with immune cells. They have been proven to be closely related to the development, immune escape, drug resistance, and prognosis of colorectal cancer. Their presence is associated with a reduced risk of recurrence and improved survival rates in almost all human solid tumors. The characteristics of tertiary lymphoid structures include density and maturity characteristics; higher density and maturity of tertiary lymphoid structures may have a protective effect against recurrence in colorectal cancer patients. Meanwhile, traditional pathological features such as vascular invasion, neural invasion, lymph node metastasis, differentiation degree, and microsatellite status have also been proven to be important factors affecting the prognosis of colorectal cancer patients, with vascular invasion and neural invasion being key indicators of poor prognosis. However, existing technologies have not yet systematically integrated the characteristics of the three-level lymphatic structure with traditional pathological features, lack a standardized process for detection, feature extraction and model building, and have not yet formed an efficient method that can accurately predict the recurrence risk of colorectal cancer patients, resulting in the inability of clinicians to identify patients with high recurrence risk in the early stage and take timely intervention measures. Summary of the Invention
[0004] Therefore, it is necessary to provide methods, systems, devices, and media for predicting colorectal cancer recurrence based on pathological characteristics to address the aforementioned technical issues. This aims to improve the accuracy of recurrence risk prediction and the personalization of clinical diagnosis and treatment, enhance the early identification of patients at high recurrence risk, and thus improve the scientificity and reliability of clinical decision-making.
[0005] Firstly, this application provides a method for predicting colorectal cancer recurrence based on pathological characteristics, including:
[0006] S1. Based on the preset inclusion and exclusion criteria, colorectal cancer subjects were screened out, and paraffin-embedded tumor tissue, clinicopathological data, and long-term follow-up data of colorectal cancer subjects were obtained. The long-term follow-up data included records of whether the colorectal cancer subjects had recurrence.
[0007] S2, Obtain a slice processing set including HE-stained slices and immunohistochemical stained slices. The immunohistochemical stained slices include CD3 immunohistochemical stained slices, CD20 immunohistochemical stained slices, CD21 immunohistochemical stained slices and CD23 immunohistochemical stained slices. The slice processing set is obtained by performing slice operations on paraffin-embedded tumor tissue blocks.
[0008] S3, based on HE-stained sections, immunohistochemical stained sections, and clinicopathological data, identifies non-dense capsule lymphocyte clusters in the tumor region and the pre-defined tumor boundary region. These non-dense capsule lymphocyte clusters are defined as tertiary lymphoid structures, and tertiary lymphoid structure features and traditional pathological features are extracted. Tertiary lymphoid structure features include tertiary lymphoid structure density features and tertiary lymphoid structure maturity features. The tertiary lymphoid structure density features, tertiary lymphoid structure maturity features, and traditional pathological features are integrated and processed to form a standardized feature set.
[0009] S4. The standardized feature set is associated and matched with long-term follow-up data to construct a feature-recurrence status corresponding dataset. The feature-recurrence status corresponding dataset is then processed by univariate logistic regression analysis and multivariate logistic regression analysis to screen out independent influencing factors of colorectal cancer recurrence.
[0010] S5. Based on independent influencing factors, construct a nomogram prediction model and perform a verification operation on the nomogram prediction model. The nomogram prediction model that passes the verification is used as the finalized nomogram prediction model.
[0011] S6. For patients who have recently undergone colorectal cancer surgery, repeat S2 and S3 sequentially to obtain a standardized feature set of patients who have recently undergone colorectal cancer surgery. Input the standardized feature set of patients who have recently undergone colorectal cancer surgery into the stereotyped nomogram prediction model to obtain the recurrence risk assessment results of patients who have recently undergone colorectal cancer surgery.
[0012] In one embodiment, based on HE-stained sections, immunohistochemically stained sections, and clinicopathological data, clusters of non-dense-capsulated lymphocytes in the tumor region and a predefined tumor boundary region are identified. These non-dense-capsulated lymphocyte clusters are defined as tertiary lymphoid structures, and tertiary lymphoid structure features and traditional pathological features are extracted. The tertiary lymphoid structure features include tertiary lymphoid structure density features and tertiary lymphoid structure maturity features. The tertiary lymphoid structure density features, tertiary lymphoid structure maturity features, and traditional pathological features are integrated to form a standardized feature set, including:
[0013] A double-blind method is used to examine HE-stained sections and immunohistochemically stained sections to obtain a first examination result and a second examination result. If the first examination result and the second examination result are inconsistent, an arbitration instruction is generated to indicate the arbitration of the first examination result and the second examination result. The arbitration instruction is sent to the preset arbitrator terminal and a unified examination result is received from the arbitrator terminal.
[0014] Based on HE staining sections, immunohistochemical staining sections, and unified slide reading results, clusters of non-dense-capsulated lymphocytes formed by the aggregation of CD20+B cells and CD3+T cells were identified in the tumor area and the pre-defined tumor boundary area. These non-dense-capsulated lymphocyte clusters were defined as tertiary lymphoid structures.
[0015] Based on HE-stained sections and unified slide reading results, the number of tertiary lymphoid structures in the preset high-power field of HE-stained sections was counted. The number of tertiary lymphoid structures was classified according to the standard that 1+ corresponds to 1-5, 2+ corresponds to 6-10, and 3+ corresponds to more than 10, so as to obtain the density characteristics of tertiary lymphoid structures.
[0016] Based on CD21 immunohistochemical staining sections, CD23 immunohistochemical staining sections, and unified slide reading results, FDC networks and germinal centers were identified. FDC networks were determined by CD21 expression, and germinal centers were determined by CD23 expression. The maturity of the three-level lymphoid structures was graded according to the criteria of E-TLS corresponding to CD21- and CD23-, PFL-TLS corresponding to CD21+ and CD23-, and SFL-TLS corresponding to CD21+ and CD23+, thus obtaining the maturity characteristics of the three-level lymphoid structures.
[0017] Based on HE-stained sections, immunohistochemical stained sections, and clinicopathological data, we extracted vascular invasion status, nerve invasion status, TNM stage, lymph node metastasis status, differentiation degree, and microsatellite status. The differentiation degree was divided into high, medium, and low differentiation, and the microsatellite status was divided into stable and unstable status. Vascular invasion status, nerve invasion status, TNM stage, lymph node metastasis status, differentiation degree, and microsatellite status were used as traditional pathological features.
[0018] The density characteristics of tertiary lymphoid structures, the maturity characteristics of tertiary lymphoid structures, and traditional pathological characteristics are integrated and processed to form a standardized feature set.
[0019] In one embodiment, a standardized feature set is correlated and matched with long-term follow-up data to construct a feature-recurrence status corresponding dataset. This dataset is then processed sequentially using univariate logistic regression analysis and multivariate logistic regression analysis to screen for independent influencing factors of colorectal cancer recurrence, including:
[0020] The features in the standardized feature set are associated and matched with the relapse status in the long-term follow-up data to construct a feature-relapse status correspondence dataset.
[0021] Univariate logistic regression analysis was performed on each matching feature in the feature-recurrence status dataset to obtain the analysis coefficient of each matching feature. Matching features with analysis coefficients less than the preset coefficient threshold were used as candidate factors associated with colorectal cancer recurrence. The matching features included age, TNM stage, lymph node metastasis status, vascular invasion status, nerve invasion status, differentiation degree, tertiary lymphoid structure density characteristics, and tertiary lymphoid structure maturity characteristics.
[0022] The candidate factors were subjected to multivariate logistic regression analysis to eliminate non-independent influencing factors, and the independent influencing factors of colorectal cancer recurrence were obtained. The independent influencing factors included the density characteristics of tertiary lymphoid structures, the maturity characteristics of tertiary lymphoid structures, the vascular invasion status, and the nerve invasion status.
[0023] In one embodiment, a nomogram prediction model is constructed based on independent influencing factors, and a validation operation is performed on the nomogram prediction model. The validated nomogram prediction model is then used as the finalized nomogram prediction model, including:
[0024] Based on independent influencing factors, a visual nomogram prediction model is constructed. Each independent influencing factor is assigned a fixed score to obtain the corresponding score, and a direct mapping relationship between the score and the probability of colorectal cancer recurrence is established.
[0025] The consistency index of the nocline plot prediction model was calculated, and the time-dependent ROC curve was plotted. The area under the curve of the time-dependent ROC curve was calculated. Based on the consistency index and the area under the curve, the discrimination of the nocline plot prediction model was verified, and the discrimination verification results were obtained.
[0026] A calibration curve is plotted, and the colorectal cancer recurrence probability predicted by the nomogram prediction model is compared with the actual recurrence rate recorded in long-term follow-up data to verify consistency. When the consistency reaches the preset standard, the calibration degree of the nomogram prediction model is determined to be up to standard, and the calibration degree verification result is obtained.
[0027] Clinical decision curves are plotted to evaluate the clinical benefits of the nomogram prediction model under different risk thresholds. When the nomogram prediction model shows positive benefits, its clinical applicability is determined to be up to standard, and the clinical applicability verification results are obtained.
[0028] Once the discrimination verification results, calibration verification results, and clinical usability verification results all meet the standards, the nomogram prediction model is determined to be a valid nomogram prediction model, and the valid nomogram prediction model is output as the finalized nomogram prediction model.
[0029] In one embodiment, the slicing operation includes slice preparation, dewaxing and hydration, antigen retrieval, and immunohistochemical staining.
[0030] Secondly, this application also provides a colorectal cancer recurrence prediction system based on pathological characteristics, including:
[0031] The data acquisition module is used to perform step 1, which includes: screening colorectal cancer subjects based on preset inclusion and exclusion criteria, and acquiring paraffin-embedded tumor tissue blocks, clinicopathological data and long-term follow-up data of colorectal cancer subjects. The long-term follow-up data includes records of whether the colorectal cancer subjects have relapsed.
[0032] The slide acquisition module is used to perform step 2, which includes: acquiring a slide processing set including HE-stained slides and immunohistochemical stained slides. The immunohistochemical stained slides include CD3 immunohistochemical stained slides, CD20 immunohistochemical stained slides, CD21 immunohistochemical stained slides, and CD23 immunohistochemical stained slides. The slide processing set is obtained by performing slide operations on paraffin-embedded tumor tissue blocks.
[0033] The feature extraction module is used to perform step 3, which includes: based on HE-stained sections, immunohistochemical stained sections, and clinicopathological data, identifying non-dense capsule lymphocyte clusters in the tumor region and the preset tumor boundary region, defining the non-dense capsule lymphocyte clusters as tertiary lymphoid structures, and extracting tertiary lymphoid structure features and traditional pathological features. The tertiary lymphoid structure features include tertiary lymphoid structure density features and tertiary lymphoid structure maturity features; integrating the tertiary lymphoid structure density features, tertiary lymphoid structure maturity features, and traditional pathological features to form a standardized feature set.
[0034] The key screening module is used to perform step 4, which includes: performing association matching processing between the standardized feature set and long-term follow-up data to construct a feature-recurrence status corresponding dataset, and then processing the feature-recurrence status corresponding dataset through univariate logistic regression analysis and multivariate logistic regression analysis to screen out independent influencing factors of colorectal cancer recurrence.
[0035] The verification module is used to execute step 5, which includes: constructing a nomogram prediction model based on independent influencing factors, performing a verification operation on the nomogram prediction model, and using the verified nomogram prediction model as the finalized nomogram prediction model.
[0036] The recurrence risk assessment module is used to perform step 6, which includes: repeating steps 2 and 3 sequentially for patients who have recently undergone colorectal cancer surgery to obtain a standardized feature set of patients who have recently undergone colorectal cancer surgery; inputting the standardized feature set of patients who have recently undergone colorectal cancer surgery into the morphological nomogram prediction model to obtain the recurrence risk assessment results of patients who have recently undergone colorectal cancer surgery.
[0037] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the first aspect.
[0038] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the first aspect.
[0039] The aforementioned method, system, equipment, and media for predicting colorectal cancer recurrence based on pathological features first screen colorectal cancer subjects and obtain paraffin-embedded tumor tissue, clinicopathological data, and long-term follow-up data, laying the foundation for subsequent feature extraction and model construction. Secondly, the paraffin-embedded tumor tissue is sectioned to obtain a slice processing set including HE-stained slices and specific immunohistochemically stained slices, providing visual materials for the identification of tertiary lymphoid structures. Subsequently, based on the stained slices and clinicopathological data, tertiary lymphoid structures are identified, and their density, maturity characteristics, and traditional pathological features are extracted and integrated to form a standardized feature set. This avoids the limitations of single-feature prediction and improves the comprehensiveness of feature expression. Simultaneously, univariate and multivariate logistic regression analysis is used to screen independent influencing factors, further ensuring the reliability of the model's core variables. Furthermore, the nomogram prediction model is constructed and validated, addressing the shortcomings of existing methods that rely on traditional staging and have insufficient prediction accuracy. Finally, the model is applied to patients who have recently undergone colorectal cancer surgery, outputting recurrence risk assessment results, improving the accuracy of clinical recurrence risk prediction, and providing support for individualized follow-up and treatment plan development. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 A flowchart of a method for predicting colorectal cancer recurrence based on pathological features, provided as an exemplary embodiment of the present invention;
[0042] Figure 2 A flowchart of a method for obtaining a typed nodal plot prediction model is provided as an exemplary embodiment of the present invention;
[0043] Figure 3 This is a schematic diagram of a colorectal cancer recurrence prediction system based on pathological features, provided as an exemplary embodiment of the present invention. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0045] In one embodiment, such as Figure 1As shown, a method for predicting colorectal cancer recurrence based on pathological features is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0046] S101, based on preset inclusion and exclusion criteria, screened colorectal cancer subjects, obtained paraffin-embedded tumor tissue, clinicopathological data and long-term follow-up data of colorectal cancer subjects, and the long-term follow-up data included records of whether the colorectal cancer subjects had recurrence.
[0047] Specifically, subjects with stage II / III colorectal cancer can be screened based on pre-defined inclusion and exclusion criteria, which can be set according to the AJCC 9th edition TNM staging and the Declaration of Helsinki's guidelines for medical research. For example, inclusion criteria could be set as follows: pathological diagnosis confirming primary colorectal cancer; undergoing radical surgical resection to ensure complete removal of tumor tissue and reduce interference from residual lesions in recurrence assessment; complete clinicopathological information, including key data such as age, sex, TNM stage, and lymph node metastasis status, providing sufficient material for subsequent feature extraction; and not having received preoperative radiotherapy, chemotherapy, immunotherapy, or other interventions to avoid the impact of treatment methods on tumor morphology and immune microenvironment. Exclusion criteria can be set as follows: patients with other malignant tumors or multiple primary colorectal cancers to avoid interference from other tumors in the recurrence risk assessment; patients with serious internal medical conditions or immune deficiency diseases, which may affect the body's immune status and thus interfere with the interpretation of the three-level lymphoid structure characteristics; patients with incomplete follow-up or data to ensure that each subject has a clear record of recurrence status; and pregnant or lactating women to avoid the influence of physiological specialties on the study results.
[0048] After screening, multidimensional data from colorectal cancer subjects can be collected. Paraffin-embedded tumor tissue blocks serve as the biological sample basis for subsequent section staining and feature extraction. Clinicopathological data can include TNM staging, lymph node metastasis status, vascular invasion status, nerve invasion status, differentiation degree, and microsatellite status, which are core sources of traditional pathological features. Long-term follow-up data can be collected through postoperative outpatient visits or telephone follow-ups, ending at the onset of recurrence or a predetermined time point. Recurrence criteria can be clearly defined as postoperative lymph node metastasis, formation of abdominal cancer nodules, organ invasion, or distant metastasis (such as liver, lung, brain, and bone metastases). This data can be used to determine whether a subject has relapsed, providing crucial labeling information for model training.
[0049] S102, Obtain a slice processing set including HE-stained slices and immunohistochemically stained slices. The immunohistochemically stained slices include CD3 immunohistochemically stained slices, CD20 immunohistochemically stained slices, CD21 immunohistochemically stained slices, and CD23 immunohistochemically stained slices. The slice processing set is obtained by performing slicing operations on paraffin-embedded tumor tissue blocks. The slicing operations include slice preparation, dewaxing and hydration, antigen retrieval, and immunohistochemical staining.
[0050] Specifically, the slide preparation process follows standard pathological procedures. For example, in slide preparation, paraffin-embedded tumor tissue is cut into continuous slides and baked to ensure close adhesion between the slides and the slides. In the dewaxing and hydration stage, a gradient reagent treatment removes paraffin and restores tissue water solubility, followed by rinsing with buffer to balance the pH. In the antigen retrieval stage, citrate buffer is used for high-temperature, high-pressure treatment of the slides to break the cross-linking bonds between antigens and proteins, exposing antigen sites. In the immunohistochemical staining stage, CD3, CD20, CD21, and CD23 specific antibodies and DAB chromogenic agent are selected. Through a series of steps including blocking, incubation, development, counterstaining, and mounting, HE-stained slides and target immunohistochemically stained slides are obtained. The final slide set obtained through these processes can include HE-stained slides and CD3, CD20, CD21, and CD23 immunohistochemically stained slides. HE-stained slides are used to observe tissue morphology and TLS density counting, while immunohistochemically stained slides are used to accurately identify T cells, B cells, FDC networks, and germinal centers, providing a reliable visual basis for subsequent TLS feature interpretation.
[0051] S103, based on HE-stained sections, immunohistochemical stained sections, and clinicopathological data, identifies non-dense capsule lymphocyte clusters in the tumor region and the pre-defined tumor boundary region. These non-dense capsule lymphocyte clusters are defined as tertiary lymphoid structures, and tertiary lymphoid structure features and traditional pathological features are extracted. Tertiary lymphoid structure features include tertiary lymphoid structure density features and tertiary lymphoid structure maturity features. The tertiary lymphoid structure density features, tertiary lymphoid structure maturity features, and traditional pathological features are integrated and processed to form a standardized feature set.
[0052] Specifically, based on HE-stained and immunohistochemically stained sections, clusters of non-dense capsule lymphocytes can be identified within the tumor region and pre-defined tumor boundary regions. These lymphocyte clusters can be defined as tertiary lymphoid structures (TLS). Subsequently, the density and maturity characteristics of TLS can be extracted. For example, TLS density can be graded based on the number of TLS within 10 high-power fields (HPF), while TLS maturity can be classified based on the presence of follicular dendritic cell networks (CD21+) and germinal centers (CD23+). These characteristics reflect the developmental degree and functional status of TLS within the tumor microenvironment and are closely related to tumor immune escape and prognosis. Simultaneously, traditional pathological features, such as TNM staging, lymph node metastasis, and vascular invasion, can be extracted. These features have been clinically proven to be associated with the prognosis of colorectal cancer. By integrating TLS features with traditional pathological features, a standardized feature set can be formed, fully utilizing information from different types of features, laying the foundation for subsequent predictive model construction, and improving the model's predictive ability.
[0053] S104. The standardized feature set is associated and matched with long-term follow-up data to construct a feature-recurrence status corresponding dataset. The feature-recurrence status corresponding dataset is then processed by univariate logistic regression analysis and multivariate logistic regression analysis to screen out independent influencing factors of colorectal cancer recurrence.
[0054] Specifically, by associating and matching standardized feature sets with long-term follow-up data, pathological features can be combined with clinical outcomes to construct a feature-relapse status correspondence dataset. Then, based on statistical principles, univariate logistic regression analysis can be used to initially screen features associated with relapse. This involves calculating the odds ratio (OR) and confidence interval (CI) for each feature to assess the strength of the association between the feature and relapse, and using the p-value to determine its statistical significance. Secondly, multivariate logistic regression analysis can be used to simultaneously consider the interactions between multiple features, screening out factors that truly have independent predictive value for relapse, providing a scientific basis for subsequent model construction.
[0055] S105. Based on independent influencing factors, construct a nomogram prediction model, perform a verification operation on the nomogram prediction model, and use the verified nomogram prediction model as the finalized nomogram prediction model.
[0056] Specifically, based on the selected independent influencing factors, a nomogram prediction model can be constructed using statistical tools such as R software. A nomogram is a visual prediction tool that quantifies and intuitively displays the contributions of multiple predictive factors. By calculating and summing the scores of each factor, a predicted relapse risk value for the patient can be obtained. The nomogram model can then be validated, including assessments of discrimination, calibration, and clinical usability, to ensure its reliability. Discrimination can be measured using the C-index and the area under the receiver operating characteristic curve (AUC), both of which reflect the model's ability to distinguish between relapsed and non-relapsed patients. Calibration can be assessed by comparing the consistency between predicted and actual values using a calibration curve, while clinical usability can be evaluated using a clinical decision curve to assess the model's benefits in practical applications. A nomogram model that passes these validations can then be used as a finalized nomogram prediction model in clinical practice.
[0057] S106. For patients who have recently undergone colorectal cancer surgery, repeat S102 and S103 sequentially to obtain a standardized feature set of patients who have recently undergone colorectal cancer surgery. Input the standardized feature set of patients who have recently undergone colorectal cancer surgery into the morphological nomogram prediction model to obtain the recurrence risk assessment results of patients who have recently undergone colorectal cancer surgery.
[0058] Specifically, for newly diagnosed colorectal cancer patients, the slicing procedure in S102 and the feature extraction process in S103 can be repeated. This involves obtaining paraffin-embedded tumor tissue blocks from the patient, preparing sections, dewaxing and hydrating them, removing antigens, and performing immunohistochemical staining to obtain a processed section set. Then, based on the sections and clinicopathological data, TLS features and traditional pathological features are extracted and integrated to form a standardized feature set. This ensures that the feature data format and extraction standards of the new patient are completely consistent with the model training data, avoiding prediction deviations due to differences in feature extraction and guaranteeing the consistency and accuracy of model application. Subsequently, the standardized feature set of the new patient is input into the nomogram prediction model. This model can automatically output the patient's recurrence risk assessment result based on preset scoring rules and the total score-recurrence probability mapping relationship. This result can be a high, medium, or low risk level or a specific recurrence probability value, providing clinicians with objective quantitative evidence of recurrence risk, replacing traditional subjective judgments based on experience. Finally, for patients at high risk of recurrence, intensive follow-up plans (such as shortening follow-up intervals and increasing the frequency of imaging examinations and tumor marker testing) or adjuvant therapy interventions (such as postoperative chemotherapy and immunotherapy) can be developed to detect recurrence early and intervene in a timely manner. For patients at low risk of recurrence, routine follow-up plans can be adopted to avoid the economic burden and adverse reactions caused by over-treatment.
[0059] The above method first addresses the lack of standardized data integration in existing prediction technologies by screening colorectal cancer subjects and obtaining their pathological and follow-up data. Secondly, it uses HE staining and immunohistochemical staining on tumor tissue to identify and extract tertiary lymphoid structures and traditional pathological features, effectively solving the problem of incomplete feature extraction in existing technologies and enhancing the characterization ability of the features. Furthermore, by integrating tertiary lymphoid structure features with traditional pathological features to form a standardized feature set, the effectiveness of the features is further improved, mitigating the limitations of feature integration. Further, univariate and multivariate logistic regression analysis is used to screen independent influencing factors, constructing and validating a nomogram prediction model, optimizing the model construction process and overcoming the lack of efficient prediction models in existing technologies. Finally, the feature set of new patients is input into the model to obtain recurrence risk assessment results. This method significantly improves the accuracy and reliability of colorectal cancer recurrence risk prediction, providing clinicians with more precise prognostic assessment and personalized treatment decision support.
[0060] In one embodiment, based on HE-stained sections, immunohistochemically stained sections, and clinicopathological data, clusters of non-densely encapsulated lymphocytes in the tumor region and a predefined tumor boundary region are identified. These non-densely encapsulated lymphocyte clusters are defined as tertiary lymphoid structures, and tertiary lymphoid structure features and traditional pathological features are extracted. The tertiary lymphoid structure features include tertiary lymphoid structure density features and tertiary lymphoid structure maturity features. The tertiary lymphoid structure density features, tertiary lymphoid structure maturity features, and traditional pathological features are integrated to form a standardized feature set, including:
[0061] A double-blind method is used to examine HE-stained sections and immunohistochemically stained sections to obtain a first examination result and a second examination result. If the first examination result and the second examination result are inconsistent, an arbitration instruction is generated to indicate the arbitration of the first examination result and the second examination result. The arbitration instruction is sent to the preset arbitrator terminal and a unified examination result is received from the arbitrator terminal.
[0062] Based on HE staining sections, immunohistochemical staining sections, and unified slide reading results, clusters of non-dense-capsulated lymphocytes formed by the aggregation of CD20+B cells and CD3+T cells were identified in the tumor area and the pre-defined tumor boundary area. These non-dense-capsulated lymphocyte clusters were defined as tertiary lymphoid structures.
[0063] Based on HE-stained sections and unified slide reading results, the number of tertiary lymphoid structures in the preset high-power field of HE-stained sections was counted. The number of tertiary lymphoid structures was classified according to the standard that 1+ corresponds to 1-5, 2+ corresponds to 6-10, and 3+ corresponds to more than 10, so as to obtain the density characteristics of tertiary lymphoid structures.
[0064] Based on CD21 immunohistochemical staining sections, CD23 immunohistochemical staining sections, and unified slide reading results, FDC networks and germinal centers were identified. FDC networks were determined by CD21 expression, and germinal centers were determined by CD23 expression. The maturity of the three-level lymphoid structures was graded according to the criteria of E-TLS corresponding to CD21- and CD23-, PFL-TLS corresponding to CD21+ and CD23-, and SFL-TLS corresponding to CD21+ and CD23+, thus obtaining the maturity characteristics of the three-level lymphoid structures.
[0065] Based on HE-stained sections, immunohistochemical stained sections, and clinicopathological data, we extracted vascular invasion status, nerve invasion status, TNM stage, lymph node metastasis status, differentiation degree, and microsatellite status. The differentiation degree was divided into high, medium, and low differentiation, and the microsatellite status was divided into stable and unstable status. Vascular invasion status, nerve invasion status, TNM stage, lymph node metastasis status, differentiation degree, and microsatellite status were used as traditional pathological features.
[0066] The density characteristics of tertiary lymphoid structures, the maturity characteristics of tertiary lymphoid structures, and traditional pathological characteristics are integrated and processed to form a standardized feature set.
[0067] Specifically, a double-blind method is used to review HE-stained and immunohistochemically stained slides. This eliminates subjective bias and information interference from the reviewers, ensuring the objectivity and reliability of the review results. As an example, two qualified pathologists, who were not involved in sample collection or experimental design, act as independent reviewers. Without informing each other of their review results, they separately observe and judge the same batch of HE-stained and immunohistochemically stained slides. During the review process, the reviewers only record the morphological characteristics of the slides and the expression of immunohistochemical markers, without obtaining additional information such as the subjects' clinicopathological data or follow-up results. Finally, a first review result and a second review result are formed. Both the first and second review results include the location markings of suspected lymphocyte clusters in the slides, judgments of cell composition, and preliminary classification conclusions. If inconsistencies are found between the first and second review results—that is, different conclusions on the identification and classification of the same suspected structure—an arbitration process is initiated. At this point, the system can generate an arbitration instruction, directing the arbitration of the first and second slide reading results, and send the arbitration instruction to a preset arbitrator terminal. A senior pathologist with many years of clinical pathological diagnosis experience can act as the arbitrator. Through this instruction, the arbitrator can simultaneously obtain the slide reading results from both readers, the corresponding stained sections, and the original image data of related markers. By carefully re-examining the cell morphology, immunohistochemical marker expression intensity, and distribution pattern of the disputed areas in the slides, and combining this with the core definition and identification criteria of the tertiary lymphoid structure, the system analyzes and judges the disputed points one by one, ultimately forming a unified slide reading result, which is then sent to the system. This unified slide reading result can serve as the basis for subsequent feature extraction, ensuring the uniqueness and accuracy of the slide reading information relied upon by subsequent steps.
[0068] Specifically, based on HE-stained sections, immunohistochemically stained sections, and standardized slide reading results, when identifying tertiary lymphoid structures within the tumor region and a pre-defined tumor boundary region, the pre-defined tumor boundary region can be defined as the area extending outward from the tumor tissue edge within a certain range, such as the area within 7 mm of the tumor boundary. This region is where the immune response is most active in the tumor microenvironment, and tertiary lymphoid structures are mostly concentrated here. The identification process is guided by standardized slide reading results, focusing on observing the lymphocyte aggregation within this region. For example, HE-stained sections can observe the overall morphology and aggregation pattern of cells. If densely aggregated lymphocytes without a dense fibrous capsule are observed, it is preliminarily identified as a suspected tertiary lymphoid structure. Subsequently, verification can be performed using CD3 and CD20 immunohistochemical staining sections. CD3 antibodies specifically bind to antigens on T cell membranes, while CD20 antibodies specifically bind to antigens on B cell membranes. If the suspected cell cluster simultaneously exhibits aggregations of both CD3-positive and CD20-positive cells, and these two cell types intertwine to form a specific spatial structure, then the non-densely encapsulated lymphocyte cluster can be confirmed as a tertiary lymphoid structure. This method of "preliminary morphological assessment + immunohistochemical specific verification" effectively avoids misclassifying simple T cell clusters, B cell clusters, or other inflammatory cell aggregations as tertiary lymphoid structures, ensuring the specificity of the identification results.
[0069] Specifically, when extracting the density features of tertiary lymphoid structures based on HE-stained sections and standardized slide reading results, the number of preset high-power fields can be determined first. By selecting multiple high-power fields for counting, sampling errors caused by single fields can be avoided, ensuring the objectivity of density assessment. For example, according to the distribution areas of tertiary lymphoid structures marked in the standardized slide reading results, within 10 high-power fields of the HE-stained section, the reader confirms the tertiary lymphoid structures in each field based on the standardized slide reading results. Following the principle of "counting if completely contained within the field or if more than half of the area is within the field," the number of tertiary lymphoid structures in each high-power field is counted. Subsequently, the count results of all preset high-power fields are summarized to obtain the total count of tertiary lymphoid structures in the sample. The total count is graded according to a preset grading standard. 1+ corresponds to a total count of 1-5, indicating low density of tertiary lymphoid structures; 2+ corresponds to a total count of 6-10, indicating moderate density; and 3+ corresponds to a total count greater than 10, indicating high density. This grading process transforms continuous count data into ordered classification features, namely, tertiary lymphoid structure density features. These features can intuitively reflect the enrichment degree of tertiary lymphoid structures in the tumor microenvironment, providing a quantitative basis for subsequent recurrence risk analysis.
[0070] Furthermore, when extracting the maturity characteristics of tertiary lymphoid structures based on CD21 immunohistochemical staining sections, CD23 immunohistochemical staining sections, and unified slide reading results, key functional components of tertiary lymphoid structures can be identified through specific immunohistochemical markers, thereby determining their developmental maturity. Specifically, CD21 antibodies can specifically recognize antigens on the surface of follicular dendritic cells (FDCs), therefore, areas positive for CD21 immunohistochemical staining indicate the location of the FDC network; CD23 antibodies can specifically recognize antigens on the surface of cells related to germinal centers, therefore, areas positive for CD23 immunohistochemical staining indicate the location of the germinal centers. For example, the location of each tertiary lymphoid structure can be determined by combining unified slide reading results, and the staining of that area can be observed on the corresponding CD21 and CD23 immunohistochemical staining sections. If the tertiary lymphoid structure is negative for both CD21 and CD23 staining, it indicates that it has not formed an FDC network and germinal centers, and is in an early developmental stage, thus being classified as E-TLS. If CD21 staining is positive and CD23 staining is negative, it indicates that a follicular network (FDC) has formed but no germinal centers, indicating a primary follicular stage, and is classified as PFL-TLS. If both CD21 and CD23 staining are positive, it indicates that both an FDC and germinal centers are present, indicating a secondary follicular stage, and is classified as SFL-TLS. Through this grading method based on the presence or absence of key functional components, the maturity characteristics of the tertiary lymphoid structures can be obtained. These characteristics directly reflect the degree of immune function perfection of the tertiary lymphoid structures; the higher the maturity of the tertiary lymphoid structures, the stronger their ability to initiate anti-tumor immune responses.
[0071] Specifically, when extracting traditional pathological features based on HE-stained sections, immunohistochemical stained sections, and clinicopathological data, vascular invasion can be determined by observing whether tumor cells have invaded blood vessels or lymphatic vessels in HE-stained sections. If tumor cells are found to form emboli or infiltrate within the blood vessels, vascular invasion is considered. Nerve invasion can be determined by observing whether tumor tissue has invaded the perineum and nerve fiber bundles. If tumor cells grow around the nerve or invade the nerve parenchyma, nerve invasion is considered. TNM staging can be determined based on preset staging criteria, combined with tumor invasion depth, lymph node metastasis, and distant metastasis. Lymph node metastasis can be determined by the presence of tumor cells in the resected lymph nodes through postoperative pathological examination. Differentiation degree can be determined based on the morphological similarity between tumor cells and normal colorectal epithelial cells. The closer the morphology is to normal cells, the higher the degree of differentiation, and vice versa. Based on this, it can be divided into highly differentiated and poorly differentiated categories. Microsatellite status can be determined by immunohistochemical detection of the expression of relevant mismatch repair proteins or by gene detection methods. If the expression of mismatch repair proteins is normal or there are no gene abnormalities, the status is stable; otherwise, it is unstable. Finally, the above-mentioned vascular invasion status, nerve invasion status, TNM stage, lymph node metastasis status, differentiation degree, and microsatellite status can be unified as traditional pathological features. These features are all key indicators that have been clinically proven to be closely related to the biological behavior and prognosis of colorectal cancer.
[0072] This approach integrates tertiary lymphoid structure density features, tertiary lymphoid structure maturity features, and traditional pathological features. For example, all features can be uniformly encoded first. For categorical features (such as 1+, 2+, and 3+ for tertiary lymphoid structure density, E-TLS, PFL-TLS, and SFL-TLS for maturity, and presence / absence of vascular invasion), numerical encoding can be used to convert them into numerical forms suitable for statistical analysis. For ordered categorical features, corresponding continuous values can be assigned according to their logical order, ensuring that the logical relationships between features are not affected by encoding. After encoding, arranging all features in a preset order forms a structured, standardized feature set. This feature set includes both tertiary lymphoid structure features reflecting the immune status of the tumor microenvironment and traditional pathological features reflecting the biological characteristics of the tumor itself, achieving a systematic integration of multi-dimensional features and providing comprehensive data input for the subsequent construction of a recurrence risk prediction model.
[0073] In one embodiment, a standardized feature set is correlated and matched with long-term follow-up data to construct a feature-recurrence status corresponding dataset. This dataset is then processed sequentially using univariate logistic regression analysis and multivariate logistic regression analysis to screen for independent influencing factors of colorectal cancer recurrence, including:
[0074] The features in the standardized feature set are associated and matched with the relapse status in the long-term follow-up data to construct a feature-relapse status correspondence dataset.
[0075] Univariate logistic regression analysis was performed on each matching feature in the feature-recurrence status dataset to obtain the analysis coefficient of each matching feature. Matching features with analysis coefficients less than the preset coefficient threshold were used as candidate factors associated with colorectal cancer recurrence. The matching features included age, TNM stage, lymph node metastasis status, vascular invasion status, nerve invasion status, differentiation degree, tertiary lymphoid structure density characteristics, and tertiary lymphoid structure maturity characteristics.
[0076] The candidate factors were subjected to multivariate logistic regression analysis to eliminate non-independent influencing factors, and the independent influencing factors of colorectal cancer recurrence were obtained. The independent influencing factors included the density characteristics of tertiary lymphoid structures, the maturity characteristics of tertiary lymphoid structures, the vascular invasion status, and the nerve invasion status.
[0077] Specifically, the standardized feature set and long-term follow-up data can be preprocessed to ensure that both contain unique subject identification information. This identification information can be the subject's medical record number or sample number, used to uniquely associate all data of the same subject. Each feature in the standardized feature set has been uniformly coded, including numerical conversions of categorical features such as tertiary lymphoid structure density, tertiary lymphoid structure maturity, and vascular invasion status. The relapse status in the long-term follow-up data is coded in a binary manner: 1 for a subject experiencing relapse and 0 for not experiencing relapse. The association matching process can be performed using data processing software to perform one-to-one matching based on the subject identification information, binding all feature data of each subject to the corresponding relapse status code, forming a feature-relapse status mapping dataset. Furthermore, a data verification mechanism can be set up during the matching process to automatically filter out missing matching identifiers, duplicate matches, or abnormal data where features do not correspond to relapse statuses, ensuring the completeness and accuracy of the dataset. The final dataset has one subject per row and one feature or relapse status per column, providing a standardized input format for subsequent regression analysis.
[0078] Specifically, when performing univariate logistic regression analysis on each matching feature in the feature-recurrence status dataset, a univariate logistic regression model can be constructed to evaluate the statistical association strength between each feature and colorectal cancer recurrence, thus screening out potential recurrence-related factors. The matching features include age, TNM stage, lymph node metastasis status, vascular invasion status, neural invasion status, differentiation degree, tertiary lymphoid structure density characteristics, and tertiary lymphoid structure maturity characteristics. These features have all been encoded and can be directly used for model calculation. The analysis process can use recurrence status as the dependent variable (binary variable) and each matching feature as an independent variable to construct a logistic regression model. The core formula of the model is:
[0079]
[0080] in, Indicates the probability of recurrence. Log-dominance ratio The value ranges from 0 to 1, representing the probability of relapse in the subject; The model intercept reflects the baseline relapse risk when this feature is not present. The regression coefficient of this feature is used to measure the strength and direction of the feature's influence on recurrence risk. A positive value indicates that the feature increases the risk of recurrence, while a negative value indicates that the feature reduces the risk of recurrence. This is the encoded value of the matching feature being analyzed.
[0081] The regression coefficients for each matching feature can be obtained through model calculation. The analysis coefficients consist of parameters including the odds ratio (OR), 95% confidence interval (95% CI), and the p-value of the statistical test. The OR is a key indicator measuring the strength of the association between a feature and recurrence; an OR > 1 indicates that the feature is a risk factor for recurrence, while an OR < 1 indicates that the feature is a protective factor against recurrence. The 95% CI is used to assess the reliability of the OR; if the 95% CI does not include 1, the association between the feature and recurrence is statistically significant. A pre-set threshold of p-value less than 0.05 was used. Matching features with p-values less than this threshold were considered candidate factors associated with colorectal cancer recurrence, meaning that these features, acting alone, showed a significant statistical association with the risk of colorectal cancer recurrence and were worthy of further inclusion in multivariate analysis.
[0082] Furthermore, performing multivariate logistic regression analysis on the candidate factors can eliminate mutual interference (i.e., confounding effects) among them, accurately identifying factors with independent predictive value for colorectal cancer recurrence. For example, all candidate factors can be simultaneously included in a multivariate logistic regression model, with recurrence status as the dependent variable. A stepwise screening method can be used to fit the model. The core logic of the stepwise screening method is to dynamically adjust the variables in the model according to preset inclusion and exclusion criteria. The inclusion criterion can be set as a p-value less than 0.05 when the variable enters the model, and the exclusion criterion can be set as a p-value greater than 0.10 after the variable enters the model. Through iterative calculation, confounding factors that have no independent impact on recurrence can be gradually removed from the model, ultimately retaining factors in the model with a p-value less than 0.05.
[0083] During model fitting, the regression coefficient, OR value, 95% CI, and P-value for each candidate factor need to be recalculated. The OR value at this point has corrected for the influence of other candidate factors and can truly reflect the independent effect of the factor on relapse risk. For example, the tertiary lymphoid structure density characteristic showed an association with relapse in univariate analysis. After being included in the multivariate model, even after correcting for the influence of other factors such as vascular and neurological invasion status, its P-value is still less than 0.05, and the OR value is less than 1, indicating that its protective effect against relapse is independent of other factors. Conversely, some candidate factors (such as age) showed an association with relapse in univariate analysis, but after being included in the multivariate model, their P-value is greater than 0.05, indicating that their influence on relapse can be explained by other factors and is not an independent influencing factor. Therefore, such factors can be identified as non-independent influencing factors and removed. The independent influencing factors for colorectal cancer recurrence can be identified as follows: tertiary lymphoid structure density, tertiary lymphoid structure maturity, vascular invasion status, and neurological invasion status. Among these, tertiary lymphoid structure density and tertiary lymphoid structure maturity are independent protective factors, while vascular invasion status and neurological invasion status are independent risk factors. These factors together constitute the core variables for constructing the subsequent predictive model, ensuring that the model focuses on the key pathological and immune features that truly affect recurrence.
[0084] In one embodiment, such as Figure 2 As shown, a nomogram prediction model is constructed based on independent influencing factors. The nomogram prediction model is then validated, and the validated model is used as the finalized nomogram prediction model. This includes:
[0085] S201: Based on independent influencing factors, a visual nomogram prediction model is constructed. Each independent influencing factor is assigned a fixed score to obtain the corresponding score, and a direct mapping relationship between the score and the probability of colorectal cancer recurrence is established.
[0086] S202: Calculate the consistency index of the nocline plot prediction model, plot the time-dependent ROC curve, calculate the area under the curve of the time-dependent ROC curve, and perform discrimination verification on the nocline plot prediction model based on the consistency index and the area under the curve to obtain the discrimination verification results.
[0087] S203: Plot a calibration curve, compare the colorectal cancer recurrence probability predicted by the nomogram prediction model with the actual recurrence rate recorded in long-term follow-up data, verify the consistency, and determine that the calibration degree of the nomogram prediction model meets the preset standard when the consistency reaches the preset standard, and obtain the calibration degree verification result.
[0088] S204: Plot clinical decision curves, evaluate the clinical benefits of the nomogram prediction model under different risk thresholds, and determine the clinical usability of the nomogram prediction model when it shows positive benefits, thus obtaining the clinical usability verification results.
[0089] S205: Once the discrimination verification results, calibration verification results, and clinical usability verification results all meet the standards, the nomogram prediction model is determined to be a valid nomogram prediction model, and the valid nomogram prediction model is output as the finalized nomogram prediction model.
[0090] Specifically, a visual nomogram prediction model is constructed based on independent influencing factors. This transforms the abstract mathematical relationships of a multivariate logistic regression model into an intuitive visualization tool, thereby lowering the barrier to clinical application. The independent influencing factors include tertiary lymphoid structure density characteristics, tertiary lymphoid structure maturity characteristics, vascular invasion status, and neural invasion status. These factors have been validated through multivariate analysis to have independent predictive value, and their impact on recurrence risk can be quantified using regression coefficients. Illustratively, a specified version of statistical analysis software and the corresponding nomogram construction toolkit can be used. First, the classification codes of each independent influencing factor are correlated with the regression coefficients in the multivariate logistic regression model. The larger the absolute value of the regression coefficient, the higher the weight of the factor's impact on recurrence risk, and the higher the assigned fixed score. For example, the regression coefficient for tertiary lymphoid structure maturity characteristics has the largest absolute value, and its corresponding scores for different grades (E-TLS, PFL-TLS, SFL-TLS) show the most significant differences. Vascular invasion status, as a risk factor, scores for the "present" state are significantly higher than those for the "absent" state. After the scoring is assigned, a direct mapping relationship between the total score and the probability of colorectal cancer recurrence can be established. The core basis of this mapping relationship is the multivariate logistic regression probability formula:
[0091]
[0092] in, This represents the recurrence probability of colorectal cancer patients, with a value ranging from 0 to 1; It is a natural constant; The model intercept reflects the baseline relapse risk when no independent influencing factors are present. to These are the regression coefficients corresponding to the density characteristics of tertiary lymphoid structures, the maturity characteristics of tertiary lymphoid structures, the vascular invasion status, and the neural invasion status, respectively. Their positive and negative values represent protective effects or risk factors, respectively. to These are the classification codes for each independent influencing factor (e.g., for the tertiary lymphoid structure density feature, 1+ is coded as 1, 2+ as 2, and 3+ as 3; for vascular invasion status, "no" is coded as 0 and "present" as 1). Using this formula, the total score obtained by summing the fixed scores corresponding to each independent influencing factor is converted into a specific recurrence probability, ultimately forming a visual nomogram. The nomogram includes a scoring axis, a grading axis for each independent influencing factor, and a recurrence probability axis. Clinical users can find the corresponding scores for each patient's characteristics on the axes, sum them, map them to the probability axis, and directly read the recurrence probability for rapid assessment.
[0093] Specifically, by calculating the concordance index of the nomogram prediction model and plotting the time-dependent ROC curve, the model's discriminative power—its ability to distinguish between relapsed and non-relapsed patients—can be verified. The concordance index (C-index) is calculated based on pairwise comparisons of all subjects. If the model correctly predicts a patient with a higher risk of relapse (i.e., a shorter relapse time or a patient who has already relapsed) within a pair, it is counted as a concordant prediction. The C-index is the ratio of the number of concordant predictions to the total number of pairwise comparisons, ranging from 0.5 to 1.0, where 0.5 indicates no discriminative ability and 1.0 indicates perfect discrimination. The calculation process is implemented using statistical software that calls a dedicated algorithm, and the 95% confidence interval of the C-index can be output to assess its reliability. The plotting of the time-dependent ROC curve considers the impact of follow-up time on relapse status to avoid bias caused by traditional ROC curves ignoring differences in follow-up time. For example, time intervals can be divided according to preset follow-up time points. For each time interval, the recurrence status within that interval is the dependent variable, and the recurrence probability predicted by the model is the independent variable. An ROC curve is plotted, and the area under the curve (AUC) is calculated. The AUC value also ranges from 0.5 to 1.0; a larger value indicates a stronger discriminative ability of the model within that time interval. When validating discriminative ability based on the concordance index and area under the curve, the preset criteria can be a concordance index ≥ 0.7 and an area under the curve ≥ 0.7. When both are met, the discriminative validation result can be considered satisfactory, indicating that the model can effectively distinguish patients with different recurrence risk levels.
[0094] Specifically, when plotting calibration curves for calibration verification, the consistency between the model's predicted recurrence probability and the actual recurrence rate can be verified to ensure that the model's predictions closely reflect clinical reality. For example, all subjects are first divided into several groups (e.g., 4-5 groups) according to the model's predicted recurrence probability from low to high. The average predicted recurrence probability for each group is calculated. Simultaneously, based on long-term follow-up data, the actual recurrence rate for each group (i.e., the ratio of the number of recurring cases in each group to the total number of cases) can be calculated. Then, a calibration curve is plotted with the average predicted recurrence probability on the horizontal axis and the actual recurrence rate on the vertical axis, and a trend line is fitted. The deviation of the trend line from the ideal diagonal (predicted probability = actual recurrence rate) is calculated. The preset consistency standard can be that the average deviation between the trend line and the ideal diagonal is less than a preset threshold, and the difference between the actual recurrence rate and the average predicted recurrence probability in each group is not statistically significant. When this standard is met, the calibration of the nomogram prediction model can be determined to be satisfactory, and the calibration verification result is qualified.
[0095] In illustrative terms, when plotting a clinical decision curve to assess clinical applicability, the core objective is to determine the model's net benefit at different clinical risk thresholds; that is, whether the model can bring actual benefits to patients when guiding clinical decisions. The vertical axis of the clinical decision curve can represent standardized net benefit, and the horizontal axis can represent the risk threshold (i.e., the lowest recurrence probability that clinicians determine a patient needs intervention). The formula for calculating standardized net benefit is:
[0096]
[0097] Wherein, TP represents the number of true positive cases (the number of patients predicted as high risk by the model and who actually relapsed), and FP represents the number of false positive cases (the number of patients predicted as high risk by the model but who did not actually relapse). The total number of patients is represented by Threshold, which represents the risk threshold. This formula quantifies the clinical value of the model by balancing the benefits of true positives with the risk of overtreatment due to false positives. During the plotting process, two reference curves can be simultaneously plotted: one for "all patients receiving intervention" and the other for "no intervention for all patients." If the decision curve of the nomogram prediction model falls within most commonly used clinical risk thresholds (e.g., 0.1-0.5), and its standardized net benefit is higher than both reference curves, showing a significant positive benefit, then the clinical applicability of the nomogram prediction model can be determined to be satisfactory, and the clinical applicability verification result is considered qualified.
[0098] When the discrimination, calibration, and clinical applicability validation results all meet the standards, it indicates that the nomogram prediction model possesses both good discrimination ability and calibration accuracy, and can provide practical value for clinical decision-making. Therefore, the nomogram prediction model can be determined as a validated nomogram prediction model. Finally, this validated nomogram prediction model can be output as a finalized nomogram prediction model. This finalized model integrates tumor microenvironment immune characteristics (density and maturity of tertiary lymphoid structures) with traditional pathological risk factors (vascular invasion and neural invasion). Through visualization design, complex prediction models can be clinically translated, thus providing an accurate and convenient tool for assessing the risk of colorectal cancer recurrence.
[0099] Based on the same inventive concept, this application also provides a pathological feature-based colorectal cancer recurrence prediction system for implementing the above-described method for predicting colorectal cancer recurrence based on pathological features. The solution provided by this system is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the pathological feature-based colorectal cancer recurrence prediction system provided below can be found in the limitations of the pathological feature-based colorectal cancer recurrence prediction method described above, and will not be repeated here.
[0100] In one exemplary embodiment, such as Figure 3 As shown, a colorectal cancer recurrence prediction system 300 based on pathological features is provided, including:
[0101] The data acquisition module 301 is used to perform step 1, which includes: screening colorectal cancer subjects based on preset inclusion and exclusion criteria, and acquiring paraffin-embedded tumor tissue blocks, clinicopathological data and long-term follow-up data of colorectal cancer subjects. The long-term follow-up data includes records of whether the colorectal cancer subjects have relapsed.
[0102] The slide acquisition module 302 is used to perform step 2, which includes: acquiring a slide processing set including HE-stained slides and immunohistochemical stained slides. The immunohistochemical stained slides include CD3 immunohistochemical stained slides, CD20 immunohistochemical stained slides, CD21 immunohistochemical stained slides and CD23 immunohistochemical stained slides; wherein, the slide processing set is obtained by performing a slide operation on the paraffin-embedded block of tumor tissue.
[0103] The feature extraction module 303 is used to perform step 3, which includes: based on HE-stained sections, immunohistochemical stained sections, and clinicopathological data, identifying non-dense capsule lymphocyte clusters in the tumor region and the preset tumor boundary region, defining the non-dense capsule lymphocyte clusters as tertiary lymphoid structures, and extracting tertiary lymphoid structure features and traditional pathological features. The tertiary lymphoid structure features include tertiary lymphoid structure density features and tertiary lymphoid structure maturity features; integrating the tertiary lymphoid structure density features, tertiary lymphoid structure maturity features, and traditional pathological features to form a standardized feature set.
[0104] The key screening module 304 is used to perform step 4, which includes: performing association matching processing between the standardized feature set and long-term follow-up data to construct a feature-recurrence status corresponding dataset, and processing the feature-recurrence status corresponding dataset through univariate logistic regression analysis and multivariate logistic regression analysis in turn to screen out independent influencing factors of colorectal cancer recurrence.
[0105] The verification module 305 is used to execute step 5, which includes: constructing a nomogram prediction model based on independent influencing factors, performing a verification operation on the nomogram prediction model, and using the verified nomogram prediction model as the finalized nomogram prediction model.
[0106] The recurrence risk assessment module 306 is used to perform step 6, which includes: repeating steps 2 and 3 sequentially for patients who have recently undergone colorectal cancer surgery to obtain a standardized feature set of patients who have recently undergone colorectal cancer surgery; inputting the standardized feature set of patients who have recently undergone colorectal cancer surgery into a morphological nomogram prediction model to obtain the recurrence risk assessment result of patients who have recently undergone colorectal cancer surgery.
[0107] In one exemplary embodiment, the present invention also provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the method for predicting recurrence of colorectal cancer based on pathological characteristics according to this application. A multi-core processor is preferred to improve the parallel processing capability of the system. The memory provides sufficient temporary storage space to support program execution and data processing. The memory capacity should be large enough to accommodate large amounts of data and computational tasks.
[0108] In one exemplary embodiment, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for predicting colorectal cancer recurrence based on pathological features of the present application. The computer-readable storage medium may include: a read-only memory, a random access memory, a solid-state drive, or an optical disk, etc.
[0109] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.
Claims
1. A method for predicting recurrence of colorectal cancer based on pathological characteristics, characterized in that, The method includes: S1. Based on preset inclusion and exclusion criteria, colorectal cancer subjects are screened out, and the paraffin-embedded tumor tissue, clinicopathological data, and long-term follow-up data of the colorectal cancer subjects are obtained. The long-term follow-up data includes records of whether the colorectal cancer subjects have relapsed. S2, Obtain a slice processing set including HE-stained slices and immunohistochemical stained slices, wherein the immunohistochemical stained slices include CD3 immunohistochemical stained slices, CD20 immunohistochemical stained slices, CD21 immunohistochemical stained slices and CD23 immunohistochemical stained slices; wherein, the slice processing set is obtained by performing a slice operation on the paraffin-embedded tumor tissue block. S3, based on the HE-stained sections, the immunohistochemically stained sections, and the clinicopathological data, clusters of non-dense-capsule lymphocytes in the tumor region and the preset tumor boundary region are identified. These non-dense-capsule lymphocyte clusters are defined as tertiary lymphoid structures, and tertiary lymphoid structure features and traditional pathological features are extracted. The tertiary lymphoid structure features include tertiary lymphoid structure density features and tertiary lymphoid structure maturity features. The tertiary lymphoid structure density features, the tertiary lymphoid structure maturity features, and the traditional pathological features are integrated to form a standardized feature set. S4, the standardized feature set is associated and matched with the long-term follow-up data to construct a feature-recurrence status corresponding dataset. The feature-recurrence status corresponding dataset is then processed by univariate logistic regression analysis and multivariate logistic regression analysis to screen out independent influencing factors of colorectal cancer recurrence. S5. Based on the independent influencing factors, construct a nomogram prediction model and perform a verification operation on the nomogram prediction model. The nomogram prediction model that passes the verification is used as the finalized nomogram prediction model. S6. For patients who have recently undergone colorectal cancer surgery, repeat S2 and S3 sequentially to obtain a standardized feature set of the patients. Input the standardized feature set of the patients who have recently undergone colorectal cancer surgery into the morphological nomogram prediction model to obtain the recurrence risk assessment results of the patients who have recently undergone colorectal cancer surgery.
2. The method according to claim 1, characterized in that, Based on the HE-stained sections, the immunohistochemical stained sections, and the clinicopathological data, clusters of non-dense-capsulated lymphocytes in the tumor region and the preset tumor boundary region are identified. These non-dense-capsulated lymphocyte clusters are defined as tertiary lymphoid structures, and tertiary lymphoid structure features and traditional pathological features are extracted. The tertiary lymphoid structure features include tertiary lymphoid structure density features and tertiary lymphoid structure maturity features. The density characteristics of the tertiary lymphoid structures, the maturity characteristics of the tertiary lymphoid structures, and the traditional pathological characteristics are integrated and processed to form a standardized feature set, including: The HE-stained sections and the immunohistochemical-stained sections are processed using a double-blind method to obtain a first review result and a second review result. If the first review result and the second review result are inconsistent, an arbitration instruction is generated to indicate arbitration processing of the first review result and the second review result. The arbitration instruction is sent to a preset arbitrator terminal, and a unified review result sent by the arbitrator terminal is received. Based on the HE-stained sections, the immunohistochemical stained sections, and the unified slide reading results, the clusters of non-dense-capsular lymphocytes formed by the aggregation of CD20+B cells and CD3+T cells were identified in the tumor region and the preset tumor boundary region, and the clusters of non-dense-capsular lymphocytes were defined as the tertiary lymphoid structures. Based on the HE-stained sections and the unified slide reading results, the number of the tertiary lymphoid structures in the preset high-power field of view of the HE-stained sections is counted, and the number of the tertiary lymphoid structures is classified according to the standard that 1+ corresponds to 1-5, 2+ corresponds to 6-10, and 3+ corresponds to more than 10, so as to obtain the density characteristics of the tertiary lymphoid structures. Based on the CD21 immunohistochemical staining sections, the CD23 immunohistochemical staining sections, and the unified slide reading results, the FDC network and germinal centers were identified. The FDC network was determined by CD21 expression, and the germinal centers were determined by CD23 expression. The maturity of the three-level lymphoid structures was graded according to the criteria of E-TLS corresponding to CD21- and CD23-, PFL-TLS corresponding to CD21+ and CD23-, and SFL-TLS corresponding to CD21+ and CD23+, to obtain the maturity characteristics of the three-level lymphoid structures. Based on the HE-stained sections, the immunohistochemically stained sections, and the clinicopathological data, the vascular invasion status, nerve invasion status, TNM stage, lymph node metastasis status, differentiation degree, and microsatellite status were extracted. The differentiation degree was divided into high, medium, and low differentiation, and the microsatellite status was divided into stable and unstable states. The vascular invasion status, nerve invasion status, TNM stage, lymph node metastasis status, differentiation degree, and microsatellite status were used as the traditional pathological features. The density characteristics of the tertiary lymphoid structures, the maturity characteristics of the tertiary lymphoid structures, and the traditional pathological characteristics are integrated and processed to form the standardized feature set.
3. The method according to claim 1, characterized in that, The standardized feature set is associated and matched with the long-term follow-up data to construct a feature-recurrence status corresponding dataset. This dataset is then processed sequentially using univariate logistic regression analysis and multivariate logistic regression analysis to screen out independent influencing factors for colorectal cancer recurrence, including: Each feature in the standardized feature set is associated and matched with the relapse status in the long-term follow-up data to construct a feature-relapse status corresponding dataset. Univariate logistic regression analysis was performed on each matching feature in the dataset corresponding to the feature-recurrence status to obtain the analysis coefficient of each matching feature. Matching features with analysis coefficients less than a preset coefficient threshold were used as candidate factors associated with the recurrence of colorectal cancer. The matching features include age, TNM stage, lymph node metastasis status, vascular invasion status, nerve invasion status, differentiation degree, density characteristics of the tertiary lymphoid structure, and maturity characteristics of the tertiary lymphoid structure. The candidate factors were subjected to multivariate logistic regression analysis to eliminate non-independent influencing factors, and the independent influencing factors of colorectal cancer recurrence were obtained. The independent influencing factors include the density characteristics of the tertiary lymphoid structure, the maturity characteristics of the tertiary lymphoid structure, the vascular invasion status, and the nerve invasion status.
4. The method according to claim 1, characterized in that, The process of constructing a nomogram prediction model based on the independent influencing factors, performing a validation operation on the nomogram prediction model, and using the validated nomogram prediction model as the finalized nomogram prediction model includes: Based on the independent influencing factors, a visual nomogram prediction model is constructed. A fixed score assignment is performed on each independent influencing factor to obtain the corresponding score, and a direct mapping relationship between the score and the probability of colorectal cancer recurrence is established. The consistency index of the nomogram prediction model is calculated, and the time-dependent ROC curve is plotted. The area under the curve of the time-dependent ROC curve is calculated. Based on the consistency index and the area under the curve, the discrimination of the nomogram prediction model is verified, and the discrimination verification result is obtained. A calibration curve is plotted, and the colorectal cancer recurrence probability predicted by the nomogram prediction model is compared with the actual recurrence rate recorded in the long-term follow-up data to verify consistency. When the consistency reaches the preset standard, the calibration degree of the nomogram prediction model is determined to be up to standard, and the calibration degree verification result is obtained. A clinical decision curve is plotted to evaluate the clinical benefit of the nomogram prediction model under different risk thresholds. When the nomogram prediction model shows a positive benefit, the clinical applicability of the nomogram prediction model is determined to be up to standard, and the clinical applicability verification result is obtained. Once the discrimination verification result, the calibration verification result, and the clinical usability verification result all meet the standards, the nomogram prediction model is determined to be the verified nomogram prediction model, and the verified nomogram prediction model is output as the finalized nomogram prediction model.
5. The method according to claim 1, characterized in that, The sectioning process includes section preparation, dewaxing and hydration, antigen retrieval, and immunohistochemical staining.
6. A colorectal cancer recurrence prediction system based on pathological characteristics, characterized in that, The system includes: The data acquisition module is used to perform step 1, which includes: screening colorectal cancer subjects based on preset inclusion and exclusion criteria, and acquiring the paraffin-embedded tumor tissue, clinicopathological data and long-term follow-up data of the colorectal cancer subjects. The long-term follow-up data includes records of whether the colorectal cancer subjects have relapsed. The slide acquisition module is used to perform step 2, which includes: acquiring a slide processing set including HE-stained slides and immunohistochemically stained slides, wherein the immunohistochemically stained slides include CD3 immunohistochemically stained slides, CD20 immunohistochemically stained slides, CD21 immunohistochemically stained slides and CD23 immunohistochemically stained slides; wherein, the slide processing set is obtained by performing a slide operation on the paraffin-embedded block of the tumor tissue. The feature extraction module is used to perform step 3, which includes: based on the HE-stained sections, the immunohistochemically stained sections, and the clinicopathological data, identifying clusters of non-dense-capsule lymphocytes in the tumor region and the preset tumor boundary region; defining the clusters of non-dense-capsule lymphocytes as tertiary lymphoid structures; and extracting tertiary lymphoid structure features and traditional pathological features, including tertiary lymphoid structure density features and tertiary lymphoid structure maturity features; integrating the tertiary lymphoid structure density features, the tertiary lymphoid structure maturity features, and the traditional pathological features to form a standardized feature set. The key screening module is used to perform step 4, which includes: performing association matching processing between the standardized feature set and the long-term follow-up data to construct a feature-recurrence status corresponding dataset, and processing the feature-recurrence status corresponding dataset through univariate logistic regression analysis and multivariate logistic regression analysis in sequence to screen out independent influencing factors of colorectal cancer recurrence. The verification module is used to execute step 5, which includes: constructing a nomogram prediction model based on the independent influencing factors, performing a verification operation on the nomogram prediction model, and using the verified nomogram prediction model as the finalized nomogram prediction model. The recurrence risk assessment module is used to perform step 6, which includes: repeating steps 2 and 3 sequentially for patients who have recently undergone colorectal cancer surgery to obtain a standardized feature set of the patients; inputting the standardized feature set of the patients who have recently undergone colorectal cancer surgery into the morphological nomogram prediction model to obtain the recurrence risk assessment result of the patients who have recently undergone colorectal cancer surgery.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.