Prostate cancer biochemical recurrence prediction system based on pathological sections and its construction method
By combining a multi-instance algorithm with a recurrent cross-attention module and a pseudo-bag strategy, the spatial position relationship and heterogeneity issues in prostate cancer pathological sections were resolved, achieving highly accurate biochemical recurrence prediction and supporting the formulation of personalized treatment plans.
Patent Information
- Application Number
- CN202510740922.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-06-05
AI Technical Summary
Existing technologies have difficulty accurately predicting the risk of biochemical recurrence of prostate cancer after surgery, mainly because they ignore the spatial position relationship and heterogeneity in pathological sections and cannot effectively capture the characteristics of key areas, resulting in limited prediction accuracy.
A multi-instance algorithm based on a recurrent cross-attention module and a pseudo-bag strategy is adopted, combined with a position and spatial information generator and multi-scale feature fusion. Pathological section features are extracted through self-supervised learning to generate high-dimensional feature representations that capture microscopic cellular abnormalities and macroscopic tissue changes.
It improves the accuracy and robustness of prostate cancer biochemical recurrence prediction, reduces diagnosis time, enhances diagnostic consistency and accuracy, and supports the development of personalized treatment plans.
Smart Images

Figure CN120260974B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of prostate cancer biochemical recurrence risk prediction and analysis, and relates to a prostate cancer biochemical recurrence prediction system based on pathological sections and a construction method. Background Art
[0002] After radical prostatectomy, approximately 30% of prostate cancer patients will experience biochemical recurrence (BCR) with elevated PSA levels. This is a key prognostic indicator for cancer recurrence and metastasis. Accurate prediction of BCR risk helps clinicians stratify management and closely follow high-risk patients, but current estimates rely primarily on subjective judgment based on clinical information and pathology reports. The most commonly used method is to assess prognosis based on the prostate-specific antigen (PSA) level at diagnosis, but its accuracy and effectiveness are suboptimal. A second approach is to use the Gleason score, but this assessment is subject to significant variability among pathologists on the same slide, significantly influenced by subjective factors, and therefore ineffective. Internationally, predictive scales are also being used to assess prognosis by integrating multiple components of information. However, these assessments are incomplete, subject to data omissions, and may not yield ideal predictions. Furthermore, pathological analysis and diagnosis is a highly labor-intensive clinical task. Millions of pathology slides are generated daily worldwide and in my country. The time required for pathological diagnosis and the shortage of specialized pathologists pose significant challenges to healthcare quality and the healthcare system. In summary, existing methods are difficult to meet the clinical risk assessment and management of BCR in patients after prostate cancer surgery, so as to make early intervention.
[0003] With the development of technology, pathology slides created from prostatectomy specimens can be digitized, meaning that information can no longer be processed solely by the naked eye. Powered by the powerful processing and analysis capabilities of computers, more meaningful information can be extracted from pathology slides that was previously unnoticed. However, most current AI pathology analysis methods have multiple key flaws when dealing with highly complex tumors such as prostate cancer, which seriously restricts their clinical application value: First, when performing pathology section image analysis, mainstream methods usually extract panoramic pathology section scan images of prostate cancer as two-dimensional features, simply storing spatially distributed image block features in a serialized form, completely ignoring the key spatial positional relationships in tissue morphology. The invasive spatial distribution of prostate cancer is jointly regulated by anatomical zoning, clonal evolution, and microenvironmental interactions, and manifests as multifocal heterogeneity, front-line invasion tendencies, and specific metastatic pathways, with a unique spatial distribution pattern; these algorithms perform poorly in multi-scale integration and are unable to effectively capture and integrate prostate cancer characteristics existing at different scales, such as microscopic changes in nuclear morphology and macroscopic tissue structural damage; second, faced with the high heterogeneity of prostate cancer, as shown by the simultaneous existence of multiple Gleason scores in the same lesion. Existing methods often use averaging to eliminate so-called "noise", but inadvertently erase a few high-level regional features that are crucial for prognostic judgment; thirdly, features closely related to neural invasion, vascular invasion and biochemical recurrence usually only exist in small areas of the entire slice. Existing algorithms often find it difficult to accurately locate these sparsely distributed but critical areas for prognostic assessment in a large amount of background tissue, resulting in limited prediction accuracy.
[0004] While cutting-edge pooling methods, such as max-pooling and mean-pooling, are simple, efficient, and widely used in pathology image analysis, they focus solely on the maximum or average value of features, ignoring the heterogeneity and crucial information of a few critical regions within prostate cancer tissue. The attention-based multi-instance learning (ABMIL) model, widely used in pathology image analysis, introduces an attention mechanism to adaptively weight the importance of different instances, offering some improvements over traditional pooling methods. However, when faced with prostate cancer pathology images characterized by high heterogeneity, multifocality, and sparse distribution of key features, it still tends to distract attention from a large amount of background regions, resulting in insufficient feature capture of prognostic key regions. These technical limitations and clinical challenges urgently require a new method that can address these pathological characteristics of prostate cancer and more efficiently extract pathological features from panoramic pathology slide scans of patients' prostate cancer, thereby accurately predicting biochemical recurrence in patients undergoing surgery. Summary of the Invention
[0005] The present invention aims to develop a prostate cancer biochemical recurrence prediction system based on pathological sections obtained immediately after radical prostatectomy surgery for prostate cancer, to assist physicians in interpreting postoperative recurrence. The system also proposes a prostate cancer biochemical recurrence prediction algorithm based on high-definition panoramic pathology section scan images of post-radical prostatectomy pathology sections. The algorithm utilizes a multi-instance algorithm based on a cyclic cross-attention module and a pseudo-bag strategy. First, a position and spatial information generator module is creatively constructed to comprehensively process image information. A pseudo-bag strategy is then proposed to address the overfitting problem. Second, a strategy for capturing and integrating image features at multiple scales is proposed for the first time in prostate cancer pathology image analysis, ultimately enabling intelligent prostate cancer pathology image analysis to predict the risk of prostate cancer biochemical recurrence. This approach is of great significance for determining patient prognosis and improving treatment methods, and has broad application prospects.
[0006] Specifically, the system of the present invention has the following technical features: it organically integrates core modules such as feature vector processing, position and spatial information generator, recurrent cross-attention, pseudo-packet construction, and multi-scale feature fusion. It efficiently extracts corresponding pathological features of panoramic pathological slice scan images according to the pathological characteristics of prostate cancer samples, thereby predicting biochemical recurrence of postoperative prostate cancer patients with high accuracy. The specific features are as follows:
[0007] (1) Input display module
[0008] Input the patient's basic information and the panoramic pathology slice scan image of the pathology slice, and display the analysis results after the analysis is completed;
[0009] (2) Feature vector processing module
[0010] The processing is performed based on a feature extraction method based on self-supervised learning: the tumor area of the pre-processed panoramic pathology slice scan image at different objective lens magnifications (preferably 10×, 20×, and 40×) is cropped into non-overlapping image blocks (patches) of a certain pixel size (such as 256×256) based on the features at each magnification, which serves as the input of the feature extraction network; then the self-supervised visual encoder UNI is used to automatically extract features from it to generate a feature representation at the panoramic pathology slice scan image level containing multiple low-dimensional feature sets at different magnifications.
[0011] Preferably, the preprocessing method of the panoramic pathology slice scan image is as follows: by adjusting the mean and standard deviation of the image block data, the data is converted into a distribution with zero mean and unit variance, thereby achieving data standardization and image intensity normalization.
[0012] 10×, 20×, and 40× represent low, medium, and high magnifications, respectively. A 10× objective magnification is used for preliminary observation of the specimen's overall structure, helping to identify the overall distribution and extent of cancer foci and their relationship to surrounding tissues. A 40× objective magnification reveals cellular morphological details, such as nuclear atypia and nuclear-cytoplasmic ratio, the tumor-nerve interface, contact between tumor cells and blood / lymphatic vessels, and intravascular micrometastases. A 20× objective magnification is a medium magnification, providing greater clarity than a low-power microscope (10×) but a wider field of view than a high-power microscope (40× and above), and is used for medium-resolution observations in histology and cytology.
[0013] The self-supervised visual encoder UNI is a universal visual encoder for pathology. Its generation process first utilizes a self-supervised learning method called DINOv2. UNI can extract meaningful features from these histological images (such as the overall distribution and range of cancer lesions and their relationship with surrounding tissues, morphological details at the cellular level, etc.) without manual labeling. After the last convolutional layer of UNI, the 1024-dimensional feature vector of each image block is extracted to generate a feature representation at the level of the panoramic pathology slice scan image. Each panoramic pathology slice scan image will generate a set of low-dimensional feature matrices at three objective lens scales, with a size of {N s ×1024}, where N represents the number of image blocks at this scale, s∈{10×, 20×, 40×}.
[0014] In order to minimize overfitting of the model and improve its generalization ability, a variety of data augmentation strategies were introduced in the feature extraction process, including horizontal flipping, vertical flipping, 45-degree rotation, and adjustment of brightness, contrast, saturation and hue in color space transformation, so that the model can learn rotation-invariant and color-invariant image representations.
[0015] For cases with multiple pathology slides, the feature vectors generated from all images of a patient at each objective magnification scale are concatenated and fused to create a representative, unified scale feature package. Each case has three feature packages representing 10×, 20×, and 40× objective magnifications, which serve as input to the subsequent classification network.
[0016] (3) Pseudo-package building module
[0017] Developing a model using a multi-instance approach is challenging, and overfitting problems often occur. In addition, there are factors in prostate cancer tissue that may interfere with judgment, such as benign prostate tissue and its hyperplasia, inflammation, and technical artifacts. The features used to predict BCR may be sparse but critical, and only exist in a few tissues, such as cell populations with high metastatic potential. Therefore, the present invention also adds a pseudo-package construction module, which uses a pseudo-package strategy to divide an original parent package into several smaller pseudo-packages according to the corresponding instance selection strategy, with the same label as the original parent package; extract the top 20% of the features in the pseudo-package ranked from large to small according to the positive probability, greatly improving the model's attention to key areas, thereby reducing the impact of noise.
[0018] The pseudo-packet strategy increases the number of packets and correspondingly reduces the number of instances within each packet. The top 20% of features ranked by positive probability from the pseudo-packet are extracted and concatenated into the model. These high-confidence features are then fed into the attention network to obtain an attention score. Instance features are weighted by the attention score to highlight the regions that are most discriminative for the final classification result. Subsequently, the selected instance features are weighted and fused using the attention score to generate a compressed pseudo-packet feature package representation, which serves as the representative feature of the pseudo-packet. Simultaneously, the original parent packet is also weighted through an independent attention network path to generate a parent packet feature package representation. This is fed into the fusion module or downstream network together with the pseudo-packet feature package for joint training and learning. This captures information about local key regions while preserving overall structural features, improving the model's ability to identify and discriminate heterogeneous lesions in images.
[0019] (4) Position and spatial information generator module
[0020] Based on the spatial distribution pattern of prostate cancer invasiveness, the input panoramic pathology slice scan image level feature representation or pseudo package and parent package feature package representation is first reconstructed into a four-dimensional tensor X reshaped , and then use convolution kernels of different sizes to process the feature data to further extract features with spatial perception and contextual information.
[0021] Preferably, the position and spatial information generator module first reconstructs the input features into a four-dimensional tensor X reshaped : , where X input is the original input feature, dimension is [R, D], four-dimensional tensor X reshaped The dimensions are [R, C, H, W], where C is the number of channels, H and W are the width and height of the image; convolution kernels with different sizes are then used to process the feature data, each kernel size corresponding to a different spatial range (7×7, 5×5, 3×3). These convolution operations produce multi-scale position-aware feature representations in the features: output = Xreshaped + Conv7(X reshaped ) + Conv5(X reshaped ) + Conv3(X reshaped ). The features are then further encoded and integrated. Convolution kernels of different scales complement each other, providing a more comprehensive and richer feature expression: X final = flatten(X onput ).
[0022] By preserving the spatial correlation in pathological images, the modeling of tissue morphological features and the extraction of spatial features are achieved.
[0023] (5) Recurrent Cross-Attention Module
[0024] This module captures dependencies at the pixel level, aggregating long-range contextual information from individual pixels in a cyclic, cross-cutting fashion. It not only preserves detailed features in local regions but also establishes global semantic associations, providing more discriminative feature inputs for subsequent prediction tasks. Convolutional kernels of different scales complement each other, and based on cross-attention, new feature maps with dense and rich contextual information are gradually generated, enabling the model to simultaneously identify microscopic cellular abnormalities and macroscopic tissue changes. Correlations between features are then calculated based on these feature representations, and attention weights are generated after normalizing these correlations.
[0025] Preferably, a new feature map is formed by an iterative cycle: in the first iteration, the feature map X final As input, the module aggregates long-range contextual information from each pixel in a cross-wise manner to produce a new output feature map X final ′, where X final and X final 'The same shape; then the generated output feature map X final ′ is input to the cross attention module again. After multiple iterative processes, the module captures the dependency at the pixel level and gradually generates a new feature map containing dense and rich context information. For any two positions (u x ,u y ) and (θ x ,θ y ), A is the attention map, and the information transfer of the two cycles is as follows:
[0026] .
[0027] Furthermore, microscopic cellular abnormalities include variability in cell size, nuclear atypia, increased chromatin, frequent mitotic figures, and abnormal cellular metabolism. Macroscopic tissue changes include irregular gland size and disordered arrangement, basement membrane disruption, perineural infiltration with associated necrosis and hemorrhage, and interstitial fibrosis. These feature representations are used to calculate inter-feature correlations—that is, the relationships between pixels in different directions—using existing methods (such as the Pearson correlation coefficient and the Spearman rank correlation coefficient). A softmax operation is used to normalize these correlations and generate attention weights. These attention weights are then applied to the calculation of bag-level feature representations, thereby prioritizing important inter-pixel correlations and enhancing the contextual information between features.
[0028] By establishing a dual association between local fine features and global context through the position and spatial information generator module and the recurrent cross-attention module, the model can simultaneously identify microscopic cellular abnormalities and macroscopic tissue structural changes. The fusion architecture simulates the thinking mode of pathologists in the diagnosis process of macroscopic scanning-suspicious area identification-microscopic confirmation, enhancing the ability to understand complex cancer tissue morphology and capturing the high heterogeneity of prostate cancer.
[0029] (6) Multi-scale feature fusion module
[0030] The characteristic information at different objective lens magnification scales is integrated, and the biochemical recurrence probabilities obtained at different objective lens magnifications are weighted averaged to obtain a recurrence or non-recurrence prediction result, which is transmitted to the input display module.
[0031] Multi-scale fusion is crucial in pathology image analysis because features at different scales complement each other and provide more comprehensive information. For example, a low-magnification objective (e.g., 10x) helps identify the overall distribution and extent of cancer lesions, as well as their relationship to surrounding tissues, such as their size, proportion, and invasion. A high-magnification objective (e.g., 40x) reveals cellular-level morphological details, such as cancer cell morphology, nuclear atypia, and nuclear-to-cytoplasmic ratio. Through multi-scale fusion, the model can better understand hierarchical structures and improve its ability to recognize complex features. Furthermore, this approach enhances model robustness, reduces the impact of noise and information loss, and can simulate the pathologist's observation method at different magnifications, improving diagnostic accuracy. This method transcends the limitations of a single scale and integrates image information at different levels, learning richer feature representations and enhancing model robustness, thereby improving prediction performance. During the testing phase, a model ensemble was used to weighted average the outputs of the three models at different scales (before the softmax operation) to generate a predicted risk score for the patient.
[0032] (7) Storage module, which stores the information processed by each module.
[0033] (8) Control module: controls the normal operation of each module.
[0034] In a second aspect, the present invention provides a method for constructing the above-mentioned prostate cancer biochemical recurrence prediction system based on pathological sections, comprising the following steps:
[0035] (1) Dataset division: The panoramic pathology slice scan images generated by the postoperative pathology slices judged by the doctor are summarized as the total dataset, and the dataset is randomly divided into five subsets. One of them is used as the validation set in turn, and the remaining four are used for training. After repeating five times, the average result is taken to evaluate the model performance;
[0036] (2) Feature vector processing module construction: The pre-processed panoramic pathology slice scan image is cropped into non-overlapping image blocks at 10×, 20×, and 40× magnifications, and the 1024-dimensional feature vector of each image block is extracted using the self-supervised learning visual encoder UNI to generate a feature representation at the panoramic pathology slice scan image level; each panoramic pathology slice scan image will generate a set of low-dimensional feature matrices at the three magnification scales, with a size of {N s ×1024}, where N represents the number of image patches at that scale, and s∈{10×, 20×, 40×}. During feature extraction, data augmentation methods are used to increase data diversity and help the model learn image representations that are invariant to rotation and color.
[0037] (3) Construction of pseudo-package construction module: An original parent package is divided into several smaller pseudo-packages according to the corresponding instance selection strategy, with the same label as the original parent package; the top 20% features ranked from large to small according to the positive probability in the pseudo-package are extracted, and these high-confidence features are spliced together and then sent to the attention network to highlight the most discriminative area for the final classification result; then, the selected instance features are weighted and fused using the attention score to generate a compressed pseudo-package feature package representation as the representative feature of the pseudo-package. At the same time, the original parent package is also feature-weighted through an independent attention network path to generate a parent package feature package representation, which is sent together with the pseudo package feature package into the fusion module or downstream network for joint training and learning to generate the feature package representation of the pseudo package and the parent package respectively;
[0038] (4) Position and spatial information generator module construction: By reconstructing the input features into four-dimensional features X reshaped : , and processed using convolution kernels of different sizes to produce multi-scale position-aware feature representations; the feature map is processed by convolution operation X output = X reshaped + Conv a (X reshaped ) + Conv b (X reshaped) + Conv c (X reshaped ) is enhanced, and finally, through encoding and integration, convolution kernels of different scales are used to provide a more comprehensive feature expression to generate the final output feature X final = flatten(X onput );
[0039] (5) Construction of cyclic cross attention module: gradually enhance the context information through multiple iterations. In the first iteration, the input feature map X final Passed to the cross attention module, the module aggregates long-distance context information from pixels and generates a new output feature map X final ′, shape and X final Same; then, X final ′ is processed again as input through the cross-attention module. After multiple iterations, the module captures the dependencies between pixels and generates a new feature map containing richer context information;
[0040] (6) A multi-scale feature fusion module is constructed to integrate feature information at different objective lens magnification scales, enabling the model to establish a hierarchical cognitive structure from global to local.
[0041] When using the system of the present invention to predict biochemical recurrence of prostate cancer, the system is first constructed based on the above method, and then trained and validated using slice images evaluated by case doctors. The trained and validated system can be integrated into the case slice image analysis machine.
[0042] Therefore, in a third aspect, the present invention provides a non-transitory computer-readable storage medium that can be installed in a computer or an image reader, and stores a computer program that implements the steps described above when executed by a processor.
[0043] Functions and effects of the invention
[0044] The prostate cancer biochemical recurrence prediction system based on pathological sections established by the present invention can effectively interpret pathological tissue through panoramic pathological section scanning images of pathological sections of prostate cancer patients, which can help doctors more accurately predict the patient's recurrence risk and thus formulate personalized treatment plans.
[0045] In terms of accuracy, the parameter combination with the highest average AUC value on the validation set across all five-fold cross-validation combinations was selected to construct the final model. This model achieved an AUC of 0.800 and an accuracy of 0.765 on the test set. Compared to existing mainstream methods, the proposed system outperformed the widely used max-pooling and mean-pooling strategies in recent years, as well as the trainable attention-based multiple instance learning (ABMIL) model, one of the best pathology image analysis models, in both AUC and accuracy. This demonstrates the model's superior ability to identify key pathological features and effectively classify them. Panoramic pathology slide scans with delineated tumor regions from 1,200 patients at Renji Hospital Affiliated to Shanghai Jiao Tong University (Renji Hospital) were used as training data. The model was then validated using panoramic pathology slide scans from 850 patients at Renji Hospital and 336 patients from the TCGA-PRAD dataset. The three-year biochemical recurrence rates were 22.9% and 11.6%, respectively.
[0046] Furthermore, the system of the present invention can be integrated into digital pathology slide reading systems. This automated tool improves diagnostic efficiency, reduces analysis time, and enhances diagnostic consistency and accuracy, ultimately contributing to improved patient treatment outcomes and survival rates. Furthermore, the system is designed to assist physicians in interpreting patient data, integrating various aspects of a patient's medical information to enable more personalized treatment. This holds great promise for determining patient prognosis and improving treatment options, and holds great promise for future advancements.
[0047] Therefore, this system eliminates the need for long-term postoperative serum prostate-specific antigen (PSA) level monitoring and clinical follow-up, enabling recurrence risk assessment based solely on routine H&E-stained pathology slides. This addresses the previously unavailable quantitative prediction of postoperative recurrence and the timing of recurrence, a challenge for pathologists. Furthermore, the system's results can inform clinicians' individualized treatment plans, enabling the deep integration of pathology information with comprehensive patient medical data and promoting the practical application of precision medicine. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 Shows the composition framework diagram of the system of the present invention.
[0049] Figure 2 The prediction process framework diagram of the system of the present invention is shown.
[0050] Figure 3 A schematic diagram of the structure of the pseudo-packet strategy is shown. DETAILED DESCRIPTION
[0051] The present invention will be described in detail below with reference to the examples and accompanying drawings. However, the following examples should not be considered as limiting the scope of the present invention.
[0052] 1. System Model Construction and Training Technology Solution
[0053] The present invention provides a prostate cancer biochemical recurrence prediction system based on pathological sections and constructs a set of intelligent prediction methods. The system first scans the prostate cancer pathological sections into high-resolution digital sections and accurately marks the tumor area at different magnifications. Subsequently, a deep feature extraction network is used to extract high-dimensional semantic features of multi-scale images, and a weakly supervised multi-instance learning model is used to generate risk representations at the level of panoramic pathological section scan images, thereby achieving a preliminary prediction of the patient's postoperative biochemical recurrence risk. On this basis, the system further integrates the model prediction results from different objective lens magnifications to construct a multi-scale integrated framework to improve the accuracy and robustness of the prediction, and ultimately outputs an intelligent judgment result on whether the patient has experienced biochemical recurrence.
[0054] Figure 1 The framework of the pathology section-based prostate cancer biochemical recurrence prediction system 100 of the present invention is shown, organically integrating an input display module 1, a feature vector processing module 2, a pseudo packet construction module 3, a position and spatial information generator module 4, a recurrent cross-attention module 5, a multi-scale feature fusion module 6, a storage module 7, and a control module 8. The core modules of the present invention are the feature vector processing module 2, the pseudo packet construction module 3, the position and spatial information generator module 4, the recurrent cross-attention module 5, and the multi-scale feature fusion module 6. These modules efficiently extract pathological features from panoramic pathology section scan images corresponding to the pathological characteristics of prostate cancer samples, thereby accurately predicting biochemical recurrence in postoperative prostate cancer patients.
[0055] Figure 2 The prediction process framework of the system of the present invention is shown. The general prediction process is as follows: First, the panoramic pathology slide scan image of the prostate cancer pathology slide is divided into tumor areas at different objective lens magnifications. The deep features of the image are extracted using a feature extractor. Then, the feature extractor is input into a weakly supervised network to generate a feature representation at the level of the panoramic pathology slide scan image. The results of the prediction of the patient's postoperative biochemical recurrence risk at different objective lens magnifications are obtained. Then, the model results at different magnification scales are integrated to predict whether the patient will eventually experience biochemical recurrence.
[0056] The prediction process of the system is similar to its training and validation process. The following describes the model construction and prediction process by using the system training process to implement the prediction algorithm, which includes the following steps:
[0057] Step (1), data set division;
[0058] Step (2), feature extraction method based on self-supervised learning and feature vector processing of multiple slices of a case;
[0059] Step (3), using the pseudo packet strategy;
[0060] Step (4), the feature vector passes through the position and spatial information generator module;
[0061] Step (5), cross attention mode iteration;
[0062] Step (6) fuses image features of different scales.
[0063] The specific steps are as follows:
[0064] (1) Dataset division
[0065] The panoramic pathology slice scan images judged by the doctor were summarized as the total data set, and the total data set was divided into a test set and a validation set. A five-fold cross-validation method was adopted to randomly divide the data set into five subsets, four subsets were used for training each time, and the remaining subset was used for validation. This process was repeated five times, and a different subset was selected as the validation set each time. Finally, the average of the five validation results was used as the performance evaluation indicator of the model. This method effectively reduced the deviation caused by data division and provided more stable and reliable model evaluation results. In addition, slices with too small tumor areas or sections where the characteristics of the cancer area could not be effectively extracted under multiple objectives, as well as samples lacking information on recurrence status or recurrence time were excluded. At the same time, patients who had received neoadjuvant therapy were also excluded to avoid the impact of therapeutic intervention on the analysis of pathological characteristics.
[0066] (2) Feature extraction method based on self-supervised learning and feature vector processing of multiple slices from one case
[0067] In the feature extraction stage, the pathologist first accurately outlines the tumor area in the prostate cancer pathology slide, and crops the outlined area into non-overlapping image blocks (patches) of size 256×256 pixels at three objective lens magnifications of 10×, 20×, and 40× as the input of the feature extraction network. In order to achieve efficient and unsupervised image representation learning, the system introduces a universal self-supervised visual encoder designed specifically for pathology images - UNI. The encoder is pre-trained based on the DINOv2 self-supervised framework and can automatically learn discriminative image features from tissue images without manual annotation. In the output of the last convolutional network layer of the UNI encoder, the system extracts a 1024-dimensional feature vector for each patch. Ultimately, each panoramic pathology slide scan image will generate a set of low-dimensional feature matrices of size {N S ×1024}, where N S Indicates the number of patches at this scale, s∈{10×, 20×, 40×}.
[0068] To enhance the model's generalization and effectively reduce the risk of overfitting, multiple data augmentation strategies were introduced during the feature extraction phase, including horizontal and vertical flips, 45-degree random rotations, and perturbations of brightness, contrast, saturation, and hue. These augmentations make the model more robust, enabling it to learn structural features that are insensitive to changes in orientation and color, thereby improving its ability to adapt to complex pathological images. It is important to note that for cases with multiple slides, the system concatenates and fuses the feature vectors extracted from all slides at each objective magnification to construct a unified scale feature package. Ultimately, each patient receives feature packages at three scales (10×, 20×, and 40×), which serve as multi-scale input to the subsequent classification network for accurate modeling of biochemical recurrence risk.
[0069] (3) Pseudo-packet strategy
[0070] Forged Packet Policy Process Reference Figure 3 First, an original parent package is divided into several smaller pseudo packages according to the corresponding instance selection strategy. Each pseudo package inherits the label information of the parent package. Then, the instance features ranked in the top 20% according to the positive class probability are selected from each pseudo package, and these high-confidence features are spliced and used as input for subsequent modeling. These filtered features are first sent to the attention network to calculate the instance-level attention score. The attention mechanism can dynamically adjust the weights of different instance features according to the attention learned by the model, thereby highlighting the areas with the most discriminative power for the final classification result ( Figure 3 The selected instance features are then weighted and fused using the attention scores to generate a compressed pseudo-packet-level representation as the representative feature of the pseudo-packet, which is then fed into the classifier for packet-level prediction. Simultaneously, the original parent package is also weighted through an independent attention network path to generate a global package representation ( Figure 3 Finally, the parent packet features and the representative features of each pseudo packet are fed into a fusion module or downstream network for joint training and learning. This allows the model to capture information about local key regions while retaining overall structural features, improving its ability to identify and discriminate heterogeneous lesions in images.
[0071] (4) The position and spatial information generator module fully integrates spatial position features and semantic representation
[0072] When the feature package is input into the spatial information generator module, it first needs to be re-dimensionalized to adapt to the subsequent spatial modeling operation. input The dimension is [R, D], which is reconstructed into a four-dimensional tensor X reshaped, with dimensions of [R, C, H, W], where C represents the number of channels, H and W are the height and width of the image, respectively. To introduce multi-scale spatial perception capabilities, the module uses three different sizes of convolution kernels (7×7, 5×5, and 3×3) to perform parallel convolution operations on the reconstructed feature maps. These convolution operations are denoted as Conv a 、Conv b and Conv c , the corresponding spatial perception features are fused in the following way: X output = X reshaped + Conv a (X reshaped ) + Conv b (X reshaped )+ Conv c (X reshaped This process can effectively combine the spatial structure information under different receptive fields, enhance the model's ability to capture local and global context, and improve the multi-scale robustness of feature expression. output Flattened into a two-dimensional form for input to subsequent modules: X final = flatten(X output ). Through this spatial information generator module, the model can more fully integrate spatial location features and semantic representations, providing more discriminative input features for downstream classification or attention mechanisms.
[0073] (5) Analysis of loop iteration of cyclic cross attention module
[0074] By iteratively taking the output of the cross-attention module as new input data, we can gradually obtain richer context information. In the first iteration, the input feature map X final is passed to the crisscross attention module, which aggregates long-range contextual information from each pixel in a crisscross manner to produce a new output feature map X final ′, where X final and X final ’ has the same shape. Then the generated output feature map X final ′ is input to the cross attention module again. After multiple iterative processes, the module captures the dependency relationship at the pixel level and gradually generates a new feature map containing dense and rich context information. x ,u y ) and (θ x ,θ y ), A is the attention map, and the information transfer of the two cycles is as follows:
[0075] .
[0076] This cross-attention-based recursive modeling approach not only improves the expressiveness of features, but also enhances the model's discriminative ability when processing pathological images with complex structures and extensive contextual dependencies.
[0077] By establishing a dual association between local fine features and global context through the position and spatial information generator module and the recurrent cross-attention module, the model can simultaneously identify microscopic cellular abnormalities and macroscopic tissue structural changes. The fusion architecture simulates the thinking mode of pathologists in the diagnosis process of macroscopic scanning-suspicious area identification-microscopic confirmation, enhancing the ability to understand complex cancer tissue morphology and capturing the high heterogeneity of prostate cancer.
[0078] (6) Multi-scale feature fusion
[0079] Multi-scale feature fusion is crucial in pathology image analysis. Image information at different magnifications is complementary and can collectively enhance the model's understanding and ability to discern tissue structure. Low-magnification objectives (e.g., 10x) help identify the overall distribution and extent of cancer lesions and their relationship to surrounding tissue, while high-magnification objectives (e.g., 40x) reveal cellular morphological details such as nuclear atypia and nuclear-cytoplasmic ratio. Fusion of feature information at different scales helps the model establish a hierarchical cognitive structure from the global to the local level, improving not only the ability to identify tumor heterogeneity but also its robustness to image noise and information loss.
[0080] This invention overcomes the limitations of traditional methods for modeling images at a single scale. By integrating image features at different magnifications, it constructs a more expressive multi-scale representation space, improving the model's ability to discriminate prostate cancer recurrence risk. During the model inference phase, an ensemble strategy is employed to perform a weighted average of the model outputs trained at three scales (the weighting operation is performed before the softmax operation), ultimately outputting an intelligent judgment of whether a patient has experienced biochemical recurrence.
[0081] 2. System Prediction Performance Evaluation and Comparative Analysis
[0082] 1. Multicenter Dataset and Model Parameter Configuration
[0083] This study used panoramic pathology slide scans with delineated tumor regions from 1,200 patients at Renji Hospital as training data. The model was validated using panoramic pathology slide scans from 850 patients at Renji Hospital and 336 patients from the TCGA-PRAD dataset. The three-year biochemical recurrence rates were 22.9% and 11.6%, respectively. The optimal model parameter configuration was determined using a five-fold cross-validation method. The training data set was randomly divided into five equal-sized subsets, one of which was used as the validation set, and the remaining four subsets were used for model training. For each training fold, a grid search method was used to systematically iterate over key hyperparameters, including the learning rate (ranging from 1e-5 to 1e-3), batch size (16 to 64), weight decay coefficient (1e-5 to 1e-3), and feature dimensionality reduction parameters. Each hyperparameter combination was trained to convergence, and the AUC was calculated on the validation set. The model was trained using the Adam optimizer with an early stopping strategy to avoid overfitting. Among all five-fold cross-validation combinations, the parameter combination with the highest average AUC value on the validation set was selected to construct the final model. The results showed that the model achieved an AUC of 0.800 and an accuracy of 0.765 on the test set, which was superior to the existing comparison methods.
[0084] 2. Performance Evaluation
[0085] In a comparative experiment for predicting biochemical recurrence in patients based on pathology images, the proposed model was trained and validated using the same samples as existing cutting-edge methods: max-pooling and mean-pooling, as well as the ABMIL model, one of the best pathology image analysis algorithms currently available. The results demonstrate significantly superior performance. Table 1 shows the results: Compared to existing mainstream methods, the proposed system outperformed the widely used max-pooling and mean-pooling strategies, as well as the ABMIL model based on a trainable attention mechanism, in terms of both AUC (0.800) and accuracy (0.765). This demonstrates the proposed model's superior ability to identify key pathological features and effectively classify them.
[0086] Furthermore, to verify the effectiveness of multi-scale image information fusion in this task, we further designed ablation experiments. The results, as shown in Table 2, show that after fusing image information at different magnifications, such as 10×, 20×, and 40×, the model performance is significantly better than using only a single scale. This demonstrates that multi-scale features play an important complementary role in modeling tumor recurrence risk, thereby improving the model's ability to represent complex tissue structures.
[0087] Table 1 Comparative results of different existing mainstream weak supervision models
[0088]
[0089] Table 2 Ablation effects at different scales
[0090]
[0091] The above results indicate that the system of the present invention has higher stability and reliability in accurately identifying the risk of biochemical recurrence, and can provide stronger support for clinical diagnosis.
[0092] Any undescribed parts of the present invention are the same as or implemented using existing technologies. The applicant declares that the present invention uses the above-mentioned embodiments to illustrate the detailed methods of the present invention, but the present invention is not limited to the above-mentioned detailed methods, that is, it does not mean that the present invention must rely on the above-mentioned detailed methods to be implemented. Those skilled in the art should understand that any improvements to the present invention, equivalent replacement of various raw materials of the product of the present invention, addition of auxiliary ingredients, selection of specific methods, etc., all fall within the scope of protection and disclosure of the present invention.
Claims
1. A prostate cancer biochemical recurrence prediction system based on pathological sections, characterized by: include: The input display module is used to input the patient's basic information and the panoramic pathology slice scan image of the pathology slice, and display the analysis results after the analysis is completed; The feature vector processing module crops the tumor region of the pre-processed panoramic pathology slide scan image at different objective lens magnifications into non-overlapping image blocks of a certain number of pixels based on the features at each magnification. This blocks serves as the input to the feature extraction network. A self-supervised visual encoder (UNI) is then used to automatically extract features from these blocks, generating a feature representation at the panoramic pathology slide scan image level containing multiple low-dimensional feature sets at different magnifications. The pseudo-package construction module refines the feature representation at the panoramic pathology slice scan image level based on the pseudo-package strategy, and generates feature package representations of pseudo packages and parent packages respectively. The pseudo-package strategy is as follows: an original parent package is divided into several smaller pseudo packages with the same label as the parent package; the features of the top 20% of the positive probability in the pseudo package are extracted, and these high-confidence features are spliced together and sent to the attention network to highlight the areas with the most discriminative power for the final classification result; the selected instance features are weightedly fused using the attention score to generate a compressed pseudo-package feature package representation as the representative feature of the pseudo package. At the same time, the original parent package generates a parent package feature package representation, which is sent together with the pseudo package feature package into the fusion module or downstream network for joint training and learning. While capturing the information of local key areas, the overall structural features are retained, thereby improving the model's ability to identify and discriminate heterogeneous lesions in the image. The position and spatial information generator module, based on the spatial distribution pattern of prostate cancer invasiveness, first reconstructs the input panoramic pathology slice scan image level feature representation or pseudo package and parent package feature package representation into a four-dimensional tensor X reshaped , and then use convolution kernels of different sizes to process feature data to further extract features with spatial perception and context information, and fuse them in the following way: output =X reshaped +Conv a (X reshaped )+Conv b (X reshaped )+Conv c (X reshaped ); then the fused four-dimensional feature X output Flattened into a two-dimensional form for input to subsequent modules: X final =flatten(X output ); The recurrent cross-attention module complements convolution kernels of different scales and iteratively generates new feature maps with dense and rich contextual information based on cross-attention, enabling the model to simultaneously identify microscopic cell abnormalities and macroscopic tissue changes. It then calculates the correlation between features based on these feature representations, and generates attention weights after normalizing these correlations. In the first iteration, the feature map X final As input, the module aggregates long-range contextual information from each pixel in a cross-wise manner to produce a new output feature map X final ', the two have the same shape; then X final ′ is input into the cross attention module again. After multiple iterative processes, the module captures the dependency at the pixel level and gradually generates a new feature map containing dense and rich context information. For any two positions (u x ,u y ) and (θ x ,θ y ), A is the attention map, and the information transfer of the two cycles is as follows: Hu″←[f(A,u x ,i y ,i x ,θy)·f(A',u x ,u y ,u x ,i y )+f(A,θ x ,u y ,i x ,i y )·f(A',u x ,u y ,i x ,u y )]; The multi-scale feature fusion module fuses feature information at different objective lens magnification scales, performs weighted averaging on the biochemical recurrence probabilities obtained at different objective lens magnifications, obtains a recurrence or non-recurrence prediction result, and transmits it to the input display module.
2. The prediction system according to claim 1, characterized in that It also includes a storage module and a control module. The storage module is used to store information processed by each module, and the control module controls the normal operation of each module.
3. The prediction system according to claim 1, wherein: The preprocessing method of panoramic pathology slide scan images is as follows: by adjusting the mean and standard deviation of the image block data, the data is converted into a distribution with zero mean and unit variance to achieve data standardization and image intensity normalization. The tumor region was cropped into non-overlapping image blocks of 256×256 pixels at 10×, 20×, and 40× objective lens magnifications and used as the input of the feature extraction network. After the last convolutional layer of UNI, the 1024-dimensional feature vector of each image block is extracted to generate the feature representation of the panoramic pathology slice scan image level. Each panoramic pathology slice scan image will generate a set of low-dimensional feature matrices at three multiple scales, with a size of {N s ×1024}, where N represents the number of image blocks at this scale, s∈{10×,20×,40×}, The feature vector processing module introduces a variety of data enhancement strategies when performing feature extraction, including horizontal flipping, vertical flipping, 45-degree random rotation, and perturbations of brightness, contrast, saturation, and hue. When a case has multiple pathological sections, the feature vectors extracted from all sections of the case at each objective lens magnification are spliced and fused to construct a unified scale feature package.
4. The prediction system according to claim 3, characterized in that The characteristic information at the 10× objective lens magnification scale includes macroscopic tissue change characteristics, and the characteristic information at the 40× objective lens magnification scale includes morphological details at the cellular level.
5. The prediction system according to claim 1, wherein: The position and spatial information generator module first reconstructs the input features into a four-dimensional tensor X reshaped :X reshaped =reshape(X input )→R 1×C×H×W , where X input is the original input feature, dimension is [R, D], four-dimensional tensor X reshaped The dimensions are [R, C, H, W], where C is the number of channels, H and W are the width and height of the image; convolution kernels with different sizes are then used to perform convolution operations on the reconstructed features. The size of each convolution kernel corresponds to a different spatial range. These convolution operations generate multi-scale spatial perception feature representations in the features and perform fusion.
6. The prediction system according to claim 1, wherein: Microscopic cellular abnormalities include variability in cell size, nuclear atypia, increased chromatin, frequent mitotic figures, and abnormal cell metabolism; macroscopic tissue changes include irregular gland size and disordered arrangement, basement membrane destruction, perineural infiltration, vascular invasion, and interstitial fibrosis; After calculating the correlation between features, the softmax operation is used to normalize these correlations and generate attention weights.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the functions of the system according to any one of claims 1 to 6 are realized.
8. The method for constructing a prostate cancer biochemical recurrence prediction system based on pathological sections according to any one of claims 1 to 6, characterized in that: The steps include: (1) Dataset division: The panoramic pathology slide scan images generated by the postoperative pathology slides judged by the doctors are summarized as the total dataset and divided into a training set and a validation set; (2) Feature vector processing module construction: The pre-processed panoramic pathology slide scan image is cropped into non-overlapping image blocks at different objective lens magnifications, and the 1024-dimensional feature vector of each image block is extracted using the self-supervised learning visual encoder UNI to generate a feature representation at the panoramic pathology slide scan image level; (3) Construction of pseudo-package construction module: An original parent package is divided into several smaller pseudo-packages according to the corresponding instance selection strategy, with the same label as the original parent package; the top 20% features ranked from large to small according to the positive probability in the pseudo-package are extracted, and these high-confidence features are spliced together and then sent to the attention network to highlight the most discriminative areas for the final classification result; the selected instance features are then weightedly fused using the attention score to generate pseudo-package representative features and parent package representative features respectively; (4) Construction of the position and spatial information generator module: The input features are reconstructed into four-dimensional features and processed using convolution kernels of different sizes to generate multi-scale position-aware feature representations; finally, through encoding and integration, convolution kernels of different scales are used to provide a more comprehensive feature expression to generate the final output feature X final =flatten(X onput ); (5) Construction of cyclic cross attention module: The context information is gradually enhanced through multiple iterations. After multiple iterations, the module captures the dependencies between pixels and generates a new feature map containing richer context information; (6) A multi-scale feature fusion module is constructed to integrate feature information at different objective lens magnification scales, enabling the model to establish a hierarchical cognitive structure from global to local.
9. The method for constructing a prostate cancer biochemical recurrence prediction system based on pathological sections according to claim 8, characterized in that: In step (1), the dataset is randomly divided into five subsets, one of which is used as the validation set in turn, and the remaining four are used for training. After repeating five times, the average result is taken to evaluate the model performance; In step (2), each panoramic pathology slide scan image generates a set of low-dimensional feature matrices at three multiple scales, with a size of {Ns×1024}, where N represents the number of image blocks at that scale, and s∈{10×,20×,40×}; During feature extraction, data augmentation methods are used to increase data diversity and help the model learn image representations that are invariant to rotation and color. In step (3), the pseudo-packet representative features and the parent-packet representative features are fed into the fusion module or downstream network for joint training and learning; In step (4), reconstruct the four-dimensional feature X reshaped The method is as follows: X reshaped =reshape(X input )→R 1×C×H×W ; The convolution kernel processing methods of different sizes are as follows: X output =X reshaped +Conv a (X reshaped )+Conv b (X reshaped )+Conv c (X reshaped ); In step (5), in the first iteration, the input feature map X final Passed to the cross attention module, the module aggregates long-distance context information from pixels and generates a new output feature map X final ′, shape and X final Same; then, X final ’ is processed again as input through the crisscross attention module.
Citation Information
Patent Citations
Primary gastric cancer digital pathological section lymph node metastasis risk prediction method
CN116705311A
Method for predicting whether breast cancer patient subjected to NAC reaches pCR or not through comparative learning and double-layer attention
CN118658005A
Pyramid vision Transform-based polyp image segmentation method
CN119399229A