Prostate cancer biochemical recurrence prediction system based on pathological section and construction method

By combining the cyclic cross attention module and a multi-instance algorithm of the pseudopacket strategy, the accuracy and efficiency of the prediction of biochemical recurrence of prostate cancer in the prior art are solved, and efficient pathological section analysis and the formulation of personalized treatment plans are achieved.

CN120260974AActive Publication Date: 2025-07-04RENJI HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE

Patent Information

Application Number
CN202510740922.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the risk of biochemical recurrence after prostate cancer, which mainly relies on PSA level and Gleason scores that are greatly affected by subjective factors, and pathological analysis is time-consuming and insufficient resources. The existing AI methods ignore spatial positional relationships and key regional characteristics when processing prostate cancer pathological sections, resulting in limited prediction accuracy.

Method used

A multi-instance algorithm based on cyclic cross attention module and pseudo-packet strategy is adopted, combined with self-supervised learning and multi-scale feature fusion, pathological features are extracted through high-definition panoramic pathological section scanning images, and a prostate cancer biochemical recurrence prediction system is constructed, and image information at different objectives is fused to enhance attention and spatial correlation of key areas.

Benefits of technology

It improves the accuracy and efficiency of biochemical recurrence prediction of prostate cancer, reduces human resources needs, achieves more accurate risk assessment and personalized treatment plans, and improves diagnostic consistency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260974A_ABST
    Figure CN120260974A_ABST
Patent Text Reader

Abstract

The invention provides a prostate cancer biochemical recurrence prediction system based on pathological sections and a construction method. Based on a high-definition panoramic pathological section scanning image of a pathological section after a radical operation, a multi-instance algorithm based on a cyclic cross attention module and a pseudo-packet strategy is adopted, deep features of tumor images with different objective lens multiples are extracted firstly, and then feature representation of a panoramic pathological section scanning image level is generated in a weak supervision network; postoperative biochemical recurrence risk prediction results under different multiples are obtained; and then integrating model results under different scales, and predicting whether the patient finally has biochemical recurrence or not through multi-center verification. According to the system, multi-center, multi-slice and multi-scale panoramic pathological section scanning images are creatively used for training and verification, the risk of prostate cancer recurrence of a patient is predicted through pathological sections after radical treatment, the risk of recurrence of the patient within 3 years and longer time after the radical treatment is accurately predicted, and more personalized treatment is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of prostate cancer biochemical recurrence risk prediction and analysis, and relates to a prostate cancer biochemical recurrence prediction system based on pathological sections and a construction method thereof. Background Art

[0002] After prostate cancer patients undergo radical resection, about 30% will experience biochemical recurrence (BCR) with elevated PSA, which is an important prognostic indicator for cancer recurrence and metastasis. Accurately predicting the BCR risk helps clinicians stratify and closely follow up high-risk patients. However, currently, it mainly relies on clinical information and pathological reports for subjective judgment. The most commonly used method is to judge the prognosis based on the level of prostate-specific antigen (PSA) at the time of diagnosis, but its accuracy and effectiveness are not ideal. Secondly, the Gleason score is used for judgment. Different pathologists have significant heterogeneity in the judgment of the same section, and it is greatly affected by subjective factors, resulting in poor results. In addition, currently, internationally, a prediction scale is also used to comprehensively consider information from multiple aspects for prognosis judgment, but the information collected is not comprehensive, there may be data missing, and the ideal prediction result cannot be achieved. Moreover, pathological analysis and diagnosis is a very labor-intensive clinical task. Currently, millions of pathological sections are generated every day worldwide and in China. Problems such as long pathological diagnosis time and insufficient number of professional pathologists pose huge pressures and challenges to the medical quality and medical system. In summary, the existing methods are difficult to meet the clinical requirements for the risk assessment and management of BCR in postoperative prostate cancer patients, and thus early intervention cannot be made.

[0003] With the development of technology, the pathological sections of specimens after radical prostatectomy can be digitized, meaning that information will no longer be limited to being processed by the naked eye. With the powerful processing and analysis capabilities of computers, more meaningful information that could not be noticed before can be obtained from pathological sections. However, most current AI pathological analysis methods have multiple key defects when dealing with highly complex tumors such as prostate cancer, severely restricting their clinical application value: First, when analyzing pathological section images, mainstream methods usually extract the panoramic pathological section scan images of prostate cancer as two-dimensional features, simply storing the features of image patches distributed in space in a serialized form, completely ignoring the key spatial position relationships in histomorphology. The invasive spatial distribution of prostate cancer is jointly regulated by anatomical partitioning, clonal evolution, and microenvironment interaction, showing multifocal heterogeneity, a tendency for frontal invasion, and specific metastasis paths, with a unique spatial distribution pattern; these algorithms are insufficient in multi-scale integration and cannot effectively capture and fuse prostate cancer features at different scales, such as nuclear morphological changes at the microscopic level and tissue structure damage at the macroscopic level at the same time; Second, in the face of the high heterogeneity of prostate cancer, such as multiple Gleason scores can exist simultaneously within the same lesion. Existing methods often use averaging to eliminate the so-called "noise", but inadvertently erase the features of a few high-grade regions that are crucial for prognosis judgment; Third, features closely related to biochemical recurrence, such as nerve invasion and vascular invasion, usually only exist in small regions of the entire section. Existing algorithms often have difficulty accurately locating these sparsely distributed but crucial regions for prognosis assessment in a large amount of background tissue, resulting in limited prediction accuracy.

[0004] Currently, pooling operations of cutting-edge new methods such as Max-pooling and Mean-pooling, although simple, efficient, and widely used in pathological image analysis, only focus on the maximum or average value of features, ignoring the heterogeneity in prostate cancer tissue and the important information in key minority regions. The attention-based multi-instance learning model (ABMIL), which is widely used in pathological image analysis, although it introduces an attention mechanism to adaptively weight the importance of different instances and is an improvement compared to traditional pooling methods, still easily disperses attention to a large number of background regions when faced with prostate cancer pathological images with high heterogeneity, multi-focality, and sparse distribution of key features, resulting in insufficient capture of features in key regions for prognosis. These technical limitations and clinical challenges urgently require a new method that can handle the above pathological characteristics of prostate cancer, more efficiently extract the pathological features of the panoramic pathological section scan images of patients' prostate cancer, and thus accurately predict the biochemical recurrence of prostate cancer patients after surgery. Summary of the Invention

[0005] The present invention aims to develop a prostate cancer biochemical recurrence prediction system based on pathological sections to assist doctors in interpreting the recurrence situation of postoperative patients. The prostate pathological section is the radical prostatectomy pathological section obtained immediately after the radical prostatectomy for prostate cancer patients. At the same time, a high-definition panoramic pathological section scanning image based on the postoperative pathological section is proposed, and a prostate cancer biochemical recurrence prediction algorithm using a multi-instance algorithm based on a cyclic cross-attention module and a pseudo-packet strategy. First, a position and spatial information generator module is creatively built to comprehensively process image information, and then a pseudo-packet strategy is proposed to solve the overfitting problem; secondly, a strategy of capturing and integrating image features at multiple scales is proposed for the first time in the analysis of prostate cancer pathological images, and finally, intelligent analysis of prostate cancer pathological images is realized to predict the risk of biochemical recurrence of prostate cancer. This is of great significance for judging the prognosis of patients and improving treatment methods and has broad application prospects.

[0006] Specifically, the system of the present invention has the following technical features: It organically integrates core modules such as feature vector processing, position and spatial information generator, cyclic cross-attention, pseudo-packet construction, and multi-scale feature fusion, and efficiently extracts the corresponding pathological features of the panoramic pathological section scanning image according to the pathological characteristics of prostate cancer samples, so as to predict the biochemical recurrence of postoperative prostate cancer patients with high accuracy, as follows:

[0007] (1) Input display module

[0008] Input the basic information of the patient and the panoramic pathological section scanning image of the pathological section, and display the analysis result after the analysis is completed;

[0009] (2) Feature vector processing module

[0010] Process based on a self-supervised learning-based feature extraction method: Cut the tumor regions of the preprocessed panoramic pathological section scanning images at different objective lens magnifications (preferably 10×, 20×, 40×) into non-overlapping image blocks (patches) of a certain pixel size (such as 256×256) as the input of the feature extraction network; then use the self-supervised visual encoder UNI to automatically extract features from it to generate a panoramic pathological section scanning image-level feature representation containing multiple low-dimensional feature sets at different magnifications.

[0011] Preferably, the preprocessing method of the panoramic pathological section scanning image is as follows: By adjusting the mean and standard deviation of the image block data, the data is converted into a distribution with zero mean and unit variance to achieve data standardization and image intensity normalization.

[0012] 10×, 20×, and 40× are low, medium, and high magnifications of the objective lens, respectively. The 10× objective lens magnification is used to preliminarily observe the overall structure of the sample, which helps to identify the overall distribution and range of the cancer lesions and their relationship with the surrounding tissues; the 40× objective lens magnification reveals morphological details at the cellular level, such as nuclear atypia and nuclear-cytoplasmic ratio, tumor-nerve interface, contact between tumor cells and blood vessels / lymphatic vessels, intravascular micrometastasis, etc.; the 20× objective lens magnification is a medium magnification, which is clearer than the low-power microscope (10×), but has a wider field of view than the high-power microscope (40× and above), and is used for medium-resolution observation of histology and cytology.

[0013] The self-supervised visual encoder UNI is a universal visual encoder for pathology. Its generation process first uses a self-supervised learning method called DINOv2. UNI can extract meaningful features from these histological images (such as the overall distribution and range of cancer lesions and their relationship with surrounding tissues, morphological details at the cellular level, etc.) without manual labeling. After the last convolutional layer of UNI, the 1024-dimensional feature vector of each image block is extracted to generate a feature representation at the level of the panoramic pathology slice scan image. Each panoramic pathology slice scan image will generate a set of low-dimensional feature matrices at three objective lens scales, with a size of {N s ×1024}, where N represents the number of image blocks at this scale, s∈{10×, 20×, 40×}.

[0014] In order to avoid overfitting of the model as much as possible and improve the generalization ability of the model, a variety of data enhancement strategies were introduced in the feature extraction process, including horizontal flipping, vertical flipping, 45-degree rotation, and adjustment of brightness, contrast, saturation and hue in color space transformation, so that the model can learn rotation-invariant and color-invariant image representations.

[0015] In the case of a case with multiple pathological sections, the feature vectors generated by all images of a patient at each objective lens magnification scale are spliced ​​and fused to generate a representative unified scale feature package. Each case has three feature packages representing 10×, 20× and 40× objective lens magnifications as the input of the subsequent classification network.

[0016] (3) Pseudo-package building module

[0017] Developing a model using the multi-instance method is somewhat challenging, and overfitting problems often occur. In addition, there are factors in prostate cancer tissues that may interfere with the judgment, such as benign prostate tissues and their hyperplasia, inflammation, and technical artifacts. The features used to predict BCR may be sparse but crucial, existing only in a small number of tissues, such as cell populations with high metastatic potential. Therefore, the present invention also adds a pseudo-package construction module, which divides an original parent package into several smaller pseudo-packages according to the corresponding instance selection strategy using the pseudo-package strategy, with the same label as the original parent package; extracts the top 20% of the features ranked from largest to smallest according to the positive probability in the pseudo-package, greatly improving the model's attention to key regions, thereby reducing the influence of noise.

[0018] The pseudo-package strategy increases the number of packages, and the number of instances in each package also decreases accordingly. Extract the top 20% of the features ranked from largest to smallest according to the positive probability in the pseudo-package and splice these high-confidence features together and refine them into the model, then send them into the attention network and obtain the attention scores. The instance features are weighted by the attention scores to highlight the regions that are most discriminative for the final classification result. Subsequently, use the attention scores to perform weighted fusion on the selected instance features to generate a compressed pseudo-package feature packet representation as the representative feature of the pseudo-package. At the same time, the original parent package also generates a parent-package feature packet representation through an independent attention network path for feature weighting, and is jointly sent into the fusion module or downstream network for joint training and learning, capturing local key region information while retaining the overall structural features, and improving the model's recognition and discrimination ability for heterogeneous lesions in the image.

[0019] (4) Location and Spatial Information Generator Module

[0020] Based on the spatial distribution pattern of prostate cancer invasiveness, first reconstruct the feature representation at the panoramic pathological slice scan image level of the input or the feature packet representations of the pseudo-package and the parent-package into a four-dimensional tensor X reshaped , and then use convolutional kernels of different sizes to process the feature data to further extract features with spatial perception and context information.

[0021] Preferably, the location and spatial information generator module first reconstructs the input features into a four-dimensional tensor X reshaped : , where X input is the original input feature, with a dimension of [R, D], and the four-dimensional tensor X reshaped has a dimension of [R, C, H, W], C is the number of channels, and H and W are the width and height of the image; then use convolutional kernels of different sizes to process the feature data, and each convolutional kernel size corresponds to a different spatial range (7×7, 5×5, 3×3), and these convolutional operations generate multi-scale location-aware feature representations in the features: X output = Xreshaped + Conv7(X reshaped ) + Conv5(X reshaped ) + Conv3(X reshaped ). Subsequently, the features are further encoded and integrated. The convolution kernels of different scales complement each other, providing a more comprehensive and richer feature representation: X final = flatten(X onput ).

[0022] By retaining the spatial correlation in the pathological images, the modeling of tissue morphological features and the extraction of spatial features are achieved.

[0023] (5) Recurrent Cross-Attention Module

[0024] This module captures dependencies at the pixel level and aggregates long-range context information from each pixel in a recurrent cross form. It not only retains the detailed features of the local region but also establishes global semantic associations, providing a more discriminative feature input for subsequent prediction tasks. The convolution kernels of different scales complement each other, and new feature maps with dense and rich context information are gradually generated based on cross-attention, enabling the model to simultaneously identify microscopic cell abnormalities and macroscopic tissue changes; then, the correlations between features are calculated based on these feature representations, and attention weights are generated after normalizing these correlations.

[0025] Preferably, a new feature map is formed in an iterative loop manner: in the first iteration, using the feature map X final as the input, the module aggregates long-range context information from each pixel in a cross form, thereby generating a new output feature map X final ′, where X final and X final ′ have the same shape; then the generated output feature map X final ′ is input into the cross-attention module again. After multiple iterative processes, the module captures dependencies at the pixel level and gradually generates a new feature map containing dense and rich context information. For any two positions (u x , u y ) and (θ x , θ y ) on the feature map, A is the attention map, and the information transfer of the two loops is as follows: .

[0026] Furthermore, microscopic cell abnormalities include uneven cell size, nuclear atypia, increased chromatin, numerous mitotic figures, and abnormal cell metabolism; the macroscopic tissue changes include irregular gland size, disordered arrangement, basement membrane disruption, perineural invasion, accompanied by necrosis and hemorrhage, and interstitial fibrosis. Using these features to represent the correlation between features calculated by existing methods (such as Pearson correlation coefficient, Spearman rank correlation coefficient), that is, the relationship between pixels in different directions. The softmax operation is used to normalize these correlations, generating attention weights. These attention weights are applied to calculate the bag-level feature representation, so that the important pixel-pixel correlations receive more attention, strengthening the context information between features.

[0027] The dual correlation between local fine features and global context is established through the position and spatial information generator module and the cyclic cross-attention module, enabling the model to simultaneously identify microscopic cell abnormalities and macroscopic tissue structure changes. The fusion architecture simulates the thinking mode of pathologists during the diagnosis process: macroscopic scanning - suspicious area identification - microscopic confirmation, enhancing the ability to understand the complex morphology of cancer tissues and capturing the high heterogeneity of prostate cancer.

[0028] (6) Multi-scale feature fusion module

[0029] Fuse the feature information at different objective lens magnification scales, perform weighted averaging on the biochemical recurrence probabilities obtained at different objective lens magnifications, obtain the recurrence or non-recurrence prediction results, and transmit them to the input display module.

[0030] Multi-scale fusion is crucial in pathological image analysis because features at different scales can complement each other, providing more comprehensive information. For example, a low-power objective lens (such as 10×) helps to identify the overall distribution, scope of cancerous lesions, and their relationship with surrounding tissues, such as the size, proportion, and invasion status of cancerous lesions. A high-power objective lens (such as 40×) reveals morphological details at the cellular level, such as the morphology of cancer cells, nuclear atypia, and nuclear-cytoplasmic ratio. Through multi-scale fusion, the model can better understand the hierarchical structure, improve the ability to identify complex features. In addition, this method also enhances the robustness of the model, reduces the impact of noise and information loss, and can simulate the observation methods of pathologists at different magnifications, improving the diagnostic accuracy. The present invention breaks through the limitation of a single scale, fuses image information at different levels, learns richer feature expressions, improves the robustness of the model, and thus improves the performance of the prediction task. During the test phase, the output results of models at three different scales were weighted averaged (before softmax) using model ensemble, and finally the predicted risk score of the patient was generated.

[0031] (7) Storage module, which stores the information processed by each module.

[0032] (8) Control module, which controls the normal operation of each module.

[0033] In the second aspect of the present invention, there is provided a method for constructing the above-mentioned prostate cancer biochemical recurrence prediction system based on pathological sections, including the following steps:

[0034] (1) Dataset division: Summarize the panoramic pathological section scan images generated from the postoperative pathological sections determined by doctors into a total dataset, randomly divide the dataset into five subsets, and use one of them as the validation set in turn, and the remaining four are used for training. After repeating five times, take the average result to evaluate the model performance;

[0035] (2) Construction of the feature vector processing module: Crop the preprocessed panoramic pathological section scan images into non-overlapping image patches at magnifications of 10×, 20×, and 40×, and use the self-supervised learning visual encoder UNI to extract 1024-dimensional feature vectors for each image patch to generate a feature representation at the panoramic pathological section scan image level; Each panoramic pathological section scan image will generate a set of low-dimensional feature matrices at three magnification scales, with a size of {N s × 1024}, where N represents the number of image patches at this scale, and s ∈ {10×, 20×, 40×}; During the feature extraction process, data augmentation methods are used to increase data diversity to help the model learn rotation- and color-invariant image representations;

[0036] (3) Construction of the pseudo-packet construction module: Divide an original parent packet into several smaller pseudo-packets according to the corresponding instance selection strategy, with the same label as the original parent packet; Extract the top 20% of the features ranked from largest to smallest in terms of positive probability in the pseudo-packet, and splice these high-confidence features together, and then send them into the attention network to highlight the regions most discriminative for the final classification result; Subsequently, use the attention scores to perform weighted fusion on the selected instance features to generate a compressed pseudo-packet feature packet representation as the representative feature of the pseudo-packet. At the same time, the original parent packet also generates a parent packet feature packet representation through an independent attention network path for feature weighting, and is jointly sent into the fusion module or downstream network for joint training and learning to generate the feature packet representations of the pseudo-packet and the parent packet respectively;

[0037] (4) Construction of the position and spatial information generator module: Reconstruct the input features into a four-dimensional feature X reshaped : and process it with convolutional kernels of different sizes to generate multi-scale position-aware feature representations; The feature map passes through the convolutional operation X output = X reshaped + Conv a (X reshaped ) + Conv b (X reshaped) + Conv c (X reshaped ) is enhanced, and finally through encoding and integration, convolutional kernels of different scales are used to provide a more comprehensive feature representation, generating the final output feature X final = flatten(X onput );

[0038] (5) Construction of the cyclic cross-attention module: The context information is gradually enhanced through multiple iterations. In the first iteration, the input feature map X final is passed to the cross-attention module, which aggregates long-range context information between pixels and generates a new output feature map X final ', with the same shape as X final ; then, X final ' is used as the input and processed again through the cross-attention module. After multiple iterations, the module captures the dependencies between pixels and generates a new feature map containing richer context information;

[0039] (6) Construction of the multi-scale feature fusion module, which fuses the feature information at different objective lens magnification scales, enabling the model to establish a hierarchical cognitive structure from global to local.

[0040] When using the system of the present invention for predicting biochemical recurrence of prostate cancer, first construct the system based on the above method, and then use the section images determined by the case doctor to train and validate the system. The trained and validated system can be integrated into the case section image analysis machine.

[0041] Therefore, in the third aspect of the present invention, a non-transitory computer-readable storage medium is provided, which can be installed in a computer or a film reader, and stores a computer program. When the computer program is executed by a processor, the above-described steps are implemented.

[0042] Functions and effects of the invention

[0043] The system for predicting biochemical recurrence of prostate cancer established by the present invention realizes effective interpretation of pathological tissues through panoramic pathological section scanning images of prostate cancer patients' pathological sections, which can help doctors more accurately predict the recurrence risk of patients, thereby formulating personalized treatment plans.

[0044] In terms of accuracy, among all five-fold cross-validation combinations, the parameter combination with the highest average AUC value in the validation set was selected to construct the final model. The AUC of this model on the test set reached 0.800, and the accuracy reached 0.765. In comparison with existing mainstream methods, the system of the present invention is superior to the widely used Max-pooling and Mean-pooling strategies in recent years, as well as the multi-instance learning model based on a trainable attention mechanism (ABMIL), which is one of the current best pathological image analysis models, in terms of both AUC and accuracy. This indicates that this model has a stronger ability to identify key pathological features and perform effective classification. Panoramic pathological section scan images with delineated tumor regions of 1,200 patients from Renji Hospital Affiliated to Shanghai Jiao Tong University (hereinafter referred to as Renji Hospital) were used as training data, and the model effect was verified on the panoramic pathological section scan images of 850 patients and 336 patients in the TCGA-PRAD dataset of Renji Hospital, where the positive rates of biochemical recurrence within three years were 22.9% and 11.6% respectively.

[0045] In addition, the system of the present invention can be integrated into a digital pathological slide reading system. Such an automated tool improves the diagnostic efficiency, reduces the analysis time, and enhances the consistency and accuracy of diagnosis, ultimately contributing to the improvement of the treatment effect and survival rate of patients. At the same time, the system aims to assist doctors in interpretation and realizes more personalized treatment by combining all aspects of patients' medical information, which is of great significance for judging the prognosis of patients and improving treatment methods and has broad prospects.

[0046] Therefore, the present invention does not need to rely on long-term postoperative monitoring of serum prostate-specific antigen (PSA) levels and clinical follow-up. It can complete the recurrence risk assessment only relying on routine H&E-stained pathological sections, achieving the difficulty of predicting whether a patient will relapse after surgery and the timing of relapse, which was impossible to quantify for past pathologists. At the same time, the system results can provide reference for clinicians to develop individualized treatment plans, realizing the deep integration of pathological information and patients' comprehensive medical data, and promoting the implementation of precision medicine. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 Shows the composition framework diagram of the system of the present invention.

[0048] Figure 2 Shows the prediction process framework diagram of the system of the present invention.

[0049] Figure 3 Shows the structural schematic diagram of the pseudo-packaging strategy. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The present invention will be described in detail below in conjunction with the embodiments and the drawings. However, the following embodiments should not be regarded as limiting the scope of the present invention.

[0051] I. System Model Construction and Training Technical Solutions

[0052] The present invention provides a prostate cancer biochemical recurrence prediction system based on pathological sections, and constructs an intelligent prediction method. First, the prostate cancer pathological sections are scanned into high-resolution digital sections, and the tumor regions are accurately marked at different magnifications. Subsequently, a deep feature extraction network is used to extract high-dimensional semantic features of multi-scale images, and a weakly supervised multi-instance learning model is used to generate a risk representation at the panoramic pathological section scan image level, so as to realize a preliminary prediction of the biochemical recurrence risk of patients after surgery. On this basis, the system further fuses the model prediction results from different objective lens magnifications, constructs a multi-scale integration framework to improve the accuracy and robustness of the prediction, and finally outputs an intelligent judgment result on whether the patient has biochemical recurrence.

[0053] Figure 1 Fig. shows the framework structure of the prostate cancer biochemical recurrence prediction system 100 based on pathological sections of the present invention, which organically integrates an input display module 1, a feature vector processing module 2, a pseudo-packet construction module 3, a position and spatial information generator module 4, a cyclic cross-attention module 5, a multi-scale feature fusion module 6, a storage module 7, and a control module 8. The core modules of the present invention are the feature vector processing module 2, the pseudo-packet construction module 3, the position and spatial information generator module 4, the cyclic cross-attention module 5, and the multi-scale feature fusion module 6, which efficiently extract the corresponding pathological features of the panoramic pathological section scan images according to the pathological characteristics of prostate cancer samples, so as to predict the biochemical recurrence of prostate cancer patients after surgery with high accuracy.

[0054] Figure 2 Fig. shows the prediction process framework diagram of the system of the present invention. The general prediction process is as follows: First, the panoramic pathological section scan images of prostate cancer pathological sections are divided into tumor regions at different objective lens magnifications, and the deep features of the images are extracted using a feature extractor. Then, the features are input into a weakly supervised network to generate a feature representation at the panoramic pathological section scan image level, and the results of predicting the biochemical recurrence risk of patients after surgery at different objective lens magnifications are obtained. Subsequently, the model results at different magnification scales are integrated using the model to predict whether the patient will ultimately have biochemical recurrence.

[0055] The prediction process of the system is similar to its training and verification processes. The following uses the system training process for implementing the prediction algorithm to illustrate the model construction and prediction processes, including the following steps:

[0056] Step (1), dataset division;

[0057] Step (2), feature extraction method based on self-supervised learning and feature vector processing of one case with multiple sections;

[0058] Step (3), use the pseudo-packaging strategy;

[0059] Step (4), the feature vector passes through the position and spatial information generator module;

[0060] Step (5), cross-attention pattern iteration;

[0061] Step (6), fuse image features of different scales.

[0062] The specific steps are as follows:

[0063] (1) Dataset division

[0064] Summarize the panoramic pathological section scan images judged by doctors into the total dataset, and divide the total dataset into a test set and a validation set. The five-fold cross-validation method is adopted, and the dataset is randomly divided into five subsets. Each time, four subsets are used for training, and the remaining one subset is used for validation. This process is repeated five times, and each time a different subset is selected as the validation set. Finally, the average value of the five validation results is used as the performance evaluation index of the model. This method effectively reduces the bias caused by data division and provides more stable and reliable model evaluation results. In addition, slices with too small tumor areas or unable to effectively extract cancer area features under multiple objective lens magnifications, as well as samples lacking recurrence status or recurrence time information, are excluded. At the same time, patients who have received neoadjuvant treatment are also excluded to avoid the influence of treatment intervention on pathological feature analysis.

[0065] (2) Feature extraction method based on self-supervised learning and feature vector processing of one case with multiple slices

[0066] In the feature extraction stage, first, a pathological expert accurately outlines the tumor area in the prostate cancer pathological section, and at three objective lens magnifications of 10×, 20×, and 40×, the outlined area is cropped into non-overlapping image patches (patches) with a size of 256×256 pixels as the input of the feature extraction network. To achieve efficient and unsupervised image representation learning, the system introduces a general self-supervised visual encoder - UNI designed specifically for pathological images. This encoder is pre-trained based on the DINOv2 self-supervised framework and can automatically learn discriminative image features from tissue images without manual annotation. In the output of the last convolutional network layer of the UNI encoder, the system extracts a 1024-dimensional feature vector for each patch. Finally, each panoramic pathological section scan image generates a set of low-dimensional feature matrices at three objective lens scales, with a size of {N S ×1024}, where N S represents the number of patches at this scale, s ∈ {10×, 20×, 40×}.

[0067] To enhance the generalization ability of the model and effectively reduce the risk of overfitting, a variety of data augmentation strategies are introduced in the feature extraction stage, including horizontal flipping, vertical flipping, random rotation by 45 degrees, and perturbations to brightness, contrast, saturation, and hue. These augmentation operations make the model more robust and able to learn structural features that are insensitive to direction changes and color transformations, thus enhancing its ability to adapt to complex pathological images. It should be particularly noted that for cases with multiple slices, at each objective magnification, the system will splice and fuse the feature vectors extracted from all slices of the case to construct a unified scale feature packet. Finally, each patient corresponds to feature packets at three scales (10×, 20×, 40×), which serve as the multi-scale input for the subsequent classification network to achieve accurate modeling of the risk of biochemical recurrence.

[0068] (3) Pseudo-packet strategy

[0069] The process of the pseudo-packet strategy is as follows Figure 3 , first divide an original parent packet into several smaller pseudo-packets according to the corresponding instance selection strategy, and each pseudo-packet inherits the label information of the parent packet; subsequently, select the top 20% of the instance features ranked according to the positive class probability from each pseudo-packet, and splice these high-confidence features as the input for subsequent modeling. These selected features are first fed into the attention network to calculate the instance-level attention scores. The attention mechanism can dynamically adjust the weights of different instance features according to the attention learned by the model, thus highlighting the regions that are most discriminative for the final classification result ( Figure 3 right figure). Then, use the attention scores to perform weighted fusion on the selected instance features to generate a compressed pseudo-packet-level representation as the representative feature of the pseudo-packet, and input it into the classifier for packet-level prediction. At the same time, the original parent packet also undergoes feature weighting through an independent attention network path to generate a global packet representation ( Figure 3 left figure). Finally, the parent packet features and the representative features of each pseudo-packet are jointly fed into the fusion module or downstream network for joint training and learning, so as to capture the local key region information while retaining the overall structural features and enhancing the model's ability to identify and discriminate heterogeneous lesions in the image.

[0070] (4) The position and spatial information generator module fully integrates spatial position features and semantic representations

[0071] After the feature packet is input into the spatial information generator module, dimensional reconstruction is first required to adapt to subsequent spatial modeling operations. The dimension of the original input feature X input is [R, D], and it is reconstructed into a four-dimensional tensor X reshaped, with dimensions [R, C, H, W], where C represents the number of channels, and H and W are the height and width of the image respectively. To introduce multi-scale spatial perception ability, the module uses three different sizes of convolutional kernels (7×7, 5×5, and 3×3 respectively) to perform parallel convolutional operations on the reconstructed feature map. These convolutional operations are denoted as Conv a , Conv b , and Conv c . The corresponding spatially-aware features are fused in the following way: X output = X reshaped + Conv a (X reshaped ) + Conv b (X reshaped ) + Conv c (X reshaped ). This process can effectively combine the spatial structure information under different receptive fields, enhance the model's ability to capture local and global contexts, and improve the multi-scale robustness of feature representation. Subsequently, the fused four-dimensional feature map X output is flattened into a two-dimensional form for input to the subsequent module: X final = flatten(X output ). Through this spatial information generator module, the model can more fully integrate spatial location features and semantic representations, providing more discriminative input features for downstream classification or attention mechanisms.

[0072] (5) Recurrent Cross-Attention Module Recursively analyze

[0073] By iteratively using the output of the cross-attention module as new input data, more abundant context information is gradually obtained. In the first iteration, the input feature map X final is passed to the cross-attention module, which aggregates long-range context information from each pixel in a cross form to generate a new output feature map X final ', where X final and X final ' have the same shape. Then the generated output feature map X final ' is input to the cross-attention module again. After multiple iterations, the module captures the dependency relationships at the pixel level and gradually generates a new feature map containing dense and rich context information. For any two positions (u x , u y ) and (θ x , θ y ) on the feature map, A is the attention map, and the information transfer for two cycles is as follows: .

[0074] This cross-attention based recursive modeling method not only improves the feature expression ability, but also enhances the discriminative ability of the model when dealing with pathological images with complex structures and extensive context dependencies.

[0075] By establishing a dual association between local fine features and global context through the position and spatial information generator module and the recurrent cross-attention module, the model can simultaneously identify microscopic cell abnormalities and macroscopic tissue structure changes. The fusion architecture simulates the thinking mode of pathological experts in the diagnosis process, namely macroscopic scanning - suspicious area identification - microscopic confirmation, enhances the understanding ability of complex cancer tissue morphology, and captures the high heterogeneity of prostate cancer.

[0076] (6) Multi-scale feature fusion

[0077] Multi-scale feature fusion is crucial in pathological image analysis. Image information at different magnifications is complementary and can jointly improve the model's understanding and discriminative ability of tissue structure. A low magnification objective lens (such as 10×) helps to identify the overall distribution, scope of cancer foci and their relationship with surrounding tissues, while a high magnification objective lens (such as 40×) reveals morphological details at the cellular level, such as nuclear atypia and nuclear-cytoplasmic ratio. Fusing feature information at different scales helps the model establish a hierarchical cognitive structure from global to local, which not only improves the ability to identify tumor heterogeneity, but also enhances the robustness to image noise and information loss.

[0078] The present invention breaks through the limitation of traditional methods for modeling single-scale images. By fusing image features at different magnifications, a more expressive multi-scale representation space is constructed, which improves the discriminative performance of the model for the recurrence risk of prostate cancer. In the model inference stage, an integration strategy is adopted to perform weighted average processing on the output results of the models trained at three scales (the weighting operation is completed before softmax), and finally an intelligent judgment result on whether the patient has biochemical recurrence is output.

[0079] II. System prediction performance evaluation and comparative analysis

[0080] 1. Multi-center dataset and model parameter configuration

[0081] The present invention uses panoramic pathological section scanning images with tumor regions delineated from 1,200 patients in Renji Hospital as training data, and validates the model's performance using panoramic pathological section scanning images of 850 patients and 336 patients in the TCGA-PRAD dataset from Renji Hospital. The positive rates of biochemical recurrence within three years are 22.9% and 11.6% respectively. The present invention adopts a five-fold cross-validation method to determine the optimal model parameter configuration. The training set data is randomly divided into five subsets of equal size. In turn, one subset is used as the validation set, and the remaining four subsets are used to train the model. For each fold of training, the grid search method is used to systematically traverse the key hyperparameters, including the learning rate (range from 1e-5 to 1e-3), batch size (16 to 64), weight decay coefficient (1e-5 to 1e-3), and feature dimensionality reduction parameters. Each combination of hyperparameters is trained until convergence, and the AUC value is calculated on the validation set. The model training uses the Adam optimizer, and the early stopping strategy is adopted to avoid overfitting. Among all the five-fold cross-validation combinations, the parameter combination with the highest average AUC value on the validation set is selected to construct the final model. The results show that the AUC of this model on the test set reaches 0.800, and the accuracy rate reaches 0.765, which is better than the existing comparative methods.

[0082] 2. Performance Evaluation

[0083] In the comparative experiment of predicting whether patients have biochemical recurrence based on pathological images, the model of the present invention, the existing cutting-edge new methods of max-pooling, mean-pooling, and the ABMIL model, which is one of the current best pathological image analysis algorithms, are trained and validated using the same samples. It shows that the model proposed by the present invention exhibits significantly superior performance. The comparison results are shown in Table 1: In the comparison with the existing mainstream methods, the system of the present invention is superior in terms of AUC (0.800) and accuracy rate (0.765) to the widely used max-pooling, mean-pooling strategies in recent years, and the multi-instance learning model (ABMIL) based on the trainable attention mechanism. This indicates that this model has a stronger ability to identify key pathological features and perform effective classification.

[0084] In addition, to verify the effectiveness of multi-scale image information fusion in this task, we further designed ablation experiments. The results shown in Table 2 indicate that after fusing image information at different magnification multiples such as 10×, 20×, and 40×, the model performance is significantly better than that of using only a single scale, indicating that multi-scale features play an important supplementary role in modeling the tumor recurrence risk, thereby enhancing the model's representation ability for complex tissue structures.

[0085] Table 1 Comparison Effects of Different Existing Mainstream Weak Supervision Models

[0086]

[0087] Table 2 Ablation effects at different scales

[0088]

[0089] The above results indicate that the system of the present invention has higher stability and reliability in accurately identifying the risk of biochemical recurrence, and can provide more powerful support for clinical diagnosis.

[0090] The parts not described in the present invention are the same as the prior art or are implemented using the prior art. The applicant declares that the present invention uses the above embodiments to illustrate the detailed method of the present invention, but the present invention is not limited to the above detailed method, that is, it does not mean that the present invention must rely on the above detailed method to be implemented. Those skilled in the art should understand that any improvement to the present invention, the equivalent substitution of each raw material of the product of the present invention, the addition of auxiliary components, the selection of specific methods, etc., all fall within the protection scope and the disclosure scope of the present invention.

Claims

1. A prostate cancer biochemical recurrence prediction system based on pathological sections, characterized in that, Comprising: An input display module, which is used to input the basic information of the patient and the panoramic pathological section scan image of the pathological section, and display the analysis result after the analysis is completed; A feature vector processing module, which cuts the tumor regions of the panoramic pathological section scan images preprocessed at different objective magnification multiples into non-overlapping image blocks of a certain number of pixels based on the features at each multiple, as the input of the feature extraction network; then uses a self-supervised vision encoder UNI to automatically extract features from it, and generates a feature representation at the panoramic pathological section scan image level containing multiple low-dimensional feature sets at different magnifications; A pseudo-packet construction module, which refines the feature representation at the panoramic pathological section scan image level based on the pseudo-packet strategy, and generates the feature packet representations of the pseudo-packet and the parent packet respectively; A location and spatial information generator module, based on the spatial distribution pattern of prostate cancer aggressiveness, first reconstructs the feature representation at the panoramic pathological section scan image level of the input or the feature packet representation of the pseudo-packet and the parent packet into a four-dimensional tensor X reshaped , and then uses convolutional kernels of different sizes to process the feature data to further extract features with spatial perception and context information; A cyclic cross-attention module, which complements convolutional kernels of different scales, and iteratively generates new feature maps with dense and rich context information based on cross-attention, enabling the model to simultaneously identify microscopic cell abnormalities and macroscopic tissue changes; then calculates the correlation between features based on these feature representations, and generates attention weights after normalizing these correlations; A multi-scale feature fusion module, which fuses the feature information at different objective magnification scales, performs weighted averaging on the biochemical recurrence probabilities obtained at different objective magnification multiples, obtains the recurrence or non-recurrence prediction result, and transmits it to the input display module.

2. The prediction system according to claim 1, wherein It also includes a storage module and a control module. The storage module is used to store the information processed by each module, and the control module controls the normal operation of each module.

3. The prediction system according to claim 1, wherein The preprocessing method of the panoramic pathological section scan image is as follows: by adjusting the mean and standard deviation of the image block data, the data is converted into a distribution with zero mean and unit variance, realizing the standardization of the data and the normalization of the image intensity, The tumor region is cut into non-overlapping image blocks of 256×256 pixels at the objective magnification multiples of 10×, 20×, and 40×, as the input of the feature extraction network, After the last convolutional layer of the UNI, 1024-dimensional feature vectors are extracted from each image patch respectively, generating a feature representation at the panoramic pathological slice scan image level. Each panoramic pathological slice scan image will generate a set of low-dimensional feature matrices at three magnification scales, with the size of {N s × 1024}, where N represents the number of image patches at this scale, and s ∈ {10×, 20×, 40×}. When the feature vector processing module performs feature extraction, it introduces a variety of data augmentation strategies, including horizontal flipping, vertical flipping, 45-degree random rotation, and perturbations to brightness, contrast, saturation, and hue, When a case has multiple pathological sections, at each objective magnification, the feature vectors extracted from all the sections of the case are spliced and fused to construct a unified scale feature packet.

4. The prediction system according to claim 3, wherein The feature information at the 10× objective magnification scale includes macroscopic tissue change features, and the feature information at the 40× objective magnification scale includes morphological details at the cell level.

5. The prediction system according to claim 1, wherein The method of the pseudo-packet strategy is as follows: Divide an original parent packet into several smaller pseudo-packets according to the corresponding instance selection strategy, with the same label as the original parent packet; Extract the top 20% of the features ranked from large to small according to the positive probability in the pseudo-packet, and splice these high-confidence features together, and then send them into the attention network to highlight the region most discriminative for the final classification result; Subsequently, use the attention scores to perform weighted fusion on the selected instance features to generate a compressed pseudo-packet feature packet representation as the representative feature of the pseudo-packet. At the same time, the original parent packet also generates a parent packet feature packet representation through an independent attention network path for feature weighting, and is jointly sent into the fusion module or downstream network for joint training and learning, while capturing local key region information and retaining the overall structural features, improving the model's recognition and discrimination ability for heterogeneous lesions in the image.

6. The prediction system according to claim 1, wherein The position and spatial information generator module first reconstructs the input features into a four-dimensional tensor X reshaped : , where X input is the original input feature with a dimension of [R, D], and the four-dimensional tensor X reshaped has a dimension of [R, C, H, W], where C is the number of channels, and H and W are the width and height of the image; then, convolution operations are performed on the reconstructed features using convolutional kernels of different sizes, and each convolutional kernel size corresponds to a different spatial range. These convolution operations generate multi-scale spatial perception feature representations in the features and are fused in the following way: X output = X reshaped + Conv a (X reshaped ) + Conv b (X reshaped ) + Conv c (X reshaped );Subsequently, the fused four-dimensional feature X output is flattened into a two-dimensional form for input to subsequent modules: X final =flatten(X output ).

7. The prediction system according to claim 6, wherein The cyclic cross-attention module forms a new feature map through an iterative loop manner: in the first iteration, using the feature map X final as the input, the module aggregates long-range context information from each pixel in a cross form, thereby generating a new output feature map X final ′, where X final and X final ′ have the same shape; then the generated output feature map X final ′ is input into the cross-attention module again. After multiple iterative processes, the module captures the dependency relationships at the pixel level and gradually generates a new feature map containing dense and rich context information. For any two positions (u x , u y ) and (θ x , θ y ) on the feature map, A is the attention map, and the information transfer of the two cycles is as follows: , Microscopic cell abnormalities include uneven cell size, nuclear atypia, increased chromatin, multiple mitotic figures, and abnormal cell metabolism; The macroscopic tissue changes include irregular gland size, disordered arrangement, basement membrane destruction, perineural invasion, vascular invasion, and interstitial fibrosis; After calculating the correlation between features, use the softmax operation to normalize these correlations to generate attention weights.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the functions of the system according to any one of claims 1 to 7.

9. A method for constructing a prostate cancer biochemical recurrence prediction system based on pathological sections according to any one of claims 1 to 7, characterized in that, It includes the following steps: (1) Dataset division: Summarize the panoramic pathological slice scan images generated from the postoperative pathological slices determined by doctors into a total dataset, and divide it into a training set and a validation set; (2) Construction of the feature vector processing module: Crop the preprocessed panoramic pathological slice scan images into non-overlapping image patches at different objective lens magnification multiples, and use the self-supervised learning visual encoder UNI to extract 1024-dimensional feature vectors for each image patch to generate a feature representation at the panoramic pathological slice scan image level; (3) Construction of the pseudo-packet construction module: Divide an original parent packet into several smaller pseudo-packets according to the corresponding instance selection strategy, with the same label as the original parent packet; Extract the top 20% of the features ranked from large to small according to the positive probability in the pseudo-packet, and splice these high-confidence features together, and then send them into the attention network to highlight the region most discriminative for the final classification result; Subsequently, use the attention scores to perform weighted fusion on the selected instance features to generate a pseudo-packet representative feature and a parent packet representative feature respectively; (4)Construction of the Location and Spatial Information Generator Module: By reconstructing the input features into four-dimensional features and processing them with convolutional kernels of different sizes to generate multi-scale location-aware feature representations; finally, through encoding and integration, convolutional kernels of different scales are used to provide a more comprehensive feature expression to generate the final output feature X final = flatten(X onput ); (5) Construction of the cyclic cross-attention module: Gradually enhance the context information through multiple iterations. After multiple iterations, the module captures the dependence relationship between pixels to generate a new feature map containing richer context information; (6) Construction of the multi-scale feature fusion module to fuse the feature information at different objective lens magnification scales, enabling the model to establish a hierarchical cognitive structure from global to local.

10. The construction method of the prostate cancer biochemical recurrence prediction system based on pathological slices according to claim 9, characterized in that: In step (1), the dataset is randomly divided into five subsets, and one of them is used as the validation set in turn, and the remaining four are used for training. After repeating five times, the average result is taken to evaluate the model performance; In step (2), a set of low-dimensional feature matrices with a size of {Ns×1024} are generated for each panoramic pathological section scan image at three magnification scales, where N represents the number of image patches at that scale, and s∈{10×, 20×, 40×}; During the feature extraction process, data augmentation methods are used to increase data diversity and help the model learn rotation- and color-invariant image representations; In step (3), the pseudo-package representative features and the parent-package representative features are jointly fed into a fusion module or a downstream network for joint training and learning; In step (4), the reconstruction into the four-dimensional feature X reshaped is as follows: ; the processing method for convolutional kernels of different sizes is as follows: X output = X reshaped + Conv a (X reshaped ) + Conv b (X reshaped ) +Conv c (X reshaped ); In step (5), in the first iteration, the input feature map X final is passed to the cross-attention module, which aggregates long-range context information between pixels to generate a new output feature map X final ', with the same shape as X final ; then, X final ' is used as the input and processed by the cross-attention module again.

Citation Information

Patent Citations

  • Pathological image classification method and system based on multi-scale domain adversarial network

    CN114299324A

  • Full-scanning pathological feature fusion extraction method and system based on multiple scales

    CN115063592A

  • Primary gastric cancer digital pathological section lymph node metastasis risk prediction method

    CN116705311A

  • Full-view digital pathological section survival prediction method based on multi-pseudo-example graph convolution

    CN117392051A

  • Cancer pathology image classification system based on multi-instance learning, medium and equipment

    CN117392468A

Cited By

  • Tumor pathological image segmentation method and system based on graph convolution contrast learning network

    CN120563847A

  • Intelligent construction method and system for chronic hepatitis B pathological image analysis model

    CN121544944A