System for predicting prostate cancer patients with homologous recombination deficiency based on magnetic resonance images and whole-slide pathology scans

Through a multimodal prediction system based on magnetic resonance images and pathological panoramic scanning sections, the problems of high cost and technical complexity of HRD detection in the prior art are solved, and economical, fast and accurate HRD status prediction is achieved.

CN117095815BActive Publication Date: 2025-05-27THE FIRST AFFILIATED HOSPITAL OF NAVAL MEDICAL UNIVERSITY OF CHINESE PEOPLES LIBERATION ARMY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310893003.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-20
Publication Date
2025-05-27
Estimated Expiration
2043-07-20

AI Technical Summary

Technical Problem

In the prior art, detecting whether prostate cancer has homologous recombination defects (HRD) has high cost and technical complexity, which makes some patients unable to bear or obtain timely test results.

Method used

A multimodal prediction system based on magnetic resonance images and pathological panoramic scanning sections is adopted. The input module collects clinical information, pathological and image data, and the quality control module performs data preprocessing. The multimodal prediction module uses machine learning and deep learning models to predict HRD state.

Benefits of technology

It significantly reduces the economic burden of HRD evaluation, improves the accuracy and reliability of diagnosis, reduces detection time, and reduces the requirements for sample quality, making HRD detection more practical and widespread.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117095815B_ABST
    Figure CN117095815B_ABST
Patent Text Reader

Abstract

The present invention provides a system for predicting prostate cancer patients with homologous recombination deficiency based on magnetic resonance images and pathological whole-slide scans, which includes: an input module, a quality control module, and a multimodal prediction module. The input module collects the clinical information, pathological whole-slide scan information, and magnetic resonance image information of patients. The quality control module preprocesses the received clinical information, pathological information, and MRI image information. The multimodal prediction module inputs the complete three-modal prostate cancer data information of the preprocessed patient into a risk prediction model to predict its HRD status. By using machine learning and image processing technologies, enhancing the prediction efficiency of the model through data augmentation, and adopting innovative data preprocessing and feature extraction strategies, the present invention successfully solves the problem of multimodal data fusion, significantly reduces the economic burden of HRD assessment, and can provide results in a shorter time to meet the needs of prostate cancer patients requiring urgent treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an artificial intelligence system for predicting homologous recombination deficiency in prostate cancer by using clinical pathological images based on multiple modalities. Background Art

[0002] Prostate cancer (PCa) is a malignant tumor mainly affecting elderly men. The incidence rate of prostate cancer is high, and the treatment and management of advanced prostate cancer are major medical challenges. According to the report of "Cancer Statistic, 2018" in the United States, it is estimated that in 2018, the number of new cases of prostate cancer in the United States was 164,690, and the number of death cases was 29,430. Among male malignancies, its incidence rate and mortality rate rank first and second respectively. In China, with the intensification of aging, the incidence rate of prostate cancer in Chinese men is rising rapidly. During the period from 2000 to 2016, its annual growth rate was 13%. In particular, once prostate cancer reaches the advanced stage, its five-year survival rate is relatively low. This is mainly due to the emergence of endocrine therapy resistance. For such patients, the only targeted therapy drug approved by the US FDA and China NMPA in the current prostate cancer treatment field is the PARP inhibitor olaparib.

[0003] However, not all patients can obtain significant efficacy by using this drug. Only those prostate cancer patients with homologous recombination deficiency (HRD) can benefit from it. Therefore, a very important clinical problem is how to find those patients with HRD.

[0004] The main method for detecting HRD in prostate cancer is genomic sequencing. Genomic sequencing is a process of analyzing genetic information by detecting DNA sequences in tumor tissues. In the context of prostate cancer and other cancers, genomic sequencing particularly focuses on finding variations that may be related to tumorigenesis, development, and treatment response. Because it can analyze genomic instability and simultaneously detect mutations in HRD-related genes, thereby comprehensively evaluating its HRD status and helping doctors determine whether patients are suitable for receiving targeted therapy.

[0005] However, there are some limitations to this method currently. First, due to its high cost and technical complexity, some patients cannot afford the HRD test. For example, currently in the United States, the reported cost of these assays ranges from $4,800 to $5,800. In addition, gene sequencing may take a long time to obtain results, which is impractical for patients in need of urgent treatment. The analysis of the gold standard loss of heterozygosity (LOH) and copy number variation (CNV) requires high-quality tumor samples, and the detection of large-scale structural variation (LST) and other genomic instabilities requires complex data analysis, which also limits its application in clinical practice. Third, the accessibility is poor. In many small cities and rural areas, due to limited medical resources and lack of necessary expertise and equipment, high-level gene testing is not easily available. Patients in remote areas may also need to travel long distances to receive the test, increasing the time and economic burden. These factors comprehensively affect the test accessibility of patients. These obstacles may limit patients from missing potential treatment opportunities. These obstacles may limit patients from missing potential treatment opportunities.

[0006] Other non-invasive emerging options such as liquid biopsy are gradually gaining favor, but their sensitivity and specificity need to be verified. For example, the amount of circulating tumor DNA (ctDNA) in many patients is insufficient, and ctDNA mutations may be below the detection limit, resulting in false negative results, and it has not been widely used.

[0007] Therefore, there is an urgent need for a more accessible HRD test method to more accurately identify prostate cancer patients who may benefit from targeted therapy. This will help improve the treatment effect and provide more treatment options and hope for patients. Summary of the Invention

[0008] In order to overcome the above technical defects, the purpose of the present invention is to provide a system for predicting prostate cancer patients with homologous recombination deficiency based on magnetic resonance images and whole slide pathology scans, which includes: an input module, a quality control module, and a multimodal prediction module;

[0009] The input module is used to collect the clinical information, whole slide pathology scan slice information, and magnetic resonance image information of the patient;

[0010] The segmentation quality control module includes a clinical information quality control sub-module, a whole slide pathology scan slice quality control sub-module, and a magnetic resonance image quality control sub-module;

[0011] Among them, the clinical information quality control sub-module is used to discard clinical data with missing key clinical information (pre-defined) exceeding a certain proportion, perform validity checks on the clinical data, and perform operations on the complete and valid clinical data;

[0012] Among them, the pathological whole-slide scanning slice quality control sub-module includes a pathological slice whole-slide scanning module, an image quality assessment module, and an image preprocessing module; the pathological slice whole-slide scanning module is used to scan the slices of patients and convert them into high-definition digital whole-slide images with 40x magnification and 1mm 3 resolution; the image quality assessment module is used to automatically detect and exclude the whole-slide scanned pathological slice images with poor quality; the image preprocessing module is used to digitally capture the whole-slide scanned pathological slice images that have passed the quality screening, and mark the regions of interest (ROIs) in a fully automatic, semi-automatic, or purely manual manner; for example, the pathological slice annotation software combines custom program modules to achieve automatic annotation of tumor regions / automatic annotation of tumor stroma / automatic annotation of benign glands to achieve fully automatic annotation; based on the above-mentioned pathological slice annotation results, custom program modules are used to add manual review on the basis of automatic annotation to achieve semi-automatic delineation; or a completely manual annotation method is adopted, that is, a pathological slice annotation software such as Qupath is used to manually delineate the ROIs by using methods such as polygons, circles, and magic wands;

[0013] Among them, the magnetic resonance image quality control sub-module includes an image receiving module, an image quality assessment module, and an image preprocessing and lesion segmentation module; the image receiving module is used to receive DICOM-format magnetic resonance images; the image quality assessment module is used to automatically detect and exclude the imaging images with quality defects; the image preprocessing and lesion segmentation module is used to convert the DICOM format into a volume image, that is, to convert the original magnetic resonance image into a two-dimensional volume image with coordinates to facilitate subsequent analysis, and mark the ROIs in a fully automatic, semi-automatic, or purely manual manner;

[0014] The multi-modal prediction module includes a clinical prediction sub-module, a pathological prediction sub-module, an imaging prediction sub-module, a multi-modal fusion module, and a multi-modal output module;

[0015] Among them, the clinical prediction sub-module includes a data filling module, a clinical feature HRD prediction module, and a clinical result display module; the data filling module is used to fill in the missing values of the patient data, and fill them with the average value and median; the clinical feature HRD prediction module is used to input the processed clinical information into a risk assessment model established by a logistic regression algorithm on the basis of preprocessing and feature engineering, and calculate the clinical risk index; the clinical result display module is used to receive the risk index obtained clinically and output and display it in the form of a percentage;

[0016] Among them, the pathological prediction sub-module includes a tile division module, a data augmentation and normalization module, a tile-level HRD status prediction module, a pathological whole-slide scan slice-level HRD status prediction module, and a pathological result display module; the tile division module is used to process the digital whole-slide scan pathological slice image, and finely divide it into multiple tiles, and further screen these tiles, only retaining the tiles with an overlap degree of more than 80% with the region of interest (ROI) to improve the calculation efficiency and accuracy; the data augmentation and normalization module is used to perform online data augmentation techniques of random horizontal and vertical flipping on the tiles, and normalize the image tiles by performing z-score normalization on the RGB channels; the tile-level HRD status prediction module is used to process the image tiles that have undergone data augmentation and normalization using a pre-trained convolutional neural network using the ResNet50 method, and calculate the probability of the HRD status for each tile; the pathological whole-slide scan slice-level HRD status prediction module is used to receive the tile-level HRD status probability as input, and use the tile likelihood histogram and bag-of-words model strategy to integrate this information, generate a feature vector at the pathological whole-slide scan slice level, and use the LightGBM method to establish a machine learning classifier model, and use the feature vector generated based on the PLH and BoW methods to predict the HRD status of the entire pathological whole-slide scan slice; the pathological result display module is used to receive the HRD status prediction at the pathological whole-slide scan slice level and present it to the user in an intuitive manner in the form of a Grad-CAM map and a pathological report;

[0017] Among them, the imaging prediction sub-module includes an imaging normalization module, a feature extraction module, an HRD risk prediction module, and an imaging result display module; the imaging normalization module is used to first sort the pixel values in the image through the intensity truncation and normalization sub-module, and truncate the intensity values so that they fall within the range of the 0.5 to 99.5 percentiles, and then through the spatial normalization sub-module, resample the image to a resolution of 1mm x 1mm x 1mm; the feature extraction module uses the PRAD-MRI-OMICS tool built by the system itself to extract the manual features required for the system to predict HRD, including geometric features, intensity features, and texture features, from the segmented and quality-controlled lesion regions; the HRD risk prediction module is used to establish a machine learning model using the support vector machine (SVM) method, predict the HRD risk according to the features selected by the model, and use the trained model to generate an HRD risk score for each sample; the imaging result display module is used to receive the HRD status prediction in radiomics and output the HRD risk score in the form of a percentage;

[0018] Among them, the multi-modal fusion module is used to integrate the outputs of the clinical result display module, the pathological result display module, and the imaging result display module by adopting pre-fusion, mid-fusion, and post-fusion strategies, and utilize multi-modal fusion technology to obtain a comprehensive evaluation result of the HRD status of prostate cancer.

[0019] Among them, the multi-modal output module is used to summarize all the results, output the overall predicted HRD risk in the form of a percentage, and simultaneously display the comprehensive evaluation result, and the predicted results of clinical, pathological, and imaging, and give reference diagnosis and treatment opinions.

[0020] Furthermore, the system for predicting prostate cancer patients with homologous recombination deficiency based on magnetic resonance images and pathological panoramic scanning slices further includes a high-throughput digital information memory and a high-speed high-frequency data processor. The high-throughput digital information memory is used to store the input clinical information, pathological images, magnetic resonance images, and their prediction results obtained by the multi-modal prediction module, as well as a computer program for storing the steps of the method for predicting the HRD status of prostate cancer. The high-speed high-frequency data processor is used to implement the steps of the method for predicting the HRD status of prostate cancer when executing the computer management program stored in the high-throughput digital information memory.

[0021] Furthermore, the system for predicting prostate cancer patients with homologous recombination deficiency based on magnetic resonance images and pathological panoramic scanning slices further includes a model evaluation module, a model reconstruction module, and a model comparison and adaptive update module. The model evaluation module is used to receive the HRD evaluation result from the multi-modal prediction module, verify the prediction result of the model using known gold standard data, and then calculate various performance metrics to evaluate the trained neural network model. The model reconstruction module is used to collect and integrate the multi-modal data obtained from the storage module, and perform full reconstruction, partial reconstruction, or manual reconstruction of the neural network to adjust the model structure and parameters, so as to identify and select the model with the optimal performance as the updated HRD evaluation model. The model comparison and adaptive update module is used to test the new model using the validation set and compare it with the performance metrics of the original model to evaluate the advantages and disadvantages of the two models. If the performance of the new model does not meet the predetermined standard or is inferior to the original model, it is possible to choose to retrain and optimize the model, or choose to retain the original model.

[0022] Furthermore, the pre-fusion strategy is the canonical correlation analysis strategy, the mid-fusion strategy is the multi-kernel learning strategy, and the post-fusion strategy is the stacking generalization strategy.

[0023] After adopting the above technical solutions, compared with the prior art, the following beneficial effects are achieved:

[0024] 1. The present invention provides an economically enhanced HRD assessment scheme: The present invention provides an economically enhanced assessment scheme for homologous recombination deficiency (HRD) in prostate cancer. Traditional methods usually rely on high-cost genomic sequencing, while the present invention makes it possible to directly extract key information from medical images by ingeniously utilizing machine learning and image processing technologies, thus significantly reducing the economic burden of HRD assessment.

[0025] 2. The present invention provides optimization of the comprehensive utilization of patient data: The present invention actively comprehensively utilizes the clinical information of patients in optimizing HRD diagnosis. The potential of this information is often not fully utilized in traditional diagnostic processes. By incorporating patient data into the training and validation of the model, the present invention increases the information richness for diagnosing HRD, thereby enhancing the accuracy and reliability of the diagnosis.

[0026] 3. The present invention enhances the prediction efficacy of the model through data augmentation: The present invention is committed to solving the problem of insufficient prediction ability of existing models due to insufficient training samples. By adopting data augmentation and precise feature extraction technologies, the present invention can continue to train samples and feature quantities after application, thereby continuously enhancing the prediction accuracy and comprehensive performance of the model.

[0027] 4. Precise integration of multi-modal data fusion: The present invention successfully solves the problem of multi-modal data fusion by adopting innovative data preprocessing and feature extraction strategies. This achievement is realized by efficiently integrating data from different modalities into a unified analysis framework, thereby enhancing the model's prediction ability.

[0028] 5. Enhancing the interpretability of multi-modal models: The present invention significantly enhances the interpretability of multi-modal models at the pathological level by adopting specific feature extraction technologies and model structures. This enhanced interpretability provides clinicians with deeper insights, helping them to more accurately understand and trust the model's predictions.

[0029] 6. Cost-effectiveness: The present invention provides a more economical HRD detection method and primary screening strategy, significantly reducing the cost of detection and enabling more patients to afford this detection.

[0030] 7. Quick results: The present invention can provide results in a shorter time by optimizing and simplifying the testing process, meeting the needs of prostate cancer patients who require urgent treatment.

[0031] 8. The present invention has lower requirements for sample quality: Compared with existing genomic sequencing methods, it does not require high-quality tumor samples, making HRD detection more practical and widely feasible.

[0032] 9. Promote personalized treatment: The HRD detection method of the present invention enables doctors to formulate treatment plans for prostate cancer patients more personalized, because it provides more information about the genetic characteristics of the patients' cancers. This enables doctors to select more targeted and effective drugs and treatment plans for patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a structural module diagram of a system for predicting prostate cancer patients with homologous recombination deficiency based on magnetic resonance images and whole-slide pathology scans in an embodiment of the present application;

[0034] Figure 2 It is to adopt Figure 1 The flowchart of predicting prostate cancer patients with homologous recombination deficiency by the prediction system in;

[0035] Figure 3 It is a structural module diagram of the multi-modal prediction module in the present application;

[0036] Figure 4 It is a structural module diagram of the model evaluation and adaptive update module in the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The advantages of the present invention are further elaborated below in conjunction with the accompanying drawings and specific embodiments. Those skilled in the art should understand that the content specifically described below is illustrative rather than restrictive, and should not be used to limit the protection scope of the present invention.

[0038] As Figure 1 shown, this embodiment provides a system for predicting prostate cancer patients with homologous recombination deficiency based on magnetic resonance images and whole-slide pathology scans, which includes: an input module, a quality control module, a multi-modal prediction module, a model evaluation and adaptive update module, and a storage module.

[0039] As Figure 2 shown, the method for predicting prostate cancer patients with homologous recombination deficiency using the above prediction system includes the following steps:

[0040] Step S1: Collect the clinical information, pathological information, and magnetic resonance imaging (MRI) image information of the patient.

[0041] The input module collects the clinical information, whole-slide pathology scan slice information, and magnetic resonance image information of the patient.

[0042] Step S2: Preprocess the received clinical information, pathological information, and MRI image information.

[0043] The quality control module preprocesses the received clinical information, pathological information, and MRI image information. The quality control module includes a clinical information quality control sub-module, a pathological panoramic scan slice quality control sub-module, and a magnetic resonance image quality control sub-module.

[0044] For clinical data, the clinical information quality control sub-module performs operations based on complete and accurate information. For example, if more than 60% of the clinical information (such as age, PSA, TNM stage, etc.) of a patient is missing, then the modality data will be discarded. At the same time, validity checks are also carried out to ensure that the clinical information is within a reasonable range.

[0045] For pathological data, the pathological panoramic scan slice quality control sub-module includes a pathological slice panoramic scan module, an image quality assessment module, and an image preprocessing module. The pathological slice panoramic scan module scans the patient's slice in a scanning instrument and uses a 40x objective lens to scan and convert the pathological slice into a high-definition digital panoramic image with 40x magnification and 1mm 3 resolution. This fine resolution and large data volume allow for in-depth and precise pathological analysis. The image quality assessment module automatically detects and excludes panoramic scan pathological slice images with poor quality, such as images with obvious folding, contamination, lack of tumor tissue, or unclear focus, and prompts for rescan. The image preprocessing module digitally captures the panoramic scan pathological slice images that have passed the quality screening. It can be done in a fully automatic, semi-automatic, or purely manual way to mark the regions of interest. For example, the pathological slice annotation software combined with a custom program module can achieve automatic annotation of the tumor region / automatic annotation of the tumor stroma / automatic annotation of benign glands to achieve full-automatic annotation; based on the above-mentioned pathological slice annotation results, a custom program module can be used to add manual review on the basis of automatic annotation to achieve semi-automatic delineation; or a completely manual annotation method can be adopted, that is, using pathological slice annotation software such as Qupath and using methods such as polygons, circles, and magic wands to manually mark the ROI in a purely manual way.

[0046] For image data, the magnetic resonance image quality control sub-module includes an image reception module, an image quality assessment module, and an image preprocessing and lesion segmentation module. The image reception module is used to receive DICOM-format magnetic resonance images from different devices. The image quality assessment module automatically detects and excludes image images with poor quality, such as those with blurring, excessive noise, insufficient contrast and brightness, artifacts, insufficient spatial resolution, or incomplete and occluded important structures. The image preprocessing and lesion segmentation module is used to convert the DICOM format into a volume image, that is, to convert the original magnetic resonance image into a two-dimensional volume image with coordinates for convenient subsequent analysis. It can be done in a fully automatic, semi-automatic, or purely manual way to mark the ROI of interest for subsequent processing.

[0047] Input the complete prostate cancer data information of the preprocessed patient in three modalities (i.e., preprocessed clinical information, pathological images, and MRI images) into the multi-modal prediction module.

[0048] Step S3: Input the complete prostate cancer data information of the preprocessed patient into the risk prediction model to predict its HRD status.

[0049] In the multi-modal prediction module, the clinical information will be processed through a logistic regression model to obtain the clinical risk index. The pathological images and MRI images will also undergo a series of processes (such as tiling, data augmentation, normalization, etc.) and be processed through a deep learning model to obtain the HRD status predictions on pathology and imaging respectively. Then, the fusion output module will combine the outputs of the clinical, pathological, and imaging sub-modules through post-fusion technology to obtain the final overall HRD assessment result and output it in the form of a percentage.

[0050] As Figure 3 shown, the multi-modal prediction module includes a clinical prediction sub-module, a pathological prediction sub-module, an imaging prediction sub-module, a multi-modal fusion module, and a multi-modal output module;

[0051] Among them, the clinical prediction sub-module includes a data filling module, a clinical feature HRD prediction module, and a clinical result display module. The data filling module is used to fill in the missing values of the patient data, using the mean and median for filling. The clinical feature HRD prediction module is used to input the processed clinical information into a risk assessment model established by the Logistic Regression (LR) algorithm on the basis of preprocessing and feature engineering, and calculate the clinical risk index. The clinical result display module is used to receive the risk index obtained clinically and output and display it in the form of a percentage.

[0052] Among them, the pathological prediction sub-module includes a tile division module, a data augmentation and normalization module, a tile-level HRD status prediction module, a pathological panoramic scan slice-level HRD status prediction module, and a pathological result display module. The tile division module is used to process the digital panoramic scan pathological slice image and finely divide it into multiple tiles. The configurable tile sizes include 256x256 pixels, 512x512 pixels, etc., to meet different computing requirements and resolution preferences. The module will further screen these tiles and only retain the tiles with an overlap of more than 80% with the region of interest (ROI) to improve the computing efficiency and accuracy. The data augmentation and normalization module is used to adopt online data augmentation techniques for the tiles, including random horizontal and vertical flips. In addition, the image tiles are normalized by performing z-score normalization on the RGB channels, thereby increasing the generalization ability of the model. The tile-level HRD status prediction module uses a pre-trained convolutional neural network of the ResNet50 method to process the image tiles that have undergone data augmentation and normalization, and calculates the probability of the HRD status for each tile. Multiple pre-trained convolutional neural networks can be used, including ResNet50, VGG16, InceptionV3, DenseNet, or Xception, etc., for flexible selection according to different needs and preferences. These networks have different structures and performance characteristics and can be used to achieve efficient and accurate tile-level HRD status prediction. The pathological panoramic scan slice-level HRD status prediction module is used to receive the HRD status probabilities at the tile level as input, and adopts the Patch Likelihood Histogram (PLH) and Bag of Words (BoW) strategies to integrate this information, generate a feature vector at the pathological panoramic scan slice level, and uses the LightGBM method to establish a machine learning classifier model, and uses the feature vector generated by the PLH and BoW methods to predict the HRD status of the entire pathological panoramic scan slice. The pathological result display module is used to receive the HRD status prediction at the pathological panoramic scan slice level and present it to the user in an intuitive way in the form of Grad-CAM diagrams and pathological reports;

[0053] Among them, the image prediction sub-module includes an image normalization module, a feature extraction module, an HRD risk prediction module, and an image result display module. The image normalization module is used to first, through the intensity truncation and normalization sub-module, sort the pixel values in the image and truncate the intensity values so that they fall within the range of the 0.5th to 99.5th percentiles, in order to reduce the influence of abnormal pixel values on the results. Then, through the spatial normalization sub-module, the image is resampled to a resolution of 1mm x 1mm x 1mm to reduce the influence of the change in voxel spacing on the results, thereby increasing the generalization ability of the model. The feature extraction module uses the PRAD-MRI-OMICS tool built by the system itself to extract the manual features required for system prediction of HRD, including geometric features, intensity features, and texture features, from the segmented and quality-controlled lesion area. The HRD risk prediction module is used to establish a machine learning model using the support vector machine (SVM) method and predict the HRD risk according to the features selected by the model. The trained model can be used to generate an HRD risk score for each sample. The image result display module is used to receive the HRD status prediction on radiomics and output the HRD risk score in the form of a percentage.

[0054] Among them, the multi-modal fusion module is used to integrate the outputs of the clinical result display module, the pathological result display module, and the image result display module by adopting pre-fusion, mid-fusion, and post-fusion strategies, and utilize multi-modal fusion technology to obtain a comprehensive evaluation result of the HRD status of prostate cancer.

[0055] This module is used to integrate the outputs of the above three sub-modules and utilize multi-modal fusion technology to obtain a comprehensive evaluation result of the HRD status of prostate cancer. This module allows three main fusion strategies to be adopted: pre-fusion, mid-fusion, and post-fusion. Various fusion methods are available for selection in the fusion strategy, including "weighted average", "voting mechanism", "Stacking", "Canonical Correlation Analysis (CCA)", and "Multiple Kernel Learning (MKL)". This module aims to provide a flexible and powerful framework to make full use of information from different modalities, so as to achieve an accurate and comprehensive evaluation of the HRD status. Users can select and adjust appropriate fusion strategies and methods according to specific application scenarios and data characteristics. In addition, the following recommended combination selections are included in the module for selection:

[0056] Pre-fusion: Under this strategy, it is recommended to use "Canonical Correlation Analysis (CCA)" to discover and integrate relevant features from different modalities. This strategy integrates data before model training and is particularly suitable for processing multi-modal data with high correlations.

[0057] Fusion in the middle: Under this strategy, "Multiple Kernel Learning (MKL)" is recommended as the fusion method. By linearly combining the data kernels of different modalities, MKL allows the model to dynamically integrate information from different modalities during training.

[0058] Fusion at the end: Under this strategy, "Stacking" is recommended to integrate the outputs from each sub-module. By training a meta-model, Stacking can learn how to effectively combine the outputs of sub-modules to optimize the prediction of HRD status.

[0059] Among them, the multi-modal output module is used to output the fused prediction results, output the overall predicted HRD risk using a set of pre-determined score cut-off values, and give reference diagnosis and treatment opinions. For example, when the comprehensive score is below 0.4, the system classifies the HRD risk as very low; when the comprehensive score is between 0.4 and 0.7, it is classified as medium risk; when the score is above 0.7, it is classified as high risk.

[0060] For example, three cases using the weighted average method under the end-fusion strategy: Patient A, Patient B, and Patient C. Patient A only provides pathological section data, and the score generated after analysis is 0.76. Since this score exceeds the cut-off value of 0.70, the system classifies Patient A as a high HRD risk according to the above criteria.

[0061] Patient B provides pathological section and clinical data. His pathological score is 0.65 and the clinical score is 0.5. The system calculates the comprehensive score by computing the simple average of these scores, which is 0.58. Since this score is between 0.4 and 0.7, the system classifies Patient B as a medium HRD risk.

[0062] Patient C provides pathological section, clinical data, and imaging data. His pathological score is 0.6, the clinical score is 0.7, and the imaging score is 0.8. The system calculates the simple average of these scores, and the comprehensive score is 0.7. Since this score is equal to the cut-off value of 0.7, the system classifies Patient C as a high HRD risk.

[0063] Preferably, the storage module is a high-throughput digital information memory that stores the input clinical information, pathological images, MRI images, the prediction results obtained by the multi-modal prediction module for the three features, and a computer program. When the computer program is executed by the processor, it implements the steps of the above method for predicting the HRD status of prostate cancer. Preferably, the processor is a high-speed high-frequency data processor.

[0064] Step S4: Model evaluation and adaptive update.

[0065] The model evaluation and adaptive update module receives the prediction results from the multimodal prediction module, validates them against the known gold standard data, and calculates performance metrics such as accuracy. If a certain amount of patient data is collected, the reconstruction module can be used to reconstruct this model. If the accuracy of the model does not meet the predetermined standard, a retraining and optimization process can be carried out, or the improved model can be rejected.

[0066] Specifically, as Figure 4 shown, the model evaluation and adaptive update module includes a model evaluation module, a model reconstruction module, and a model comparison and adaptive update module.

[0067] The model evaluation module This module is mainly responsible for evaluating the performance of the model. It uses the HRD evaluation results from the multimodal prediction module and the known gold standard data, such as loss of heterozygosity (LOH), copy number variation (CNV), and large-scale genomic instability (LST), to verify the prediction results of the model. After that, it calculates various performance metrics, such as accuracy, recall, F1-score, and area under the curve (AUC). It also includes the analysis of the confusion matrix to gain an in-depth understanding of the model's performance on different classes, as well as the evaluation of the sensitivity and specificity for specific problems.

[0068] The model reconstruction module This module mainly involves adjusting the model structure and parameters. Its purpose is to ensure that the accuracy and reliability of the HRD evaluation continuously improve over time and with the accumulation of data.

[0069] According to different needs and situations, there can be different levels of reconstruction:

[0070] A. Full reconstruction: The module fetches and integrates new multimodal data from the storage module, then re-performs feature extraction and selection, and comprehensively constructs a new neural network model with unchanged hyperparameters.

[0071] B. Partial reconstruction: In this case, only the end nodes of the neural network are adjusted, while the previous model is used as a pre-trained model, keeping the hyperparameters unchanged. This is a faster adjustment method.

[0072] C. Manual reconstruction: Based on full reconstruction, this module allows manual adjustment of the hyperparameters, structure options, and other settings for model construction to achieve more refined optimization. Hyperparameters include changing the learning rate, optimizer, activation function, etc.

[0073] The model comparison and adaptive update module is used to test the new model with the validation set in this module and compare it with the performance metrics such as the accuracy and AUC of the original model. This can evaluate the advantages and disadvantages of the two models. If the performance of the new model does not meet the predetermined standard or is inferior to the original model, one can choose to retrain and optimize the model or choose to retain the original model. It aims to identify and select the model with the best performance as the updated HRD evaluation model to improve the prediction effect of the model.

[0074] It should be noted that the embodiments of the present invention have better implementability and are not any form of limitation to the present invention. Any person skilled in the art may use the technical content disclosed above to change or modify it into equivalent effective embodiments. However, as long as it does not depart from the content of the technical solution of the present invention, any modification, equivalent change or modification made to the above embodiments according to the technical essence of the present invention still falls within the scope of the technical solution of the present invention.

Claims

1. A system for predicting prostate cancer patients with homologous recombination deficiency based on magnetic resonance images and pathological whole-slide scans, characterized in that, it includes: an input module, a quality control module, and a multimodal prediction module; the input module is used to collect the clinical information, pathological whole-slide scan slice information, and magnetic resonance image information of the patient; the quality control module includes a clinical information quality control sub-module, a pathological whole-slide scan slice quality control sub-module, and a magnetic resonance image quality control sub-module; wherein, the clinical information quality control sub-module is used to discard clinical data with key clinical information missing by more than a certain proportion, perform validity checks on the clinical data, and perform calculations on the complete and valid clinical data; Among them, the pathological whole-slide scanning slice quality control sub-module includes a pathological slice whole-slide scanning module, an image quality assessment module, and an image preprocessing module; the pathological slice whole-slide scanning module is used to scan the slices of patients and convert them into high-definition digital whole-slide images with 40x magnification and 1mm 3 resolution; the image quality assessment module is used to automatically detect and exclude panoramic scanning pathological slice images with poor quality; the image preprocessing module is used to digitally capture the panoramic scanning pathological slice images that have passed quality screening and mark the ROI in a fully automatic, semi-automatic, or purely manual manner; wherein, the magnetic resonance image quality control sub-module includes an image receiving module, an image quality assessment module, and an image preprocessing and lesion segmentation module; the image receiving module is used to receive DICOM-format magnetic resonance images; the image quality assessment module is used to automatically detect and exclude image images with quality defects; the image preprocessing and lesion segmentation module is used to convert the DICOM format into a volume image, that is, convert the original magnetic resonance image into a two-dimensional volume image with coordinates, and mark the ROI in a fully automatic, semi-automatic, or purely manual manner; the multimodal prediction module includes a clinical prediction sub-module, a pathological prediction sub-module, an imaging prediction sub-module, a multimodal fusion module, and a multimodal output module; wherein, the clinical prediction sub-module includes a data filling module, a clinical feature HRD prediction module, and a clinical result display module; the data filling module is used to fill in the missing values of the patient data, using the average value and median for filling; the clinical feature HRD prediction module is used to input the processed clinical information into a risk assessment model established by a logistic regression algorithm on the basis of preprocessing and feature engineering, and calculate the clinical risk index; the clinical result display module is used to receive the risk index obtained clinically and output and display it in the form of a percentage; Among them, the pathological prediction sub-module includes a tile division module, a data augmentation and normalization module, a tile-level HRD status prediction module, a pathological whole-slide scan slice-level HRD status prediction module, and a pathological result display module; the tile division module is used to process the digitized whole-slide scan pathological slice image, and finely divide it into multiple tiles, and further screen these tiles, only retaining the tiles with an overlap degree with the ROI exceeding 80% to improve the calculation efficiency and accuracy; the data augmentation and normalization module is used to perform online data augmentation techniques of random horizontal and vertical flipping on the tiles, and normalize the image tiles by performing z-score normalization on the RGB channels; the tile-level HRD status prediction module is used to process the image tiles that have undergone data augmentation and normalization using a pre-trained convolutional neural network using the ResNet50 method, and calculate the probability of the HRD status for each tile; the pathological whole-slide scan slice-level HRD status prediction module is used to receive the HRD status probability at the tile level as input, and use the tile likelihood histogram and bag-of-words model strategy to integrate this information, generate a feature vector at the pathological whole-slide scan slice level, and use the LightGBM method to establish a machine learning classifier model, and use the feature vector generated based on the PLH and BoW methods to predict the HRD status of the entire pathological whole-slide scan slice; the pathological result display module is used to receive the HRD status prediction at the pathological whole-slide scan slice level and present it to the user in an intuitive manner in the form of a Grad-CAM map and a pathological report; Among them, the imaging prediction sub-module includes an imaging normalization module, a feature extraction module, an HRD risk prediction module, and an imaging result display module; the imaging normalization module is used to first sort the pixel values in the image through an intensity truncation and normalization sub-module, and truncate the intensity values so that they fall within the range of the 0.5th to 99.5th percentiles, and then through a spatial normalization sub-module, resample the image to a resolution of 1mm x 1mm x 1mm; the feature extraction module is used to use the PRAD-MRI-OMICS tool to extract the manual features required for systematic prediction of HRD, including geometric features, intensity features, and texture features, from the segmented and quality-controlled lesion area; the HRD risk prediction module is used to establish a machine learning model using the support vector machine method, predict the HRD risk according to the features selected in the system, and use the trained model to generate an HRD risk score for each sample; the imaging result display module is used to receive the HRD status prediction on the imaging omics and output the HRD risk score in the form of a percentage; Among them, the multi-modal fusion module is used to integrate the outputs of the clinical result display module, the pathological result display module, and the imaging result display module using the pre-fusion, mid-fusion, and post-fusion strategies, and use multi-modal fusion technology to obtain a comprehensive evaluation result of the HRD status of prostate cancer; Among them, the multi-modal output module is used to summarize all results, output the overall predicted HRD risk in the form of a percentage, display the comprehensive evaluation results, as well as the predicted results of clinical, pathological, and imaging simultaneously, and give reference diagnosis and treatment opinions.

2. The system for predicting prostate cancer patients with homologous recombination deficiency based on magnetic resonance images and pathological whole-slide scans according to claim 1, characterized in that it further includes a high-throughput digital information memory and a high-speed high-frequency data processor. The high-throughput digital information memory is used to store the input clinical information, pathological images, magnetic resonance images, and their predicted results obtained by the multi-modal prediction module, as well as a computer program storing the steps of the method for predicting the HRD status of prostate cancer. The high-speed high-frequency data processor is used to implement the steps of the method for predicting the HRD status of prostate cancer when executing the computer management program stored in the high-throughput digital information memory.

3. The system for predicting prostate cancer patients with homologous recombination deficiency based on magnetic resonance images and pathological whole-slide scans according to claim 2, characterized in that it further includes a model evaluation module, a model reconstruction module, and a model comparison and adaptive update module; the model evaluation module is used to receive the HRD evaluation results from the multi-modal prediction module, verify the predicted results of the model using known gold standard data, and then calculate various performance indicators to evaluate the trained neural network model; the model reconstruction module is used to collect and integrate the multi-modal data obtained from the storage module, and perform complete reconstruction, partial reconstruction, or manual reconstruction of the neural network to adjust the model structure and parameters, so as to identify and select the model with the optimal performance as the updated HRD evaluation model; the model comparison and adaptive update module is used to test the new model using the validation set, and compare the performance indicators with those of the original model to evaluate the advantages and disadvantages of the two models. If the performance of the new model does not meet the predetermined standard or is inferior to the original model, it is possible to choose to retrain and optimize the model, or choose to retain the original model.

4. The system for predicting prostate cancer patients with homologous recombination deficiency based on magnetic resonance images and pathological whole-slide scans according to claim 1, characterized in that the pre-fusion strategy is a canonical correlation analysis strategy, the mid-fusion strategy is a multi-kernel learning strategy, and the post-fusion strategy is a stacking generalization strategy.

Citation Information

Patent Citations

  • Therapeutic effect prediction method based on multi-modal fusion model and terminal equipment

    CN115036002A

  • MRI image multi-omics evaluation method for breast cancer homologous recombination repair defect

    CN116030261A