Prediction method and system for breast cancer immunohistochemical index and typing
By combining a deep learning model with breast ultrasound and magnetic resonance images, the problems of invasiveness, long time, sampling limitations and resource imbalance in traditional breast cancer diagnosis have been solved, and non-invasive, rapid and accurate breast cancer immunohistochemistry indicators and typing predictions have been achieved.
Patent Information
- Application Number
- CN202510980902.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-12
AI Technical Summary
Traditional breast cancer diagnostic methods have problems such as invasive examination risks, long diagnosis time, sampling limitations, strong subjectivity, and uneven medical resources. In addition, the prediction accuracy of a single imaging modality is limited, making it difficult to fully characterize the biological behavior of the tumor.
A deep learning model is combined with breast ultrasound and magnetic resonance images to extract and fuse multimodal imaging features, build a machine learning model, and predict breast cancer immunohistochemical indicators and typing.
It provides non-invasive, rapid and accurate breast cancer immunohistochemistry indicators and typing predictions, improves diagnostic efficiency and accuracy, reduces dependence on medical resources, and enhances the consistency of diagnostic results.
Smart Images

Figure CN120636779A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image analysis and breast cancer diagnosis, and in particular to a method and system for predicting breast cancer immunohistochemical indicators and typing. Background Art
[0002] Traditional preoperative diagnosis of breast cancer mainly relies on histopathological examination, which requires obtaining tissue samples through puncture biopsy for immunohistochemical staining analysis to determine the expression levels of key biomarkers such as ER, PR, Her2, and Ki-67, and then perform molecular typing to guide the selection of treatment options.
[0003] However, traditional breast cancer diagnosis methods have many limitations. Specifically, traditional methods face the following five problems.
[0004] 1. Risks of invasive examinations and patient acceptance issues.
[0005] Traditional breast cancer immunohistochemistry testing requires obtaining tissue samples through a puncture biopsy. This invasive examination method has multiple risks and limitations. First, the puncture biopsy itself has certain medical risks, including complications such as bleeding, hematoma formation, infection, and vascular damage. Although the incidence is relatively low (about 1-3%), these risks may be significantly increased for patients with abnormal coagulation function, those taking anticoagulants, or those with severe cardiopulmonary diseases. Secondly, the puncture process will cause obvious pain and psychological pressure to the patient, especially for patients who are older, sensitive to pain, or have weak psychological tolerance, which may lead to poor compliance with the examination and affect the smooth implementation of the diagnosis and treatment plan. Thirdly, the invasive nature of the puncture biopsy limits its feasibility of repeated examinations and cannot meet the clinical needs of dynamically monitoring changes in tumor biological characteristics. Especially in the process of neoadjuvant therapy, it is difficult to evaluate the treatment response and adjust the treatment plan in real time.
[0006] 2. Long diagnosis period and low efficiency.
[0007] The traditional pathological diagnosis process is complicated and time-consuming, which seriously affects the efficiency of breast cancer diagnosis and treatment and the psychological state of patients. From the puncture biopsy to the final acquisition of a complete immunohistochemistry test report, it usually takes 5-10 working days, or even longer. The specific process includes: tissue fixation (24-48 hours), dehydration and embedding (6-8 hours), slice preparation (2-3 hours), HE staining and preliminary pathological diagnosis (1-2 days), immunohistochemistry staining (2-3 days), result interpretation and report writing (1-2 days). This long waiting process puts tremendous psychological pressure on patients. Emotional problems such as anxiety, fear, and insomnia are common, which not only affects the quality of life of patients, but may also have a negative impact on the immune system.
[0008] 3. Sampling limitations and insufficient assessment of tumor heterogeneity.
[0009] A puncture biopsy can only obtain local tissue samples of the tumor. Usually, the tissue strip obtained by a 14G needle core biopsy is only 1-2 cm long and about 2 mm in diameter. This small sample size is negligible relative to the entire tumor volume. Because breast cancer has significant intra-tumor heterogeneity and inter-tumor heterogeneity, cancer cells in different regions may exhibit different molecular characteristics and biological behaviors. Local sampling may not fully reflect the true biological characteristics of the entire tumor, and there is a risk of sampling bias. Studies have shown that the expression of ER, PR, and Her2 may vary in different regions of the same tumor, and the Ki-67 proliferation index may also show spatial heterogeneity. This sampling limitation may lead to inaccurate immunohistochemistry test results, which in turn affects the accuracy of molecular typing and the choice of treatment options.
[0010] 4. Problems of strong subjectivity and significant differences between observers.
[0011] Traditional immunohistochemistry interpretation relies primarily on the pathologist's subjective judgment and microscopic morphological observation, a method subject to significant subjectivity and variability. Pathologists vary in their experience, professional background, and understanding of interpretation criteria, potentially leading to differing diagnostic results for the same specimen. In particular, assessment of the Ki-67 proliferation index requires manual calculation of the percentage of positive nuclei, a process that relies heavily on subjective judgment. Even experienced pathologists can experience 10-20% variability in repeated counts. Interpretation of HER2 status presents similar challenges, particularly in borderline 2+ cases, where interobserver agreement is relatively poor. Counting the percentage of ER and PR-positive cells also exhibits interobserver and intraobserver variability. This subjectivity not only compromises diagnostic accuracy and consistency but can also lead to inappropriate treatment options for patients.
[0012] 5. Uneven distribution of medical resources and low level of technical standardization.
[0013] Accurate diagnosis of breast cancer requires advanced pathology technology platforms, experienced pathologists, and standardized laboratory operating procedures, but the distribution of these resources is extremely uneven among medical institutions in different regions and at different levels. In economically developed areas, tertiary hospitals are usually equipped with complete pathology equipment, standardized immunohistochemistry platforms, and experienced professional teams to provide high-quality diagnostic services. However, in primary medical institutions, county hospitals, or hospitals in remote areas, the necessary equipment and technical personnel may be lacking, and standardized immunohistochemistry tests cannot be carried out. Patients often need to be referred to higher-level hospitals, increasing the cost and time of diagnosis and treatment. Even in different hospitals in the same city, differences in equipment models, reagent selection, operating procedures, etc. may lead to inconsistent test results.
[0014] Currently, some studies have explored methods for predicting breast cancer biomarkers based on a single imaging modality. For example, some studies have used texture features from breast ultrasound images to predict Her2 status, and dynamic enhancement features from magnetic resonance imaging to predict ER status. However, most of these studies are limited to a single imaging modality, extracting relatively limited feature information, and the prediction accuracy needs to be improved. Single-modality imaging often only reflects a single aspect of tumor characteristics and cannot fully capture the complex biological behavior of tumors. Summary of the Invention
[0015] In view of the above problems, the present invention is proposed to provide a method and system for predicting breast cancer immunohistochemical indicators and typing, which overcomes the above problems or at least partially solves the above problems.
[0016] According to one aspect of the present invention, a method for predicting breast cancer immunohistochemical indicators and typing is provided, the method comprising:
[0017] Obtain breast ultrasound images and breast magnetic resonance images of patients;
[0018] performing preprocessing and quality control on the breast ultrasound image and the breast magnetic resonance image;
[0019] Using a deep learning model to perform breast lesion region segmentation on the preprocessed breast ultrasound image and the breast magnetic resonance image to obtain a segmented lesion region;
[0020] extracting multimodal radiomics features of the breast ultrasound image and the breast magnetic resonance image based on the segmented lesion area;
[0021] fusing the multimodal imaging omics features of the breast ultrasound image and the breast magnetic resonance image, constructing a machine learning model, performing immunohistochemical index prediction, and obtaining immunohistochemical index prediction results;
[0022] Breast cancer molecular typing analysis is performed based on the prediction results of the immunohistochemical indicators.
[0023] Optionally, the preprocessing and quality control of the breast ultrasound image specifically includes:
[0024] Convert raw ultrasound images into DICOM standard format;
[0025] Image dimensions were normalized to uniform pixel spacing and image size;
[0026] Grayscale values are normalized to the range of [0,255];
[0027] Use bilateral filter for speckle noise suppression;
[0028] Apply histogram equalization to enhance image contrast;
[0029] Establish quality control standards and eliminate images with blurred images, severe artifacts, and unclear lesion visualization.
[0030] Optionally, the preprocessing and quality control of the breast magnetic resonance image specifically includes: using motion artifact correction;
[0031] N4 offset field correction;
[0032] Multi-sequence image registration to maintain spatial consistency among T1-weighted imaging (T1WI), T2-weighted imaging (T2WI), and diffusion-weighted imaging (DWI) sequences;
[0033] Normalize signal strength;
[0034] The image is resampled to a uniform voxel size;
[0035] Magnetic resonance image quality assessment was performed.
[0036] Optionally, the step of performing breast lesion region segmentation on the pre-processed breast ultrasound image using a deep learning model specifically includes: constructing a segmentation network based on an improved U-Net, comprising 4 downsampling layers and 4 upsampling layers;
[0037] Use a weighted combination of Dice loss function and Focal loss function for training;
[0038] Introducing attention gating mechanism for lesion detection;
[0039] Output the binary segmentation mask of the lesion area;
[0040] Morphological post-processing is performed on the segmentation results, including hole filling and boundary smoothing.
[0041] Optionally, the step of performing breast lesion region segmentation on the pre-processed breast magnetic resonance image using a deep learning model specifically includes:
[0042] Use the improved U-Net segmentation network to perform joint segmentation of multiple sequence magnetic resonance images;
[0043] Fusion of complementary information from T1-weighted imaging (T1WI), T2-weighted imaging (T2WI), diffusion-weighted imaging (DWI), and dynamic contrast-enhanced sequences;
[0044] Use deep supervision mechanism for segmentation;
[0045] Use conditional random fields to perform post-processing optimization of segmentation results;
[0046] Generate 3D lesion segmentation masks and volume measurements.
[0047] Optionally, the extracting of multimodal radiomics features of the breast ultrasound image based on the segmented lesion area specifically includes:
[0048] Morphological feature extraction to calculate the area, perimeter, roundness, ellipticity, and irregularity index of the lesion;
[0049] First-order statistical feature extraction, including mean, variance, skewness, kurtosis and entropy;
[0050] Second-order texture feature extraction, calculating contrast, correlation, energy, and homogeneity based on the gray-level co-occurrence matrix; high-order feature extraction, including wavelet transform features and local binary pattern features;
[0051] Echo characteristic analysis to evaluate the internal echo pattern and posterior acoustic shadow characteristics of the lesion.
[0052] Optionally, the extracting of multimodal radiomics features of the breast magnetic resonance image based on the segmented lesion area specifically includes:
[0053] Morphological feature extraction to calculate the volume, surface area, sphericity, and compactness of the lesion;
[0054] Signal intensity feature extraction, analysis of signal features of T1-weighted imaging (T1WI), T2-weighted imaging (T2WI), and diffusion-weighted imaging (DWI) sequences;
[0055] Texture feature extraction, based on gray-level co-occurrence matrix, gray-level run-length matrix and gray-level size area matrix;
[0056] Dynamic feature extraction, analyzing the time-signal intensity curve characteristics of dynamic contrast-enhanced sequences;
[0057] Diffusion feature extraction, calculation of statistical parameters and histogram features of apparent diffusion coefficient.
[0058] Optionally, fusing the multimodal imaging omics features of the breast ultrasound image and the breast magnetic resonance image to construct a machine learning model and perform immunohistochemical index prediction specifically includes:
[0059] fusing multimodal imaging omics features of the breast ultrasound image and the breast magnetic resonance image to obtain a fusion feature;
[0060] Performing dimensionality reduction processing on the fused features, and performing feature selection using principal component analysis and least absolute shrinkage selection operators;
[0061] Build multi-classification machine learning models, including support vector machines, random forests, and gradient boosted trees;
[0062] Use cross-validation and leave-one-out methods to evaluate model performance;
[0063] Establish immunohistochemical index prediction models to predict the expression status of ER, PR, Her2, and Ki-67 respectively; construct a breast cancer molecular typing decision tree model based on the immunohistochemical prediction results;
[0064] Provides predicted probabilities and confidence intervals for clinical decision making.
[0065] Optionally, the prediction method further comprises: visualizing the immunohistochemical index prediction result to generate a probability distribution graph of the immunohistochemical index;
[0066] Draw classification probability maps and decision boundaries for molecular classification of breast cancer;
[0067] Generate a three-dimensional visualization display of the lesion segmentation results;
[0068] Get feature importance analysis and model interpretability reports.
[0069] The present invention also provides a prediction system for breast cancer immunohistochemical indicators and typing, which uses the above-mentioned prediction method for breast cancer immunohistochemical indicators and typing. The prediction system specifically includes: a data acquisition module for acquiring a patient's breast ultrasound image and breast magnetic resonance image; an image preprocessing module for preprocessing and quality control of the breast ultrasound image and the breast magnetic resonance image; a deep learning segmentation module for using a deep learning model to segment the preprocessed breast ultrasound image and the breast magnetic resonance image into breast lesion areas to obtain segmented lesion areas; a feature extraction module for extracting multimodal imaging genomics features of the breast ultrasound image and the breast magnetic resonance image based on the segmented lesion areas; a machine learning module for fusing the multimodal imaging genomics features of the breast ultrasound image and the breast magnetic resonance image, constructing a machine learning model, performing immunohistochemical indicator prediction, and obtaining immunohistochemical indicator prediction results; and a result analysis module for performing breast cancer molecular typing analysis based on the immunohistochemical indicator prediction results.
[0070] Optionally, the prediction system further includes: a visualization display module for visualizing the prediction results of the immunohistochemistry indicators.
[0071] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for predicting breast cancer immunohistochemistry and typing. The present invention provides a method and system for predicting breast cancer immunohistochemistry indicators and typing, the prediction method comprising: obtaining a patient's breast ultrasound image and breast magnetic resonance image; preprocessing and quality control of the breast ultrasound image and the breast magnetic resonance image; using a deep learning model to segment the preprocessed breast ultrasound image and the breast magnetic resonance image to obtain a segmented lesion area; extracting multimodal imaging genomics features of the breast ultrasound image and the breast magnetic resonance image based on the segmented lesion area; fusing the multimodal imaging genomics features of the breast ultrasound image and the breast magnetic resonance image to construct a machine learning model, perform immunohistochemistry indicator prediction, and obtain an immunohistochemistry indicator prediction result; and perform breast cancer molecular typing analysis based on the immunohistochemistry indicator prediction result. By fusing multimodal imaging information, tumor characteristics can be more comprehensively portrayed, and the accuracy of biomarker prediction can be improved.
[0072] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0074] Figure 1 A flowchart of a method for predicting breast cancer immunohistochemical indicators and typing provided by an embodiment of the present invention;
[0075] Figure 2 This is a detailed flow chart for implementing a method for predicting breast cancer immunohistochemical indicators and typing provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0076] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0077] The terms "including" and "having," as well as any variations thereof, in the description, embodiments, claims, and drawings of the present invention are intended to cover non-exclusive inclusions, such as, for example, inclusion of a series of steps or units. The technical solutions of the present invention are further described in detail below with reference to the drawings and embodiments.
[0078] Example 1
[0079] The purpose of this invention is to provide a method and system for predicting breast cancer immunohistochemical markers and molecular typing based on multimodal ultrasound and magnetic resonance imaging. By integrating the characteristic information of breast ultrasound and magnetic resonance imaging and applying advanced machine learning algorithms, this method accurately predicts immunohistochemical markers and molecular typing in breast cancer patients, providing clinicians with noninvasive, rapid, and accurate diagnostic support. This system, combined with the development of a complete clinical application system, improves analytical efficiency and clinical practicality.
[0080] like Figure 1 As shown, a method for predicting breast cancer immunohistochemical indicators and typing includes:
[0081] Obtain breast ultrasound images and breast magnetic resonance images of patients;
[0082] performing preprocessing and quality control on the breast ultrasound image and the breast magnetic resonance image;
[0083] Using a deep learning model to perform breast lesion region segmentation on the preprocessed breast ultrasound image and the breast magnetic resonance image to obtain a segmented lesion region;
[0084] extracting multimodal radiomics features of the breast ultrasound image and the breast magnetic resonance image based on the segmented lesion area;
[0085] fusing the multimodal imaging omics features of the breast ultrasound image and the breast magnetic resonance image, constructing a machine learning model, performing immunohistochemical index prediction, and obtaining immunohistochemical index prediction results;
[0086] Breast cancer molecular typing analysis is performed based on the prediction results of the immunohistochemical indicators.
[0087] The technical solutions of the present invention are as follows:
[0088] The present invention provides a breast cancer immunohistochemical index and molecular typing prediction method based on ultrasound and magnetic resonance multimodal imaging. The method implements a complete breast cancer immunohistochemical index and molecular typing prediction analysis process through 10 core steps:
[0089] Breast ultrasound image acquisition. A high-frequency linear array probe (7-15 MHz) was used to perform breast ultrasound examinations, acquiring two-dimensional grayscale ultrasound images encompassing the lesion area. Image acquisition was performed in strict accordance with the Breast Imaging Reporting and Data System (BI-RADS) standards to ensure image quality. Acquisition parameters, including gain, time gain compensation, dynamic range, and focus position, were standardized. Images were acquired from at least two orthogonal planes for each lesion, including the maximum radial image and the perpendicular radial image.
[0090] Breast MRI images were acquired using a 1.5T or 3.0T MRI machine with a dedicated breast coil. Scanning sequences included T1-weighted imaging (T1WI), T2-weighted imaging (T2WI), T1-weighted dynamic contrast-enhanced imaging (DCE-MRI), and diffusion-weighted imaging (DWI). Scanning parameters were standardized: slice thickness ≤ 3 mm, slice spacing ≤ 1 mm, and matrix ≥ 256 × 256. The contrast agent used was gadopentetate dimeglumine at a dose of 0.1 mmol / kg body weight and an injection rate of 2–3 ml / s.
[0091] Ultrasound image preprocessing and quality control. The acquired ultrasound images were subjected to standardized preprocessing, including: (1) image format conversion and standardization; (2) noise filtering, using a combination of Gaussian filtering and median filtering; (3) contrast enhancement, using the contrast-limited adaptive histogram equalization (CLAHE) algorithm; (4) image size normalization, uniformly adjusted to 512 × 512 pixels; (5) quality assessment, establishing an image quality scoring standard, and excluding images with unqualified quality.
[0092] MRI image preprocessing and quality control. MRI images were systematically preprocessed: (1) motion artifact correction using a rigid body registration algorithm; (2) bias field correction to eliminate the effects of magnetic field inhomogeneity; (3) intensity normalization using the N4 bias field correction algorithm; (4) image registration to register different sequence images to T1WI space; (5) resampling to a uniform spatial resolution of 1 × 1 × 1 mm. 3 ; (6) Quality control, establish a multi-sequence image quality assessment system.
[0093] Breast ultrasound image lesion segmentation. A deep learning approach is used to achieve automatic lesion segmentation: (1) Construct a segmentation network based on the U-Net architecture and integrate an attention mechanism to improve segmentation accuracy; (2) Use a large amount of labeled data for model training to establish a robust segmentation model; (3) Achieve pixel-level accurate segmentation of the lesion area; (4) Manual verification and correction are performed to ensure the accuracy of the segmentation results; (5) Extract the geometric features of the lesion, including area, perimeter, major diameter, minor diameter, shape index, etc.
[0094] Breast MRI lesion segmentation. A specialized segmentation algorithm was developed based on the characteristics of multi-sequence MRI images: (1) Preliminary lesion localization based on DCE-MRI early enhancement features; (2) Build a 3D segmentation model by combining multi-sequence information; (3) Use deep learning methods to achieve accurate 3D segmentation; (4) Consider tumor heterogeneity to achieve tumor subregion segmentation; (5) Generate a 3D model of the lesion and calculate 3D geometric features such as volume and surface area.
[0095] Breast ultrasound image feature extraction. Multi-dimensional features are extracted from the segmented lesion area: (1) Morphological features: including geometric parameters such as area, perimeter, major diameter, minor diameter, roundness, ellipticity, convexity, and compactness; (2) Texture features: extracting first-order statistical features and second-order statistical features based on methods such as gray-level co-occurrence matrix (GLCM), gray-level range matrix (GLRLM), and gray-level zone size matrix (GLSZM); (3) Edge features: analyzing morphological features such as smoothness, burr sign, and lobulation sign of the lesion boundary; (4) Echo features: extracting acoustic features such as echo pattern and echo intensity distribution inside the lesion; (5) Blood flow features: analyzing the blood flow distribution pattern of the lesion in combination with color Doppler information.
[0096] Breast MRI feature extraction. Fully exploit the information of multiple MRI sequences: (1) Morphological features: Extract three-dimensional geometric features such as volume, surface area, sphericity, and fractal dimension based on three-dimensional segmentation results; (2) Signal intensity features: Analyze the signal intensity distribution characteristics on T1WI and T2WI; (3) Dynamic enhancement features: Extract the time-signal intensity curve characteristics of DCE-MRI, including parameters such as peak enhancement, time to peak, and washout rate; (4) Diffusion features: Calculate the apparent diffusion coefficient (ADC) value and its distribution characteristics based on DWI sequences; (5) Texture features: Extract texture features based on statistics, models, and transformations on multiple sequences; (6) Perfusion features: Analyze the vascular permeability and hemodynamic characteristics of the tumor.
[0097] Statistical analysis and modeling of breast cancer classification. Construction of a multi-level prediction model system: (1) Feature selection: univariate analysis, multivariate analysis, LASSO regression and other methods were used to screen the optimal feature subset; (2) Feature fusion: Multimodal feature fusion strategies were designed, including early fusion, late fusion and hybrid fusion methods; (3) Immunohistochemical index prediction: Binary or multi-classification models were constructed for ER, PR, Her2 and Ki-67 status respectively; (4) Molecular classification prediction: Based on the results of immunohistochemical indicators, molecular subtypes such as Luminal A, Luminal B, Her2 positive and triple negative were further predicted; (5) Model optimization: Hyperparameter tuning was performed using grid search, Bayesian optimization and other methods; (6) Cross-validation: K-fold cross-validation was used to evaluate the generalization performance of the model.
[0098] System clinical validation, explanation, and optimization. Establish a complete clinical validation and continuous optimization mechanism: (1) Clinical validation study: Design prospective and retrospective study plans to verify system performance in a multicenter environment; (2) Performance evaluation: Use accuracy, sensitivity, specificity, AUC, F1-score and other indicators to comprehensively evaluate model performance; (3) Interpretability analysis: Use methods such as SHAP and LIME to analyze feature importance and provide interpretability of model decisions; (4) Clinical applicability evaluation: Evaluate the stability of the system in different devices, different operators, and different patient groups; (5) Continuous learning: Establish an online learning mechanism to continuously optimize model performance based on new data; (6) Quality management: Establish a complete quality control system to ensure the reliability of system output; (7) User training: Develop standardized operating procedures and training programs; (8) Effect evaluation: Evaluate the system's improvement effect on clinical diagnosis and treatment efficiency and accuracy.
[0099] The present invention also provides a breast cancer immunohistochemistry index and typing prediction system, which includes seven core components: data acquisition module, image preprocessing module, deep learning segmentation module, feature extraction module, machine learning module, result output module and quality control module.
[0100] The data acquisition module interfaces with the hospital's existing ultrasound and MRI equipment, supporting the acquisition of image data in the DICOM standard format. Compatible with devices from multiple vendors, the module automatically identifies and analyzes image parameters from different devices, ensuring standardized and consistent data collection. Furthermore, the module includes built-in patient information management capabilities, automatically linking image data with basic patient information to create a complete data archive.
[0101] The image preprocessing module uses advanced image processing algorithms to standardize raw image data. For ultrasound images, this module focuses on addressing noise suppression, contrast enhancement, and size normalization. For magnetic resonance images, it primarily addresses technical challenges such as motion artifact correction, bias field correction, and multi-sequence registration. This module utilizes GPU-accelerated computing to significantly improve image processing efficiency, reducing the processing time for a single image to under 5 seconds.
[0102] Lesion segmentation module. Based on a deep convolutional neural network architecture, this system automatically and accurately segments lesions. Using an improved U-Net network, the system integrates an attention mechanism and multi-scale feature extraction technology, achieving segmentation accuracy exceeding 95%. To ensure segmentation quality, the system provides a manual interface that allows physicians to quickly verify and correct segmentation results.
[0103] The feature extraction module is the core technical component of the system, capable of extracting high-dimensional, multi-layered imaging features from segmented lesion areas. Ultrasound feature extraction focuses on features such as morphology, texture, edges, and echogenicity; magnetic resonance feature extraction covers multiple dimensions, including morphology, signal, dynamics, diffusion, and texture. The system extracts over 2,500+ dimensional feature parameters, providing a rich data foundation for subsequent machine learning modeling. The machine learning module integrates a variety of advanced machine learning algorithms, including support vector machines, random forests, gradient boosting trees, and deep neural networks. The system employs an ensemble learning strategy, combining multiple base learners to improve prediction accuracy and stability. Dedicated prediction models are constructed for different prediction tasks (ER, PR, Her2, and Ki-67 status prediction and molecular typing) to ensure optimal predictive performance. The result output module provides an intuitive and user-friendly result display interface, presenting prediction results in various formats, including charts, numerical values, and text. The system output includes the predicted probability, confidence interval, molecular typing results, and corresponding clinical recommendations for each immunohistochemical indicator. At the same time, the system provides detailed analysis reports, including key feature analysis, explanation of prediction basis, etc., to help doctors understand and apply prediction results.
[0104] The quality control module monitors data quality, algorithm performance, and result reliability throughout the entire system operation. This module establishes a multi-level quality assessment system, including image quality scoring, segmentation accuracy verification, feature validity verification, and prediction consistency checks, to ensure the reliability and clinical applicability of system outputs.
[0105] Example 2
[0106] like Figure 2 As shown, the present invention provides a breast cancer immunohistochemical index and typing prediction method based on ultrasound and magnetic resonance multimodal imaging, the specific steps of which include:
[0107] Step 101: Acquiring breast ultrasound images
[0108] Breast ultrasound examinations were performed using a high-frequency linear array probe with the following technical parameters: probe frequency 7-15 MHz, penetration depth 4-6 cm, near-field resolution ≤ 0.3 mm, and far-field resolution ≤ 0.8 mm. The examinations were conducted in strict accordance with the Breast Imaging Reporting and Data System (BI-RADS) version 5 standards. Technical implementation details:
[0109] (1) Patient position: supine with the affected limb raised to fully expose the breast and axilla;
[0110] (2) Probe placement: Use gentle contact to avoid excessive pressure that may affect blood flow display;
[0111] (3) Scanning method: A combination of radial and concentric scanning is used to ensure full breast coverage;
[0112] (4) Image acquisition: At least six standard sections were obtained for each lesion, including the transverse and longitudinal sections of the maximum diameter, the transverse and longitudinal sections of the perpendicular diameter, the section showing the relationship between the lesion and the chest wall, and the section showing the relationship between the lesion and the nipple;
[0113] (5) Parameter setting: Gain is adjusted until the noise just disappears, the time gain compensation curve is straight or slightly tilted to the right, the dynamic range is set to 60-80dB, and the focus is set at the center of the lesion or slightly deeper;
[0114] (6) Color Doppler parameters: PRF was set to 1000-1500 Hz, wall filter was set to 50-100 Hz, and color gain was adjusted to a level where noise just appeared and then slightly decreased;
[0115] (7) Image storage: stored in DICOM format, pixel matrix ≥512×512, bit depth ≥8 bits;
[0116] (8) Quality control: Each image must be labeled with key information such as lesion location, size, probe frequency, etc. Step 102 Breast MRI image acquisition
[0117] Scans are performed using a 1.5T or 3.0T MRI machine equipped with a dedicated 8-channel or 16-channel breast phased array coil. The patient lies prone, with the breasts naturally suspended within the coil aperture to avoid compression and deformation.
[0118] Detailed scanning plan:
[0119] (1) Positioning image: three-plane positioning to ensure that the scanning range covers both breasts and chest wall;
[0120] (2) T2-weighted imaging (T2WI): fast spin echo sequence, TR / TE = 4000-6000 ms / 80-120 ms, slice thickness 3 mm, slice spacing 0.3 mm, matrix 256 × 256 or 512 × 512, field of view (FOV) 320-380 mm, 2-3 excitations, fat suppression using STIR or frequency-selective fat suppression technology;
[0121] (3) Diffusion-weighted imaging (DWI): A single-shot echo planar imaging sequence was used, with TR / TE = 6000–8000 ms / shortest, slice thickness 4 mm, slice spacing 0.4 mm, matrix 128 × 128, and b values set to 0, 50, 800, and 1000 s / mm 2 , diffusion-sensitive gradient directions ≥ 3, fat suppression using STIR technology;
[0122] (4) T1-weighted dynamic contrast-enhanced imaging (DCE-MRI): using a three-dimensional gradient echo sequence, TR / TE = 4-6 ms / 1.8-2.4 ms, flip angle 10-15°, slice thickness 1.2-2 mm, matrix 256 × 256 or 512 × 512, temporal resolution 60-90 seconds, total scanning time 8-10 minutes, including 1 plain scan period and 6-8 enhanced phases;
[0123] (5) Contrast agent injection: Use gadopentetate dimeglumine (Gd-DTPA) or gadobutrol (Gadobutrol) at a dose of 0.1 mmol / kg body weight, via cubital vein bolus injection at a rate of 2-3 ml / s. Immediately after the injection, inject 20 ml of normal saline at the same rate for flushing;
[0124] (6) Image post-processing: Automatically generate subtraction images, maximum intensity projection (MIP) images, and time-intensity curves (TIC), and calculate pharmacokinetic parameters including Ktrans, Kep, and Ve.
[0125] Step 103 of ultrasound image preprocessing and quality control includes: establishing a standardized ultrasound image preprocessing process to ensure consistency and comparability of image quality.
[0126] Specific preprocessing steps:
[0127] (1) Image format standardization: convert ultrasound images generated by various manufacturers' equipment into the standard DICOM format and extract image metadata including equipment model, probe parameters, scanning parameters, etc.
[0128] (2) Image cropping and region of interest (ROI) extraction: Automatically identify and crop non-anatomical structures such as text annotations, measurement scales, and color bars in the image to extract pure breast tissue image areas;
[0129] (3) Noise filtering: An improved non-local mean filtering algorithm combined with Gaussian filtering is used, with a filter kernel size of 3×3 and a standard deviation of σ = 1.0, to effectively remove speckle noise and electronic noise;
[0130] (4) Contrast enhancement: The contrast-limited adaptive histogram equalization (CLAHE) algorithm is used with clipLimit set to 2.0 and tileGridSize set to 8×8 to improve image contrast and detail visibility;
[0131] (5) Gamma correction: Gamma correction is performed according to the characteristics of different devices, with a correction coefficient of γ = 0.8-1.2 to unify the image brightness characteristics;
[0132] (6) Image size normalization: bilinear interpolation method is used to uniformly adjust the size of all images to 512 × 512 pixels, maintaining the aspect ratio;
[0133] (7) Intensity normalization: Normalize the pixel intensity value to the interval [0, 1], the formula is I_norm = (I-I_min) / (I_max-I_min);
[0134] (8) Quality assessment indicators: establish image quality scoring standards, including signal-to-noise ratio (SNR ≥ 25dB), contrast-to-noise ratio (CNR ≥ 3), image clarity index, artifact degree score, etc., and automatically exclude images with unqualified quality;
[0135] (9) Data enhancement (optional): For training data, random rotation (±15°), scaling (0.9-1.1 times), and flipping methods are used for data enhancement to increase model robustness.
[0136] Step 104: Magnetic resonance image preprocessing and quality control.
[0137] According to the characteristics of multi-sequence magnetic resonance images, a systematic preprocessing process is established.
[0138] Detailed processing flow:
[0139] (1) DICOM parsing and sequence recognition: Automatically parse DICOM file header information, identify different sequence types (T1WI, T2WI, DWI, DCE-MRI), and extract scan parameters;
[0140] (2) Motion artifact correction: A registration-based motion correction algorithm is used to correct the registration error caused by the patient's breathing and body movement using a rigid body transformation model, with a registration accuracy of ≤1mm;
[0141] (3) Bias field correction: The N4ITK bias field correction algorithm was used to eliminate RF field inhomogeneity, with 50 iterations and a convergence threshold of 0.001.
[0142] (4) Meningeal suppression and background removal: Morphological opening and connected domain analysis are used to remove the image background and retain the breast tissue area;
[0143] (5) Inter-sequence registration: Using the T1WI plain scan image as a reference, other sequence images are registered to the same space, using mutual information as the similarity measure, and the registration algorithm uses Powell optimization;
[0144] (6) Resampling and interpolation: resample all sequence images to a uniform spatial resolution of 1×1×1mm 3 , using B-spline interpolation method;
[0145] (7) Intensity normalization: For T1WI and T2WI images, the Z-score normalization method is used, with the formula I_norm = (I-μ) / σ, where μ is the mean and σ is the standard deviation; for DWI images, the apparent diffusion coefficient (ADC) map is calculated, with the formula ADC = -ln(S / S0) / b, where S is the diffusion-weighted signal and S0 is the signal when b = 0;
[0146] (8) Dynamic enhancement post-processing: Calculate the signal enhancement ratio using the formula ER = (S_post - S_pre) / S_pre × 100% to generate a pharmacokinetic parameter map;
[0147] (9) Quality control indicators: Evaluate image signal-to-noise ratio, motion artifact level, bias field correction effect, registration accuracy, etc., and establish a multi-dimensional quality scoring system;
[0148] (10) Outlier detection: The 3σ criterion is used to detect and process abnormal pixel values to prevent extreme values from affecting subsequent analysis.
[0149] Step 105: segmenting breast ultrasound image lesions.
[0150] Develop an intelligent segmentation algorithm based on deep learning to achieve accurate and automatic segmentation of the lesion area.
[0151] Technical implementation plan:
[0152] (1) Network architecture design: An improved U-Net architecture is adopted. The encoder uses the ResNet-50 backbone network for feature extraction, and the decoder uses upsampling and skip connections to restore spatial resolution.
[0153] (2) Attention mechanism integration: Channel Attention and Spatial Attention are added at the jump connection. The weight calculation formula is Attention = σ(W2*ReLU(W1*X)), where σ is the Sigmoid function and W1 and W2 are learning parameters.
[0154] (3) Multi-scale feature fusion: The feature pyramid network (FPN) structure is used to fuse feature information of different scales to improve the segmentation ability of lesions of different sizes;
[0155] (4) Loss function design: A combination of Dice Loss and Binary Cross Entropy Loss is used with a weight ratio of 1:1. The formula is Loss = -log(Dice) + BCE.
[0156] (5) Data annotation: Two senior sonographers independently performed pixel-level annotation, and the inconsistent areas were determined through consultation to establish a gold standard dataset;
[0157] (6) Training strategy: Adam optimizer is used, with an initial learning rate of 0.001, a learning rate decay strategy of 0.1 times the original rate every 30 epochs, a batch size of 16, and a training epoch number of 100;
[0158] (7) Data enhancement: random rotation (-30° to 30°), random scaling (0.8 to 1.2 times), random flipping, random brightness adjustment (±20%), random contrast adjustment (±20%);
[0159] (8) Post-processing optimization: Morphological opening operation is used to remove small noise areas, the largest connected domain is retained as the lesion area, and the active contour model is used to further refine the boundary;
[0160] (9) Quality evaluation: Calculate the Dice coefficient, IoU, Hausdorff distance and other indicators to evaluate the segmentation quality. A Dice coefficient ≥ 0.85 is considered a successful segmentation.
[0161] (10) Human interaction interface: Provides a visual interface for doctors to verify and manually correct segmentation results, supporting brush tools, eraser tools, magic wand tools, etc.
[0162] Step 106: Breast MRI Lesion Segmentation
[0163] Based on the characteristics of three-dimensional magnetic resonance images, a special 3D segmentation algorithm is developed.
[0164] Specific implementation plan:
[0165] (1) Multi-sequence information fusion: T1WI, T2WI, DWI, and DCE-MRI sequences are used as multi-channel inputs, and the network architecture adopts 3D U-Net, which can simultaneously utilize spatial and temporal dimension information;
[0166] (2) Dynamic enhancement-guided segmentation: DCE-MRI early enhancement features are used to preliminarily locate the lesion, calculate the early enhancement rate threshold (>50%), and generate candidate regions;
[0167] (3) Network structure design: The encoder contains five downsampling blocks, each of which contains two 3×3×3 convolutional layers and one 2×2×2 maximum pooling layer. The decoder is symmetrically designed and uses transposed convolution for upsampling.
[0168] (4) Multi-scale loss function: Losses are calculated at different resolution levels, deep supervision improves training results, and the total loss is the weighted sum of losses at each layer;
[0169] (5) Region growing algorithm: Using high-confidence pixels as seed points, region growing is performed based on the similarity criterion, and the similarity threshold is determined by cross-validation;
[0170] (6) Graph cut algorithm optimization: The improved maximum flow-minimum cut algorithm is used for fine segmentation. The energy function includes data terms and smoothing terms. The weight parameters are optimized through grid search.
[0171] (7) Time series analysis: DCE-MRI time series were analyzed, time-intensity curve (TIC) features were calculated, and segmentation was assisted based on enhancement patterns;
[0172] (8) Post-processing strategy: Use 3D morphological operations to remove isolated small areas and retain a volume ≥ 0.5 cm 3 In the area, 3D active contour model is used to refine the boundary;
[0173] (9) Multi-observer validation: Three senior radiologists independently annotated the data, and the STAPLE algorithm was used to fuse the multi-observer results to generate a probabilistic segmentation map;
[0174] (10) Segmentation quality control: Calculate indicators such as 3D Dice coefficient, surface distance, and volume correlation coefficient, and establish an automatic quality assessment system.
[0175] Step 107: Breast Ultrasound Image Feature Extraction
[0176] Extract multi-dimensional and multi-level image features from the segmented lesion area.
[0177] Detailed feature extraction scheme:
[0178] (1) Geometric features (30 dimensions): including area, perimeter, major axis, minor axis, aspect ratio, circularity (4πArea / Perimeter) 2 ), Ellipticity, Convexity (Convexity = ConvexArea / Area), Compactness (Compactness = Area / ConvexArea), Shape Factor (Shape Factor = 4πArea / Perimeter 2 ), rectangularity, irregularity index, etc.
[0179] (2) First-order statistical features (20 dimensions): Calculate mean, standard deviation, variance, skewness, kurtosis, energy, entropy, minimum, maximum, range, uniformity, coefficient of variation, etc. based on pixel intensity distribution;
[0180] (3) Second-order statistical features (88 dimensions): Based on the gray-level co-occurrence matrix (GLCM) calculation, the distance is set to 1, 2, and 3 pixels, and the direction is set to 0°, 45°, 90°, and 135°. Features such as contrast, correlation, energy, entropy, homogeneity, variance, sum average, sum variance, sum entropy, difference variance, and difference entropy are extracted.
[0181] (4) High-order statistical features (32 dimensions): Short run emphasis (SRE), long run emphasis (LRE), grayscale nonuniformity (GLN), run length nonuniformity (RLN), run percentage (RP), etc. are calculated based on the grayscale run matrix (GLRLM); small zone emphasis (SZE), large zone emphasis (LZE), low grayscale zone emphasis (LGZE), high grayscale zone emphasis (HGZE), etc. are calculated based on the grayscale zone size matrix (GLSZM);
[0182] (5) Gradient features (16 dimensions): Calculate the first-order moment, second-order moment, gradient direction entropy, gradient intensity distribution, etc. of the image gradient;
[0183] (6) Frequency domain features (24 dimensions): Spectral features are extracted through Fourier transform and wavelet transform, including main frequency, spectral center of gravity, spectral width, wavelet energy distribution, etc.
[0184] (7) Boundary features (18 dimensions): Analyze the roughness, fractal dimension, boundary gradient intensity, boundary contrast, burr sign quantitative index, lobulation sign quantitative index, etc. of the lesion boundary;
[0185] (8) Texture directional features (12 dimensions): Calculate the main direction, directional strength, anisotropy coefficient, etc. of the texture;
[0186] (9) LBP feature (59 dimensions): Local binary pattern (LocalBinaryPattern) is used to describe local texture. The radius is set to 1, 2, and 3, the number of neighborhood points is set to 8 and 16, and the LBP histogram feature is calculated.
[0187] (10) Feature standardization: The Z-score standardization method is used to ensure that different features have the same numerical range.
[0188] Step 108 is to extract breast MRI image features to fully exploit the rich information of multi-sequence MRI images. Comprehensive feature extraction strategy:
[0189] (1) Three-dimensional geometric features (25 dimensions): Volume, surface area, sphericity (Sphericity = π^(1 / 3) × (6V)^(2 / 3) / SA), compressibility (Compactness = SA / (36πV 2 )^(1 / 3)), maximum 3D radial line, fractal dimension, convex hull volume ratio, etc.;
[0190] (2) T1WI signal features (32 dimensions): Extract the signal intensity statistical features of T1-weighted images, including mean, median, standard deviation, 25% quantile, 75% quantile, skewness, kurtosis, signal intensity ratio (lesion / normal breast tissue), etc.
[0191] (3) T2WI signal features (32 dimensions): Similar to T1WI, it extracts the signal intensity distribution features of T2-weighted images, paying special attention to the high and low changes of T2 signals;
[0192] (4) Dynamic enhancement features (45 dimensions): Time-signal intensity curve features are calculated based on the DCE-MRI sequence, including baseline signal intensity, peak enhancement ratio (Peak Enhancement Ratio = (S_peak - S_baseline) / S_baseline), time to peak, initial enhancement rate (Initial Enhancement Rate), delayed slope (Delayed Slope), washout rate (Washout Rate), area under the curve (AUC), enhancement heterogeneity index, etc.
[0193] (5) Pharmacokinetic characteristics (24 dimensions): calculation of statistical parameters such as the mean, standard deviation, skewness, kurtosis, 10% quantile, and 90% quantile of Ktrans (vascular permeability parameter), Kep (extravascular extracellular space reflux constant), Ve (extravascular extracellular space volume fraction), and Vp (plasma volume fraction);
[0194] (6) Diffusion characteristics (28 dimensions): Calculate the distribution characteristics of ADC (apparent diffusion coefficient) values based on DWI sequences, including ADC mean, standard deviation, minimum, maximum, skewness, kurtosis, ADC histogram parameters, diffusion heterogeneity index, etc.
[0195] (7) Multi-sequence texture features (180 dimensions): GLCM, GLRLM, and GLSZM texture features are extracted on T1WI, T2WI, and ADC images, respectively, with 60-dimensional texture features for each sequence;
[0196] (8) Morphological characteristics (36 dimensions): Analyze the three-dimensional morphological characteristics of the lesion, including the lengths of the three principal axes obtained by principal component analysis, surface roughness, concavity, boundary clarity, shape complexity, etc.
[0197] (9) Perfusion characteristics (20 dimensions): Calculation of perfusion parameters such as cerebral blood flow (CBF), cerebral blood volume (CBV), and mean transit time (MTT) based on DSC-PWI or ASL sequences (if available);
[0198] (10) Multi-parameter correlation features (15 dimensions): Calculate the correlation between different sequences, such as the correlation coefficient between T1WI and ADC, the correlation between DCE and DWI, etc.
[0199] (11) Time series features (30 dimensions): Fourier analysis is performed on the DCE-MRI time series to extract frequency domain features such as main frequency, phase, and amplitude;
[0200] (12) Spatial heterogeneity characteristics (22 dimensions): Analyze the spatial heterogeneity within the lesion, including the spatial variation coefficient of signal intensity, texture spatial distribution entropy, and spatial consistency of enhancement pattern.
[0201] Step 109: Statistical analysis and modeling of breast cancer classification, building a multi-level, multi-task prediction model system. Detailed modeling process:
[0202] (1) Data preprocessing: The ultrasound and magnetic resonance features were combined to form a high-dimensional feature vector (with a total dimension of approximately 800-1000 dimensions). K-nearest neighbor imputation was used to process missing values, and the isolation forest algorithm was used to detect outliers.
[0203] (2) Feature selection strategy: A multi-stage feature selection was adopted. In the first stage, single-factor screening was performed using analysis of variance (ANOVA), with a p-value threshold of <0.05. In the second stage, recursive feature elimination (RFE) combined with cross-validation was used to select the optimal number of features. In the third stage, LASSO regression (L1 regularization) was used for sparse feature selection, and the regularization parameter λ was optimized through 10-fold cross-validation.
[0204] (3) Feature engineering: Box-Cox transformation is performed on continuous features to improve their distribution; feature interaction terms are constructed, including the ratio and product of ultrasound features and MRI features; principal component analysis (PCA) is used for dimensionality reduction, retaining 95% of the variance contribution rate.
[0205] (4) Multimodal fusion strategy: Three fusion methods are implemented: early fusion (feature-level fusion) directly connects the features of two modalities; mid-term fusion (decision-level fusion) trains the single-modal models separately and then fuses the prediction results; late fusion (model-level fusion) uses a meta-learner to integrate multiple base classifiers.
[0206] (5) Immunohistochemical index prediction model: A binary classification model was constructed for ER, PR, and Her2 status, and a regression model (continuous value prediction) and a classification model (high and low expression classification, threshold 14%) were constructed for Ki-67; the algorithm selection included support vector machine (SVM), random forest (RF), gradient boosting decision tree (GBDT), extreme gradient boosting (XGBoost), lightweight gradient boosting (LightGBM), and deep neural network (DNN). (6) Molecular typing prediction: Based on the immunohistochemical prediction results, a five-classification model was constructed to predict Luminal A type (ER+ / PR+ / Her2- / Ki-67<14%), Luminal B type Her2 negative (ER+ / PR+ / - / Her2- / Ki-67≥14%), Luminal B type Her2 positive (ER+ / PR+ / - / Her2+), Her2 positive type (ER- / PR- / Her2+), and triple negative type (ER- / PR- / Her2-). (7) Hyperparameter optimization: The Bayesian optimization algorithm was used to search for hyperparameters. The search space included learning rate (0.001-0.1), tree depth (3-15), regularization parameter (0.001-10), etc. The objective function was the cross-validation AUC value, and the number of optimization iterations was 100.
[0207] (8) Model integration: Three integration strategies are adopted: Voting, Stacking, and Blending. Voting includes hard voting and soft voting. Stacking uses logistic regression as a meta-learner. Blending uses a validation set for weight learning.
[0208] (9) Imbalanced data processing: Use SMOTE oversampling, Borderline-SMOTE, ADASYN and other methods to deal with category imbalance problems; adjust category weights and use balanced parameters for automatic calculation.
[0209] (10) Model validation: Stratified k-fold cross-validation (k = 5 or 10) was used to ensure that the proportion of each category in each fold was consistent; nested cross-validation was used for unbiased performance evaluation; and 95% confidence intervals were calculated to evaluate the stability of the results.
[0210] (11) Performance indicators: For classification tasks, accuracy, precision, recall, F1-score, AUC, and specificity are calculated; for regression tasks, mean square error (MSE), mean absolute error (MAE), and coefficient of determination (R 2 ); macro-average and micro-average indicators are calculated for multi-classification tasks.
[0211] Step 110: Systematic clinical validation interpretation and optimization, establishing a complete clinical validation system and continuous optimization mechanism. Comprehensive validation plan:
[0212] (1) Clinical study design: A prospective multicenter clinical validation study was designed. Inclusion criteria included patients aged 18-80 years, with imaging findings of breast space-occupying lesions, and patients who were scheduled for pathological biopsy. Exclusion criteria included patients in pregnancy and lactation, those with severe cardiopulmonary diseases, and those with contraindications to MRI. Sample size calculation was based on power analysis, with a test level of α = 0.05, a test power of 1-β = 0.80, an expected AUC of 0.85, and a minimum detectable difference of 0.10. The required sample size was calculated.
[0213] (2) Multi-center coordination: Establish a unified standard operating procedure (SOP) for data collection, including equipment calibration, scanning parameters, image quality control, etc.; regularly organize operator training to ensure operational consistency; establish a centralized data management platform to achieve standardized data storage and transmission.
[0214] (3) Reference standard establishment: Histopathological results were used as the gold standard, and immunohistochemical staining was standardized according to the ASCO / CAP guidelines; the criterion for ER and PR positivity was >1% positive nuclear staining; Her2 was assessed using immunohistochemistry combined with fluorescence in situ hybridization (FISH); and Ki-67 was counted using a digital image analysis system.
[0215] (4) Performance evaluation system: Calculate sensitivity, specificity, positive predictive value, negative predictive value, accuracy, and Kappa consistency coefficient; draw ROC curves, calculate AUC and 95% confidence intervals; perform subgroup analysis, including performance differences among different age groups, different pathological types, and different lesion sizes.
[0216] (5) Interpretability analysis: Use SHAP (SHapley Additive exPlanations) values to analyze feature contributions and generate feature importance rankings; use LIME (Local Interpretable Model-agnostic Explanations) for local interpretation to explain individual prediction results; build a decision rule extraction algorithm to convert black box models into understandable rules; and visualize the distribution differences and decision boundaries of important features.
[0217] (6) Clinical practicality evaluation: Evaluate the performance of the system in a real clinical environment, including the applicability to different equipment models, different operators, and different patient populations; test the robustness of the system, including the impact of image quality changes and parameter setting differences on the results; analyze failure cases and summarize the system limitations and improvement directions.
[0218] (7) Diagnostic efficacy analysis: calculate the degree of improvement in diagnostic accuracy and compare it with traditional imaging diagnosis; analyze the impact of the system on clinical decision-making, including changes in biopsy rates, shortened diagnostic time, etc.; conduct cost-benefit analysis to evaluate the economic value of the system.
[0219] (8) Continuous learning mechanism: Establish an online learning framework to support incremental learning and parameter updates of the model; design a data quality monitoring system to automatically detect data drift and concept drift; implement model version management to support model rollback and A / B testing.
[0220] (9) Quality assurance system: Establish a multi-level quality control mechanism, including data quality inspection, algorithm performance monitoring, and result consistency verification; design an anomaly detection system to automatically identify abnormal inputs and abnormal results; establish a user feedback mechanism to collect problems and suggestions during clinical use.
[0221] (10) Security and privacy protection: Implement end-to-end data encryption to ensure patient privacy; establish an access control mechanism to implement role-based permission management; design an audit log system to record all operation traces; and comply with HIPAA, GDPR and other regulatory requirements.
[0222] (11) User training and support: Develop tiered training programs, including system administrator training, clinician training, and technician training; develop online training systems and knowledge bases, and provide 24 / 7 technical support; establish a user community and experience sharing platform.
[0223] Through the systematic implementation of the above ten key steps, the present invention constructs a complete breast cancer immunohistochemistry and typing prediction system based on multimodal imaging, realizes a full-process technical solution from image acquisition to clinical application, and provides strong technical support for the precise diagnosis and treatment of breast cancer.
[0224] Beneficial Effects: Rapid, non-invasive testing. Based on conventional ultrasound and magnetic resonance imaging, this method is completely non-invasive, highly acceptable to patients, and repeatable, making it particularly suitable for patients who cannot tolerate a biopsy. The system can provide a prediction within 30 minutes of completing the imaging examination, significantly shortening diagnosis time, facilitating timely treatment planning, significantly improving medical efficiency, and reducing patient anxiety and waiting time.
[0225] Technical advantages of multimodal fusion. This invention innovatively integrates two complementary imaging modalities: ultrasound and magnetic resonance imaging. Ultrasound is good at displaying the morphological boundary characteristics of tumors, while magnetic resonance imaging is more sensitive in reflecting tumor vascularization and metabolic activity. Multimodal fusion can more comprehensively characterize the biological characteristics of tumors and significantly improve the prediction accuracy. Compared with a single modality method, the prediction accuracy is increased by 15-25%. The system simultaneously predicts four key immunohistochemical indicators: ER, PR, Her2, and Ki-67, and further performs molecular typing, providing a complete evaluation of the biological characteristics of breast cancer.
[0226] High-precision, objective diagnosis. Through large-sample training and multi-center validation, this system has achieved high accuracy in predicting various immunohistochemical indicators: ER prediction accuracy is over 90%, PR prediction accuracy is over 88%, Her2 prediction accuracy is over 85%, and Ki-67 prediction accuracy is over 86%. Traditional pathological diagnosis is subject to inter-observer and intra-observer variability, especially in the interpretation of indicators such as Ki-67, which is highly subjective. Based on objective imaging features and machine learning algorithms, this system eliminates the influence of human subjective factors, ensures the objectivity and consistency of diagnostic results, and approaches or even exceeds the diagnostic level of pathologists.
[0227] This invention analyzes images of the entire tumor, better reflecting its overall biological behavior and overcoming sampling bias. The system is reusable, with minimal marginal cost, significantly reducing the cost of a single test. Rapid and accurate diagnosis helps reduce unnecessary repeated examinations and the additional costs associated with delayed treatment.
[0228] Wide applicability and continuous optimization capabilities. This system is suitable for breast cancer patients of different age groups and different pathological types and has good universality. The system design fully considers factors such as equipment differences and operator differences. It has strong robustness and can operate stably in medical institutions of different levels. The system has an online learning function and can continuously optimize the model performance based on new clinical data. With the extension of usage time and the increase of data accumulation, the prediction accuracy of the system will continue to improve. In addition to providing prediction results, the system also provides detailed analysis reports and clinical recommendations, including treatment plan recommendations, prognosis assessment, etc., to provide clinicians with comprehensive decision-making support, and establish a complete quality control system and safety assurance mechanism to ensure the reliability of the results and the security of patient data.
[0229] The above specific implementation methods further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific implementation methods of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for predicting breast cancer immunohistochemical indicators and typing, characterized in that: The prediction method comprises: Obtain breast ultrasound images and breast magnetic resonance images of patients; performing preprocessing and quality control on the breast ultrasound image and the breast magnetic resonance image; Using a deep learning model to perform breast lesion region segmentation on the preprocessed breast ultrasound image and the breast magnetic resonance image to obtain a segmented lesion region; extracting multimodal radiomics features of the breast ultrasound image and the breast magnetic resonance image based on the segmented lesion area; fusing the multimodal imaging omics features of the breast ultrasound image and the breast magnetic resonance image, constructing a machine learning model, performing immunohistochemical index prediction, and obtaining immunohistochemical index prediction results; Breast cancer molecular typing analysis is performed based on the prediction results of the immunohistochemical indicators.
2. The method for predicting breast cancer immunohistochemical indicators and typing according to claim 1, wherein: The preprocessing and quality control of the breast ultrasound image specifically includes: Convert raw ultrasound images into DICOM standard format; Image dimensions were normalized to uniform pixel spacing and image size; Grayscale values are normalized to the range of [0,255]; Use bilateral filter for speckle noise suppression; Apply histogram equalization to enhance image contrast; Establish quality control standards and eliminate images with blurred images, severe artifacts, and unclear lesion visualization.
3. The method for predicting breast cancer immunohistochemical indicators and typing according to claim 1, wherein: The preprocessing and quality control of the breast magnetic resonance image specifically includes: Motion artifact correction was employed; N4 offset field correction; Multi-sequence image registration to maintain spatial consistency among T1-weighted imaging (T1WI), T2-weighted imaging (T2WI), and diffusion-weighted imaging (DWI) sequences; Normalize signal strength; The image is resampled to a uniform voxel size; Magnetic resonance image quality assessment was performed.
4. The method for predicting breast cancer immunohistochemical indicators and typing according to claim 1, wherein: The use of the deep learning model to segment the breast lesion area of the pre-processed breast ultrasound image specifically includes: Construct a segmentation network based on the improved U-Net, which includes 4 downsampling layers and 4 upsampling layers; Use a weighted combination of Dice loss function and Focal loss function for training; Introducing attention gating mechanism for lesion detection; Output the binary segmentation mask of the lesion area; Morphological post-processing is performed on the segmentation results, including hole filling and boundary smoothing.
5. The method for predicting breast cancer immunohistochemical indicators and typing according to claim 3, characterized in that: The use of a deep learning model to segment the breast lesion region on the pre-processed breast magnetic resonance image specifically includes: Use the improved U-Net segmentation network to perform joint segmentation of multiple sequence magnetic resonance images; Fusion of complementary information from T1-weighted imaging (T1WI), T2-weighted imaging (T2WI), diffusion-weighted imaging (DWI), and dynamic contrast-enhanced sequences; Use deep supervision mechanism for segmentation; Use conditional random fields to perform post-processing optimization of segmentation results; Generate 3D lesion segmentation masks and volume measurements.
6. The method for predicting breast cancer immunohistochemical indicators and typing according to claim 1, wherein: The step of extracting the multimodal radiomics features of the breast ultrasound image based on the segmented lesion area specifically includes: Morphological feature extraction to calculate the area, perimeter, roundness, ellipticity, and irregularity index of the lesion; First-order statistical feature extraction, including mean, variance, skewness, kurtosis and entropy; Second-order texture feature extraction, calculating contrast, correlation, energy and homogeneity based on the gray-level co-occurrence matrix; High-order feature extraction, including wavelet transform features and local binary pattern features; Echo characteristic analysis to evaluate the internal echo pattern and posterior acoustic shadow characteristics of the lesion.
7. The method for predicting breast cancer immunohistochemical indicators and typing according to claim 1, wherein: The step of extracting the multimodal imaging omics features of the breast magnetic resonance image based on the segmented lesion area specifically includes: Morphological feature extraction to calculate the volume, surface area, sphericity, and compactness of the lesion; Signal intensity feature extraction, analysis of signal features of T1-weighted imaging (T1WI), T2-weighted imaging (T2WI), and diffusion-weighted imaging (DWI) sequences; Texture feature extraction, based on gray-level co-occurrence matrix, gray-level run-length matrix and gray-level size area matrix; Dynamic feature extraction, analyzing the time-signal intensity curve characteristics of dynamic contrast-enhanced sequences; Diffusion feature extraction, calculation of statistical parameters and histogram features of apparent diffusion coefficient.
8. The method for predicting breast cancer immunohistochemical indicators and typing according to claim 1, characterized in that: The fusing of the multimodal imaging omics features of the breast ultrasound image and the breast magnetic resonance image, constructing a machine learning model, and performing immunohistochemical index prediction specifically includes: fusing multimodal imaging omics features of the breast ultrasound image and the breast magnetic resonance image to obtain a fusion feature; Performing dimensionality reduction processing on the fused features, and performing feature selection using principal component analysis and least absolute shrinkage selection operators; Build multi-classification machine learning models, including support vector machines, random forests, and gradient boosted trees; Use cross-validation and leave-one-out methods to evaluate model performance; Immunohistochemical index prediction models were established to predict the expression status of ER, PR, Her2 and Ki-67 respectively; Based on the immunohistochemistry prediction results, a breast cancer molecular typing decision tree model was constructed; Provides predicted probabilities and confidence intervals for clinical decision making.
9. The method for predicting breast cancer immunohistochemical indicators and typing according to claim 1, wherein: The prediction method further comprises: Visualizing the immunohistochemical index prediction results to generate a probability distribution graph of the immunohistochemical index; Draw classification probability maps and decision boundaries for molecular classification of breast cancer; Generate a three-dimensional visualization display of the lesion segmentation results; Get feature importance analysis and model interpretability reports.
10. A system for predicting breast cancer immunohistochemical indices and typing, using a method for predicting breast cancer immunohistochemical indices and typing according to any one of claims 1 to 9, characterized in that: The prediction system specifically includes: A data acquisition module, used for acquiring breast ultrasound images and breast magnetic resonance images of patients; An image preprocessing module, configured to perform preprocessing and quality control on the breast ultrasound image and the breast magnetic resonance image; A deep learning segmentation module is used to segment the breast lesion area on the preprocessed breast ultrasound image and the breast magnetic resonance image using a deep learning model to obtain a segmented lesion area; a feature extraction module, configured to extract multimodal imaging features of the breast ultrasound image and the breast magnetic resonance image based on the segmented lesion area; a machine learning module, configured to fuse the multimodal imaging omics features of the breast ultrasound image and the breast magnetic resonance image, construct a machine learning model, perform immunohistochemical index prediction, and obtain immunohistochemical index prediction results; The result analysis module is used to perform molecular typing analysis of breast cancer based on the prediction results of the immunohistochemical indicators.
11. A breast cancer immunohistochemical index and classification prediction system according to claim 10, characterized in that: The prediction system further includes: The visualization display module is used to visualize the prediction results of the immunohistochemical indicators.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the breast cancer immunohistochemistry and typing prediction method according to any one of claims 1 to 9 are implemented.
Citation Information
Cited By
Bone metastasis detecting and positioning method and system based on SPECT bone imaging
CN121213539A
Breast cancer multi-classification method and device based on bimodal image, medium and product
CN121236074A
Biomarker auxiliary screening system based on radiomics conjoint analysis
CN121259814A
Colorectal cancer MRI image segmentation method and system based on multi-dimensional feature fusion
CN121280359A
Colorectal cancer MRI image segmentation method and system based on multi-dimensional feature fusion
CN121280359B