Method, device and equipment for differentiating benign and malignant breast calcification
By using generative adversarial networks and target segmentation models to perform data augmentation and segmentation on breast calcification images, and extracting and fusing deep learning and radiomics features, the problem of relying on experience-based judgment for the identification of breast calcification foci has been solved, achieving more efficient and accurate breast cancer screening.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY)
- Filing Date
- 2026-03-26
- Publication Date
- 2026-08-04
AI Technical Summary
In current technologies, the differentiation between benign and malignant breast calcifications relies heavily on the experience of radiologists, resulting in problems such as large differences in diagnostic results, low efficiency, and a high rate of missed diagnoses of microcalcifications. Furthermore, traditional methods lack quantitative interpretation, which affects the accuracy of early breast cancer screening.
Generative adversarial networks are used to augment medical image data, generating synthetic samples that conform to pathological characteristics. Calcification lesion regions are segmented using a target segmentation model, and deep learning features and radiomics features are extracted and fused to determine the benign or malignant differentiation of breast calcifications.
It significantly improves the accuracy and interpretability of early breast cancer screening, provides reliable support for clinical diagnosis, and enhances the accuracy and reliability of differentiating between benign and malignant calcifications.
Smart Images

Figure CN122510149A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus and equipment for differentiating benign and malignant breast calcifications. Background Technology
[0002] Breast cancer is one of the most common malignant tumors among women worldwide, and early detection and accurate diagnosis are crucial for improving cure and survival rates. Mammography, as the preferred method for breast cancer screening, can clearly show microcalcifications and early lesion structures. However, the differentiation between benign and malignant calcifications relies heavily on the experience and subjective judgment of radiologists, resulting in problems such as large discrepancies in diagnostic results, low efficiency, and a high rate of missed diagnoses of microcalcifications. With the development of artificial intelligence technology, especially the application of deep learning in medical image analysis, the automated identification and classification of mammogram images has become a research hotspot.
[0003] In related technologies, most focus on single-modal image analysis, lacking quantitative interpretation of diagnostic criteria and the accuracy of imaging in early breast cancer screening. Summary of the Invention
[0004] This application provides a method, apparatus, and device for differentiating between benign and malignant breast calcifications, as well as a computer program product, to improve the accuracy and reliability of differentiating between benign and malignant breast calcifications, and to provide a more effective reference for the diagnosis and treatment of breast diseases.
[0005] In a first aspect, embodiments of this application provide a method for differentiating benign from malignant breast calcifications, comprising: performing data augmentation processing on medical image data using a generative adversarial network to obtain augmented medical image data, the medical image data including mammogram data; segmenting the calcification foci region in the augmented medical image data using a target segmentation model to obtain a segmentation mask; extracting deep learning features and radiomics features from the segmentation mask, and performing feature fusion processing on the deep learning features and radiomics features to obtain a fused feature vector; and determining the benign or malignant breast calcification differentiation result corresponding to the medical image data based on the fused feature vector.
[0006] In one possible implementation, the generative adversarial network includes a generator that performs data augmentation on medical image data to obtain augmented medical image data. This includes: transforming the medical image data based on generator parameters corresponding to the generator to generate synthetic samples that conform to pathological characteristics; and obtaining augmented medical image data based on the synthetic samples.
[0007] In one possible implementation, acquiring enhanced medical image data based on synthetic samples includes: determining target samples in the synthetic samples, wherein the target samples include synthetic samples that conform to pathological characteristics and have a confidence level greater than a preset confidence threshold; and obtaining enhanced medical sample data based on the target samples and medical image data.
[0008] In one possible implementation, the generative adversarial network further includes a discriminator, and the method further includes: inputting synthetic samples and original medical image data into the discriminator for adversarial training to obtain gradient information generated by the adversarial training; and optimizing the generator parameters based on the gradient information.
[0009] In one possible implementation, the enhanced medical image data includes cephalic and endoscopic oblique views. The enhanced medical image data is segmented into calcified lesion regions using a target segmentation model to obtain a segmentation mask. This process includes: extracting joint features from the cephalic and endoscopic oblique views using the target segmentation model to obtain feature maps from different perspectives; and fusing the feature maps from different perspectives using a target attention gating mechanism to generate a segmentation mask for the calcified lesion regions. The target attention gating mechanism is determined based on the calcification features of the calcified lesion regions.
[0010] In one possible implementation, before performing joint feature extraction on the cephaloscopic and oblique views using the target segmentation model, the method further includes: determining the medical image quantification features corresponding to the cephaloscopic and oblique views, wherein the medical image quantification features include at least one or more of voxel spacing, size, number of categories, and organ scale; determining the task configuration parameters for calcification region segmentation based on the medical image quantification features; and configuring the target segmentation model according to the task configuration parameters.
[0011] In one possible implementation, extracting deep learning features and radiomics features from the segmentation mask includes: encoding the segmentation mask based on a pre-acquired convolutional neural network model, and extracting high-dimensional feature vectors from the front of the fully connected layer as deep learning features based on the encoding results; the convolutional neural network model is trained using a medical image dataset; and extracting multi-dimensional radiomics features from the segmentation mask, wherein the multi-dimensional radiomics features include at least one of morphological features, texture features, intensity features, and spatial distribution features.
[0012] In one possible implementation, determining the benign or malignant identification result of breast calcifications corresponding to medical imaging data based on the fused feature vector includes: performing feature transformation on the fused feature vector through the convolutional layer of a pre-acquired benign or malignant identification model of breast calcifications to obtain the hidden feature representation corresponding to the fused feature vector; mapping the hidden feature representation to the output space, and determining the malignancy probability of the calcification lesion region based on the mapping result; and determining the benign or malignant identification result of breast calcifications corresponding to medical imaging data based on the malignancy probability.
[0013] Secondly, embodiments of this application provide a device for differentiating benign and malignant breast calcifications, comprising: an enhancement module for performing data enhancement processing on medical image data through a generative adversarial network to obtain enhanced medical image data, the medical image data including mammogram data; a segmentation module for segmenting the calcification foci region on the enhanced medical image data through a target segmentation model to obtain a segmentation mask; a fusion module for extracting deep learning features and radiomics features from the segmentation mask, and performing feature fusion processing on the deep learning features and radiomics features to obtain a fused feature vector; and an identification module for determining the benign or malignant breast calcification identification result corresponding to the medical image data based on the fused feature vector.
[0014] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0016] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0017] The method, apparatus, and device for differentiating benign and malignant breast calcifications provided in this application utilize generative adversarial networks to augment medical imaging data, especially mammogram data. This effectively expands the dataset size and increases data diversity, mitigating overfitting issues caused by limited data volume. A target segmentation model is used to segment the calcification region of the augmented data to obtain a segmentation mask, which accurately locates the calcification region and reduces interference from irrelevant regions in feature extraction. Deep learning features and radiomics features are extracted from the segmentation mask and fused. Deep learning features capture complex nonlinear relationships in the image, while radiomics features quantitatively reflect texture, morphology, and other information. The fusion of these two features integrates multi-dimensional information to obtain a more representative fused feature vector. Finally, the benign or malignant breast calcification differentiation result is determined based on this fused feature vector, improving the accuracy and reliability of the differentiation and providing a more effective reference for the diagnosis and treatment of breast diseases. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] Figure 1 A schematic diagram illustrating a scenario for differentiating benign and malignant breast calcifications according to an embodiment of this application;
[0020] Figure 2 A flowchart illustrating the method for differentiating benign and malignant breast calcifications provided in this application embodiment. Figure 1 ;
[0021] Figure 3 This is a schematic diagram of the structure of the moving inverse residual convolution and the fused inverse residual convolution block provided in the embodiments of this application;
[0022] Figure 4 A diagram of a generative adversarial network architecture provided in the embodiments of this application;
[0023] Figure 5 A schematic diagram of the skip connection structure of the adaptive medical image segmentation network provided in the embodiments of this application;
[0024] Figure 6 A schematic diagram of a two-dimensional U-shaped convolutional network architecture provided in an embodiment of this application;
[0025] Figure 7 A schematic diagram of a three-dimensional U-shaped convolutional network architecture provided in an embodiment of this application;
[0026] Figure 8 This is a schematic diagram of the system architecture for differentiating benign and malignant breast calcifications provided in this embodiment;
[0027] Figure 9 A schematic diagram of the overall process for differentiating benign and malignant breast calcifications provided in the embodiments of this application;
[0028] Figure 10 A schematic diagram of the module composition of the breast calcification benign and malignant differentiation system provided in the embodiments of this application;
[0029] Figure 11 A schematic diagram of the device for differentiating benign and malignant breast calcifications provided in this application;
[0030] Figure 12 A schematic diagram of the structure of the electronic device provided in this application.
[0031] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0032] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0033] First, let me explain the terms used in this application:
[0034] Generative Adversarial Networks (GANs) are deep learning model architectures consisting of a generator and a discriminator. The generator is responsible for generating realistic new data samples, such as lifelike images, based on input random noise or specific data. The discriminator, on the other hand, is responsible for identifying whether the input data comes from a real dataset or is fake data generated by the generator. During training, the generator and discriminator compete against each other, constantly playing a game. The generator strives to improve the quality of the generated data to fool the discriminator, while the discriminator continuously improves its discrimination ability to accurately distinguish between real and fake data. Through this dynamic adversarial training, the generator can eventually generate high-quality, near-realistic data samples, finding wide applications in many fields such as image generation and data augmentation.
[0035] Segmentation masks are binary images used in image processing and analysis, especially in object detection and segmentation tasks, to accurately identify specific target regions in an image. They are presented at the same size as the original image, where pixels in the target region are assigned a specific value (usually 1, representing the foreground), while pixels in the non-target region take another value (usually 0, representing the background). By clearly distinguishing between the foreground and background, the position, shape, and extent of the target in the image can be defined intuitively and accurately, providing crucial basic information for subsequent operations such as feature extraction, target recognition, and medical image analysis.
[0036] In related technologies, the differentiation between benign and malignant calcifications highly relies on the experience and subjective judgment of radiologists, resulting in problems such as large discrepancies in diagnostic results, low efficiency, and a high rate of missed diagnoses of microcalcifications. With the development of artificial intelligence technology, especially the application of deep learning in medical image analysis, the automated identification and classification of mammogram images has become a research hotspot. Existing technologies, through methods such as convolutional neural networks (CNNs), object detection algorithms, and transfer learning, attempt to solve the challenges of calcification detection and classification, but still face bottlenecks such as high data annotation costs, insufficient model generalization ability, and limited effectiveness in identifying micro and atypical calcifications. In addition, traditional computer-aided diagnosis (CAD) systems rely on manually designed feature extraction methods, focusing mostly on single-modal image analysis and lacking quantitative interpretation of diagnostic criteria, affecting the accuracy of early breast cancer screening.
[0037] To address the aforementioned issues, this application provides a scheme for differentiating benign from malignant breast calcifications. It enhances raw medical image data using a generative adversarial network (GAN) to generate synthetic samples that conform to pathological characteristics, thus alleviating the scarcity of malignant calcification samples. Secondly, a segmentation model is used to locate and segment the calcification region in the enhanced image, outputting a pixel-level segmentation mask to provide structured input for subsequent feature extraction. Finally, deep learning features (such as texture and morphological features) and radiomics features (such as aggregation degree and grayscale contrast) are extracted based on the segmentation mask. The two types of features are then fused to differentiate between benign and malignant breast calcifications, significantly improving the accuracy and interpretability of early breast cancer screening and providing reliable support for clinical diagnosis.
[0038] The breast calcification differentiation scheme provided in this application is applicable to automated screening scenarios using mammography (mammography). The target users are radiologists, hospital radiology departments, and public health screening institutions, etc., and this application does not impose any restrictions on them.
[0039] First, examples illustrate the application scenarios of this application. Please refer to... Figure 1 , Figure 1 This is a schematic diagram of a scenario provided for an embodiment of this application, such as... Figure 1 As shown, the implementation environment includes a terminal device 110 and a server 120. The terminal device 110 and the server 120 communicate with each other via a wired or wireless network. The terminal device 110 can upload its own data to the server 120, and can also retrieve data from the server 120.
[0040] Among them, terminal device 110 may include, but is not limited to, mobile phones, tablets, laptops, computers, voice interaction devices, etc.; server 120 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. This document does not restrict the specific form of terminal devices and servers.
[0041] It should be noted that, Figure 1 The number of terminal devices 110 and servers 120 is merely illustrative; any number of terminal devices 110 and servers 120 can be used as needed.
[0042] In an exemplary embodiment, the method for differentiating benign and malignant breast calcifications provided in the embodiments of this application can be executed by a terminal device 110. For example, the terminal device 110 can perform data augmentation processing on medical image data using a generative adversarial network to obtain augmented medical image data, including mammogram data. Then, the augmented medical image data is segmented into calcification regions using a target segmentation model to obtain a segmentation mask. Deep learning features and radiomics features are extracted from the segmentation mask, and these features are fused to obtain a fused feature vector. Finally, the benign or malignant breast calcification differentiation result corresponding to the medical image data is determined based on the fused feature vector, thereby improving the effectiveness of benign and malignant breast calcification differentiation.
[0043] In another exemplary embodiment, server 120 may have functions similar to terminal device 110 to execute the method for identifying benign and malignant breast calcifications provided in this application embodiment. For example, server 120 may perform data augmentation processing on medical image data using a generative adversarial network to obtain augmented medical image data, including mammogram data. Then, the augmented medical image data is segmented into calcification regions using a target segmentation model to obtain a segmentation mask. Deep learning features and radiomics features are extracted from the segmentation mask, and these are fused to obtain a fused feature vector. Finally, the benign and malignant breast calcification identification result corresponding to the medical image data is determined based on the fused feature vector, thereby improving the effectiveness of identifying benign and malignant breast calcifications.
[0044] In another exemplary embodiment, the terminal device 110 and the server 120 may also jointly execute the image processing method provided in the embodiments of this application. The specific implementation process is as described in the foregoing embodiments and will not be repeated here.
[0045] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0046] Figure 2 The flowchart provided in this application illustrates the process for differentiating between benign and malignant breast calcifications. Figure 1 .like Figure 2 As shown, this method for differentiating benign from malignant breast calcifications includes steps S201 to S204, which are described in detail below:
[0047] Step S201: Perform data augmentation processing on the medical image data using a generative adversarial network to obtain augmented medical image data, which includes mammogram data.
[0048] For example, considering the limited medical imaging data and the uneven distribution of calcifications, an online data augmentation mechanism can be introduced to improve the model's generalization ability and anti-overfitting performance. Data augmentation operations include: random rotation (within ±15°), horizontal or vertical flipping, scaling, brightness and contrast adjustment, gamma transformation, and adding Gaussian noise. Simultaneously, to maintain the spatial relationships and realistic features of the calcifications...
[0049] In this embodiment, reasonable constraints are set on the enhancement parameters to ensure that the morphology and density information of the lesions are not distorted. The above enhancement strategy is dynamically executed during the training process, enabling the model to learn the characteristics of calcifications under different imaging angles, brightness conditions, and noise levels, thereby improving its recognition accuracy and stability in real clinical scenarios.
[0050] Optionally, considering the inherent imbalance in the number of benign and malignant breast calcification samples during actual clinical data collection, simply relying on traditional geometric data augmentation techniques (such as rotation and flipping) is insufficient to fundamentally address the bias defects in model training. Traditional supervised learning models, when the sample distribution is significantly skewed, are prone to overfitting to the majority class (benign samples), leading to insufficient weight updates in the feature extraction network when faced with scarce malignant calcifications, making it difficult to fully capture the complex visual semantic features of malignant lesions in terms of microscopic texture and macroscopic morphology. In this embodiment, a high-dimensional fitting distribution model for malignant samples can be constructed using a Generative Adversarial Network (GAN) with its generative and discriminative adversarial mechanism. The core logic of this model lies in approximating the true data distribution of malignant calcifications through mathematical means. .
[0051] Specifically, the generator does not simply copy existing images, but rather learns various structural features of malignant calcifications in the latent feature space, including the clustering patterns of microcalcifications, linear branching morphology, and coarse, heterogeneous grayscale textures. The generator can map random noise into synthetic image data that is highly similar to real malignant samples in terms of texture detail, grayscale distribution, and spatial topology. Meanwhile, the discriminator, acting as the opposing side in the adversarial process, is responsible for strictly distinguishing between real clinical samples and synthetic samples, and continuously guides the generator to optimize its parameters through gradient information generated during adversarial training. This dynamically balanced adversarial process ensures that the generated samples not only fill the quantitative gap in malignant data but also greatly enrich the diversity of training data and the completeness of pathological features, thereby significantly improving the final classification model's ability to capture the features of malignant lesions and its diagnostic robustness.
[0052] Optionally, in one feasible embodiment, before performing data augmentation on the medical image data, standard craniotomy and oblique mammograms can be obtained from mammogram data in the hospital's Picture Archiving and Communication System (PACS) or local storage devices, and image standardization can be performed on them, for example, grayscale normalization, spatial resolution normalization, and contrast enhancement, to eliminate imaging differences between different devices. Furthermore, basic data augmentation techniques such as random rotation, flipping, and brightness adjustment can be applied to further enhance data diversity. Then, the preprocessed medical image data undergoes data augmentation processing via a generative adversarial network.
[0053] Step S202: The enhanced medical image data is segmented into calcified lesion regions using a target segmentation model to obtain a segmentation mask.
[0054] For example, a target segmentation model can be used to jointly extract features from head-to-tail and inner / outer oblique views to obtain feature maps from different perspectives. For instance, the segmentation model can extract deep features from two views separately, and then perform weighted fusion of the feature maps from different perspectives based on a target attention gating mechanism determined by features of the calcification region. The attention mechanism can learn the importance of features from different perspectives for the current pixel segmentation task; for example, for some calcification points, which may be clearer in the head-to-tail view, features from that view are given higher weights. After decoding and upsampling, the fused feature map finally generates an accurate segmentation mask for the calcification region.
[0055] Optionally, after standardizing and enhancing the image data, a segmentation model (nnU-Net structure) can be used to automatically detect and accurately segment calcification areas in breast images. The model locates calcifications through region proposal and multi-scale feature fusion mechanisms, and outputs corresponding bounding boxes and pixel-level segmentation masks. This module can accurately identify small, dense, or irregularly shaped calcifications against a complex breast tissue background, providing structured input for subsequent quantitative analysis and classification.
[0056] Step S203: Extract deep learning features and radiomics features from the segmentation mask, and perform feature fusion processing on the deep learning features and radiomics features to obtain a fused feature vector.
[0057] For example, based on the segmentation mask of the segmented calcification region, the original mammogram can be cropped according to the detected bounding boxes to obtain local image patches containing the calcifications. To extract deep semantic features, these image patches are input into a pre-trained EfficientNetV2-B0 convolutional neural network model for feature encoding. This model has been pre-trained on a large-scale natural image dataset and has strong feature extraction and transfer capabilities, effectively capturing subtle morphological differences and texture variations in calcifications.
[0058] Please see Figure 3 , Figure 3 This diagram illustrates the structure of the moving inverse residual convolution and the fused inverse residual convolution block provided in this embodiment. The EfficientNetV2-B0 convolutional neural network model uses two types of convolutional modules at different stages: the front end employs a fused inverse residual convolution block (Fused-MBConv) to improve training speed. This structure effectively expands the receptive field and accelerates gradient propagation, allowing early features to contain more local texture and primary morphological information, suitable for modeling dense granular textures in calcification images. The back end of the network uses the classic moving inverse residual convolution (MBConv) structure to enhance the model's high-dimensional semantic representation capabilities. Its inverse bottleneck structure includes dimensionality-increasing convolution, depthwise convolution, and a squeeze-and-excitation (SE) mechanism, which enhances the network's sensitivity to local morphological changes and texture combination patterns, exhibiting good feature capture capabilities for subtle edge abrupt changes and irregular density textures present in calcifications.
[0059] Step S204: Determine the benign or malignant identification result of breast calcification corresponding to the medical imaging data based on the fused feature vector.
[0060] For example, the benign or malignant identification result of breast calcifications corresponding to the medical image data can be determined based on the fused feature vector. This can be achieved by performing feature transformation on the fused feature vector using the convolutional or fully connected layers of a pre-acquired breast calcification benign / malignant identification model, such as a CNN classification module, to obtain the hidden feature representation corresponding to the fused feature vector. This identification model can be a small convolutional neural network or a fully connected network. The hidden feature representation is then mapped to the output space, resulting in a malignancy probability for the breast calcification corresponding to the fused feature vector. The final identification result is then determined based on this malignancy probability. For instance, by setting a malignancy probability threshold, if the malignancy probability is greater than or equal to the threshold, the breast calcification result is considered malignant; if the malignancy probability is less than the threshold, the breast calcification result is considered benign.
[0061] In summary, the embodiments provided in this application, by combining multiple network architectures such as Generative Adversarial Networks (GAN), nnU-Net, EfficientNetV2, and Convolutional Neural Networks (CNN), and through a series of structural improvements and feature fusion strategies, significantly enhance the extraction, fusion, and classification capabilities of calcification features. Specifically:
[0062] First, to address the scarcity of malignant calcification samples, this invention employs a GAN network for data augmentation. A generator synthesizes calcification images, producing high-quality, realistic malignant samples with lifelike pathological features. A multi-scale augmentation strategy is used to particularly highlight the morphology and distribution characteristics of microcalcifications, resulting in more realistic and diagnostically valuable images. This data synthesis module effectively expands the training dataset and enhances the model's generalization ability.
[0063] Next, this embodiment utilizes nnU-Net to perform precise region of interest (ROI) segmentation on the enhanced calcification image. nnU-Net automatically selects the optimal network structure and employs adaptive convolutional modules and skip connections to retain details that may be lost during segmentation, thereby improving segmentation accuracy. The segmentation results at this stage provide accurate region localization information for subsequent feature extraction and classification.
[0064] Subsequently, the segmented image regions are input into the EfficientNetV2 network for deep feature extraction. To further improve the model's classification ability, this embodiment also introduces a dual-path network structure: one path focuses on the morphological features of calcifications (such as granular, rod-like, etc.), and the other path focuses on their distribution features (such as clustered, linear, segmental, etc.). Through a feature fusion module, the network merges the feature information extracted by the two paths to generate a highly discriminative fused feature vector, thereby improving the ability to distinguish between benign and malignant calcifications.
[0065] Finally, the fused features are input into the CNN classification module, and the loss function (such as cross-entropy loss) is optimized through supervised learning, enabling the network to establish a stable and highly generalizable decision boundary in the high-dimensional feature space, thus achieving accurate classification of benign and malignant calcifications.
[0066] Based on the above embodiments, in one exemplary embodiment provided in this application, the generative adversarial network includes a generator, and the specific implementation process of performing data augmentation processing on medical image data through the generative adversarial network to obtain augmented medical image data may further include steps S301 and S302, which are described in detail below:
[0067] Step S301: Based on the generator parameters corresponding to the generator, the medical image data is transformed to generate a synthetic sample that conforms to pathological characteristics.
[0068] Step S302: Based on the synthetic sample, obtain enhanced medical image data.
[0069] For example, the generator does not simply copy existing images, but rather learns in-depth the various structural features of malignant calcifications in the latent feature space, including the clustering patterns of microcalcifications, linear branching morphology, and coarse, heterogeneous grayscale textures. The generator can map random noise into synthetic image data that is highly similar to real malignant samples in terms of texture detail, grayscale distribution, and spatial topology.
[0070] Optionally, the generator parameters need to be clearly defined. These parameters are key factors controlling the characteristics of the generated samples. They are carefully designed and adjusted to precisely define various attributes of the generated samples. After obtaining the generator parameters, they are applied to medical image data. Through the generator's complex algorithms and model architecture, a series of transformation operations are performed on the medical image data. These transformations may involve multiple aspects such as image texture, shape, and color distribution. The goal is to generate synthetic samples that conform to pathological characteristics. These synthetic samples should simulate the pathological features of lesions in real medical images as realistically as possible, such as the morphology, density, and edge features of tumors, so as to provide valuable references for subsequent medical research and diagnosis. After successfully generating synthetic samples that conform to pathological characteristics, these synthetic samples are fused with the original medical image data. Through reasonable algorithms and strategies, the effective information from the synthetic samples is added to the original data to obtain enhanced medical image data. This enhanced data can increase the richness and diversity of the data while retaining the main features of the original data, which helps to improve the accuracy and reliability of medical image analysis and provide doctors with more comprehensive and accurate diagnostic basis.
[0071] In the embodiments provided in this application, the medical image data is transformed according to the generator parameters to generate synthetic samples that conform to pathological characteristics, which can enrich the diversity and complexity of the dataset. The enhanced medical image data obtained in this way can effectively alleviate the problem of data scarcity, provide more sufficient and representative data for model training, and thus improve the model's learning ability and diagnostic accuracy of medical image features.
[0072] Based on the above embodiments, in one exemplary embodiment provided in this application, the specific implementation process of obtaining enhanced medical image data based on synthetic samples may further include steps S401 and S402, which are described in detail below:
[0073] Step S401: Determine the target sample in the synthetic sample. The target sample includes synthetic samples that meet the pathological characteristics and have a confidence level greater than a preset confidence threshold.
[0074] Step S402: Based on the target sample and medical image data, the enhanced medical sample data is obtained.
[0075] For example, after generating numerous synthetic samples, the target sample needs to be selected. This requires first defining a pre-set reliability threshold, which is determined after integrating professional medical knowledge and a large amount of experimental data. Next, each synthetic sample can be meticulously analyzed using appropriate algorithms to calculate its confidence level in conforming to pathological characteristics. The calculated confidence level is compared with the pre-set reliability threshold; only synthetic samples with a confidence level greater than this threshold are selected as target samples, as they can more realistically and accurately simulate pathological characteristics. After determining the target samples, the next step is to obtain the enhanced medical sample data. Specifically, based on the characteristics of the target sample and the original medical imaging data, key pathological feature information in the target sample can be accurately extracted and then cleverly integrated into the medical imaging data. This allows the enhanced medical sample data to retain the main information of the original data while adding new pathological features, resulting in higher quality and richer information.
[0076] In the embodiments provided in this application, by selecting target samples that meet pathological characteristics and have a confidence level greater than a preset confidence threshold, the quality of the synthesized samples can be guaranteed. Combining these samples with the original medical image data yields enhanced medical sample data, which increases both the amount of data and the reliability of the data. This helps the model learn medical image features more accurately and improves the performance of tasks such as disease diagnosis.
[0077] Based on the above embodiments, in one exemplary embodiment provided in this application, the generative adversarial network further includes a discriminator, and the specific implementation process of the above-mentioned method for differentiating benign and malignant breast calcifications may further include steps S501 and S502, which are described in detail below:
[0078] Step S501: Input the synthetic sample and the original medical image data into the discriminator for adversarial training to obtain the gradient information generated by the adversarial training.
[0079] Step S502: Optimize generator parameters based on gradient information.
[0080] For example, in generative adversarial networks (GANs), the discriminator, as the opposing side in the game, is responsible for strictly distinguishing between real clinical samples and synthetic samples. It continuously guides the generator to optimize its parameters through gradient information generated during adversarial training. This dynamically balanced adversarial process not only fills the gap in the number of malignant data but also greatly enriches the diversity of training data and the completeness of pathological features, thereby significantly improving the final classification model's ability to capture the features of malignant lesions and its diagnostic robustness.
[0081] Optional, please refer to Figure 4 , Figure 4 The generative adversarial network architecture diagram provided in the embodiments of this application is as follows: Figure 4 As shown, a generative adversarial network consists of a generator (G) and a discriminator (D), where the generator (G) is composed of parameters. The neural network implementation, whose input follows a distribution random vectors The output is the generated sample. It can be regarded as coming from the generating distribution. The actual data distribution is denoted as... The purpose of training is to make the generated distribution approximate the true distribution. In order to... near We can sample m samples from the generated distribution and maximize the likelihood:
[0082]
[0083] Its optimal parameters satisfy:
[0084]
[0085] In theory, when the discriminator reaches its optimal state:
[0086]
[0087] Substituting this into the optimization objective yields the Jensen-Shannon divergence (JSD) between the generated and true distributions. Therefore, the problem can be transformed into minimizing the JSD between the two distributions. The standard objective function for a generator (GAN) is then:
[0088]
[0089] The objective is essentially binary cross-entropy, aiming to make the generated distribution approximate the true distribution. During training, if the discriminator (D) is too strong, it will... When the generator approaches saturation, the gradient vanishes. To alleviate the gradient problem, the optimization objective of the generator is often changed to maximizing... Although gradient vanishing cannot be completely avoided, training can also lead to mode collapse due to the limited expressive power of the generator (G) and discriminator (D). Therefore, subsequent research has focused on improving the original generator (GAN) by modifying the loss function, introducing new structures, or employing stabilization techniques.
[0090] Furthermore, during the model training phase, the generated high-fidelity malignant calcification samples can be used together with real data for training. This significantly balances the ratio of benign to malignant samples while ensuring data authenticity, effectively mitigating the impact of data imbalance on model performance. This method not only enhances the model's ability to characterize malignant calcifications but also improves the stability and generalization performance of the overall classifier across different categories of samples, ensuring higher robustness and diagnostic accuracy in clinical applications.
[0091] In the embodiments provided in this application, by inputting synthetic samples and original medical image data into the discriminator for adversarial training, the discriminator can be prompted to accurately distinguish between real and fake samples. The resulting gradient information can be fed back to the generator. Based on this information, the generator parameters can be optimized, enabling the generator to generate more realistic samples that better match the distribution of real medical image features, thereby improving the data augmentation effect and model performance.
[0092] Based on the above embodiments, in one exemplary embodiment provided in this application, the enhanced medical image data includes cephalothorax and oblique views. The specific implementation process of segmenting the calcification region of the enhanced medical image data using a target segmentation model to obtain a segmentation mask may further include steps S601 and S602, which are described in detail below:
[0093] Step S601: Use the target segmentation model to perform joint feature extraction on the head and tail positions and the inner and outer oblique views to obtain feature maps from different perspectives.
[0094] Step S602: Based on the target attention gating mechanism, feature maps from different perspectives are fused to generate a segmentation mask for the calcification foci region. The target attention gating mechanism is determined based on the calcification features of the calcification foci region.
[0095] For example, enhanced medical image data typically includes cephalic and endoscopic oblique views. Then, a target segmentation model can perform joint feature extraction on these views to obtain feature maps from different perspectives. For instance, the model can extract deep features from each view separately, and then perform weighted fusion of the feature maps from different perspectives based on a target attention gating mechanism determined by calcification region features. The attention mechanism can learn the importance of features from different perspectives for the current pixel segmentation task; for example, for some calcifications that may be clearer in the cephalic and endoscopic views, features from that view are given higher weights. The fused feature maps undergo decoding and upsampling operations to ultimately generate an accurate segmentation mask for the calcification region.
[0096] In the embodiments provided in this application, by jointly extracting features from head-to-tail and inner-outer oblique views, more comprehensive feature maps can be obtained by making full use of information from different perspectives. The feature maps are fused based on the target attention gating mechanism determined by calcification features, which can focus on key calcification information and suppress irrelevant interference, thereby generating a more accurate calcification foci region segmentation mask and improving segmentation accuracy and reliability.
[0097] Based on the above embodiments, in one exemplary embodiment provided in this application, before performing joint feature extraction on the head-tail and internal-external oblique views using the target segmentation model, the specific implementation process of the above-mentioned method for differentiating benign and malignant breast calcifications may further include steps S701 and S702, which are detailed below:
[0098] Step S701: Determine the quantitative features of medical images corresponding to the head-tail and oblique views. The quantitative features of medical images include at least one or more of the following: voxel spacing, size, number of categories, and organ scale.
[0099] Step S702: Determine the task configuration parameters for calcification region segmentation based on the quantitative features of medical images, and configure the target segmentation model according to the task configuration parameters.
[0100] For example, an adaptive medical image segmentation network (nnU-Net) can be used as the backbone network to obtain the segmentation mask. nnU-Net, as an automatically configured medical image segmentation framework, automatically determines whether to use a 2D U-shaped convolutional network (2DU-Net), a 3D full-resolution U-Net, or a 3D cascaded U-Net, based on information such as voxel spacing, size, number of categories, and organ scale. It also adjusts the number of channels, network depth, kernel size, and number of pooling / upsampling layers for each layer. Furthermore, it adjusts preprocessing methods (resampling, normalization, patch size, etc.) and post-processing rules to achieve optimal segmentation performance under different data conditions. The skip connection structure used by nnU-Net is an improvement and extension of the traditional U-Net skip connection mechanism. (See also...) Figure 5 , Figure 5 This is a schematic diagram of the skip connection structure of the medical image segmentation network provided in the embodiments of this application, as shown below. Figure 5 As shown, the core of nnU-Net lies in adaptively adjusting the skip connections.
[0101] Optionally, the system determines the quantitative features of the medical images corresponding to the cephalic, endoscopic, and oblique views. These features include, but are not limited to, voxel spacing, image size, number of segmentation categories (e.g., background, calcifications), and the scale of the target organ (calcifications). Then, based on these quantitative features, the system automatically determines the task configuration parameters for calcification region segmentation, such as deciding whether to use 2DU-Net, a 3D U-shaped convolutional network (3DU-Net), or a 3D cascaded U-Net, and configuring the number of channels, kernel size, and pooling strategy for each layer. Finally, the target segmentation model is initialized or configured according to the task configuration parameters. nnU-Net achieves "plug-and-play" functionality in this way, automatically finding the optimal segmentation strategy for medical image data from different sources and formats, thus ensuring segmentation accuracy and robustness.
[0102] For example, for a two-dimensional mammogram, a 2DU-Net might be automatically configured. Its encoder extracts features through progressive downsampling using multiple convolutional and pooling layers, while the decoder recovers spatial details by upsampling through transposed convolutions and combining them with skip connections. Finally, a 1x1 convolution outputs the probability that each pixel belongs to a calcification, forming a segmentation mask. Skip connections fuse the shallow detail features of the encoder with the deep semantic features of the decoder, effectively preventing detail loss, which is crucial for segmenting tiny calcifications.
[0103] nnU-Net automatically determines whether to use 2DU-Net, 3D full-resolution U-Net, or 3D cascaded U-Net. U-Net is a convolutional neural network (CNN) for image segmentation. Its network architecture has a symmetrical encoder-decoder structure (i.e., a U-shaped structure). During input image processing, initial feature extraction is performed through convolutional layers (Conv3×3), followed by pooling layers (Max Pooling2×2) to progressively reduce the spatial resolution of the image while enhancing its semantic information. The encoder extracts hierarchical features of the image through multiple convolution and pooling operations, enabling the network to capture local texture and global morphological information. In the decoder stage, transposed convolutions upsample the feature maps, restoring the low-resolution feature maps to the original image's spatial dimensions. After each upsampling operation, the decoder concatenates the feature maps from the same layers in the encoder, helping the network better recover details and improve segmentation accuracy. See also... Figure 6 , Figure 6This is a schematic diagram of a two-dimensional U-shaped convolutional network architecture provided in an embodiment of this application, as shown below. Figure 6 As shown, the network ultimately maps the fused feature map to the category prediction of each pixel through a 1×1 convolutional layer. During training, optimization is performed using joint loss functions (such as Dice loss and cross-entropy loss) to achieve accurate segmentation of the target region. In particular, the skip connection design in the network effectively addresses the problem of image detail loss, improving the accuracy of the segmentation results.
[0104] In 3DU-Net, the encoder extracts multi-scale features from the input image through 3D convolution operations. Unlike 2DU-Net, 3DU-Net can directly process three-dimensional data, extracting features in three spatial dimensions (height, width, and depth) through convolution. Each convolution operation not only operates on slices of the image but also captures features in the depth direction, thus effectively extracting morphological and textural features from volumetric data. The encoding process of each layer consists of 3D convolution and three-dimensional max pooling, enabling the feature map to obtain a higher level of representation while reducing resolution. Its calculation method can be represented as follows:
[0105]
[0106] The decoder restores spatial resolution through 3D transposed convolution and performs skip connections with features at the same scale as the encoder to fuse shallow details and deep semantics, thereby improving segmentation accuracy. The fused feature map is represented as follows:
[0107]
[0108] The final output stage uses 1×1×1 convolutions to classify each voxel, generating segmentation predictions:
[0109]
[0110] Please see Figure 7 , Figure 7 This is a schematic diagram of the three-dimensional U-shaped convolutional network architecture provided in the embodiments of this application, as shown below. Figure 7 As shown, relying on the three-dimensional convolution, skip connections and hierarchical feature fusion mechanism, 3DU-Net has significant advantages in restoring volumetric data details and representing complex three-dimensional structures, and is suitable for accurate segmentation of three-dimensional target regions in medical images.
[0111] Unlike 3DU-Net, 3D-Cascaded U-Net employs a two-stage segmentation strategy: first, a coarser, low-resolution network extracts approximate features, and then a high-resolution network refines the segmentation of details. This design helps reduce computational overhead while improving segmentation accuracy. In the first stage, 3D-Cascaded U-Net uses a low-resolution convolutional neural network for initial region segmentation, quickly obtaining approximate target regions with minimal computational resources. This process gradually extracts features through 3D convolution and pooling operations, simplifying the complexity of the segmentation task. In the second stage, the network refines the initial segmentation results using the high-resolution 3DU-Net. This stage uses transposed convolutions and skip connections to gradually restore the low-resolution feature maps to the original image resolution, further enhancing the ability to recover local details. Skip connections fuse features at different levels, enabling the network to simultaneously process both global semantics and local details of the image, thus improving segmentation accuracy.
[0112] In segmenting medical images of calcifications, 2DU-Net effectively extracts local features through its encoder-decoder structure and preserves detailed information through skip connections, making it suitable for simpler two-dimensional structures. 3DU-Net extends to three dimensions, fully utilizing the spatial information of volumetric data to provide finer segmentation when processing complex three-dimensional structures, especially suitable for three-dimensional medical images. 3D Cascaded U-Net further optimizes this by employing a two-stage segmentation strategy: first coarse segmentation followed by fine-tuning, saving computational resources while improving detail recovery. This architecture is particularly suitable for processing multi-layered, highly detailed calcification images, improving efficiency while maintaining accuracy. nnU-Net leverages the different advantages of the three models, achieving high accuracy, efficiency, and robustness in calcification segmentation tasks.
[0113] In the embodiments provided in this application, by determining the quantitative features of medical images in the cephalic and endoscopic oblique views, the basic information of the images can be fully grasped. Based on these features, the task configuration parameters for calcification region segmentation can be determined, which allows the target segmentation model to be accurately configured according to the actual image situation, improving the model's adaptability and targeting to different images, thereby improving the accuracy and effectiveness of calcification region segmentation.
[0114] Based on the above embodiments, in one exemplary embodiment provided in this application, the specific implementation process of extracting deep learning features and image omics features from the segmentation mask may further include steps S801 and S802, which are described in detail below:
[0115] Step S801: Based on the pre-acquired convolutional neural network model, the segmentation mask is feature-encoded, and high-dimensional feature vectors are extracted from the front of the fully connected layer based on the encoding results as deep learning features. The convolutional neural network model is trained using a medical image dataset.
[0116] For example, for the segmented calcification region, the segmentation mask (or the calcification region image patch cropped from the original image based on the segmentation mask) can be feature-encoded based on a pre-acquired convolutional neural network model. This convolutional neural network model is preferably an EfficientNetV2-B0 model pre-trained on a large-scale natural image dataset (such as ImageNet) and fine-tuned on a medical image dataset. The calcification region image is input into this model, and the network performs multi-level feature extraction through the front-end Fused-MBConv module and the back-end MBConv module (containing the SE attention mechanism). Finally, based on the encoding results, a high-dimensional feature vector is extracted from the front of the fully connected layer (i.e., after the last convolutional layer or global pooling layer) as deep learning features. For example, a feature vector with a dimension of 1280 can be extracted, which contains abstract morphological and textural semantic information of the calcification.
[0117] Optionally, during feature extraction, this embodiment can select the high-dimensional feature vector preceding the fully connected layer of the convolutional neural network as the deep representation result. This feature vector contains multi-level visual semantic information of the calcifications, including local edge features, spatial texture distribution, and overall morphological features. Unlike traditional manually defined features, this deep feature is automatically learned by the network, which can adaptively capture the potential complex morphological patterns in the calcifications, thereby providing high-dimensional, abstract, and highly discriminative input features for subsequent classification models.
[0118] Step S802: Extract multi-dimensional image omics features from the segmentation mask. The multi-dimensional image omics features include at least one of morphological features, texture features, intensity features, and spatial distribution features.
[0119] In addition to deep learning features, this invention further utilizes Python and YAML (fixed parameters) scripts to extract multi-dimensional radiomics features based on the pixel-level segmentation mask output by nnU-Net, thereby achieving refined quantification of the physical and statistical characteristics of calcifications. The radiomics features are divided into the following four categories:
[0120] (1) Morphological characteristics:
[0121] The geometric shape indices of calcifications are calculated based on segmentation masks, including area, perimeter, roundness, aspect ratio, and fractal dimension. These features reflect the morphological complexity and edge irregularity of calcifications; malignant calcifications are typically characterized by irregular shapes and blurred boundaries.
[0122] (2) Texture features:
[0123] By constructing statistical models such as the gray-level co-occurrence matrix and gray-level run-length matrix, the texture distribution characteristics of calcified lesions, such as contrast, correlation, energy, entropy, and homogeneity, are calculated. These characteristics reflect the complexity and structural uniformity of gray-level changes within the calcified lesion region, while malignant calcification typically exhibits higher texture heterogeneity.
[0124] (3) Intensity characteristics:
[0125] Statistical analysis was performed on the grayscale values within the segmented regions to extract indicators such as average grayscale value, maximum value, minimum value, standard deviation, skewness, and kurtosis. These features describe the brightness and grayscale distribution characteristics of calcification foci and can be used to reflect the differences in density and composition of calcified tissue.
[0126] (4) Spatial distribution characteristics:
[0127] When calcifications are clustered, this embodiment can calculate indicators such as the number of calcification clusters, spatial aggregation, average distance between centroids, and nearest neighbor distance. These features are used to describe the arrangement pattern and spatial relationship of calcifications in breast tissue. Malignant calcifications are usually densely distributed or linearly aggregated.
[0128] In the embodiments provided in this application, by simultaneously extracting deep features and radiomics features, where deep features provide global semantics and high-dimensional abstract information, and radiomics features provide interpretable quantitative indicators, multi-dimensional feature fusion from the semantic layer to the physical layer is achieved. The combination of the two not only ensures the model's discriminative ability, but also enhances its clinical interpretability.
[0129] Based on the above embodiments, in one exemplary embodiment provided in this application, the specific implementation process of determining the benign or malignant breast calcification identification result corresponding to medical image data based on the fused feature vector may further include steps S901 to S903, which are described in detail below:
[0130] Step S901: The fused feature vector is transformed by the convolutional layer of the pre-acquired breast calcification benign and malignant identification model to obtain the hidden feature representation corresponding to the fused feature vector;
[0131] Step S902: Map the hidden feature representation to the output space, and determine the malignancy probability of the calcified lesion region based on the mapping result;
[0132] Step S903: Based on the probability of malignancy, determine the benign or malignant differentiation result of breast calcification corresponding to the medical imaging data.
[0133] For example, high-dimensional features obtained from deep learning via convolutional neural networks are concatenated with quantized features extracted from radiomics in the feature space to form a unified fused feature vector. This fused feature vector contains both high-level semantic information of calcifications and retains low-level physical properties such as morphology, texture, intensity, and spatial distribution, thus achieving information complementarity at both the global semantic and local structural levels. To ensure consistency in the numerical range of different features, all feature components are standardized and normalized before fusion to avoid differences in feature scale affecting the stability and convergence performance of model training.
[0134] The classification task employs a classification module of a Convolutional Neural Network (CNN). First, the fused feature vector is transformed using convolutional or fully connected layers to obtain a more discriminative abstract representation. This process can be represented as:
[0135]
[0136] Where W1 and b1 are trainable parameters, h is the fused feature vector, and σ is a non-linear activation function (such as ReLU). Subsequently, the classifier maps the hidden representation u to the output space and calculates the probability that the calcification is malignant using the sigmoid function.
[0137]
[0138] This allows for the construction of a nonlinear combination of hidden feature representations in a high-dimensional space, enabling the model to capture semantic changes corresponding to subtle morphological differences in calcifications. During the training phase, the model uses benign and malignant calcification samples verified by pathological biopsy as labeled data. Supervised learning is used to optimize the loss function (such as cross-entropy or log-likelihood loss) to establish a nonlinear mapping between features and categories. This training process enables the model to form a stable and generalizable decision boundary in a high-dimensional fused feature space, achieving effective differentiation of different categories of calcifications.
[0139] Then, during the inference phase, the classifier discriminates the input fused features and outputs a continuous probability value to represent the confidence level that the calcification belongs to the malignant category (e.g., malignancy probability).
[0140] .in This represents the confidence level that the calcifications are malignant. This embodiment uses cross-entropy loss for supervised optimization of the classifier.
[0141]
[0142] This loss function is used to update the parameters in the classification network through backpropagation, enabling the network to gradually form a decision boundary with stable generalization ability in the fused feature space. When the probability value is higher than a set threshold (e.g., 0.5), the system determines the calcification to be malignant; when it is lower than the threshold, it is determined to be benign. This probability output can be dynamically adjusted according to the actual clinical scenario to achieve an optimal balance between sensitivity and specificity.
[0143] Furthermore, this embodiment can output the classification results and confidence scores in the form of visual annotations on mammogram images, providing doctors with intuitive auxiliary diagnostic basis and realizing the quantitative expression and interpretable presentation of diagnostic results.
[0144] In the embodiments provided in this application, by utilizing the convolutional layer of the breast calcification benign and malignant differentiation model to transform the fused feature vector, a deeper level of hidden feature representation can be mined, which is then mapped to the output space to determine the probability of malignancy. This allows for a more accurate measurement of the likelihood of malignancy in calcifications, and ultimately, the differentiation result is derived based on the probability of malignancy, effectively improving the accuracy and reliability of the diagnosis of benign and malignant breast calcifications.
[0145] Please see Figure 8 , Figure 8 This is a schematic diagram of the system architecture for differentiating benign and malignant breast calcifications provided in this embodiment. Figure 8 As shown, the system architecture includes at least the following modules:
[0146] (1) Image input module
[0147] This module is used for line imaging data, supporting the reading and parsing of mammograms from hospital image archiving and communication systems (PACS) or local storage devices in the Digital Imaging and Communications in Medicine (DICOM) standard format. The module features image import, format conversion, and metadata verification functions, and can automatically identify the image acquisition perspective (cephalopelvic or oblique) and related imaging parameters, providing complete input information for subsequent processing.
[0148] (2) Preprocessing module
[0149] This module is used to standardize and enhance input images to eliminate the effects of different imaging devices, exposure conditions, and noise levels. Its main functions include grayscale normalization, spatial resolution normalization, contrast enhancement, and noise suppression. This module incorporates various image enhancement algorithms (such as adaptive histogram equalization and nonlocal mean denoising) and can automatically select the optimal processing strategy based on image features, thereby ensuring the clarity and consistency of the input image and providing a high-quality data foundation for subsequent feature extraction.
[0150] (3) Calcification detection and segmentation module
[0151] This module incorporates a pre-trained deep learning instance segmentation model (nnU-Net structure) for automatic detection and accurate segmentation of calcification regions in breast images. The model locates calcifications through region proposal and multi-scale feature fusion mechanisms, outputting corresponding bounding boxes and pixel-level segmentation masks. This module can accurately identify small, dense, or irregularly shaped calcifications against complex breast tissue backgrounds, providing structured input for subsequent quantitative analysis and classification.
[0152] (4) Multimodal feature extraction module
[0153] This module comprises two sub-modules: a deep learning feature extraction unit and a radiomics feature computation unit. The deep learning feature extraction unit, based on a pre-trained convolutional neural network (EfficientNetV2-B0 structure), encodes features of detected calcification foci regions, extracting high-dimensional deep semantic features. The radiomics feature computation unit, based on a segmentation mask, automatically calculates the morphological features (area, roundness, fractal dimension, etc.), texture features (contrast, entropy, correlation, etc.), intensity features (average gray value, skewness, kurtosis, etc.), and spatial distribution features (calcification quantity, clustering degree, nearest neighbor distance, etc.) of calcification foci. This module can generate quantitative and interpretable feature parameters, providing multidimensional input for the classification model.
[0154] (5) Integration Classification Module
[0155] This module is responsible for fusing deep features and radiomics features, and using a trained classification model to distinguish between benign and malignant calcifications. First, the two types of features are concatenated into a fused feature vector, which is then standardized and input into the classifier. The classifier can employ structures such as fully connected neural networks, support vector machines, or gradient boosting decision trees. By optimizing the supervised learning loss function, high-precision benign / malignant classification is achieved. The model's output is the probability value of a calcification belonging to the malignant category, and then it can automatically classify benign / malignant lesions based on a set threshold (e.g., 0.5).
[0156] (6) Results Output and Visualization Module
[0157] This module presents the results of Artificial Intelligence (AI) analysis in various formats to assist doctors in decision-making. The system can automatically generate structured diagnostic reports, including the location coordinates, number, malignancy probability, and key radiomics features of calcifications. The visualization unit highlights benign and malignant calcification areas in different colors on the original image and supports zooming, overlaying, and switching between multiple views. The doctor review interface provides interactive functions, allowing doctors to confirm, modify, or reject the AI's automatic identification results, achieving human-machine collaborative diagnosis.
[0158] In addition, please see Figure 9 , Figure 9 This is a schematic diagram of the overall process for differentiating benign and malignant breast calcifications provided in the embodiments of this application, as shown below. Figure 9 As shown, in the front-end processing stage, the image receiving module is responsible for receiving mammogram images from external systems (such as PACS) and performing standardized preprocessing. The calcification region segmentation module uses nnU-Net to accurately segment calcifications, providing accurate regional localization. The feature extraction and classification module extracts the morphological and distribution features of calcifications using the EfficientNetV2 model and classifies them as benign or malignant. The result generation and visualization module uses an improved version of Gradient-weighted Class Activation Mapping (Grad-CAM) technology to highlight key regions supporting the classification conclusion, enhancing interpretability. The report output module generates a detailed diagnostic report, providing segmentation results, classification information, and visual annotations to facilitate physician decision-making. Through modular design, this system improves the accuracy and efficiency of diagnosis while enhancing the interpretability of AI diagnostic results, increasing physician trust, and enhancing its clinical application value.
[0159] Please see Figure 10 , Figure 10 This is a block diagram of the breast calcification benign / malignant differentiation system provided in the embodiments of this application, as shown below. Figure 10 As shown in an exemplary embodiment, the breast calcification benign / malignant differentiation system provided in this application includes a breast image input unit, an image normalization and quality enhancement unit, a calcification detection and segmentation unit, a multi-source feature analysis and risk grading unit, and a diagnostic conclusion annotation and report generation unit connected in sequence. The units are closely linked through sequential data flow and progressively advancing functions: the output of the upstream unit serves as the input of the downstream unit, sequentially completing the entire process from raw image input, standardized preprocessing, precise lesion region segmentation, multi-dimensional feature extraction and risk analysis, to the final generation of a structured diagnostic report, forming an end-to-end automated assisted diagnostic process.
[0160] In summary, the embodiments provided in this application have at least the following technical effects:
[0161] 1. High precision and high robustness
[0162] This application's embodiments achieve multi-dimensional representation of image information by fusing two types of features—deep learning features and radiomics features—within the same framework. The deep learning features, derived from the high-level semantic space of a pre-trained convolutional neural network, automatically capture the texture and morphological patterns of calcifications. The radiomics features quantify attributes such as area, roundness, entropy, and aggregation density of calcifications from physical and statistical perspectives. This combination endows the model with both powerful abstract recognition capabilities and precise perception of the physical characteristics of lesions, significantly improving classification accuracy and generalization performance. The model maintains stable recognition performance under different imaging devices, breast density, and background noise conditions, demonstrating high robustness.
[0163] 2. Fully automated process
[0164] This application's embodiments construct an end-to-end automated workflow from image input, preprocessing, calcification detection and segmentation, feature extraction, classification and discrimination to result visualization. The system uses a deep instance segmentation network (such as nnU-Net) to automatically locate and segment calcifications, and a feature fusion classification module to determine benign or malignant lesions. The entire process requires no manual intervention. This automated workflow significantly improves diagnostic efficiency, enabling the analysis of an entire image in seconds, reducing repetitive work for doctors in large-scale breast cancer screening tasks, and effectively increasing the workload and screening coverage of medical institutions.
[0165] 3. Highly interpretable
[0166] This embodiment differs from traditional "black box" deep learning models. In addition to deep features, it introduces a radiomics feature calculation module, providing clear quantitative evidence for the AI's diagnostic process. Doctors can understand the key factors in the model's judgment through the morphological, textural, and spatial distribution feature parameters generated by the system, thus rationally reviewing the AI's output. This design balances algorithm performance with medical interpretability, enhancing doctors' trust in AI-assisted diagnostic results and facilitating the formation of a traceable and verifiable chain of evidence in clinical decision-making.
[0167] 4. Assist in early diagnosis
[0168] This application's embodiments leverage the ability of deep learning models to extract high-dimensional features at the pixel level, enabling the identification of minute, atypical calcifications that are difficult to detect with the naked eye. These early lesions are often in the initial stages of breast cancer development and are easily overlooked by traditional manual image reading. The introduction of this system can significantly improve the detection rate of early malignant calcifications, providing a basis for early clinical intervention and possessing significant public health and medical application value.
[0169] 5. Standardization and Consistency
[0170] The algorithm model in this application is based on unified data preprocessing, feature extraction, and classification standards, enabling it to output consistent analysis results across different hospitals and equipment conditions. Through quantitative feature representation and a threshold-based decision-making mechanism, the system avoids discrepancies arising from subjective physician judgment, thus improving the objectivity and reproducibility of diagnostic results. This standardized analysis model can serve as an auxiliary reference system for breast calcification assessment, providing technical support for the standardization and quality control of radiology diagnostic processes.
[0171] In summary, this application presents an innovative method for differentiating between benign and malignant breast calcifications, significantly improving the model's interpretability and physician trust. Unlike traditional end-to-end deep learning models, this invention processes tasks step-by-step, performing segmentation, feature extraction, and classification separately, enhancing the transparency of each processing step. This allows physicians to clearly understand the operation and results of each step during use, thereby increasing their trust in the diagnostic conclusions.
[0172] Furthermore, the innovation of this embodiment lies in the introduction of visualization technology. It not only provides classification results for benign and malignant calcifications but also highlights key areas on the image that support this classification conclusion. Through improved Grad-CAM technology, the most diagnostically valuable calcification particles or regions can be automatically identified. These key areas are typically associated with malignant characteristics of calcifications, such as densely distributed microcalcifications or branching morphology. When reviewing the classification results, doctors can directly see the image areas most relevant to the classification, thus gaining a more intuitive understanding of the basis for AI diagnosis, enhancing trust in the model, and improving the accuracy of assisted decision-making.
[0173] Figure 11 This is a schematic diagram of the device for differentiating benign and malignant breast calcifications provided in this application, as shown below. Figure 11As shown, the breast calcification benign / malignant identification device 1100 provided in this embodiment includes: an enhancement module 1110, used to perform data enhancement processing on medical image data through a generative adversarial network to obtain enhanced medical image data, the medical image data including breast X-ray image data; a segmentation module 1120, used to segment the calcification foci region of the enhanced medical image data through a target segmentation model to obtain a segmentation mask; a fusion module 1130, used to extract deep learning features and radiomics features from the segmentation mask, and perform feature fusion processing on the deep learning features and radiomics features to obtain a fused feature vector; and an identification module 1140, used to determine the benign / malignant identification result of the breast calcification corresponding to the medical image data based on the fused feature vector.
[0174] In one possible implementation, the enhancement module 1110 is specifically used to transform medical image data based on the generator parameters corresponding to the generator to generate a synthetic sample that conforms to pathological characteristics; and to obtain enhanced medical image data based on the synthetic sample.
[0175] In one possible implementation, the enhancement module 1110 is further configured to: determine the target sample in the synthetic sample, the target sample including the synthetic sample with a confidence level greater than a preset confidence threshold that conforms to pathological characteristics; and obtain enhanced medical sample data based on the target sample and medical image data.
[0176] In one possible implementation, the enhancement module 1110 is further configured to input the synthetic sample and the original medical image data into the discriminator for adversarial training to obtain gradient information generated by the adversarial training; and optimize the generator parameters based on the gradient information.
[0177] In one possible implementation, the aforementioned enhancement module 1120 is specifically used to extract joint features from the head-tail and inner-outer oblique views using a target segmentation model to obtain feature maps from different perspectives; and to generate a segmentation mask for the calcification region by fusing the feature maps from different perspectives according to a target attention gating mechanism, wherein the target attention gating mechanism is determined based on the calcification features of the calcification region.
[0178] In one possible implementation, the enhancement module 1120 is further configured to: determine the medical image quantification features corresponding to the head-tail and oblique views, the medical image quantification features including at least one or more of voxel spacing, size, number of categories and organ scale; determine the task configuration parameters for calcification region segmentation based on the medical image quantification features; and configure the target segmentation model according to the task configuration parameters.
[0179] In one possible implementation, the fusion module 1130 is specifically used to encode the segmentation mask based on a pre-acquired convolutional neural network model, and extract high-dimensional feature vectors from the front of the fully connected layer as deep learning features based on the encoding results. The convolutional neural network model is trained through a medical image dataset. Multi-dimensional radiomics features are extracted from the segmentation mask. The multi-dimensional radiomics features include at least one of morphological features, texture features, intensity features, and spatial distribution features.
[0180] In one possible implementation, the identification module 1140 is specifically used to: perform feature transformation on the fused feature vector through the convolutional layer of the pre-acquired breast calcification benign and malignant identification model to obtain the hidden feature representation corresponding to the fused feature vector; map the hidden feature representation to the output space and determine the malignant probability of the calcification lesion region based on the mapping result; and determine the breast calcification benign and malignant identification result corresponding to the medical imaging data based on the malignant probability.
[0181] The breast calcification benign or malignant identification device provided in this embodiment can perform the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0182] Figure 12 A schematic diagram of the structure of the electronic device provided in this application. Figure 12 As shown, the electronic device 1200 provided in this embodiment includes at least one processor 1210 and a memory 1220. Optionally, the device 1200 further includes a communication component 1230. The processor 1210, the memory 1220, and the communication component 1230 are connected via a bus 1240.
[0183] In a specific implementation, at least one processor 1210 executes computer execution instructions stored in memory 1220, causing at least one processor 1210 to perform the above-described method.
[0184] The specific implementation process of processor 1210 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0185] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0186] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0187] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0188] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0189] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0190] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0191] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an application-specific integrated circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0192] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0193] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0194] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0195] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0196] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0197] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A method for differentiating benign from malignant breast calcifications, characterized in that, include: The medical image data, including mammogram data, is augmented by a generative adversarial network. The enhanced medical image data is segmented into calcified lesion regions using a target segmentation model to obtain a segmentation mask. Deep learning features and radiomics features are extracted from the segmentation mask, and feature fusion processing is performed on the deep learning features and the radiomics features to obtain a fused feature vector; The benign or malignant differentiation result of breast calcification corresponding to the medical imaging data is determined based on the fused feature vector.
2. The method according to claim 1, characterized in that, The generative adversarial network includes a generator, and the process of performing data augmentation on the medical image data through the generative adversarial network to obtain augmented medical image data includes: Based on the generator parameters corresponding to the generator, the medical image data is transformed to generate a synthetic sample that conforms to pathological characteristics. Based on the synthesized sample, enhanced medical image data is obtained.
3. The method as described in claim 2, characterized in that, The process of obtaining enhanced medical image data based on the synthesized sample includes: The target samples in the synthetic samples are determined, and the target samples include synthetic samples that meet the pathological characteristics and have a confidence level greater than a preset confidence threshold; Based on the target sample and the medical image data, enhanced medical sample data is obtained.
4. The method according to claim 2, characterized in that, The generative adversarial network further includes a discriminator, and the method further includes: The synthetic sample and the original medical image data are input into the discriminator for adversarial training to obtain gradient information generated by the adversarial training. The generator parameters are optimized based on the gradient information.
5. The method according to any one of claims 1 to 4, characterized in that, The enhanced medical image data includes cephalic and endoscopic oblique views. The step of segmenting the enhanced medical image data into calcified lesion regions using a target segmentation model to obtain a segmentation mask includes: The target segmentation model is used to extract joint features from the head and tail positions and the inner and outer oblique views to obtain feature maps from different perspectives. The feature maps from different perspectives are fused according to the target attention gating mechanism to generate a segmentation mask for the calcification foci region. The target attention gating mechanism is determined based on the calcification features of the calcification foci region.
6. The method as described in claim 5, characterized in that, Before performing joint feature extraction on the head-tail and inner / outer oblique views using the target segmentation model, the method further includes: Determine the medical image quantification features corresponding to the cephalic and endoscopic oblique views, wherein the medical image quantification features include at least one or more of voxel spacing, size, number of categories, and organ scale; Based on the quantitative features of the medical images, the task configuration parameters for segmenting calcified lesion regions are determined, and the target segmentation model is configured according to the task configuration parameters.
7. The method according to any one of claims 1 to 4, characterized in that, The extraction of deep learning features and image omics features from the segmentation mask includes: The segmentation mask is feature-encoded based on a pre-acquired convolutional neural network model, and high-dimensional feature vectors are extracted from the front of the fully connected layer based on the encoding results as deep learning features. The convolutional neural network model is trained using a medical image dataset. Extract multi-dimensional radiomics features from the segmentation mask. The multi-dimensional radiomics features include at least one of morphological features, texture features, intensity features, and spatial distribution features.
8. The method as described in claim 7, characterized in that, The determination of the benign or malignant differentiation result of breast calcifications corresponding to the medical imaging data based on the fused feature vector includes: The fused feature vector is transformed by the convolutional layer of the pre-acquired breast calcification benign and malignant identification model to obtain the hidden feature representation corresponding to the fused feature vector; The hidden feature representation is mapped to the output space, and the malignancy probability of the calcified lesion region is determined based on the mapping result; Based on the malignancy probability, the benign or malignant differentiation result of the breast calcification corresponding to the medical imaging data is determined.
9. A device for differentiating benign and malignant breast calcifications, characterized in that, include: An enhancement module is used to perform data enhancement processing on medical image data through a generative adversarial network to obtain enhanced medical image data, including mammogram data. The segmentation module is used to segment the calcification region of the enhanced medical image data using a target segmentation model to obtain a segmentation mask. The fusion module is used to extract deep learning features and radiomics features from the segmentation mask, and to perform feature fusion processing on the deep learning features and the radiomics features to obtain a fused feature vector; The identification module is used to determine the benign or malignant identification result of breast calcifications corresponding to the medical imaging data based on the fused feature vector.
10. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1 to 8.