Mammary gland X-ray photography auxiliary diagnosis system based on feature image selective area fusion

By using deep learning-based feature image selection and fusion technology, the problems of low efficiency, insufficient accuracy, and missing localization in mammography diagnosis are solved, achieving efficient and accurate breast cancer diagnosis and reducing misdiagnosis rate and storage costs.

CN121789919APending Publication Date: 2026-04-03BOCE BIOMEDICAL (TIANJIN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Current mammography diagnostic techniques are inefficient, lack diagnostic accuracy, lack localization assistance, and are fragmented throughout the entire process. They fail to effectively combine feature recognition, image synthesis, and stereotactic localization, resulting in high misdiagnosis rates and high diagnostic costs.

Method used

Employing deep learning-based feature image selection and fusion technology, through image acquisition, feature recognition, selection and fusion, and stereo localization modules, the deep learning model identifies features such as calcifications in breast cancer, generates efficient and accurate synthetic images, and performs three-dimensional stereo localization to provide auxiliary diagnostic information.

Benefits of technology

It significantly reduces the misdiagnosis rate, reduces the number of images by more than 50%, shortens the doctor's image reading time, reduces radiation exposure time, improves diagnostic efficiency and reduces storage costs, and provides sub-millimeter-level precision lesion localization assistance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789919A_ABST
    Figure CN121789919A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical image processing, and discloses a breast X-ray photography auxiliary diagnosis system based on feature image selective area fusion. Comprising an image acquisition module used for acquiring a mammary gland X-ray tomographic image; the feature recognition module is used for performing feature recognition and marking on breast cancer calcification points, radial structures and circular focus high-density shadows in the breast X-ray tomographic image by using a deep learning method; the selected area fusion module is used for carrying out image synthesis on the key selected area according to the feature recognition and marking result, obtaining a synthesized image, reducing the absolute number of the images and keeping the diagnosis accuracy and sensitivity; and the three-dimensional positioning module is used for performing three-dimensional coordinate positioning on the synthesized image, outputting a focus position, dynamically optimizing 3D occupation of a focus in the synthesized image according to the focus position, analyzing and deeply learning a surrounding blood vessel and connective tissue density shadow, and outputting auxiliary diagnosis information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical image processing technology, and in particular to a mammography-assisted diagnostic system based on feature image selection region fusion. Background Technology

[0002] In the field of computer-aided diagnosis of mammography, current technologies mainly rely on manual page-by-page review of tomographic images and traditional image learning methods based on maximum intensity projection (MIP). These traditional methods have the following significant drawbacks:

[0003] 1. Inefficiency: Breast tomographic synthetic imaging typically contains dozens to hundreds of images. Diagnosing each image page by page takes a lot of time and requires a large amount of storage, which leads to a significant increase in the time that doctors and patients are exposed to X-ray radiation.

[0004] 2. Insufficient diagnostic accuracy: Image synthesis methods based on maximum density projection cannot specifically extract key lesion features such as calcification points and radial structures, and are easily affected by background noise, resulting in a high misdiagnosis rate.

[0005] 3. Lack of localization assistance: For lesions distributed across multiple layers, current technology lacks systematic three-dimensional localization optimization methods, making it difficult to accurately present the spatial relationship between the lesion and surrounding tissues (such as blood vessels and connective tissue), which affects the accuracy of subsequent biopsy interventions.

[0006] 4. Fragmented process: Image generation, transmission, storage and diagnosis are independent of each other, lacking an integrated optimization solution, making it difficult to achieve a synergistic improvement in diagnostic efficiency and cost.

[0007] Therefore, a system is urgently needed to solve at least one of the above problems. Summary of the Invention

[0008] This application provides a mammography-assisted diagnostic system based on feature image selection and fusion, aiming to address the shortcomings of existing technologies, which have not yet proposed a deep learning-based feature image selection and fusion technology, nor have they integrated feature recognition, image synthesis, stereo localization, and cloud imaging systems into a comprehensive optimization scheme. Traditional methods do not address the problem of cross-layer lesion localization by dynamically optimizing the 3D lesion location and deepening the learning of surrounding tissue features.

[0009] In a first aspect, this application provides a mammography-assisted diagnostic system based on feature image selection region fusion, comprising:

[0010] The image acquisition module is used to acquire mammograms.

[0011] The feature recognition module is used to identify and label breast cancer calcifications, radial structures, and high-density circular lesions in the breast X-ray tomographic images using deep learning methods.

[0012] The selected area fusion module is used to synthesize images of key selected areas based on the results of feature recognition and labeling, thereby reducing the absolute number of images and preserving diagnostic accuracy and sensitivity.

[0013] The stereo positioning module is used to locate the synthesized image using stereo coordinates, output the lesion location, dynamically optimize the 3D occupancy of the lesion in the synthesized image based on the lesion location, analyze and deepen the learning of the density shadows of surrounding blood vessels and connective tissue, and output auxiliary diagnostic information.

[0014] In some embodiments, the feature recognition and labeling of breast cancer calcifications, radial structures, and high-density circular lesions in the mammogram using deep learning methods includes: inputting the preprocessed mammogram into a pre-trained convolutional neural network model, which is trained on a breast image dataset containing annotations of calcifications, radial structures, and circular lesions; the neural network model extracts features from pixel regions in the mammogram per region, and identifies the shape, edge, and texture features of the high-density lesions through multiple convolutional and pooling layers; generating feature bounding boxes based on feature matching results, labeling the location and type of calcifications, radial structures, and circular lesions, and outputting a feature coordinate map.

[0015] In some embodiments, the step of synthesizing images of key selected regions based on the results of feature recognition and labeling to obtain synthesized images includes: determining the smallest image region containing feature-marked boxes as key selected regions based on the feature coordinate map; inputting multi-layer tomographic images corresponding to the key selected regions into a generative adversarial network model, wherein the generative adversarial network model learns the differences in pixel distribution between normal breast tissue and lesion regions through adversarial training; and performing weighted fusion of multi-layer images within the key selected regions according to the density threshold and spatial distribution of feature markers, retaining high-density shadow details corresponding to lesion features, suppressing background noise, and generating a synthesized image containing lesion features, thereby reducing the number of synthesized images compared to the original tomographic images.

[0016] In some embodiments, the step of locating the lesion location from the synthesized image using stereo coordinates includes: calibrating the synthesized image using multi-view camera parameters to obtain image pairs taken from different angles; calculating the disparity information of the lesion features in images from different viewpoints using a stereo matching algorithm, and constructing a three-dimensional coordinate system by combining the calibrated camera intrinsic and extrinsic parameters; calculating the three-dimensional coordinates of the center point of the lesion feature marker box using triangulation to generate lesion location data containing X, Y, and Z axis coordinate values.

[0017] In some embodiments, the step of dynamically optimizing the 3D occupancy of the lesion in the synthesized image based on the lesion location, analyzing and deepening the learning of the density shadows of surrounding blood vessels and connective tissue, and outputting auxiliary diagnostic information includes: inputting the three-dimensional coordinates of the lesion location into a three-dimensional convolutional neural network model, extracting three-dimensional voxel features from the tomographic images within a preset radius around the lesion; the three-dimensional convolutional neural network model identifies the tubular structural features of blood vessels and the grid-like density distribution of connective tissue, generating a spatial relationship heatmap of the lesion and surrounding tissues; adjusting the 3D occupancy boundary of the lesion in the synthesized image based on the blood vessel distribution density and the degree of connective tissue adhesion in the spatial relationship heatmap, and outputting a three-dimensional model of the lesion containing the feature weights of surrounding tissues and interventional path suggestions.

[0018] In some embodiments, an adaptive threshold segmentation algorithm is introduced into the feature recognition module. The adaptive threshold segmentation algorithm automatically calculates the optimal segmentation threshold based on the overall grayscale distribution of the image through a machine learning model. The optimal segmentation threshold is used to binarize the mammogram X-ray tomographic image to highlight the contrast between high-density shadows and normal tissue. The processed image is then input into the feature recognition model to improve the detection sensitivity of small lesions such as calcifications.

[0019] In some embodiments, the selected region fusion module employs a multi-task learning neural network to simultaneously achieve lesion feature preservation and image noise reduction; the network sets a feature preservation loss function and a noise suppression loss function, and optimizes the network parameters through backpropagation; when synthesizing images, the weights of the loss function are dynamically adjusted according to the priority of feature labels.

[0020] In some embodiments, the stereo positioning module integrates an incremental learning mechanism. When a new biopsy pathology result is obtained, the image features corresponding to the pathology result are associated with the positioning coordinates and annotated. The newly annotated data is input into the stereo positioning model for iterative training, and the three-dimensional coordinate calculation parameters are updated so that the subsequent lesion positioning error is reduced compared with the initial training.

[0021] In some embodiments, the feature recognition module employs an attention mechanism neural network, adding a spatial attention module and a channel attention module after the convolutional layer; the spatial attention module focuses on the region where the high-density shadow is located by generating a two-dimensional attention map, and the channel attention module enhances the response intensity of calcification points and radial structures by adjusting the feature channel weights; the interference of background noise such as adipose tissue is suppressed through the dual attention mechanism.

[0022] In some embodiments, the system includes a real-time feedback optimization module. When the lesion display effect of the synthesized image is marked, the real-time feedback optimization module collects the marking position and feedback type to generate feedback data. The feedback data is then input into a reinforcement learning model to generate a selection area fusion parameter adjustment strategy for the case. Based on the selection area fusion parameter adjustment strategy, the key selection area range and image synthesis weight are dynamically adjusted to generate an optimized synthesized image.

[0023] This application reduces the number of images by more than 50% through feature selection and fusion technology, shortening doctors' image reading time and reducing radiation exposure time for both doctors and patients. It utilizes a deep learning model to specifically extract high-risk features such as calcifications and radial structures, suppressing background noise and significantly reducing the misdiagnosis rate compared to traditional maximum density projection methods. Through stereoscopic coordinate localization and 3D feature learning, it dynamically optimizes the 3D lesion location, clearly presenting the spatial relationship between the lesion and surrounding tissues, providing sub-millimeter-level precision localization assistance for biopsy interventions. Combined with a cloud imaging system, it achieves end-to-end digital optimization from image generation, transmission, storage to diagnosis, reducing storage and transmission costs and creating a demand-side innovation of "fewer images, faster diagnosis, and lower cost."

[0024] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0025] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a schematic block diagram of a mammography-assisted diagnostic system based on feature image selection and fusion provided in an embodiment of this application;

[0027] Figure 2 This is a schematic diagram of the principle of a mammography-assisted diagnostic system based on feature image selection and fusion provided in an embodiment of this application.

[0028] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Detailed Implementation

[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0030] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0031] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present invention, the terms "first" and "second" are used in the embodiments of the present invention to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.

[0032] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0033] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0034] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0035] In the field of computer-aided diagnosis of mammography, current technologies mainly rely on manual page-by-page review of tomographic images and traditional image learning methods based on maximum intensity projection (MIP). These traditional methods have the following significant drawbacks:

[0036] 1. Inefficiency: Breast tomographic synthetic imaging typically contains dozens to hundreds of images. Diagnosing each image page by page takes a lot of time and requires a large amount of storage, which leads to a significant increase in the time that doctors and patients are exposed to X-ray radiation.

[0037] 2. Insufficient diagnostic accuracy: Image synthesis methods based on maximum density projection cannot specifically extract key lesion features such as calcification points and radial structures, and are easily affected by background noise, resulting in a high misdiagnosis rate.

[0038] 3. Lack of localization assistance: For lesions distributed across multiple layers, current technology lacks systematic three-dimensional localization optimization methods, making it difficult to accurately present the spatial relationship between the lesion and surrounding tissues (such as blood vessels and connective tissue), which affects the accuracy of subsequent biopsy interventions.

[0039] 4. Fragmented process: Image generation, transmission, storage and diagnosis are independent of each other, lacking an integrated optimization solution, making it difficult to achieve a synergistic improvement in diagnostic efficiency and cost.

[0040] Current technologies have not yet proposed a feature image selection and fusion technology based on deep learning, nor have they combined feature recognition, image synthesis, stereo localization, and cloud imaging systems to form a complete optimization solution. Traditional methods do not address the challenge of locating cross-layer lesions by dynamically optimizing the 3D lesion footprint and deepening the learning of surrounding tissue features.

[0041] To solve the above problem, please refer to Figures 1 to 2 This application provides a mammography-assisted diagnostic system based on feature image selection and fusion, comprising: an image acquisition module for acquiring mammography tomographic images; a feature recognition module for identifying and labeling breast cancer calcifications, radial structures, and high-density circular lesions in the mammography tomographic images using deep learning methods; a selection and fusion module for synthesizing images of key selected areas based on the feature recognition and labeling results, obtaining synthesized images, reducing the absolute number of images, and preserving diagnostic accuracy and sensitivity; and a stereo positioning module for locating the synthesized images using stereo coordinates, outputting the lesion location, dynamically optimizing the 3D occupancy of the lesion in the synthesized image based on the lesion location, analyzing and deepening the learning of surrounding blood vessels and connective tissue density shadows, and outputting auxiliary diagnostic information.

[0042] Specifically, this system aims to completely revolutionize the traditional mammography diagnostic process by deeply integrating deep learning, intelligent image synthesis, and three-dimensional stereoscopic positioning technologies to build a full-link, integrated intelligent solution from image input to diagnostic assistance output.

[0043] The core idea of ​​this system is to "intelligently extract key information from massive amounts of data." Instead of directly presenting users with raw tomographic images of dozens to hundreds of layers, it first processes, filters, and enhances the data through an intelligent front-end, ultimately presenting the most crucial and intuitive diagnostic information. The system receives the raw data sequence of mammograms via an image acquisition module. Using a deep learning model, the feature recognition module detects and marks key lesion features in parallel and automatically across the entire set of images.

[0044] Based on the markings from the previous step, the selective area fusion module performs intelligent image synthesis only on key areas containing lesions or suspected areas, generating a very small number of highly concentrated "essential" images. The stereoscopic localization module then performs three-dimensional spatial analysis on the synthesized images to accurately locate the lesions and analyze their spatial relationship with surrounding tissues, generating the final auxiliary diagnostic report.

[0045] The image acquisition module is responsible for interfacing with PACS or mammography equipment. Its core task is to receive, parse, and preprocess raw breast tomographic image sequences (usually in DICOM format) in a standardized manner. Preprocessing includes image normalization, noise reduction, and necessary coordinate system unification, laying a high-quality data foundation for subsequent deep learning analysis.

[0046] The interface adapter receives image data transmitted from imaging equipment or PACS servers via a standard DICOM network interface (such as DICOM C-Store SCP).

[0047] Data parsing extracts key patient information, scanning parameters, and pixel data and spatial location information for each image layer by parsing the DICOM file header.

[0048] Image preprocessing includes: grayscale normalization, which maps the grayscale values ​​of images acquired from different devices and under different exposure conditions to a unified standard range, eliminating device differences. Lightweight noise reduction applies algorithms such as nonlocal means or wavelet transform to suppress background noise without losing detail.

[0049] Volume data construction: Based on the layer thickness and interlayer spacing information of each image layer, all two-dimensional slices are reconstructed into a three-dimensional volume data to prepare for three-dimensional positioning.

[0050] The feature recognition module uses a pre-trained deep learning model to perform parallel analysis on the reconstructed three-dimensional volume data or all two-dimensional slices, automatically and accurately identifying and delineating key imaging features related to breast cancer.

[0051] The backbone network employs advanced 3D convolutional neural networks or Transformer architectures (such as 3D U-Net, ViT). These models can utilize contextual information both within and between slices, giving them a natural advantage in detecting lesions distributed across layers.

[0052] The training data was generated using a large-scale dataset of breast tomographic images precisely annotated by experienced radiologists. The annotations included: calcifications (marked as dots or small regions), radial structures (delineated with their star-shaped outlines), and round lesions / masses (delineated with their boundaries).

[0053] Multi-task recognition includes: Calcification point detection: The model outputs a heatmap highlighting all tiny clusters of calcification points and performing preliminary classification of their morphology (e.g., fine polymorphism, casting, etc.). Radial structure detection: The model can identify linear structures radiating outward from the center point and accurately segment their core and spur regions. Mass detection and segmentation: For circular or near-circular high-density shadows, the model can not only detect their presence but also accurately segment their 3D boundaries. Output: The output of this module is a set of "feature markers" with 3D spatial coordinates. Each marker is associated with the feature type, confidence score, and its specific bounding box or mask in the 3D volumetric data.

[0054] The region fusion module is key to solving the problems of "low efficiency" and "storage pressure." It abandons the traditional "one-size-fits-all" synthesis method of MIP and instead performs "feature-based intelligent selective fusion." It only synthesizes images from local areas and key layers marked by the feature recognition module that contain important information, thereby greatly reducing the number of images while preserving or even enhancing diagnostic information to the maximum extent.

[0055] Region selection: Based on the output of the feature recognition module, the system calculates one or more three-dimensional "regions of interest" that encompass all marked lesions.

[0056] Multiplanar reconstruction (MPR) automatically generates thin-slice MPR images in the coronal, sagittal, and transverse planes passing through the center of the major lesion. This provides physicians with a familiar view, similar to traditional three-dimensional anatomy.

[0057] Slab MIP generates a thin MIP layer, 5-10 mm thick, centered on the calcification center, for each calcification cluster. Compared to full-layer MIP, this method effectively eliminates interference from overlapping tissues, allowing calcifications to be more clearly visible.

[0058] Feature-weighted fusion can use a weighted fusion algorithm to assign higher weights to lesion areas during synthesis, making them more prominent in the final image.

[0059] Output: The output of this module is no longer hundreds of original images, but may only be a few to a dozen highly condensed, multi-angle synthetic images (such as several key-level MPR images and several slab MIP images for different lesions). These images constitute the core basis for subsequent diagnosis.

[0060] The stereo positioning module aims to solve the problem of "lack of positioning assistance." Based on the synthetic image generated by the selected area fusion module, it performs precise positioning and relationship analysis in three-dimensional space, providing direct and reliable navigation information for clinical biopsies or surgeries.

[0061] The 3D coordinate registration system maps each pixel in the synthesized image back to the original 3D volume data coordinate system, ensuring that every point on the 2D synthesized image has accurate 3D spatial coordinates (X,Y,Z).

[0062] The dynamic optimization of 3D lesion location begins with the system generating a coarse 3D model of the lesion based on the segmentation mask provided by the feature recognition module. Then, the system uses 3D segmentation algorithms such as graph cut or level set to refine the region within the original volume data, dynamically optimizing the lesion's morphology and boundaries to obtain a more accurate 3D location model.

[0063] Deep learning of surrounding tissue relationships involves defining an expanded three-dimensional region centered on the lesion, and then using another deep learning model (such as a relationship network) to analyze the density, course, and morphological changes of blood vessels and connective tissue within that region.

[0064] The model determines whether the lesion pulls on or infiltrates surrounding blood vessels, and whether it causes thickening or tortuosity of connective tissue (i.e., indirect signs such as the "funnel sign"), thereby outputting high-level auxiliary information about the malignancy and invasiveness of the lesion.

[0065] The system outputs auxiliary diagnostic information including: a precise location report, which outputs the specific quadrant, clock face position, and depth (in millimeters from the skin surface) of the lesion in the breast through text and graphic overlay; a 3D visualization model, which generates an interactive 3D model in which lesions, key blood vessels, and Cooper's ligaments are highlighted in different colors, allowing doctors to rotate and section them at will to intuitively understand their spatial relationships; and a BI-RADS grading suggestion, which, based on the comprehensive analysis results (lesion morphology, calcification type, surrounding relationships, etc.), provides an AI-based BI-RADS density and category grading suggestion for doctors' reference. This system achieves end-to-end optimization of "cloud-edge-device" collaboration by seamlessly integrating the above four modules and connecting with the cloud-based PACS system.

[0066] During image generation, the device can perform preliminary feature recognition and region fusion, reducing the amount of data uploaded. Transmission and storage are significantly reduced by requiring only the uploading and storage of a small amount of fused key images and related location data, greatly saving network bandwidth and cloud storage costs. In the diagnostic process, doctors receive a comprehensive report generated by the system at their workstation, containing intelligently synthesized images, 3D localization maps, and structured diagnostic suggestions, fundamentally improving diagnostic efficiency and quality.

[0067] The system provided in this application, through deep learning-driven feature recognition and intelligent region selection fusion technology, upgrades the diagnosis of breast tomography from a labor-intensive, page-by-page review to a highly precise, intelligent guidance system. It not only solves the inherent shortcomings of traditional methods in terms of efficiency, accuracy, and positioning, but also brings revolutionary progress to the entire breast imaging diagnostic process through its integrated design.

[0068] In some embodiments, the feature recognition and labeling of breast cancer calcifications, radial structures, and high-density circular lesions in the mammogram using deep learning methods includes: inputting the preprocessed mammogram into a pre-trained convolutional neural network model, which is trained on a breast image dataset containing annotations of calcifications, radial structures, and circular lesions; the neural network model extracts features from pixel regions in the mammogram per region, and identifies the shape, edge, and texture features of the high-density lesions through multiple convolutional and pooling layers; generating feature bounding boxes based on feature matching results, labeling the location and type of calcifications, radial structures, and circular lesions, and outputting a feature coordinate map.

[0069] This embodiment focuses on the core implementation details of the feature recognition module. Its core idea is to utilize a deep convolutional neural network trained on a large amount of professional data to automatically detect and locate three key lesions in mammograms with pixel-level precision: calcifications, radial structures, and circular lesions. This network simulates the human eye's visual perception mechanism, extracting image features layer by layer from low to high level, ultimately achieving accurate lesion identification and spatial labeling.

[0070] The model architecture is chosen and constructed using an encoder-decoder network structure, such as U-Net or its variants. The encoder part (downsampling path) uses deep convolutional networks (such as convolutional blocks of VGG or ResNet) to extract and abstract image features. The decoder part (upsampling path) then restores the abstract features to high resolution through deconvolution or upsampling operations, enabling class prediction for each pixel.

[0071] The last layer of the network uses the Softmax or Sigmoid activation function to output a probability value for each pixel, representing the probability that it belongs to "calcification point", "radial structure", "circular lesion" or "background".

[0072] Data preparation and training include: Training dataset: Collecting thousands to tens of thousands of breast computed tomography (CT) scans. Each dataset requires meticulous annotation by multiple senior radiologists under strict quality control. The annotation is in the form of pixel-level segmentation masks, precisely outlining each cluster of calcifications, the spiculations and core of radial structures, and the boundaries of circular masses using different colored regions. Training process: The preprocessed images and corresponding annotation masks are input into the network. By optimizing the loss function (such as a hybrid loss function combining Dice Loss and cross-entropy), the backpropagation algorithm is used to continuously adjust the millions of weight parameters of the network, making the network's predictions increasingly closer to the gold standard annotations of doctors.

[0073] The inference and labeling process includes: Input: A preprocessed mammogram is input into the trained network. Region-by-region feature extraction: The image passes through multiple convolutional and pooling layers. Shallow networks identify basic edges and corners; deep networks integrate this basic information to identify complex shapes (such as circles), textures (such as spiky textures), and patterns (such as clustered distributions).

[0074] Generate Feature Boxes: The network outputs a segmentation map of the same size as the input image. The system performs connected component analysis on this segmentation map, calculating the minimum bounding box (MUB) for each individual region predicted as a lesion. Output Feature Map: The system generates a new image (feature map) that overlays these bounding boxes onto the original image. Each bounding box is associated with the following metadata: Type: calcification / radial structure / circular lesion. Location Coordinates: The coordinates of the top-left and bottom-right corners of the bounding box in the 2D image (x1, y1, x2, y2). Confidence Score: The confidence score of the network predicting that the region is a lesion.

[0075] In some embodiments, the step of synthesizing images of key selected regions based on the results of feature recognition and labeling to obtain synthesized images includes: determining the smallest image region containing feature-marked boxes as key selected regions based on the feature coordinate map; inputting multi-layer tomographic images corresponding to the key selected regions into a generative adversarial network model, wherein the generative adversarial network model learns the differences in pixel distribution between normal breast tissue and lesion regions through adversarial training; and performing weighted fusion of multi-layer images within the key selected regions according to the density threshold and spatial distribution of feature markers, retaining high-density shadow details corresponding to lesion features, suppressing background noise, and generating a synthesized image containing lesion features, thereby reducing the number of synthesized images compared to the original tomographic images.

[0076] This embodiment aims to optimize the image synthesis quality of the selected region fusion module. It introduces the advanced paradigm of generative adversarial networks (GANs), whose core innovation lies in abandoning fixed fusion rules and instead allowing the model to learn how to generate an ideal synthesized image that is both clear and preserves lesions while being clean and noise-free. Through the "adversarial game" between the generator and the discriminator, a synthesized image that surpasses the performance of traditional fusion algorithms is ultimately obtained.

[0077] The system receives the feature coordinate map, calculates the union of all feature bounding boxes, and appropriately extends it outward by a safety boundary to form one or more three-dimensional selection areas covering all suspected lesions.

[0078] Building and training a GAN model involves: Generator: The input is a multi-layered original tomographic image within a selected region. The generator's goal is to output a single, high-quality synthetic image. The generator is typically a U-Net-like structure responsible for learning a complex mapping function. Discriminator: The input is an image (either a "fake image" synthesized by the generator or a "real image" approved by a doctor—i.e., an ideal image manually fused from the same selected region by an expert). The discriminator's goal is to correctly determine whether the input image is real or fake.

[0079] Adversarial training involves the generator trying to "deceive" the discriminator into believing that the synthesized image is "real." The discriminator then continuously evolves to better distinguish between real and fake images. This game forces the generator to synthesize images that are infinitely close to expert level in terms of detail, texture, and overall appearance.

[0080] During the generator fusion process, feature labeling information is explicitly utilized. For example, pixels within the feature labeling bounding box are given higher weights in the loss function to ensure that the details of these regions are perfectly preserved.

[0081] Through training, the model learned how to adjust the contribution of different layers of images during fusion (i.e., weighted fusion) based on the density threshold of features (such as calcification points usually having the highest density) and spatial distribution (such as calcification point clusters needing to be displayed as a whole), while instinctively suppressing noise from background fat and glandular tissue.

[0082] Ultimately, for a set of raw data with hundreds of layers, the system may only output 5-10 composite images for different selection areas and different perspectives. The number is drastically reduced, but the key information is not lost.

[0083] In some embodiments, the step of locating the lesion location from the synthesized image using stereo coordinates includes: calibrating the synthesized image using multi-view camera parameters to obtain image pairs taken from different angles; calculating the disparity information of the lesion features in images from different viewpoints using a stereo matching algorithm, and constructing a three-dimensional coordinate system by combining the calibrated camera intrinsic and extrinsic parameters; calculating the three-dimensional coordinates of the center point of the lesion feature marker box using triangulation to generate lesion location data containing X, Y, and Z axis coordinate values.

[0084] This embodiment details how the stereo positioning module achieves precise conversion from two-dimensional synthetic images to three-dimensional spatial coordinates. Its technical inspiration comes from human binocular vision and photogrammetry; by simulating "viewing" the lesion from different angles and calculating its parallax, its true three-dimensional position is deduced.

[0085] In the acquisition of multi-view image pairs and camera calibration, the system does not actually use multiple cameras to take pictures, but rather utilizes digital reconstruction radiographic imaging technology. That is, from the three-dimensional volume data, two or more two-dimensional projection images are virtually generated from different angles (such as 5-10 degrees to the left and right in the front direction), forming an "image pair".

[0086] Camera parameter calibration is a crucial step. The system precisely records the internal parameters (such as focal length and principal point) and external parameters (i.e., the virtual camera's position and orientation in the 3D world coordinate system) of each virtual camera. These parameters are known and accurate during DRR generation.

[0087] Stereo matching and disparity calculation: For the same lesion feature point in two virtual images, the system uses a stereo matching algorithm (such as Semi-Global Matching, SGM) to find its corresponding point.

[0088] After finding the corresponding points, calculate the difference in their horizontal coordinates on the image plane, which is the parallax. Parallax is inversely proportional to depth (Z coordinate): the greater the parallax, the closer the object is to the "camera".

[0089] The triangulation method calculates three-dimensional coordinates by performing geometric calculations based on the calibrated camera parameters and the calculated parallax value, using the principles of triangulation.

[0090] Specifically, by substituting the pixel (u_l, v_l) in the left image and the matching point (u_r, v_r) in the right image, along with the camera parameters, into the collinearity equation or the direct linear transformation formula, the precise coordinates (X, Y, Z) of the point in the three-dimensional world coordinate system can be calculated.

[0091] The system outputs structured lesion location data, for example: {Lesion ID:1, Type:"Round Lesion", Coordinates:(X:125.3mm,Y:67.8mm,Z:45.2mm). Here, the Z coordinate typically represents the depth of the lesion from the detector surface or skin surface.

[0092] In some embodiments, the step of dynamically optimizing the 3D occupancy of the lesion in the synthesized image based on the lesion location, analyzing and deepening the learning of the density shadows of surrounding blood vessels and connective tissue, and outputting auxiliary diagnostic information includes: inputting the three-dimensional coordinates of the lesion location into a three-dimensional convolutional neural network model, extracting three-dimensional voxel features from the tomographic images within a preset radius around the lesion; the three-dimensional convolutional neural network model identifies the tubular structural features of blood vessels and the grid-like density distribution of connective tissue, generating a spatial relationship heatmap of the lesion and surrounding tissues; adjusting the 3D occupancy boundary of the lesion in the synthesized image based on the blood vessel distribution density and the degree of connective tissue adhesion in the spatial relationship heatmap, and outputting a three-dimensional model of the lesion containing the feature weights of surrounding tissues and interventional path suggestions.

[0093] This embodiment enhances the functionality of the stereotactic localization module, enabling it to not only provide the coordinates of a center point but also depict the three-dimensional morphology of the lesion (3D occupancy) and intelligently analyze its spatial relationship with surrounding key tissues (blood vessels, connective tissue), thereby providing interventional guidance with greater clinical value.

[0094] The three-dimensional feature extraction uses the calculated lesion center coordinates (X,Y,Z) as the center of a sphere and extracts a three-dimensional sub-block with a preset radius (e.g., 30mm) from the original three-dimensional volume data.

[0095] The 3D sub-block is input into a 3D convolutional neural network. The convolutional kernels of a 3D CNN can move in three dimensions, thus it can capture the three-dimensional structural information of the lesion and its three-dimensional spatial relationship with the surrounding tissues extremely effectively.

[0096] The generated spatial relationship heatmap is trained using a 3D CNN and can identify and segment blood vessels (which appear as tubular, continuous low to medium density shadows) and connective tissue (such as Cooper's ligaments, which appear as linear or lattice-like high density shadows) within sub-blocks.

[0097] The network output is a spatial heatmap. On this map, different colors or brightness levels represent different tissue types or their "proximity" to the lesion. For example, blood vessels that are stretched and twisted by the lesion will be highlighted in bright red.

[0098] Dynamic optimization of 3D occupancy and output includes: Boundary optimization: The system uses heat map information to correct the boundaries of the lesions initially segmented in Example 1. For example, if the heat map shows that one side of the lesion is tightly adhered to normal blood vessels, the system may slightly shrink the boundary of the lesion on that side to distinguish the tumor tissue from the pushed-out normal tissue.

[0099] The output auxiliary information includes: a 3D lesion model: a three-dimensional mesh model in which lesions, blood vessels, and connective tissue are rendered with different colors and transparency. This model includes weights of surrounding tissue features, meaning that doctors can clearly see which blood vessels are "dangerous" and need to be avoided during surgery. Interventional path suggestion: Based on this 3D model, the system can automatically calculate an optimal puncture biopsy path from the skin surface to the lesion. This path ensures avoidance of large blood vessels and dense glandular areas, improving the safety and accuracy of the biopsy.

[0100] In some embodiments, an adaptive threshold segmentation algorithm is introduced into the feature recognition module. The adaptive threshold segmentation algorithm automatically calculates the optimal segmentation threshold based on the overall grayscale distribution of the image through a machine learning model. The optimal segmentation threshold is used to binarize the mammogram X-ray tomographic image to highlight the contrast between high-density shadows and normal tissue. The processed image is then input into the feature recognition model to improve the detection sensitivity of small lesions such as calcifications.

[0101] This embodiment enhances the feature recognition module to address the issue of missed detection of microcalcifications caused by uneven image contrast. It adds a preprocessing step, adaptive thresholding, before the image is fed into the deep learning model, aiming to pre-highlight potential lesions, especially microcalcifications, from the complex background.

[0102] Adaptive thresholding no longer uses a single global threshold. Instead, it employs a machine learning model based on local image characteristics (such as using Gaussian weighting or the mean and standard deviation of local neighboring pixels) to compute a "locally optimal threshold" for each pixel region in the image. For example, a variant of Otsu's Method can be used to compute the optimal segmentation threshold separately within multiple local windows of the image.

[0103] Binarization is performed on the original grayscale image using a calculated local threshold: pixels above the local threshold are set to white (foreground) and pixels below the threshold are set to black (background).

[0104] After this processing, all relatively high-density areas in the image (including lesions and some normal tissues) will be initially extracted to form a high-contrast binary image.

[0105] Input Feature Recognition Model: This binary image can be used as an additional channel, concatenated with the original grayscale image, and then fed into the CNN. In this way, the CNN not only receives the original texture information but also a strong cue about "where the problem might be." This is equivalent to giving the network a "focusing lens," making it easier to activate when faced with tiny, low-contrast calcifications, thereby significantly improving detection sensitivity (i.e., reducing false negatives).

[0106] In some embodiments, the selected region fusion module employs a multi-task learning neural network to simultaneously achieve lesion feature preservation and image noise reduction; the network sets a feature preservation loss function and a noise suppression loss function, and optimizes the network parameters through backpropagation; when synthesizing images, the weights of the loss function are dynamically adjusted according to the priority of feature labels.

[0107] This embodiment optimizes the network structure of the selected region fusion module by adopting a multi-task learning framework. Its core idea is to enable a network to simultaneously learn the seemingly contradictory yet crucial tasks of "preserving lesions" and "suppressing noise," and through a clever design of the loss function, allow the network to find the optimal balance point on its own.

[0108] The network architecture design involves constructing a network with a shared encoder and two independent decoders.

[0109] The shared encoder is responsible for extracting common, low-level features from the input multi-layered image.

[0110] The dual decoder and loss functions include: Task 1: Feature Preservation. The first decoder generates a feature-preserving map, aiming to make lesion areas clearly visible. Dynamic Weight Adjustment: During training and inference, the system dynamically adjusts the weights (λ1 and λ2) of the two loss functions based on the features of the currently processed case. For example, for a case filled with scattered calcifications, the system increases the weight of the feature preservation loss function (λ1) to ensure that all calcifications are fused in. Conversely, for an image with strong background noise, the system increases the weight of the noise suppression loss function (λ2). This dynamic adjustment mechanism makes the fusion strategy more adaptive and intelligent.

[0111] In some embodiments, the stereo positioning module integrates an incremental learning mechanism. When a new biopsy pathology result is obtained, the image features corresponding to the pathology result are associated with the positioning coordinates and annotated. The newly annotated data is input into the stereo positioning model for iterative training, and the three-dimensional coordinate calculation parameters are updated so that the subsequent lesion positioning error is reduced compared with the initial training.

[0112] This embodiment introduces a continuously self-evolving capability into the stereotactic localization module. Traditional models remain fixed after training, while this system, through the integration of an incremental learning mechanism, can utilize new data constantly generated in clinical work (especially the gold standard of biopsy pathology results), enabling the lesion localization accuracy to continuously improve with the increase of usage time.

[0113] Data association occurs when a doctor performs a biopsy on a case diagnosed with system assistance. The system then obtains the pathological results of that case (e.g., "invasive ductal carcinoma" or "benign fibroadenoma"). These pathological results are automatically associated and labeled with the previously generated lesion location coordinates and image features, forming a high-quality new training sample with final validation.

[0114] Iterative training involves periodic (e.g., quarterly) or, after accumulating a certain number of new samples, initiating an offline training process. Newly labeled data is mixed with existing training data to incrementally train the stereo localization model (including 3D coordinate calculation and the 3D CNN model). During training, techniques to prevent catastrophic forgetting are employed to ensure that the model learns new knowledge without forgetting a large amount of previously acquired knowledge.

[0115] After iterative training, the parameters in the model are updated. The new model is able to more accurately understand the true 3D location corresponding to different image features. The longer the system is used and the more biopsy feedback data is accumulated, the lower the system's localization error will be, forming a positive feedback loop that gets smarter with use.

[0116] In some embodiments, the feature recognition module employs an attention mechanism neural network, adding a spatial attention module and a channel attention module after the convolutional layer; the spatial attention module focuses on the region where the high-density shadow is located by generating a two-dimensional attention map, and the channel attention module enhances the response intensity of calcification points and radial structures by adjusting the feature channel weights; the interference of background noise such as adipose tissue is suppressed through the dual attention mechanism.

[0117] This embodiment represents another in-depth optimization of the feature recognition module. It embeds an attention mechanism into the CNN, mimicking the attention allocation when a doctor reads images, allowing the network to actively and selectively "focus" on suspected lesion areas while ignoring irrelevant background interference.

[0118] Network module integration involves adding two parallel attention modules after the convolutional layers of the CNN. The spatial attention module learns to generate a two-dimensional weight map with the same spatial dimensions as the feature map. Each pixel value in this map represents the importance of that location. Regions where lesions are likely to occur (such as high-density areas) are assigned high weights, while homogeneous fat regions are assigned low weights. The original feature map is then multiplied point-by-point with this weight map. As a result, the features of lesion regions are enhanced, while the features of background regions are suppressed.

[0119] Channel attention modules (such as SENet) analyze the importance of each feature channel. For example, some convolutional kernels may be specifically responsible for detecting "star spikes," and these channels will be given high weights.

[0120] The dual attention mechanism involves spatial attention telling the network "where to look" and channel attention telling the network "what features to look at." Through this mechanism, the network can: enhance calcification responses: even in noisy environments, channels associated with tiny point-like features are amplified. Suppress background noise: background elements such as adipose tissue are suppressed in both spatial and channel dimensions, making it easier to correctly identify genuine lesions in subsequent classification and segmentation.

[0121] In some embodiments, the system includes a real-time feedback optimization module. When the lesion display effect of the synthesized image is marked, the real-time feedback optimization module collects the marking position and feedback type to generate feedback data. The feedback data is then input into a reinforcement learning model to generate a selection area fusion parameter adjustment strategy for the case. Based on the selection area fusion parameter adjustment strategy, the key selection area range and image synthesis weight are dynamically adjusted to generate an optimized synthesized image.

[0122] This embodiment adds a real-time feedback optimization module at the system level, forming a closed loop. It allows doctors to evaluate the system's output in real time, and the system uses this evaluation to dynamically adjust the treatment strategy for the current case through reinforcement learning algorithms, achieving personalized "diagnosis and optimization" services.

[0123] Feedback data collection is achieved through a simple interactive interface on the workstation where doctors review the synthesized images. Doctors can click on the image, mark an area as "unclear" or "needs more context," and specify the desired type of optimization.

[0124] The reinforcement learning framework is constructed as follows: Agent: The region fusion module of this system. Environment: The current breast tomographic image data and the doctor's feedback. Action: Adjusting the parameters of the region fusion, such as the range of the key selection area, the fusion weight of different layers of images, and the parameters of the synthesis algorithm. Reward: The system generates reward signals based on the doctor's feedback. For example, a negative reward is given if the doctor marks "unclear"; a positive reward is given if the doctor confirms "image satisfactory".

[0125] Policy generation and dynamic adjustment are achieved by using a reinforcement learning model to learn and generate a parameter-adjusting policy based on the current "state" (image features and feedback) and the obtained "reward".

[0126] The optimized image is generated, and the selection and fusion process is immediately re-executed based on the new strategy. For example, if a doctor points out that the edges of a cluster of calcifications are unclear, the reinforcement learning model might decide to expand the current selection area and increase the fusion weight of the layer containing the edges. The system then pushes the optimized synthetic image back to the doctor. This process can iterate rapidly until the doctor is satisfied with the image quality. This real-time optimization capability greatly enhances the system's usability and user experience.

[0127] It should be noted that the acquisition of any information mentioned in the system is in accordance with relevant regulations and with the user's consent, and will not infringe on the user's privacy or violate relevant laws and regulations.

[0128] This application provides an embodiment of a mammography-assisted diagnostic method based on feature image selection region fusion. The execution device of the method is a computer device deployed in the mammography-assisted diagnostic system based on feature image selection region fusion provided in any embodiment of this application.

[0129] The provided method includes steps S101 to S103. The computer device can be a handheld terminal, a laptop computer, a wearable device, or a robot, etc. This is used to implement steps S101 to S103 and their corresponding embodiments.

[0130] It should be noted that the acquisition of any information mentioned in the provided methods is in compliance with relevant regulations and is carried out with the user's consent, and will not infringe on the user's privacy or violate relevant laws and regulations.

[0131] Step S101. Obtain breast X-ray tomographic images; for breast cancer calcifications, radial structures, and high-density circular lesions in the breast X-ray tomographic images, use deep learning methods for feature recognition and labeling;

[0132] Step S102. The selection area fusion module is used to synthesize images of key selection areas based on the results of feature recognition and labeling, obtain synthesized images, reduce the absolute number of images, and retain diagnostic accuracy and sensitivity;

[0133] Step S103. Stereoscopic localization module, used to locate the synthesized image using stereoscopic coordinates, output the lesion location, dynamically optimize the 3D occupancy of the lesion in the synthesized image based on the lesion location, analyze and deepen the learning of the density shadows of surrounding blood vessels and connective tissue, and output auxiliary diagnostic information.

[0134] In some embodiments, the feature recognition and labeling of breast cancer calcifications, radial structures, and high-density circular lesions in the mammogram using deep learning methods includes: inputting the preprocessed mammogram into a pre-trained convolutional neural network model, which is trained on a breast image dataset containing annotations of calcifications, radial structures, and circular lesions; the neural network model extracts features from pixel regions in the mammogram per region, and identifies the shape, edge, and texture features of the high-density lesions through multiple convolutional and pooling layers; generating feature bounding boxes based on feature matching results, labeling the location and type of calcifications, radial structures, and circular lesions, and outputting a feature coordinate map.

[0135] In some embodiments, the step of synthesizing images of key selected regions based on the results of feature recognition and labeling to obtain synthesized images includes: determining the smallest image region containing feature-marked boxes as key selected regions based on the feature coordinate map; inputting multi-layer tomographic images corresponding to the key selected regions into a generative adversarial network model, wherein the generative adversarial network model learns the differences in pixel distribution between normal breast tissue and lesion regions through adversarial training; and performing weighted fusion of multi-layer images within the key selected regions according to the density threshold and spatial distribution of feature markers, retaining high-density shadow details corresponding to lesion features, suppressing background noise, and generating a synthesized image containing lesion features, thereby reducing the number of synthesized images compared to the original tomographic images.

[0136] In some embodiments, the step of locating the lesion location from the synthesized image using stereo coordinates includes: calibrating the synthesized image using multi-view camera parameters to obtain image pairs taken from different angles; calculating the disparity information of the lesion features in images from different viewpoints using a stereo matching algorithm, and constructing a three-dimensional coordinate system by combining the calibrated camera intrinsic and extrinsic parameters; calculating the three-dimensional coordinates of the center point of the lesion feature marker box using triangulation to generate lesion location data containing X, Y, and Z axis coordinate values.

[0137] In some embodiments, the step of dynamically optimizing the 3D occupancy of the lesion in the synthesized image based on the lesion location, analyzing and deepening the learning of the density shadows of surrounding blood vessels and connective tissue, and outputting auxiliary diagnostic information includes: inputting the three-dimensional coordinates of the lesion location into a three-dimensional convolutional neural network model, extracting three-dimensional voxel features from the tomographic images within a preset radius around the lesion; the three-dimensional convolutional neural network model identifies the tubular structural features of blood vessels and the grid-like density distribution of connective tissue, generating a spatial relationship heatmap of the lesion and surrounding tissues; adjusting the 3D occupancy boundary of the lesion in the synthesized image based on the blood vessel distribution density and the degree of connective tissue adhesion in the spatial relationship heatmap, and outputting a three-dimensional model of the lesion containing the feature weights of surrounding tissues and interventional path suggestions.

[0138] In some embodiments, an adaptive threshold segmentation algorithm is introduced into the feature recognition module. The adaptive threshold segmentation algorithm automatically calculates the optimal segmentation threshold based on the overall grayscale distribution of the image through a machine learning model. The optimal segmentation threshold is used to binarize the mammogram X-ray tomographic image to highlight the contrast between high-density shadows and normal tissue. The processed image is then input into the feature recognition model to improve the detection sensitivity of small lesions such as calcifications.

[0139] In some embodiments, the selected region fusion module employs a multi-task learning neural network to simultaneously achieve lesion feature preservation and image noise reduction; the network sets a feature preservation loss function and a noise suppression loss function, and optimizes the network parameters through backpropagation; when synthesizing images, the weights of the loss function are dynamically adjusted according to the priority of feature labels.

[0140] In some embodiments, the stereo positioning module integrates an incremental learning mechanism. When a new biopsy pathology result is obtained, the image features corresponding to the pathology result are associated with the positioning coordinates and annotated. The newly annotated data is input into the stereo positioning model for iterative training, and the three-dimensional coordinate calculation parameters are updated so that the subsequent lesion positioning error is reduced compared with the initial training.

[0141] In some embodiments, the feature recognition module employs an attention mechanism neural network, adding a spatial attention module and a channel attention module after the convolutional layer; the spatial attention module focuses on the region where the high-density shadow is located by generating a two-dimensional attention map, and the channel attention module enhances the response intensity of calcification points and radial structures by adjusting the feature channel weights; the interference of background noise such as adipose tissue is suppressed through the dual attention mechanism.

[0142] In some embodiments, the system includes a real-time feedback optimization module. When the lesion display effect of the synthesized image is marked, the real-time feedback optimization module collects the marking position and feedback type to generate feedback data. The feedback data is then input into a reinforcement learning model to generate a selection area fusion parameter adjustment strategy for the case. Based on the selection area fusion parameter adjustment strategy, the key selection area range and image synthesis weight are dynamically adjusted to generate an optimized synthesized image.

[0143] It should be noted that those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the mammography-assisted diagnostic device and its modules based on feature image selection fusion described above can be referred to the corresponding process in the embodiment of the mammography-assisted diagnostic system based on feature image selection fusion described in any embodiment of this application, and will not be repeated here.

[0144] The provided mammography-assisted diagnostic method based on feature image selection fusion can be implemented as a computer program that can run on the provided device.

[0145] The computer device provided in this application includes a processor, a memory, and a network interface connected via a device bus, wherein the memory may include a storage medium and internal memory.

[0146] The storage medium may store operating devices and computer programs. The computer program includes program instructions that, when executed, cause the processor to perform any embodiment of a mammography-assisted diagnostic method based on feature image selection region fusion.

[0147] The processor provides computing and control capabilities, supporting the operation of the entire computer device.

[0148] Internal memory provides an environment for the execution of computer programs in non-volatile storage media. When executed by a processor, the computer program enables the processor to execute any method of a mammography-assisted diagnostic system based on feature image selection region fusion.

[0149] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that a specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0150] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0151] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps:

[0152] Obtain mammograms; use deep learning methods to identify and label features such as calcifications, radial structures, and high-density circular lesions in the mammograms.

[0153] Based on the results of feature recognition and labeling, image synthesis is performed on key selected areas to obtain synthesized images, reducing the absolute number of images while preserving diagnostic accuracy and sensitivity.

[0154] The system locates the lesion in the synthesized image using stereoscopic coordinates, dynamically optimizes the 3D occupancy of the lesion in the synthesized image based on the lesion location, analyzes and deepens the learning of the density shadows of surrounding blood vessels and connective tissue, and outputs auxiliary diagnostic information.

[0155] In some embodiments, the feature recognition and labeling of breast cancer calcifications, radial structures, and high-density circular lesions in the mammogram using deep learning methods includes: inputting the preprocessed mammogram into a pre-trained convolutional neural network model, which is trained on a breast image dataset containing annotations of calcifications, radial structures, and circular lesions; the neural network model extracts features from pixel regions in the mammogram per region, and identifies the shape, edge, and texture features of the high-density lesions through multiple convolutional and pooling layers; generating feature bounding boxes based on feature matching results, labeling the location and type of calcifications, radial structures, and circular lesions, and outputting a feature coordinate map.

[0156] In some embodiments, the step of synthesizing images of key selected regions based on the results of feature recognition and labeling to obtain synthesized images includes: determining the smallest image region containing feature-marked boxes as key selected regions based on the feature coordinate map; inputting multi-layer tomographic images corresponding to the key selected regions into a generative adversarial network model, wherein the generative adversarial network model learns the differences in pixel distribution between normal breast tissue and lesion regions through adversarial training; and performing weighted fusion of multi-layer images within the key selected regions according to the density threshold and spatial distribution of feature markers, retaining high-density shadow details corresponding to lesion features, suppressing background noise, and generating a synthesized image containing lesion features, thereby reducing the number of synthesized images compared to the original tomographic images.

[0157] In some embodiments, the step of locating the lesion location from the synthesized image using stereo coordinates includes: calibrating the synthesized image using multi-view camera parameters to obtain image pairs taken from different angles; calculating the disparity information of the lesion features in images from different viewpoints using a stereo matching algorithm, and constructing a three-dimensional coordinate system by combining the calibrated camera intrinsic and extrinsic parameters; calculating the three-dimensional coordinates of the center point of the lesion feature marker box using triangulation to generate lesion location data containing X, Y, and Z axis coordinate values.

[0158] In some embodiments, the step of dynamically optimizing the 3D occupancy of the lesion in the synthesized image based on the lesion location, analyzing and deepening the learning of the density shadows of surrounding blood vessels and connective tissue, and outputting auxiliary diagnostic information includes: inputting the three-dimensional coordinates of the lesion location into a three-dimensional convolutional neural network model, extracting three-dimensional voxel features from the tomographic images within a preset radius around the lesion; the three-dimensional convolutional neural network model identifies the tubular structural features of blood vessels and the grid-like density distribution of connective tissue, generating a spatial relationship heatmap of the lesion and surrounding tissues; adjusting the 3D occupancy boundary of the lesion in the synthesized image based on the blood vessel distribution density and the degree of connective tissue adhesion in the spatial relationship heatmap, and outputting a three-dimensional model of the lesion containing the feature weights of surrounding tissues and interventional path suggestions.

[0159] In some embodiments, an adaptive threshold segmentation algorithm is introduced into the feature recognition module. The adaptive threshold segmentation algorithm automatically calculates the optimal segmentation threshold based on the overall grayscale distribution of the image through a machine learning model. The optimal segmentation threshold is used to binarize the mammogram X-ray tomographic image to highlight the contrast between high-density shadows and normal tissue. The processed image is then input into the feature recognition model to improve the detection sensitivity of small lesions such as calcifications.

[0160] In some embodiments, the selected region fusion module employs a multi-task learning neural network to simultaneously achieve lesion feature preservation and image noise reduction; the network sets a feature preservation loss function and a noise suppression loss function, and optimizes the network parameters through backpropagation; when synthesizing images, the weights of the loss function are dynamically adjusted according to the priority of feature labels.

[0161] In some embodiments, the stereo positioning module integrates an incremental learning mechanism. When a new biopsy pathology result is obtained, the image features corresponding to the pathology result are associated with the positioning coordinates and annotated. The newly annotated data is input into the stereo positioning model for iterative training, and the three-dimensional coordinate calculation parameters are updated so that the subsequent lesion positioning error is reduced compared with the initial training.

[0162] In some embodiments, the feature recognition module employs an attention mechanism neural network, adding a spatial attention module and a channel attention module after the convolutional layer; the spatial attention module focuses on the region where the high-density shadow is located by generating a two-dimensional attention map, and the channel attention module enhances the response intensity of calcification points and radial structures by adjusting the feature channel weights; the interference of background noise such as adipose tissue is suppressed through the dual attention mechanism.

[0163] In some embodiments, the system includes a real-time feedback optimization module. When the lesion display effect of the synthesized image is marked, the real-time feedback optimization module collects the marking position and feedback type to generate feedback data. The feedback data is then input into a reinforcement learning model to generate a selection area fusion parameter adjustment strategy for the case. Based on the selection area fusion parameter adjustment strategy, the key selection area range and image synthesis weight are dynamically adjusted to generate an optimized synthesized image.

[0164] It should be noted that those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the processor described above can be referred to the corresponding process in the method embodiments of the above embodiments, and will not be repeated here.

[0165] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement the steps of the mammography-assisted diagnosis method based on feature image selection region fusion provided in the above embodiments of this application.

[0166] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.

[0167] It should be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. It should be understood that when an element or layer is referred to as “on,” “adjacent to,” “connected to,” or “coupled to” other elements or layers, it may be directly on, adjacent to, connected to, or coupled to other elements or layers, or there may be intervening elements or layers. Conversely, when an element is referred to as “directly on,” “directly adjacent to,” “directly connected to,” or “directly coupled to” other elements or layers, there are no intervening elements or layers. It should be understood that although the terms first, second, third, etc., may be used to describe various elements, components, areas, layers, and / or portions, these elements, components, areas, layers, and / or portions should not be limited by these terms. These terms are merely used to distinguish one element, component, area, layer, or portion from another element, component, area, layer, or portion. Therefore, without departing from the teachings of this application, the first element, component, area, layer, or portion discussed below may be referred to as a second element, component, area, layer, or portion.

[0168] Spatial relation terms such as “below,” “under,” “below,” “under,” “above,” “above,” etc., are used herein for convenience of description to describe the relationship between one element or feature shown in the figure and other elements or features. It should be understood that, in addition to the orientation shown in the figure, spatial relation terms are intended to also include different orientations of the device in use and operation. For example, if the device in the figure is flipped, then the element or feature described as “below,” “under,” or “below” other elements or features will be oriented “above” other elements or features. Therefore, the exemplary terms “below” and “under” can include both above and below orientations. The device may be otherwise oriented (rotated 90 degrees or otherwise) and the spatial descriptive terms used herein will be interpreted accordingly.

[0169] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. When used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising” and / or “including,” when used in this specification, identify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups. When used herein, the term “and / or” includes any and all combinations of the associated listed items.

[0170] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0171] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A mammography-assisted diagnostic system based on feature image region fusion, characterized in that, include: The image acquisition module is used to acquire mammograms. The feature recognition module is used to identify and label breast cancer calcifications, radial structures, and high-density circular lesions in the breast X-ray tomographic images using deep learning methods. The selected area fusion module is used to synthesize images of key selected areas based on the results of feature recognition and labeling, thereby reducing the absolute number of images and preserving diagnostic accuracy and sensitivity. The stereo positioning module is used to locate the synthesized image using stereo coordinates, output the lesion location, dynamically optimize the 3D occupancy of the lesion in the synthesized image based on the lesion location, analyze and deepen the learning of the density shadows of surrounding blood vessels and connective tissue, and output auxiliary diagnostic information.

2. The system according to claim 1, characterized in that, The method of using deep learning to identify and label breast cancer calcifications, radial structures, and high-density circular lesions in the mammogram includes: The preprocessed mammogram X-ray tomographic images are input into a pre-trained convolutional neural network model, which is trained using a mammogram dataset containing annotations of calcification points, radial structures, and circular lesions. The neural network model extracts features from pixel regions in mammograms by region, identifies the shape, edge, and texture features of high-density shadows through multiple convolutional and pooling layers, generates feature bounding boxes based on feature matching results, marks the location and type of calcifications, radial structures, and circular lesions, and outputs feature coordinate maps.

3. The system according to claim 1, characterized in that, The step of performing image synthesis on the key selected area based on the results of feature recognition and labeling to obtain a synthesized image includes: Based on the feature coordinate map, the smallest image region containing the feature marker box is determined as the key selection area; The multi-layer tomographic images corresponding to the key selected areas are input into the generative adversarial network model, and the generative adversarial network model learns the differences in pixel distribution between normal breast tissue and lesion areas through adversarial training. Based on the density threshold and spatial distribution of feature markers, multi-layer images within the key selected area are weighted and fused to preserve the high-density shadow details corresponding to the lesion features, suppress background noise, and generate a synthetic image containing lesion features, thereby reducing the number of synthetic images compared to the original tomographic images.

4. The system according to claim 1, characterized in that, The step of locating the lesion position from the synthesized image using stereoscopic coordinates includes: The synthesized image is calibrated using multi-view camera parameters to obtain image pairs taken from different angles; The disparity information of lesion features in images from different viewpoints is calculated using a stereo matching algorithm, and a three-dimensional coordinate system is constructed by combining the calibrated camera intrinsic and extrinsic parameters. The three-dimensional coordinates of the center point of the lesion feature marker box are calculated by triangulation, generating lesion location data containing X, Y and Z axis coordinate values.

5. The system according to claim 1, characterized in that, The process involves dynamically optimizing the 3D occupancy of the lesion in the synthesized image based on its location, analyzing and deepening the learning of surrounding blood vessels and connective tissue density shadows, and outputting auxiliary diagnostic information, including: The three-dimensional coordinates of the lesion location are input into a three-dimensional convolutional neural network model to extract three-dimensional voxel features from the tomographic images within a preset radius around the lesion; the three-dimensional convolutional neural network model identifies the tubular structural features of blood vessels and the grid-like density distribution of connective tissue, and generates a spatial relationship heatmap between the lesion and surrounding tissues. Based on the vascular distribution density and connective tissue adhesion in the spatial relationship heatmap, the 3D occupancy boundary of the lesion in the synthetic image is adjusted, and a 3D model of the lesion containing the feature weights of surrounding tissues and intervention path suggestions are output.

6. The system according to claim 1, characterized in that, By introducing an adaptive threshold segmentation algorithm into the feature recognition module, the adaptive threshold segmentation algorithm automatically calculates the optimal segmentation threshold based on the overall grayscale distribution of the image through a machine learning model; the optimal segmentation threshold is used to binarize the mammogram X-ray tomographic image to highlight the contrast between high-density shadows and normal tissue; the processed image is then input into the feature recognition model to improve the detection sensitivity of small lesions such as calcifications.

7. The system according to claim 1, characterized in that, The selected area fusion module employs a multi-task learning neural network to simultaneously preserve lesion features and reduce image noise. The network is configured with a feature preservation loss function and a noise suppression loss function, and the network parameters are optimized through backpropagation. When synthesizing images, the weights of the loss functions are dynamically adjusted according to the priority of feature labels.

8. The system according to claim 1, characterized in that, The stereo positioning module integrates an incremental learning mechanism. When a new biopsy pathology result is obtained, the image features corresponding to the pathology result are associated with the positioning coordinates and labeled. The newly labeled data is input into the stereo positioning model for iterative training, and the three-dimensional coordinate calculation parameters are updated, so that the subsequent lesion positioning error is reduced compared with the initial training.

9. The system according to claim 1, characterized in that, The feature recognition module employs an attention mechanism neural network, adding a spatial attention module and a channel attention module after the convolutional layer. The spatial attention module focuses on the region where high-density shadows are located by generating a two-dimensional attention map, while the channel attention module enhances the response intensity of calcification points and radial structures by adjusting the feature channel weights. The dual attention mechanism suppresses the interference of background noise such as adipose tissue.

10. The system according to claim 1, characterized in that, The system is equipped with a real-time feedback optimization module. When the lesion display effect of the synthesized image is marked, the real-time feedback optimization module collects the marking position and feedback type to generate feedback data. The feedback data is input into the reinforcement learning model to generate a selection region fusion parameter adjustment strategy for the case. The selection area range and image synthesis weight are dynamically adjusted according to the selection area fusion parameter adjustment strategy to generate an optimized synthesized image.