Medical image segmentation method and system based on machine learning
Through the methods of multimodal medical image preprocessing, federated learning and attention mechanism optimization, the problems of insufficient robustness and automation in medical image segmentation are solved, and high precision, small sample adaptability and real-time segmentation are achieved, which complies with data privacy protection standards.
Patent Information
- Application Number
- CN202510812657.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-26
AI Technical Summary
Existing medical image segmentation methods have deficiencies in robustness, generalization ability and automation, especially in the problems of multimodal image fusion, data scarcity and high computational complexity.
A machine learning-based approach, including multimodal medical image preprocessing, federated learning small sample model construction, multi-scale feature extraction, progressive decoder optimization, and morphological post-processing, is used in combination with attention mechanism and loss function optimization to achieve high-precision segmentation.
It improves segmentation accuracy, enhances the generalization ability and automation level of the model, reduces doctors' manual operation time, ensures data privacy protection, and is suitable for small sample scenarios and real-time diagnosis.
Smart Images

Figure CN120707575A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and machine learning, and in particular to a medical image segmentation method and system based on machine learning. Background Art
[0002] Medical image segmentation is a key technology for extracting regions of interest (ROIs) from medical images, and its accuracy directly affects the reliability of clinical decision-making. Traditional medical image segmentation methods mainly include image processing-based algorithms such as threshold segmentation, region growing, and edge detection. However, these methods have significant limitations, including the following defects: (1) Lack of robustness: Affected by factors such as medical image noise, grayscale inhomogeneity (such as field strength deviation in MRI), and low tissue contrast, traditional algorithms find it difficult to stably segment the target region in complex scenarios. For example, in lung CT images, the grayscale difference between ground-glass nodules and normal lung tissue is subtle, and traditional threshold methods can easily lead to blurred or missed segmentation boundaries; (2) Weak generalization ability: There are individual differences in the anatomical structures of different patients (such as the morphological diversity of brain tumors). Traditional algorithms rely on manually designed features (such as texture and shape), which are difficult to adapt to the changing characteristics of medical images. They need to be repeatedly adjusted for different cases, which is inefficient; (3) Low degree of automation: In clinical practice, doctors often need to manually outline the lesion contours, which is time-consuming and laborious (such as tumor segmentation in liver MRI takes more than 30 minutes), and the segmentation results are significantly affected by subjective experience, making it difficult to meet the needs of large-scale image screening.
[0003] In recent years, machine learning technology based on deep learning has shown potential in the field of medical image segmentation, but it still faces the following challenges: (1) Strong data dependence: high-quality medical annotation data is scarce (for example, pathological sections require experts to annotate pixel by pixel), and the model is prone to overfitting in small sample scenarios and has insufficient generalization ability; (2) Difficulty in multimodal fusion: Clinical diagnosis often combines multimodal images such as CT (displaying anatomical structure) and PET (reflecting metabolic activity), but the effectiveness of existing algorithms for cross-modal feature fusion needs to be improved; (3) Balance between real-time and accuracy: The amount of data in three-dimensional medical images (such as whole-body CT) is huge, and high-precision segmentation models (such as 3D U-Net) have high computational complexity, which makes it difficult to meet real-time diagnosis needs.
[0004] Based on this, it is of great practical significance to provide a segmentation method and system that combines the advantages of machine learning and adapts to the characteristics of medical images. Summary of the Invention
[0005] In view of this, the present invention proposes a medical image segmentation method and system based on machine learning, aiming to solve the problems of low segmentation accuracy, weak generalization ability and insufficient automation in current technology.
[0006] The present invention proposes a medical image segmentation method based on machine learning, comprising the following steps:
[0007] S1: Preprocessing the multimodal medical image to obtain preprocessed image data;
[0008] S2: Build a small sample model based on federated learning;
[0009] S3: Perform multi-scale feature extraction and use the attention mechanism to optimize feature weight distribution;
[0010] S4: Optimize using a progressive decoder and loss function;
[0011] S5: Output the segmentation results and perform post-processing through morphological optimization.
[0012] Preferably, the multimodal medical images include CT, MRI and PET original images;
[0013] The preprocessing includes: performing grayscale normalization on the multimodal medical images to eliminate imaging differences between devices; using non-local mean filtering or bilateral filtering to reduce the noise of the multimodal medical images while retaining edge details; performing rigid or non-rigid registration on the multimodal medical images, and constructing a joint feature space through feature-level fusion.
[0014] Preferably, the construction of a small sample model based on federated learning includes:
[0015] Data partitioning: The labeled medical image data is partitioned into a support set and a query set. The support set contains labeled samples, and the query set contains samples to be segmented.
[0016] Build a small sample model based on federated learning: The client uses local small sample data to train a lightweight segmentation model and only uploads model parameter updates to the server. The server aggregates multi-center model parameters to generate a global model, which is returned to each client for a new round of training, and the cycle continues until convergence.
[0017] Meta-learning optimization: A meta-learning algorithm is introduced to enable the model to quickly adapt to the segmentation task of new types of lesions.
[0018] Preferably, performing multi-scale feature extraction and optimizing feature weight distribution using an attention mechanism includes:
[0019] An improved ResNet-50 is used as the encoder, which outputs multi-scale feature maps at different levels (Conv2-Conv5) to capture features from local details to global structures.
[0020] A spatial attention map is generated through convolution operations to suppress background noise and enhance the feature response of the target area; global average pooling and fully connected layers are used to automatically calibrate the importance of each channel feature.
[0021] Preferably, the optimization using a progressive decoder and a loss function includes:
[0022] It uses skip connections to fuse shallow detail features and deep semantic features of the encoder, gradually restores the image resolution through transposed convolution, and finally outputs a segmentation mask of the same size as the input image, with pixel values 0 / 1 representing background / foreground.
[0023] The loss function optimization combines cross entropy loss and Dice loss.
[0024] Preferably, the morphological optimization includes: performing operations such as dilation and erosion on the segmentation results to eliminate isolated noise points and smooth boundaries.
[0025] Preferably, the post-processing further comprises: performing uncertainty assessment and clinical verification;
[0026] The uncertainty assessment is as follows: estimating the uncertainty of the segmentation result through Monte Carlo Dropout (MC Dropout), generating a confidence map, and assisting doctors in judging the reliability of the segmentation.
[0027] Preferably, the clinical verification is as follows: comparing the segmentation results with the manual annotation results of experts, calculating evaluation indicators such as the Dice similarity coefficient DSC and the Jaccard index, and outputting the final segmentation result if DSC ≥ 0.9; otherwise, returning to the model for retraining.
[0028] The present invention also provides a medical image segmentation system based on machine learning, comprising:
[0029] Data preprocessing module: performs standardization, noise reduction, registration and fusion operations on multimodal medical images and outputs preprocessed image data;
[0030] Federated learning training module: supports distributed training of small sample data from multiple centers, and generates segmentation models with strong generalization capabilities through meta-learning and model parameter aggregation;
[0031] Feature extraction and segmentation module: Based on the improved ResNet-UNet network, it integrates multi-scale feature extraction and attention mechanism to achieve pixel-level segmentation prediction;
[0032] Post-processing and evaluation module: performs morphological optimization and uncertainty analysis on the segmentation results, and outputs clinically applicable accurate segmentation masks and evaluation reports.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] (1) The present invention can achieve high-precision segmentation. Through multimodal feature fusion and attention mechanism, it can effectively capture the subtle differences between tumors and surrounding tissues. Compared with the traditional U-Net algorithm, the DSC is improved by 8%-12%, especially in the segmentation of small lesions.
[0035] (2) The present invention also has small sample adaptability. Based on federated learning and meta-learning, the model can converge quickly with 5-10 labeled samples, alleviating the problem of medical data scarcity and being suitable for the diagnosis of rare diseases.
[0036] (3) In addition, the method described in the present invention integrates automated preprocessing, real-time segmentation and uncertainty assessment functions, reducing doctors' manual operation time by more than 70%, while reducing the risk of misdiagnosis through confidence prompts; further, the federated learning framework supports model training under cross-hospital data privacy protection, avoiding the leakage of sensitive medical data, and complying with GDPR and HIPAA regulations. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0038] Figure 1 This is a flow chart of the medical image segmentation method based on machine learning according to the present invention;
[0039] Figure 2 Schematic diagram of the medical image segmentation system based on machine learning described in the present invention. DETAILED DESCRIPTION
[0040] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that, unless there is a conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0041] In the description of this application, it should be understood that the terms "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on this application.
[0042] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature specified as "first" or "second" may explicitly or implicitly include one or more of such features. Throughout this application, unless otherwise specified, "plurality" means two or more.
[0043] In the description of this application, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in this application based on the specific circumstances.
[0044] The present invention provides a medical image segmentation method based on machine learning, comprising the following steps:
[0045] S1: Preprocessing the multimodal medical image to obtain preprocessed image data;
[0046] S2: Build a small sample model based on federated learning;
[0047] S3: Perform multi-scale feature extraction and use the attention mechanism to optimize feature weight distribution;
[0048] S4: Optimize using a progressive decoder and loss function;
[0049] S5: Output the segmentation results and perform post-processing through morphological optimization.
[0050] In the present invention, the multimodal medical images include original images such as CT, MRI, and PET;
[0051] The preprocessing includes: performing grayscale normalization on the multimodal medical images to eliminate imaging differences between devices; using non-local mean filtering or bilateral filtering to reduce the noise of the multimodal medical images while retaining edge details; performing rigid or non-rigid registration on the multimodal medical images, and constructing a joint feature space through feature-level fusion.
[0052] The preprocessing described in the present invention can solve the spatial misalignment and noise interference of cross-modal images. It adopts a combination of bilateral filtering and non-local mean filtering. Bilateral filtering uses dual weights of spatial distance and grayscale difference to suppress MRI magnetic susceptibility artifacts while retaining the tumor-normal tissue boundary (edge blurring is reduced by 40%); non-local mean filtering uses image block similarity to maintain the edge clarity of small lung nodules (diameter ≤ 5mm) in low-dose CT noise reduction. Traditional Gaussian filtering easily causes blurring of small targets.
[0053] In the present invention, the construction of a small sample model based on federated learning includes:
[0054] Data partitioning: The labeled medical image data is partitioned into a support set and a query set. The support set contains labeled samples, and the query set contains samples to be segmented.
[0055] Build a small sample model based on federated learning: The client uses local small sample data to train a lightweight segmentation model and only uploads model parameter updates to the server. The server aggregates multi-center model parameters to generate a global model, which is returned to each client for a new round of training, and the cycle continues until convergence.
[0056] Meta-learning optimization: The introduction of meta-learning algorithms enables the model to quickly adapt to the segmentation tasks of new types of lesions and reduce dependence on large-scale labeled data.
[0057] This invention adopts a federated learning architecture to solve the problem of data authenticity and privacy protection of medical annotation. Hospitals in various places only upload model parameter updates and do not share original patient imaging data. This complies with HIPAA / GDPR regulations and solves the privacy leakage risks of traditional centralized training.
[0058] The meta-learning algorithm can quickly adapt to new tasks, so that the model can converge with only 1 to 2 gradients under few samples.
[0059] In the present invention, the multi-scale feature extraction and the optimization of feature weight distribution using the attention mechanism include:
[0060] An improved ResNet-50 is used as the encoder, which outputs multi-scale feature maps at different levels (Conv2-Conv5) to capture features from local details to global structures.
[0061] A spatial attention map is generated through convolution operations to suppress background noise and enhance the feature response of the target area; global average pooling and fully connected layers are used to automatically calibrate the importance of each channel feature.
[0062] The present invention performs multi-scale feature extraction and adopts attention mechanism optimization to solve the problems of missed detection of small lesions and background noise interference.
[0063] In the present invention, the optimization using a progressive decoder and a loss function includes:
[0064] It uses skip connections to fuse shallow detail features and deep semantic features of the encoder, gradually restores the image resolution through transposed convolution, and finally outputs a segmentation mask of the same size as the input image, with pixel values 0 / 1 representing background / foreground.
[0065] The loss function optimization combines cross entropy loss and Dice loss.
[0066] The formula of the loss function of the present invention is:
[0067]
[0068] Among them, α is the balance parameter, y is the true mask, is the prediction mask.
[0069] In the present invention, the morphological optimization preferably includes: performing operations such as dilation and erosion on the segmentation results to eliminate isolated noise points and smooth the boundaries.
[0070] In the present invention, the post-processing further includes: performing uncertainty assessment and clinical verification;
[0071] The uncertainty assessment is as follows: estimating the uncertainty of the segmentation result through Monte Carlo Dropout (MC Dropout), generating a confidence map, and assisting doctors in judging the reliability of the segmentation.
[0072] In the present invention, the clinical verification is as follows: comparing the segmentation results with the manual annotation results of experts, calculating evaluation indicators such as the Dice similarity coefficient DSC and the Jaccard index, and outputting the final segmentation results if DSC ≥ 0.9; otherwise, returning to the model for retraining.
[0073] The present invention verifies the reliability of segmentation results and improves the efficiency of manual review through post-processing and uncertainty assessment.
[0074] The present invention also provides a medical image segmentation system based on machine learning, comprising:
[0075] Data preprocessing module: performs standardization, noise reduction, registration and fusion operations on multimodal medical images and outputs preprocessed image data;
[0076] Federated learning training module: supports distributed training of small sample data from multiple centers, and generates segmentation models with strong generalization capabilities through meta-learning and model parameter aggregation;
[0077] Feature extraction and segmentation module: Based on the improved ResNet-UNet network, it integrates multi-scale feature extraction and attention mechanism to achieve pixel-level segmentation prediction;
[0078] Post-processing and evaluation module: performs morphological optimization and uncertainty analysis on the segmentation results, and outputs clinically applicable accurate segmentation masks and evaluation reports.
[0079] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or a combination of software and hardware embodiments. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0080] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0081] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0082] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. A medical image segmentation method based on machine learning, characterized in that: The following steps are involved: S1: Preprocessing the multimodal medical image to obtain preprocessed image data; S2: Build a small sample model based on federated learning; S3: Perform multi-scale feature extraction and use the attention mechanism to optimize feature weight distribution; S4: Optimize using a progressive decoder and loss function; S5: Output the segmentation results and perform post-processing through morphological optimization.
2. The medical image segmentation method based on machine learning according to claim 1, characterized in that: The multimodal medical images include CT, MRI and PET original images; The preprocessing includes: performing grayscale normalization on the multimodal medical images to eliminate imaging differences between devices; using non-local mean filtering or bilateral filtering to reduce the noise of the multimodal medical images while retaining edge details; performing rigid or non-rigid registration on the multimodal medical images, and constructing a joint feature space through feature-level fusion.
3. The medical image segmentation method based on machine learning according to claim 1, characterized in that The construction of a small sample model based on federated learning includes: Data partitioning: The labeled medical image data is partitioned into a support set and a query set. The support set contains labeled samples, and the query set contains samples to be segmented. Build a small sample model based on federated learning: The client uses local small sample data to train a lightweight segmentation model and only uploads model parameter updates to the server. The server aggregates multi-center model parameters to generate a global model, which is returned to each client for a new round of training, and the cycle continues until convergence. Meta-learning optimization: A meta-learning algorithm is introduced to enable the model to quickly adapt to the segmentation task of new types of lesions.
4. The medical image segmentation method based on machine learning according to claim 1, characterized in that The multi-scale feature extraction and the optimization of feature weight distribution using the attention mechanism include: An improved ResNet-50 is used as the encoder to output multi-scale feature maps at different levels, capturing features from local details to global structures. A spatial attention map is generated through convolution operations to suppress background noise and enhance the feature response of the target area; global average pooling and fully connected layers are used to automatically calibrate the importance of each channel feature.
5. The medical image segmentation method based on machine learning according to claim 1, characterized in that: The optimization using a progressive decoder and a loss function includes: It uses skip connections to fuse shallow detail features and deep semantic features of the encoder, gradually restores the image resolution through transposed convolution, and finally outputs a segmentation mask of the same size as the input image, with pixel values 0 / 1 representing background / foreground. The loss function optimization combines cross entropy loss and Dice loss.
6. The medical image segmentation method based on machine learning according to claim 1, characterized in that: The morphological optimization includes: performing expansion and corrosion operations on the segmentation results, eliminating isolated noise points, and smoothing boundaries.
7. The medical image segmentation method based on machine learning according to claim 1, characterized in that: The post-processing also includes: performing uncertainty assessment and clinical validation; The uncertainty assessment is as follows: estimating the uncertainty of the segmentation result through Monte Carlo Dropout (MC Dropout), generating a confidence map, and assisting doctors in judging the reliability of the segmentation.
8. The medical image segmentation method based on machine learning according to claim 7, characterized in that: The clinical verification is as follows: the segmentation results are compared with the manual annotation results of experts, and the Dice similarity coefficient DSC and Jaccard index evaluation indicators are calculated. If DSC ≥ 0.9, the final segmentation result is output; otherwise, the model is returned to retraining.
9. A medical image segmentation system based on machine learning, characterized in that: include: Data preprocessing module: performs standardization, noise reduction, registration and fusion operations on multimodal medical images and outputs preprocessed image data; Federated learning training module: supports distributed training of small sample data from multiple centers, and generates segmentation models with strong generalization capabilities through meta-learning and model parameter aggregation; Feature extraction and segmentation module: Based on the improved ResNet-UNet network, it integrates multi-scale feature extraction and attention mechanism to achieve pixel-level segmentation prediction; Post-processing and evaluation module: performs morphological optimization and uncertainty analysis on the segmentation results, and outputs clinically applicable accurate segmentation masks and evaluation reports.
Citation Information
Cited By
Waste slag field functional area zero sample identification method based on image processing and product
CN121170608A
Image processing-based zero sample identification method for functional areas of a spoil site and product
CN121170608B