Medical image multi-organ automatic segmentation large model method and system
By constructing a deep neural network model, the encoder-decoder bidirectional feature interaction mechanism and semantic perception mechanism are adopted, the category interference problem caused by the difference in the labeling standards of multi-center medical imaging data sets is solved, and the efficiency and accuracy of automatic segmentation of multiple organs is achieved, reducing doctor dependence and enhancing the value of clinical application.
Patent Information
- Application Number
- CN202510719877.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-18
AI Technical Summary
The existing technology has problems such as large demand for high-quality labeling, large differences in labeling standards, and insufficient accuracy and repeatability caused by manual outline reliance on doctor experience in the construction of multi-center medical imaging data sets, which limits the application effect of multi-organ segmentation models.
A deep neural network model is constructed, the encoder-decoder bidirectional feature interaction mechanism is adopted, and the semantic perception mechanism is introduced. Through multi-scale encoder and decoder, semantic guidance encoding module, dynamic filter encoding module and semantic guidance dynamic decoding module, automatic organ segmentation and output organ masks.
It significantly reduces the time and learning costs of doctors during multi-organ segmentation, improves segmentation efficiency and accuracy, and provides reliable technical support for clinical diagnosis.
Smart Images

Figure CN120339263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image organ segmentation, and in particular to a method and system for automatically segmenting multiple organs in medical images using a large model. Background Art
[0002] Accurate organ segmentation technology plays a crucial role in the early diagnosis of diseases, precise staging, and the optimization of radiotherapy plans. However, in actual clinical applications, the construction of multi-center medical image datasets faces significant challenges: Firstly, high-quality annotation requires a large amount of professional human resources, including experienced radiologists and imaging experts; Secondly, there are significant differences in disease spectra, image acquisition protocols, and annotation standards among different medical institutions. These factors together lead to a partially annotated pattern in the dataset. This partially annotated characteristic will cause serious category interference problems during the model training process. These limitations severely restrict the application effect of general mask segmentation models in multi-organ segmentation tasks. There is an urgent need to develop new technical solutions to solve this key problem in order to promote the clinical application and innovative development of medical image analysis technology.
[0003] The literature (Ma J, He Y, Li F, et al. Segment anything in medical images[J]. Nature Communications, 2024, 15(1): 654.) provides a general large model for medical organ segmentation and explores its clinical value. The article proposes an innovative basic model for medical image segmentation, MedSAM, aiming to break through the limitations of existing methods in cross-modal and cross-disease type applications. The development of this model is based on a large-scale medical image dataset containing 1,570,263 image mask pairs, covering 10 different imaging modalities and more than 30 cancer types, ensuring the wide applicability of the model. Through a systematic 146 validation tasks (including 86 internal validations and 60 external validations), MedSAM is significantly superior to traditional modality-based professional models in terms of accuracy and robustness.
[0004] The method has the following main limitations in clinical applications: First, its core operation relies on doctors to manually delineate the category information of different organs in a professional software platform. This human dependence causes the accuracy of the evaluation results to be directly restricted by two key factors: the degree of doctors' clinical experience accumulation and the precision control of the delineation position, which inevitably reduces the repeatability and objectivity of the method. Second, from the operational level, the manual delineation process requires doctors to invest significant learning costs and time costs, including: proficiently mastering software interface operations, understanding the anatomical features of various organs, and precisely controlling professional skills such as delineation. More importantly, there are individual differences among different doctors in terms of delineation habits, judgment criteria, and experience levels. This subjective factor amplifies the deficiencies of the method in terms of consistency and reliability. These limiting factors are superimposed on each other, restricting its promotion and application value and large-scale application prospects in clinical practice. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method and system for automatically segmenting multiple organs in medical images in view of the above-mentioned deficiencies of the prior art. By constructing a deep neural network model, automatic extraction and precise positioning of the category information of the organs to be segmented are realized.
[0006] To solve the above technical problems, the technical solutions adopted by the present invention are as follows: On the one hand, the present invention provides a method for automatically segmenting multiple organs in medical images, including:
[0007] Obtain CT medical images;
[0008] Perform standardization processing of medical image parameters on the obtained CT medical images to construct a CT medical image dataset;
[0009] Divide the CT medical image dataset into a training set, a validation set, and a test set according to a ratio;
[0010] Construct a multi-organ segmentation network model; the multi-organ segmentation network model realizes precise organ segmentation through an encoder-decoder bidirectional feature interaction mechanism;
[0011] Train the multi-organ segmentation network model with the training set, and use the validation set and the test set to verify and test the trained multi-organ segmentation network model;
[0012] Use the trained multi-organ segmentation network model to perform automatic segmentation of multiple organs in medical images and output the delineation results of organ masks.
[0013] Furthermore, the CT medical image is a dcm file compliant with the DICOM (Digital Imaging and Communications in Medicine) protocol, or multi-frame images parsed from the dcm file and their corresponding labeled tags.
[0014] Furthermore, the multi-organ segmentation network model uses the U-Net architecture as the benchmark framework, introduces a semantic perception mechanism, and constructs five core modules, including a multi-scale encoder, a multi-scale decoder, a semantic-guided encoding module, a dynamic filtering encoding module, and a semantic-guided dynamic decoding module;
[0015] The multi-scale encoder uses a multi-scale convolution layer to extract multi-scale features from the input 3D-CT image, constructs a multi-scale feature map, and inputs it into the multi-scale decoder and the semantic guidance encoding module; the multi-scale decoder uses multi-scale upsampling to generate a feature map to be identified that matches the spatial scale of the input image for the multi-scale feature map constructed by the multi-scale encoder; the semantic guidance encoding module applies multi-scale sine and cosine encoding to the multi-scale feature map extracted by the multi-scale encoder to generate a multi-scale position encoding; the multi-scale feature map is combined with the corresponding position encoding to generate a multi-scale feature information sequence; at the same time, a set of learnable parameters is assigned to each learnable organ category as preset organ information; the cross-attention layer is used to cross-fuse the multi-scale feature information sequence with the preset organ information to obtain cross-level fused CT image multi-organ features as multi-organ semantic guidance information; the dynamic filtering encoding module encodes the guidance information extracted by the semantic guidance encoding module into dynamic convolution kernel parameters through parameter partitioning encoding operation; the semantic guidance dynamic decoding module uses the dynamic convolution kernel generated by the dynamic filtering encoding module to perform a convolution operation on the feature map to be identified generated by the multi-scale decoder to achieve semantically guided dynamic decoding and complete the delineation of the organ mask.
[0016] Furthermore, the multi-scale encoder adopts a progressive downsampling strategy and realizes multi-scale feature extraction through multi-level feature extraction units; each level of feature extraction unit includes a residual convolution layer with a 3×3 convolution kernel, a batch normalization layer, a ReLU activation layer and a maximum pooling downsampling layer; the multi-scale encoder extracts multi-scale feature maps from high-resolution detail features to deep semantic features step by step through a hierarchical structure.
[0017] Furthermore, the multi-scale decoder adopts a progressive upsampling strategy to generate a feature map to be identified that matches the input spatial scale through a multi-level feature extraction unit; each level of feature extraction unit includes a residual convolution layer with a 3×3 convolution kernel, a batch normalization layer, a ReLU activation layer, and a transposed convolution upsampling layer; the multi-scale decoder extracts multi-scale feature maps from deep semantic features to high-resolution detail features step by step through a hierarchical structure; and the multi-scale feature map is gradually reconstructed through cross-layer feature splicing and fusion operations.
[0018] Furthermore, the semantically guided encoding module first uses an embedding layer to assign a set of learnable parameters as preset organ information to each learnable organ category, and then uses a three-layer cascaded cross-attention layer to apply multi-scale sine and cosine encoding to the multi-scale feature map extracted by the multi-scale encoder to generate multi-scale position encoding; the multi-scale feature map is combined with the corresponding position encoding to generate a multi-scale feature information sequence; the multi-scale feature information sequence is then cross-fused with the preset organ information to obtain cross-level CT image multi-organ features as multi-organ semantic guidance information; the semantically guided encoding module extracts unified representation information of different organs in different CT images, decouples traditional image feature extraction and organ feature representation tasks, and then obtains semantic guidance information for multiple organs.
[0019] Furthermore, the cross-attention layer includes a cascaded deformable attention fusion layer and a position-aware attention layer; the deformable attention fusion layer utilizes a deformable convolution mechanism to perform deformable convolution offset learning on features of different scales extracted by the multi-scale coding layer, adaptively adjusts the receptive field range, and normalizes the extracted multi-layer features to achieve weighted adaptive fusion of features; the position-aware attention layer extracts information specific to multiple organs in CT images according to their respective key positions in combination with the spatial prior of medical anatomical atlases, thereby enhancing the network's perception of the spatial distribution of organs.
[0020] Furthermore, the dynamic filtering coding module encodes the guidance information extracted by the semantic guidance coding module into dynamic convolution kernel parameters through parameter division coding operation, and the specific method is as follows:
[0021] The semantic guidance information of multiple organs extracted by the semantic guidance encoding module is regarded as a collection of specific information of different organs. First, the input semantic guidance information is linearly transformed in channel dimension through a linear layer to generate dynamic filtering parameters. Then, the generated dynamic filtering parameters are divided into a set of 1×1 convolution kernel parameters used by multiple convolution layers through a parameter partitioning layer to realize dynamic adjustment of the input and output dimensions of the dynamic convolution kernel parameters.
[0022] Furthermore, the semantically guided dynamic decoding module uses a dynamic convolution strategy to construct three dynamic convolution layers, each layer uses a 1×1 convolution kernel parameter generated by the dynamic filter coding module, and performs a dynamic convolution operation on a feature map to be identified that is generated by a multi-scale decoder and matches the spatial scale of the input image; a multi-organ segmentation mask image is generated through three dynamic convolution layers.
[0023] On the other hand, the present invention provides a large model system for automatic segmentation of multiple organs in medical images, including a medical image acquisition module, a data set construction module, a model construction module, a training module and a segmentation module;
[0024] The medical image acquisition module is used to acquire CT medical images;
[0025] The dataset construction module is used to perform standardization processing of medical image parameters on the acquired CT medical images, construct a CT medical image dataset; and divide the CT medical image dataset into a training set, a validation set and a test set according to a ratio;
[0026] The model construction module is used to construct a multi-organ segmentation network model; the multi-organ segmentation network model realizes accurate organ segmentation through an encoder-decoder bidirectional feature interaction mechanism;
[0027] The training module uses the training set to train the multi-organ segmentation network model, and uses the validation set and the test set to verify and test the trained multi-organ segmentation network model;
[0028] The segmentation module uses the trained multi-organ segmentation network model to perform automatic segmentation of multiple organs in medical images and outputs the drawing results of organ masks.
[0029] The beneficial effects produced by adopting the above technical solutions are as follows: A method and system for a large model of automatic multi-organ segmentation of medical images provided by the present invention uses the automatically obtained organ category information as the guiding information of the neural network, effectively solving the problem of category interference caused by partial annotations in traditional methods. Compared with the traditional manual drawing method, the present invention significantly reduces the time cost and learning cost of doctors in the use process, improves the efficiency and accuracy of multi-organ segmentation, and provides more reliable technical support for clinical diagnosis and treatment. Description of the Drawings
[0030] Figure 1 It is a flowchart of a method for a large model of automatic multi-organ segmentation of medical images provided by an embodiment of the present invention;
[0031] Figure 2 It is a structural schematic diagram of a multi-organ segmentation network model provided by an embodiment of the present invention;
[0032] Figure 3 It is an example of training data input into the multi-organ segmentation network model provided by an embodiment of the present invention, wherein (a) is the input CT image sequence, and (b) is the output organ segmentation mask label;
[0033] Figure 4 It is a visualization diagram of the multi-organ drawing result provided by an embodiment of the present invention, wherein (a) is the cross-sectional view, (b) is the sagittal view, (c) is the coronal view, and (d) is the 3D reconstruction view. Detailed Embodiments
[0034] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0035] In this embodiment, a method for a large multi-organ automatic segmentation model of medical images uses chest and abdomen medical images to automatically detect all organs to be delineated in the entire image, so as to delineate the organs matching the current image type, such as Figure 1 shown, and includes the following steps:
[0036] Step S1: Obtain CT medical images; the CT medical images are dcm files that comply with the Digital Imaging and Communications in Medicine (DICOM) protocol, or multi-frame images and their corresponding annotated labels after parsing the dcm files.
[0037] In this embodiment, 3D-CT images of the chest and abdomen are input; the input 3D-CT images of the chest and abdomen are dcm files that comply with the Digital Imaging and Communications in Medicine (DICOM) protocol, or multi-frame images (in nii.gz format) and their corresponding annotated labels (partially annotated or fully annotated) after parsing the dcm files.
[0038] Step S2: Perform standardization processing on the obtained CT medical images, including unified correction of image display parameters such as window width and window level, and spatial attribute parameters such as voxel size and image orientation, to construct a CT medical image dataset;
[0039] Step S3: Divide the CT medical image dataset into a training set, a validation set, and a test set according to a ratio;
[0040] The present invention divides two types of dataset:
[0041] (1) Training and validation datasets: Include CT images with unified information and corresponding labels, and are used to train and validate the deep neural network model;
[0042] (2) Test dataset: Include CT images with unified information and corresponding labels, and are used to clinically evaluate the segmentation DSC index;
[0043] Step S4: Construct a multi-organ segmentation network model; the multi-organ segmentation network model realizes precise organ segmentation through an encoder-decoder bidirectional feature interaction mechanism;
[0044] The multi-organ segmentation network model is based on the U-Net architecture, introduces a semantic perception mechanism, and constructs five core modules, such as Figure 2As shown, it includes a multi-scale encoder (Multi-scale Encoder), a multi-scale decoder (Multi-scale Decoder), a semantic-guided encoding module (Semantic-guided Encoder), a dynamic filtering encoding module (Dynamic Filtering Encoder) and a semantic-guided dynamic decoding module;
[0045] The multi-scale encoder uses a multi-scale convolution layer to extract multi-scale features from the input 3D-CT image, constructs a multi-scale feature map, and inputs it into the multi-scale decoder and the semantic guidance encoding module; the multi-scale decoder uses multi-scale upsampling to generate a feature map to be identified that matches the spatial scale of the input image for the multi-scale feature map constructed by the multi-scale encoder; the semantic guidance encoding module applies multi-scale sine and cosine encoding to the multi-scale feature map extracted by the multi-scale encoder to generate a multi-scale position encoding; the multi-scale feature map is combined with the corresponding position encoding to generate a feature information sequence; at the same time, a set of learnable parameters is assigned to each learnable organ category as preset organ information; the cross-attention layer is used to cross-fuse the multi-scale feature information sequence with the preset organ information to obtain cross-level fused CT image multi-organ features as multi-organ semantic guidance information; the dynamic filtering encoding module encodes the guidance information extracted by the semantic guidance encoding module into dynamic convolution kernel parameters through parameter partitioning encoding operation; the semantic guidance dynamic decoding module uses the dynamic convolution kernel generated by the dynamic filtering encoding module to perform a convolution operation on the feature map to be identified generated by the multi-scale decoder to achieve semantically guided dynamic decoding and complete the delineation of the organ mask.
[0046] In this embodiment, the multi-scale encoder adopts a progressive downsampling strategy and realizes multi-scale feature extraction through multi-level feature extraction units; each level of feature extraction unit includes a residual convolution layer with a 3×3 convolution kernel, a batch normalization layer, a ReLU activation layer and a maximum pooling downsampling layer; the multi-scale encoder extracts multi-scale feature maps from high-resolution detail features to deep semantic features step by step through a hierarchical structure.
[0047] The multi-scale decoder adopts a progressive upsampling strategy and generates a feature map to be identified that matches the input spatial scale through a multi-level feature extraction unit; each level of feature extraction unit includes a residual convolution layer with a 3×3 convolution kernel, a batch normalization layer, a ReLU activation layer, and a transposed convolution upsampling layer; the multi-scale decoder extracts multi-scale feature maps from deep semantic features to high-resolution detail features step by step through a hierarchical structure; and the multi-scale feature map is gradually reconstructed through cross-layer feature splicing and fusion operations.
[0048] The semantically guided encoding module first uses an embedding layer to assign a set of learnable parameters as preset organ information to each learnable organ category, and then uses a three-layer cascaded cross-attention layer to apply multi-scale sine and cosine encoding to the multi-scale feature map extracted by the multi-scale encoder to generate a multi-scale position code; the multi-scale feature map is combined with the corresponding position code to generate a multi-scale feature information sequence; the multi-scale feature information sequence is then cross-fused with the preset organ information to obtain cross-level CT image multi-organ features as multi-organ semantic guidance information; the feature information of each organ in the current input image is extracted through a deformable attention operation; the semantically guided encoding module extracts unified representation information of different organs in different CT images, decouples traditional image feature extraction and organ feature representation tasks, and then obtains semantic guidance information of multiple organs.
[0049] In this embodiment, the cross-attention layer includes a cascaded deformable attention fusion layer and a position-aware attention layer; the deformable attention fusion layer uses a deformable convolution mechanism to perform deformable convolution offset learning on features of different scales extracted by the multi-scale coding layer, adaptively adjusts the receptive field range, and normalizes the extracted multi-layer features to achieve weighted adaptive fusion of features; the position-aware attention layer combines the spatial prior of the medical anatomical atlas to extract the information specific to multiple organs in the CT image according to their respective key positions, thereby enhancing the network's perception of the spatial distribution of organs.
[0050] In this embodiment, the dynamic filtering coding module encodes the guidance information extracted by the semantic guidance coding module into dynamic convolution kernel parameters through parameter division coding operation, and the specific method is as follows:
[0051] The semantic guidance information of multiple organs extracted by the semantic guidance encoding module is regarded as a collection of specific information of different organs. First, the input semantic guidance information is linearly transformed in channel dimension through a linear layer to generate dynamic filtering parameters. Then, the generated dynamic filtering parameters are divided into a set of 1×1 convolution kernel parameters used by multiple convolution layers through a parameter partitioning layer to realize dynamic adjustment of the input and output dimensions of the dynamic convolution kernel parameters.
[0052] The semantic-guided dynamic decoding module uses a dynamic convolution strategy to construct three dynamic convolution layers. Each layer uses a 1×1 convolution kernel parameter generated by a dynamic filter encoding module to perform a dynamic convolution operation on a feature map to be identified that matches the spatial scale of the input image and is generated by a multi-scale decoder. A multi-organ segmentation mask image is generated through the three dynamic convolution layers.
[0053] The innovation of the multi-organ segmentation network model of the present invention lies in the proposal of an organ-level semantic perception mechanism, which systematically decouples the two tasks of image feature extraction and organ feature representation in traditional medical images. This mechanism consists of three parts working together: the semantic-guided encoding module: assigns a set of independent and learnable pre-set organ information to each learnable organ category; through a three-cascade cross-attention mechanism, fuses multi-scale encoded features and pre-set organ information to extract cross-level features with organ distinctiveness; combines deformable convolution and position-aware attention to enhance the perception of the spatial distribution and local fine-grained structure of organs; realizes the extraction of unified and robust organ representation under different image inputs. The dynamic filtering encoding module: encodes the organ features extracted by the semantic-guided encoding module into dynamic convolution kernel parameters; through linear mapping and parameter partitioning, converts multi-organ guidance information into dynamic 1×1 convolution kernels adapted to different convolution layers to achieve dynamic adjustment of the input and output channels of the convolution parameters. The semantic-guided dynamic decoding module: designs a decoding structure based on dynamic convolution, and uses the dynamically generated convolution kernels to perform layer-by-layer convolution on the feature map to be recognized; accurately models the segmentation masks of multiple organs, improving the spatial matching and organ distinctiveness.
[0054] By introducing organ-level semantic perception, the multi-organ segmentation network model of the present invention breaks through the limitation of previous feature extraction based only on low-level information such as image intensity and texture, and can effectively decouple and model between organ category consistency and inter-image differences, overcoming the limitation that some annotation characteristics will cause serious category interference during model training, thereby greatly improving the accuracy and robustness of multi-organ segmentation in complex medical images.
[0055] Step S5: Train the multi-organ segmentation network model using the training set, as Figure 3 shown, and use the validation set and test set to verify and test the trained multi-organ segmentation network model;
[0056] Perform the network parameter initialization operation, and set the running state of the deep neural network according to the characteristics of the input data (such as the data type it belongs to), including three basic running states: training mode (training), validation mode (validation), and testing mode (testing);
[0057] In the training mode, if it is necessary to load specific initial weights in special task situations, then load the set specific pre-trained model weights for model training;
[0058] Step S6: Use the trained multi-organ segmentation network model to automatically segment multiple organs in medical images and output the drawing results of organ masks, as Figure 4 shown.
[0059] The trained model weights are loaded through the framework's model persistence interface (such as PyTorch's load_state_dict()), the network architecture is instantiated and parameter injection is completed, and the model is switched to the inference-ready state. To perform inference calculations, the multi-scale encoder extracts the multi-scale features of the 3D-CT image through multi-scale convolution and inputs them into the multi-scale decoder and semantic guidance module. The semantic guidance module uses sine and cosine encoding to generate position encoding, and combines the learnable parameters corresponding to each organ to fuse the multi-scale features with the preset organ information through cross attention to form semantic guidance features. The dynamic filtering module encodes the guidance features into dynamic convolution kernel parameters. The multi-scale decoder upsamples to generate the feature map to be identified, and the semantic guidance dynamic decoding module then convolves it with the dynamic convolution kernel, finally completing the segmentation of the organ mask and outputting the accurate delineation result of the organ mask.
[0060] The method of the present invention addresses the following problems existing in previous methods: (1) manual tracing is required, and the calculation results are affected by the doctor's clinical experience and have poor repeatability; (2) category interference is prone to occur when using partially annotated data sets for training; (3) general large models lack organ-specific information when performing mask segmentation tasks, and still require manual subsequent annotation guidance. By constructing a deep neural network and innovatively introducing organ semantic information as a guide, the category interference problem caused by partially annotated data sets is effectively overcome. This method significantly reduces data interference and the time cost of subsequent annotation by doctors, and realizes the automatic calculation of multi-organ segmentation. By integrating organ semantic information, this method not only improves segmentation accuracy, but also enhances the model's ability to learn the characteristics of different organs, providing more reliable technical support for medical image analysis.
[0061] The method of the present invention uses deep learning technology to achieve accurate automatic delineation of multiple organs in medical images. To verify the effectiveness and robustness of the method, this embodiment was systematically tested on multiple groups of CT image data sets containing different scanning parameters and imaging conditions. The experimental results show that this method can stably generate high-precision organ segmentation masks on CT images from different sources, and the single segmentation operation time meets the clinical real-time requirements. The present invention has the following significant advantages: (1) The algorithm is simple and efficient, and easy to deploy; (2) The operation speed is fast, and real-time processing can be achieved on ordinary GPU devices; (3) The whole process is automated, without the need for manual intervention, which greatly reduces the clinical workload. These advantages enable the present invention to fully meet the practical application needs of modern medical image processing.
[0062] In this embodiment, a large model system for automatic segmentation of multiple organs in medical images includes a medical image acquisition module, a data set construction module, a model construction module, a training module and a segmentation module;
[0063] The medical image acquisition module is used to acquire CT medical images;
[0064] The data set construction module is used to perform standardization processing on the acquired CT medical images and construct a CT medical image data set; and divide the CT medical image data set into a training set, a validation set and a test set in proportion;
[0065] The model building module is used to build a multi-organ segmentation network model; the multi-organ segmentation network model realizes accurate organ segmentation through an encoder-decoder bidirectional feature interaction mechanism;
[0066] The training module uses a training set to train a multi-organ segmentation network model, and uses a validation set and a test set to verify and test the trained multi-organ segmentation network model;
[0067] The segmentation module uses the trained multi-organ segmentation network model to automatically segment multiple organs in medical images and outputs the delineation results of organ masks.
[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.
Claims
1. A method for a large model of automatic multi-organ segmentation of medical images, characterized in that include: Obtain CT medical images; Performing standardized processing of medical imaging parameters on the acquired CT medical images to construct a CT medical imaging data set; The CT medical image dataset is divided into training set, validation set and test set in proportion; Constructing a multi-organ segmentation network model; the multi-organ segmentation network model realizes accurate organ segmentation through an encoder-decoder bidirectional feature interaction mechanism; The training set is used to train the multi-organ segmentation network model, and the validation set and test set are used to verify and test the trained multi-organ segmentation network model; The trained multi-organ segmentation network model is used to automatically segment multiple organs in medical images and output the outline results of the organ mask.
2. The method for a large model of automatic multi-organ segmentation of medical images according to claim 1, wherein, The CT medical image is a dcm file that complies with the medical digital image transmission protocol, or a plurality of frames of images and their corresponding annotated labels after the dcm file is parsed.
3. A method for a large model of automatic multi-organ segmentation of medical images according to claim 1, characterized in that, The multi-organ segmentation network model uses the U-Net architecture as the benchmark framework, introduces a semantic perception mechanism, and constructs five core modules, including a multi-scale encoder, a multi-scale decoder, a semantic-guided encoding module, a dynamic filtering encoding module, and a semantic-guided dynamic decoding module; The multi-scale encoder uses a multi-scale convolutional layer to extract multi-scale features from the input 3D-CT image, constructs a multi-scale feature map, and inputs it into the multi-scale decoder and the semantic guidance encoding module; the multi-scale decoder uses multi-scale upsampling to generate a feature map to be identified that matches the spatial scale of the input image for the multi-scale feature map constructed by the multi-scale encoder; The semantically guided encoding module applies multi-scale sine and cosine encoding to the multi-scale feature map extracted by the multi-scale encoder to generate a multi-scale position code; the multi-scale feature map is combined with the corresponding position code to generate a multi-scale feature information sequence; at the same time, a set of learnable parameters is assigned to each learnable organ category as the preset organ information; the multi-scale feature information sequence is cross-fused with the preset organ information using a cross-attention layer to obtain cross-level fused CT image multi-organ features as multi-organ semantic guidance information; the dynamic filtering encoding module encodes the guidance information extracted by the semantically guided encoding module into dynamic convolution kernel parameters through parameter partitioning encoding operations; the semantically guided dynamic decoding module uses the dynamic convolution kernel generated by the dynamic filtering encoding module to perform a convolution operation on the feature map to be identified generated by the multi-scale decoder to achieve semantically guided dynamic decoding and complete the delineation of the organ mask.
4. A method for a large model of automatic multi-organ segmentation of medical images according to claim 3, characterized in that The multi-scale encoder adopts a progressive downsampling strategy and realizes multi-scale feature extraction through multi-level feature extraction units; each level of feature extraction unit includes a residual convolution layer with a 3×3 convolution kernel, a batch normalization layer, a ReLU activation layer and a maximum pooling downsampling layer; the multi-scale encoder extracts multi-scale feature maps from high-resolution detail features to deep semantic features step by step through a hierarchical structure.
5. A method for a large model of automatic multi-organ segmentation of medical images according to claim 4, characterized in that, The multi-scale decoder adopts a progressive upsampling strategy to generate a feature map to be identified that matches the input spatial scale through a multi-level feature extraction unit; Each level of feature extraction unit includes a residual convolution layer with a 3×3 convolution kernel, a batch normalization layer, a ReLU activation layer, and a transposed convolution upsampling layer; the multi-scale decoder extracts multi-scale feature maps from deep semantic features to high-resolution detail features step by step through a hierarchical structure; and the multi-scale feature maps are gradually reconstructed through cross-layer feature splicing and fusion operations.
6. The method for a large model of automatic multi-organ segmentation of medical images according to claim 5, wherein, The semantically guided encoding module first uses an embedding layer to assign a set of learnable parameters as preset organ information to each learnable organ category, and then uses a three-layer cascaded cross-attention layer to apply multi-scale sine and cosine encoding to the multi-scale feature map extracted by the multi-scale encoder to generate a multi-scale position code; the multi-scale feature map is combined with the corresponding position code to generate a multi-scale feature information sequence; the multi-scale feature information sequence is then cross-fused with the preset organ information to obtain cross-level CT image multi-organ features as multi-organ semantic guidance information; the semantically guided encoding module extracts unified representation information of different organs in different CT images, decouples traditional image feature extraction and organ feature representation tasks, and then obtains semantic guidance information of multiple organs.
7. A method for a large model of automatic multi-organ segmentation of medical images according to claim 6, characterized in that, The cross-attention layer includes a cascaded deformable attention fusion layer and a position-aware attention layer; the deformable attention fusion layer uses a deformable convolution mechanism to perform deformable convolution offset learning on features of different scales extracted by the multi-scale coding layer, adaptively adjusts the receptive field range, and extracts multi-layer features through normalization to achieve weighted adaptive fusion of features; the position-aware attention layer combines the spatial prior of medical anatomical atlases to extract the information specific to multiple organs in CT images according to their respective key positions, thereby enhancing the network's perception of the spatial distribution of organs.
8. A method for a large model of automatic multi-organ segmentation of medical images according to claim 7, characterized in that, The dynamic filtering coding module encodes the guidance information extracted by the semantic guidance coding module into dynamic convolution kernel parameters through parameter division coding operation, and the specific method is as follows: The semantic guidance information of multiple organs extracted by the semantic guidance encoding module is regarded as a collection of specific information of different organs. First, the input semantic guidance information is linearly transformed in channel dimension through a linear layer to generate dynamic filtering parameters. Then, the generated dynamic filtering parameters are divided into a set of 1×1 convolution kernel parameters used by multiple convolution layers through a parameter partitioning layer to realize dynamic adjustment of the input and output dimensions of the dynamic convolution kernel parameters.
9. A method for a large model of automatic multi-organ segmentation of medical images according to claim 8, characterized in that, The semantic-guided dynamic decoding module uses a dynamic convolution strategy to construct three dynamic convolution layers. Each layer uses a 1×1 convolution kernel parameter generated by a dynamic filter encoding module to perform a dynamic convolution operation on a feature map to be identified that matches the spatial scale of the input image and is generated by a multi-scale decoder. A multi-organ segmentation mask image is generated through the three dynamic convolution layers.
10. A large model system for automatic multi-organ segmentation of medical images, implemented based on the method for automatic multi-organ segmentation of medical images described in claim 1, characterized in that, It includes medical image acquisition module, data set construction module, model construction module, training module and segmentation module; The medical image acquisition module is used to acquire CT medical images; The data set construction module is used to perform standardization processing on the acquired CT medical images and construct a CT medical image data set; and divide the CT medical image data set into a training set, a validation set and a test set in proportion; The model building module is used to build a multi-organ segmentation network model; the multi-organ segmentation network model realizes accurate organ segmentation through an encoder-decoder bidirectional feature interaction mechanism; The training module uses a training set to train a multi-organ segmentation network model, and uses a validation set and a test set to verify and test the trained multi-organ segmentation network model; The segmentation module uses the trained multi-organ segmentation network model to automatically segment multiple organs in medical images and outputs the delineation results of organ masks.