Abdominal CT image segmentation method and system based on deep learning, terminal and medium

By using a deep learning-based cascaded convolutional decoder to segment the network model, extracting and fusing feature maps, and using attention gating and edge enhancement modules, the problem of insufficient segmentation accuracy of abdominal CT images is solved, and diagnostic efficiency and accuracy are improved.

CN119992097APending Publication Date: 2025-05-13CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510155388.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art lacks segmentation accuracy in abdominal CT image segmentation, resulting in limited diagnostic efficiency and accuracy, especially when the boundaries of different organs are blurred, which increases the difficulty of observation for doctors.

Method used

The network model is segmented by a cascaded convolutional decoder based on deep learning. By extracting low-dimensional spatial feature maps and high-dimensional semantic feature maps, local and global context extraction is performed, combining attention gating modules and edge enhancement decoder modules, feature extraction and edge features fusion are enhanced.

Benefits of technology

It significantly improves the segmentation accuracy of abdominal CT images, assists doctors in better observation of abdominal CT images, reduces doctors' observation experience requirements, and improves diagnosis efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992097A_ABST
    Figure CN119992097A_ABST
Patent Text Reader

Abstract

The invention discloses an abdomen CT image segmentation method and system based on deep learning, a terminal and a medium, and relates to the technical field of image processing, and the method comprises the steps: obtaining a to-be-segmented abdomen multi-organ CT image; preprocessing the abdomen multi-organ CT image to obtain a preprocessed image; inputting the preprocessed image into a trained cascade convolution decoder segmentation network model for organ segmentation to obtain an abdominal organ CT image segmentation label image; and outputting the segmentation label image of the abdominal organ CT image. According to the method, the feature extraction capability of the segmentation model is enhanced, the segmentation precision of the abdominal CT image is improved, a doctor can be assisted to better observe the abdominal CT image, the observation experience requirement on the doctor is reduced, and the diagnosis efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a deep learning-based abdominal CT image segmentation method, system, terminal and medium. Background Art

[0002] With the development of social economy and the improvement of science and technology, CT (Computed Tomography) imaging technology has been widely used in the field of human medical diagnosis, which has brought about earth-shaking changes in traditional medical diagnosis methods and improved the diagnostic efficiency of doctors. However, the slice images produced by CT imaging equipment are all two-dimensional tomographic images. Doctors often rely on these images for diagnosis with subjectivity. The diagnosis results are highly dependent on the doctor's diagnostic experience and professional level, and even the diagnosis results between different experts may be inconsistent. Especially for abdominal CT images, the boundaries between different organs may be blurred, which greatly increases the difficulty of doctors' observation and diagnosis.

[0003] If the region of interest can be clearly segmented directly from the CT image, it will help doctors to visually observe the region of interest, which can greatly improve the doctor's diagnostic efficiency and the accuracy of the diagnostic results. In recent years, image segmentation methods based on deep learning have shown strong competitiveness compared to traditional segmentation methods. Among them, a series of U-shaped image segmentation models represented by U-Net are often used in medical image segmentation tasks including CT images, and have achieved certain results. However, the segmentation accuracy needs to be further improved, and better feature extraction methods need to be studied to improve the problems of mis-segmentation and over-segmentation. Summary of the invention

[0004] The purpose of the present invention is to provide a deep learning-based abdominal CT image segmentation method, system, terminal and medium, which enhances the segmentation accuracy of abdominal CT images, can assist doctors to better observe abdominal CT images, and improve the efficiency and accuracy of diagnosis.

[0005] The present invention is achieved through the following technical solutions:

[0006] In a first aspect, an embodiment of the present invention provides a method for segmenting abdominal CT images based on deep learning, comprising:

[0007] Acquire a multi-organ CT image of the abdomen to be segmented;

[0008] Preprocessing the abdominal multi-organ CT image to obtain a preprocessed image;

[0009] The preprocessed image is input into the trained cascade convolution decoder segmentation network model for organ segmentation to obtain the abdominal organ CT image segmentation label image;

[0010] Output abdominal organ CT image segmentation label image.

[0011] Furthermore, the cascaded convolutional decoder segmentation network model extracts a low-dimensional spatial feature map from the encoder and a high-dimensional semantic feature map from the decoder, performs local context extraction on the low-dimensional spatial feature map to obtain a first feature map, performs global context extraction on the high-dimensional semantic feature map to obtain a second feature map, adds and fuses the first feature map and the second feature map to obtain a first fused feature map, and then uses 1×1 convolution to project the first fused feature map to two channels, calculates the spatial selection attention feature map through the Softmax of the channel dimension, and then splits the spatial selection attention feature map according to the channel dimension, generating two spatial selection versions. New feature map, multiply one of the new feature maps with the low-dimensional space feature map and then add and fuse them with the low-dimensional space feature map to obtain a second fused feature map, multiply another new feature map with the high-dimensional semantic feature map and then add and fuse them with the high-dimensional semantic feature map to obtain a third fused feature map, perform feature mapping on the second fused feature map and the third fused feature map through 1×1 convolution respectively and then add and fuse them to obtain a fourth fused feature map, use the ReLU activation function to perform nonlinear transformation on the fourth fused feature map, and finally use the sigmoid function to generate an attention weight and multiply it with the second fused feature map to obtain the first output feature map.

[0012] Furthermore, the specific method of obtaining the first feature map after performing local context extraction on the low-dimensional spatial feature map includes: performing depth convolution, depth expansion convolution and point-wise convolution on the input low-dimensional spatial feature map to obtain the first feature map after spatial feature refinement;

[0013] The specific method of obtaining the second feature map after performing global context extraction on the high-dimensional semantic feature map includes:

[0014] The input high-dimensional semantic feature map is subjected to global average pooling and global maximum pooling to obtain the first pooling result and the second pooling result. The first pooling result and the second pooling result are spliced, and after splicing, convolution processing and batch normalization processing are performed in sequence to obtain the second feature map.

[0015] Furthermore, the method further includes: the cascaded convolutional decoder segmentation network model enhances the edge features of the first output feature map by fusing high-frequency features, specifically including:

[0016] Extract high-frequency features from the original input image by Laplacian pyramid operation, multiply the high-frequency features with the first output feature map, and then perform concatenation and convolution processing with the first output feature map to obtain a second output feature map;

[0017] The second output feature map is subjected to convolution and sigmoid activation function to generate attention weights, which are then multiplied with the second output feature map to obtain a third output feature map. The third output feature map is residually connected with the first output feature map to obtain a fourth output feature map.

[0018] Furthermore, the method also includes training the cascaded convolutional decoder segmentation network model, specifically including:

[0019] Obtaining an original abdominal CT image dataset;

[0020] Preprocess and annotate the original abdominal CT images;

[0021] Divide the data set into training set and test set according to the set ratio;

[0022] Input the training set into the cascade convolution decoder segmentation network model for training, and obtain the trained cascade convolution decoder segmentation network model;

[0023] The test set is input into the trained cascade convolution decoder segmentation network model, and the label image is output. The organ distribution of the slice is confirmed according to the label distribution of each slice, and the abdominal CT image segmentation is completed to obtain the trained cascade convolution decoder segmentation network model.

[0024] Furthermore, the specific method of preprocessing includes:

[0025] Adjust the window level and width of abdominal CT images;

[0026] The values ​​of the pixels in the abdominal CT images were normalized to the range of 0 to 1;

[0027] Convert abdominal CT images into NIFTI format and save them.

[0028] In a second aspect, another embodiment of the present invention provides an abdominal CT image segmentation system based on deep learning, comprising: an image acquisition module, an image preprocessing module, a segmentation network model module and an output module;

[0029] The image acquisition module is used to acquire the abdominal multi-organ CT image to be segmented;

[0030] The image preprocessing module is used to preprocess the abdominal multi-organ CT image to obtain a preprocessed image;

[0031] The segmentation network model module uses the trained cascade convolution decoder segmentation network model to perform organ segmentation on the preprocessed image to obtain an abdominal organ CT image segmentation label image;

[0032] The output module is used to output the abdominal organ CT image segmentation label image.

[0033] Further, the cascaded convolutional decoder segmentation network model includes an attention gating module, an edge enhancement decoder module and an upsampling module;

[0034] The attention gating module extracts a low-dimensional spatial feature map from the encoder and a high-dimensional semantic feature map from the decoder, performs local context extraction on the low-dimensional spatial feature map to obtain a first feature map, performs global context extraction on the high-dimensional semantic feature map to obtain a second feature map, adds and fuses the first feature map and the second feature map to obtain a first fused feature map, then uses 1×1 convolution to project the first fused feature map to two channels, calculates the spatial selection attention feature map through the Softmax of the channel dimension, and then splits the spatial selection attention feature map according to the channel dimension, generating two new feature maps of the spatial selection version, and converts them into One of the new feature maps is multiplied by the low-dimensional space feature map and then added and fused with the low-dimensional space feature map to obtain a second fused feature map. Another new feature map is multiplied by the high-dimensional semantic feature map and then added and fused with the high-dimensional semantic feature map to obtain a third fused feature map. The second fused feature map and the third fused feature map are respectively feature mapped by 1×1 convolution and then added and fused to obtain a fourth fused feature map. The fourth fused feature map is nonlinearly transformed using the ReLU activation function. Finally, an attention weight is generated by the sigmoid function and multiplied with the second fused feature map to obtain the first output feature map.

[0035] The edge enhancement decoder module includes a high-frequency feature fusion unit, which extracts high-frequency features from the original input image by Laplacian pyramid operation, multiplies the high-frequency features with the first output feature map, and then performs splicing and convolution processing with the first output feature map to obtain a second output feature map;

[0036] The second output feature map is subjected to convolution and sigmoid activation function to generate attention weights, which are then multiplied with the second output feature map to obtain a third output feature map. The third output feature map is residually connected with the first output feature map to obtain a fourth output feature map.

[0037] The upsampling module is used to upsample the feature maps obtained by the encoder in multiple stages.

[0038] In a third aspect, another embodiment of the present invention provides an intelligent terminal, comprising a processor, an input device, an output device and a memory, wherein the processor is connected to the input device, the output device and the memory respectively, the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions to execute the method described in the above embodiment.

[0039] In a fourth aspect, another embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method described in the above embodiment.

[0040] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0041] The embodiments of the present invention provide a deep learning-based abdominal CT image segmentation method, system, terminal and medium, which enhance the feature extraction capability of the segmentation model and improve the segmentation accuracy of abdominal CT images, can assist doctors in better observing abdominal CT images, reduce the observation experience requirements for doctors, and improve the efficiency and accuracy of diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without creative work. In the drawings:

[0043] Figure 1 A flowchart of an abdominal CT image segmentation method based on deep learning provided in the first embodiment of the present invention;

[0044] Figure 2 It is an overall architecture diagram of the cascaded convolutional decoder segmentation network model in the first embodiment of the present invention;

[0045] Figure 3 for Figure 2 Flowchart of attention gating;

[0046] Figure 4 for Figure 3 Flowchart of global context extraction in [5];

[0047] Figure 5 is a flow chart of an edge enhancement decoder module in the first embodiment of the present invention;

[0048] Figure 6 This is a flow chart of edge features of the first output feature map fused with high-frequency features in the first embodiment of the present invention;

[0049] Figure 7 The segmentation results and comparison diagrams obtained by using the abdominal CT image segmentation method based on deep learning provided by the first embodiment of the present invention;

[0050] Figure 8A structural block diagram of an abdominal CT image segmentation system based on deep learning provided by another embodiment of the present invention;

[0051] Fig. 9 A structural block diagram of an intelligent terminal provided in another embodiment of the present invention. DETAILED DESCRIPTION

[0052] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with embodiments and drawings. The exemplary embodiments of the present invention and their description are only used to explain the present invention and are not intended to limit the present invention.

[0053] Example 1

[0054] like Figure 1 As shown, a deep learning-based abdominal CT image segmentation method provided by the first embodiment of the present invention includes the following steps:

[0055] Acquire a multi-organ CT image of the abdomen to be segmented;

[0056] Preprocessing the abdominal multi-organ CT image to obtain a preprocessed image;

[0057] The preprocessed image is input into the trained cascade convolution decoder segmentation network model for organ segmentation to obtain the abdominal organ CT image segmentation label image;

[0058] Output abdominal organ CT image segmentation label image.

[0059] Figure 2 It is an architecture diagram of the cascaded convolutional decoder segmentation network model, and its encoder end is composed of a Pyramid Vision Transformer (PVT). The cascaded convolutional decoder segmentation network model includes an attention gating module, an edge enhancement decoder module, and an upsampling module. The method for improving the cascaded convolutional decoder segmentation network model is: improving the large-core grouped attention gating module on its decoder end. The original large-core grouped attention gate module mainly uses group convolution to process low-dimensional spatial information and high-dimensional semantic information from the skip connection stage, by directly adding two feature maps and subsequent batch normalization, convolution and activation functions, and residual connection output gated weight scores. An embodiment of the present invention proposes an attention gating module based on spatial selection and cross-modulation. The data processing flow chart of the attention gating module is shown in FIG. Figure 3As shown in the figure. After the image is input, the high-resolution spatial detail information of the low-dimensional spatial feature map from the encoder and the semantic information of the high-dimensional semantic feature map from the decoder are extracted through the group convolution and global context feature extraction modules respectively. The first feature map is obtained after local context extraction on the low-dimensional spatial feature map, and the second feature map is obtained after global context extraction on the high-dimensional semantic feature map. The first feature map and the second feature map are added and fused to obtain the first fused feature map. The first fused feature map is then projected to two channels using 1×1 convolution. The spatial selective attention feature map is calculated by the Softmax of the channel dimension, and then the spatial selective attention feature map is split according to the channel dimension, generating two new feature maps of the spatial selective version, emphasizing the key areas, and further refining the segmentation through precise context-aware attention.

[0060] In order to further refine the features, one of the new feature maps is multiplied by the low-dimensional space feature map and then added and fused with the low-dimensional space feature map to obtain a second fused feature map. Another new feature map is multiplied by the high-dimensional semantic feature map and then added and fused with the high-dimensional semantic feature map to obtain a third fused feature map. The second fused feature map and the third fused feature map are respectively feature mapped by 1×1 convolution and then added and fused to obtain a fourth fused feature map. The ReLU activation function is used to perform a nonlinear transformation on the fourth fused feature map. Finally, a sigmoid function is used to generate an attention weight and multiply it with the second fused feature map to obtain the first output feature map.

[0061] Among them, the specific method of obtaining the first feature map after performing local context extraction on the low-dimensional spatial feature map includes: performing deep convolution, deep expansion convolution and point-wise convolution on the input low-dimensional spatial feature map to obtain the first feature map after spatial feature refinement.

[0062] The specific method flow chart of obtaining the second feature map after extracting the global context of the high-dimensional semantic feature map is as follows: Figure 4 As shown. The input high-dimensional semantic feature map is subjected to global average pooling and global maximum pooling to obtain the first pooling result and the second pooling result, the first pooling result and the second pooling result are spliced, and the convolution process and batch normalization process are performed in sequence after the splicing to obtain the second feature map.

[0063] like Figure 5 As shown, the edge enhancement decoder module uses high-frequency feature fusion to enhance the first output feature map and the feature map obtained by upsampling, and adds high-frequency edge features.

[0064] like Figure 6 As shown, an abdominal CT image segmentation method based on deep learning provided by an embodiment of the present invention also includes: a cascaded convolution decoder segmentation network model enhances edge features of a first output feature map by fusing high-frequency features, specifically including:

[0065] Extract high-frequency features from the original input image by Laplacian pyramid operation, multiply the high-frequency features with the first output feature map, and then perform concatenation and convolution processing with the first output feature map to obtain a second output feature map;

[0066] The second output feature map is subjected to convolution and sigmoid activation function to generate attention weights, which are then multiplied with the second output feature map to obtain a third output feature map. The third output feature map is residually connected with the first output feature map to obtain a fourth output feature map.

[0067] The same method is used to process the upsampled decoder feature map, fuse the edge features, and obtain the fifth output feature map. The edge enhancement decoder module concatenates the fourth output feature map and the fifth output feature map and performs convolution processing to obtain the sixth output feature map. The cascade convolution decoder segmentation network model is trained, specifically including:

[0068] Obtaining an original abdominal CT image dataset;

[0069] Preprocess and annotate the original abdominal CT images;

[0070] Divide the data set into training set and test set according to the set ratio;

[0071] Input the training set into the cascade convolution decoder segmentation network model for training, and obtain the trained cascade convolution decoder segmentation network model;

[0072] The test set is input into the trained cascade convolution decoder segmentation network model, and the label image is output. The organ distribution of the slice is confirmed according to the label distribution of each slice, and the abdominal CT image segmentation is completed to obtain the trained cascade convolution decoder segmentation network model.

[0073] Among them, the method for obtaining the original abdominal CT image dataset is: first, obtain CT images in Dicom format from hospitals and other medical institutions, and professional doctors are required to manually segment the liver, spleen, pancreas, left kidney, right kidney and other organs in each slice to obtain labeled images.

[0074] In the preprocessing stage, the window level and width of the CT image are adjusted to increase the contrast between the abdominal organs and the background. Then the pixel values ​​are standardized to the range of 0 to 1. Finally, the CT image and the label image are converted into NIFTI format and saved.

[0075] The data set is divided into training and test sets in a 3:2 ratio.

[0076] In this embodiment, the Synapse abdominal multi-organ CT image dataset is used as the original dataset, and the CT images and label images in the dataset are compressed into nii.gz format files. The dataset contains 30 multi-slice contrast abdominal CT images, of which 8 organs (aorta, gallbladder, left kidney, right kidney, liver, pancreas, spleen, stomach) are manually segmented and annotated for abdominal multi-organ segmentation, and are divided into training set and test set in a ratio of 3:2, that is, 18 training sets and 12 test sets.

[0077] Set the training parameters: batch size is 6, learning rate is 0.0001, optimizer type is AdamW, and maximum number of iterations is 200.

[0078] During the model training process, the traditional iterative optimization algorithm is used to optimize the model's loss function based on the gradient descent method. After each batch of training samples is input into the model, the model performs forward propagation to generate a predicted output. The loss function calculates the difference between the predicted output and the actual value, that is, the loss value. After obtaining the loss value, the model adjusts the parameters through the back-propagation process to reduce the error between the predicted value and the actual value. Training continues until the model's loss value tends to stabilize, and it is considered that the training has achieved the predetermined goal.

[0079] In the test phase, the training weights are loaded, the model is set to evaluation mode, and only forward reasoning is performed. The abdominal multi-organ CT images to be segmented are input into the trained cascade convolution decoder segmentation network model, and the abdominal organ CT image segmentation label image is output. The segmentation results and comparison diagrams are shown in Figure 1. Figure 7 shown.

[0080] An abdominal CT image segmentation method based on deep learning provided by an embodiment of the present invention better extracts and fuses contextual information from low-dimensional spatial feature maps and high-dimensional semantic feature maps at jump connections, and uses high-frequency feature fusion to further enhance edge features, thereby improving the feature extraction capability of the algorithm and improving segmentation accuracy. In the training phase, the decoder fuses the low-dimensional spatial features and high-dimensional semantic features, calculates the loss function with the input label, and updates the weight parameters of the model through back propagation. In the testing phase, the model directly outputs the segmentation label image as the result.

[0081] An abdominal CT image segmentation method based on deep learning provided by an embodiment of the present invention enhances the feature extraction capability of the segmentation model and improves the segmentation accuracy of abdominal CT images. It can assist doctors in better observing abdominal CT images, reduce the observation experience requirements for doctors, and improve the efficiency and accuracy of diagnosis.

[0082] Example 2

[0083] like Figure 8As shown, another embodiment of the present invention provides an abdominal CT image segmentation system based on deep learning, including: an image acquisition module, an image preprocessing module, a segmentation network model module and an output module; the image acquisition module is used to acquire the abdominal multi-organ CT image to be segmented; the image preprocessing module is used to preprocess the abdominal multi-organ CT image to obtain a preprocessed image; the segmentation network model module uses a trained cascade convolutional decoder segmentation network model to perform organ segmentation on the preprocessed image to obtain an abdominal organ CT image segmentation label image; the output module is used to output the abdominal organ CT image segmentation label image.

[0084] Among them, the cascade convolution decoder segmentation network model includes an attention gating module, an edge enhancement decoder module and an upsampling module. The attention gating module extracts the low-dimensional spatial feature map from the encoder and the high-dimensional semantic feature map from the decoder, performs local context extraction on the low-dimensional spatial feature map to obtain the first feature map, performs global context extraction on the high-dimensional semantic feature map to obtain the second feature map, adds and fuses the first feature map and obtains the first fused feature map, then uses 1×1 convolution to project the first fused feature map to two channels, calculates the spatial selection attention feature map by Softmax of the channel dimension, and then decomposes the spatial selection attention feature map according to the channel dimension. The second fused feature map is obtained by multiplying the other new feature map with the high-dimensional semantic feature map and then adding and fusion with the high-dimensional semantic feature map to obtain a third fused feature map. The second fused feature map and the third fused feature map are respectively feature mapped by 1×1 convolution and then added and fused to obtain a fourth fused feature map. The ReLU activation function is used to perform nonlinear transformation on the fourth fused feature map. Finally, the sigmoid function is used to generate an attention weight which is multiplied by the second fused feature map to obtain the first output feature map.

[0085] The edge enhancement decoder module uses high-frequency feature fusion to enhance the first output feature map and the feature map obtained by upsampling, adding high-frequency edge features.

[0086] The edge enhancement decoder module includes a high-frequency feature fusion unit, which extracts high-frequency features from the original input image by Laplacian pyramid operation, multiplies the high-frequency features with the first output feature map, and then performs splicing and convolution processing with the first output feature map to obtain a second output feature map;

[0087] The second output feature map is subjected to convolution and sigmoid activation function to generate attention weights, which are then multiplied with the second output feature map to obtain a third output feature map. The third output feature map is residually connected with the first output feature map to obtain a fourth output feature map.

[0088] The upsampling module is used to upsample the feature maps obtained by the encoder in multiple stages.

[0089] The attention gating module includes a local context feature extraction unit and a global context feature extraction unit. The local context feature extraction unit is used to perform deep convolution and deep dilation convolution on the input low-dimensional spatial feature map to obtain a first feature map after spatial feature refinement. The global context feature extraction unit performs global average pooling and global maximum pooling on the input high-dimensional semantic feature map to obtain a first pooling result and a second pooling result, and splices the first pooling result and the second pooling result, and then performs convolution processing and batch normalization processing in sequence to obtain a second output feature map.

[0090] An abdominal CT image segmentation system based on deep learning provided by an embodiment of the present invention enhances the feature extraction capability of the segmentation model and improves the segmentation accuracy of abdominal CT images. It can assist doctors in better observing abdominal CT images, reduce the observation experience requirements for doctors, and improve the efficiency and accuracy of diagnosis.

[0091] Example 3

[0092] like Fig. 9 As shown, a structural block diagram of an intelligent terminal provided in another embodiment of the present invention, the terminal includes a processor, an input device, an output device and a memory, the processor is connected to the input device, the output device and the memory respectively, the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method described in the first embodiment above.

[0093] It should be understood that in the embodiments of the present invention, the processor referred to may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0094] Input devices may include a touchpad, a fingerprint collection sensor (for collecting the user's fingerprint information and fingerprint direction information), a microphone, etc., and output devices may include a display (LCD, etc.), a speaker, etc.

[0095] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0096] In a specific implementation, the processor, input device, and output device described in the embodiments of the present invention may execute the implementation described in the method embodiment provided in the embodiments of the present invention, or may execute the implementation described in the system embodiment described in the embodiments of the present invention, which will not be described in detail here.

[0097] Example 4

[0098] Another embodiment of the present invention further provides an embodiment of a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method described in the first embodiment.

[0099] The computer-readable storage medium may be an internal storage unit of the terminal described in the foregoing embodiments, such as a hard disk or memory of the terminal. The computer-readable storage medium may also be an external storage device of the device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Further, the computer-readable storage medium may also include both an internal storage unit of the terminal and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the terminal. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.

[0100] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0101] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the terminals and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.

[0102] In the several embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or it can be an electrical, mechanical or other form of connection.

[0103] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A deep learning-based abdominal CT image segmentation method, characterized in that: include: Acquire a multi-organ CT image of the abdomen to be segmented; Preprocessing the abdominal multi-organ CT image to obtain a preprocessed image; The preprocessed image is input into the trained cascade convolution decoder segmentation network model for organ segmentation to obtain the abdominal organ CT image segmentation label image; Output abdominal organ CT image segmentation label image.

2. The method according to claim 1, characterized in that The cascaded convolution decoder segmentation network model extracts a low-dimensional spatial feature map from the encoder and a high-dimensional semantic feature map from the decoder, performs local context extraction on the low-dimensional spatial feature map to obtain a first feature map, performs global context extraction on the high-dimensional semantic feature map to obtain a second feature map, adds and fuses the first feature map and the second feature map to obtain a first fused feature map, then uses 1×1 convolution to project the first fused feature map to two channels, calculates the spatial selection attention feature map through the Softmax of the channel dimension, and then splits the spatial selection attention feature map according to the channel dimension, generating two new feature maps of the spatial selection version. Figure, multiply one of the new feature maps with the low-dimensional space feature map and then add and fuse them with the low-dimensional space feature map to obtain a second fused feature map, multiply the other new feature map with the high-dimensional semantic feature map and then add and fuse them with the high-dimensional semantic feature map to obtain a third fused feature map, perform feature mapping on the second fused feature map and the third fused feature map through 1×1 convolution respectively, and then add and fuse them to obtain a fourth fused feature map, use the ReLU activation function to perform nonlinear transformation on the fourth fused feature map, and finally use the sigmoid function to generate an attention weight and multiply it with the second fused feature map to obtain the first output feature map.

3. The method according to claim 2, characterized in that The specific method of extracting local context from the low-dimensional spatial feature map to obtain the first feature map includes: performing depth convolution, depth expansion convolution and point-wise convolution on the input low-dimensional spatial feature map to obtain the first feature map after spatial feature refinement; The specific method of obtaining the second feature map after performing global context extraction on the high-dimensional semantic feature map includes: The input high-dimensional semantic feature map is subjected to global average pooling and global maximum pooling to obtain the first pooling result and the second pooling result. The first pooling result and the second pooling result are spliced, and after splicing, convolution processing and batch normalization processing are performed in sequence to obtain the second feature map.

4. The method according to claim 3, characterized in that The method further includes: the cascaded convolutional decoder segmentation network model enhances the edge features of the first output feature map by fusing high-frequency features, specifically including: Extract high-frequency features from the original input image by Laplacian pyramid operation, multiply the high-frequency features with the first output feature map, and then perform concatenation and convolution processing with the first output feature map to obtain a second output feature map; The second output feature map is subjected to convolution and sigmoid activation function to generate attention weights, which are then multiplied with the second output feature map to obtain the third output feature map. The third output feature map is residually connected with the first output feature map to obtain the fourth output feature map.

5. The method according to claim 4, characterized in that The method further includes: training the cascaded convolutional decoder segmentation network model, specifically including: Obtaining an original abdominal CT image dataset; Preprocess and annotate the original abdominal CT images; Divide the data set into training set and test set according to the set ratio; Input the training set into the cascade convolution decoder segmentation network model for training, and obtain the trained cascade convolution decoder segmentation network model; The test set is input into the trained cascade convolution decoder segmentation network model, and the label image is output. The organ distribution of the slice is confirmed according to the label distribution of each slice, and the abdominal CT image segmentation is completed to obtain the trained cascade convolution decoder segmentation network model.

6. The method according to claim 5, characterized in that The specific method of the pretreatment includes: Adjust the window level and width of abdominal CT images; The values ​​of the pixels in the abdominal CT images were normalized to the range of 0 to 1; Convert abdominal CT images into NIFTI format and save them.

7. An abdominal CT image segmentation system based on deep learning, characterized in that: include: Image acquisition module, image preprocessing module, segmentation network model module and output module; The image acquisition module is used to acquire the abdominal multi-organ CT image to be segmented; The image preprocessing module is used to preprocess the abdominal multi-organ CT image to obtain a preprocessed image; The segmentation network model module uses the trained cascade convolution decoder segmentation network model to perform organ segmentation on the preprocessed image to obtain an abdominal organ CT image segmentation label image; The output module is used to output the abdominal organ CT image segmentation label image.

8. The system according to claim 7, characterized in that The cascaded convolutional decoder segmentation network model includes an attention gating module, an edge enhancement decoder module and an upsampling module; The attention gating module extracts a low-dimensional spatial feature map from the encoder and a high-dimensional semantic feature map from the decoder, performs local context extraction on the low-dimensional spatial feature map to obtain a first feature map, performs global context extraction on the high-dimensional semantic feature map to obtain a second feature map, adds and fuses the first feature map and the second feature map to obtain a first fused feature map, then uses 1×1 convolution to project the first fused feature map to two channels, calculates the spatial selection attention feature map through the Softmax of the channel dimension, and then splits the spatial selection attention feature map according to the channel dimension, generating two new feature maps of the spatial selection version, and converts them into One of the new feature maps is multiplied by the low-dimensional space feature map and then added and fused with the low-dimensional space feature map to obtain a second fused feature map. Another new feature map is multiplied by the high-dimensional semantic feature map and then added and fused with the high-dimensional semantic feature map to obtain a third fused feature map. The second fused feature map and the third fused feature map are respectively feature mapped by 1×1 convolution and then added and fused to obtain a fourth fused feature map. The fourth fused feature map is nonlinearly transformed using the ReLU activation function. Finally, an attention weight is generated by the sigmoid function and multiplied with the second fused feature map to obtain the first output feature map. The edge enhancement decoder module includes a high-frequency feature fusion unit, which extracts high-frequency features from the original input image by Laplacian pyramid operation, multiplies the high-frequency features with the first output feature map, and then performs splicing and convolution processing with the first output feature map to obtain a second output feature map; The second output feature map is subjected to convolution and sigmoid activation function to generate attention weights, which are then multiplied with the second output feature map to obtain a third output feature map. The third output feature map is residually connected with the first output feature map to obtain a fourth output feature map. The upsampling module is used to upsample the feature maps obtained by the encoder in multiple stages.

9. An intelligent terminal, characterized in that: It includes a processor, an input device, an output device and a memory, the processor is connected to the input device, the output device and the memory respectively, the memory is used to store a computer program, the computer program includes program instructions, and is characterized in that the processor is configured to call the program instructions to execute the method according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Ultrasonic breast cancer focus self-attention identification method based on multi-dimensional deep convolution

    CN121599995A