CT neck lymph node automatic segmentation method and system

By using U-Net Efficient network and multi-scale feature fusion technology in the automatic segmentation algorithm of CT cervical lymph nodes, the problem of unclear and fine segmentation of CT cervical lymph nodes is solved, and the precise segmentation accuracy and model performance are achieved.

CN119963580APending Publication Date: 2025-05-09THE SECOND HOSPITAL OF TIANJIN MEDICAL UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510442970.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing CT cervical lymph node automatic segmentation algorithm fails to fully consider the subgroup characteristics of lymph nodes, resulting in the segmentation results being unclear and refined enough, and can only basically meet the preliminary assessment needs for the overall distribution and size of lymph nodes.

Method used

Deep convolutional neural network (DCNN) is used to replace it with U-Net Efficient network, and the multi-scale feature fusion and hollow space pyramid pooling module is combined with dual-branch cascade fusion and weighted loss function optimization model to achieve more refined lymph node segmentation.

Benefits of technology

The precise segmentation of CT neck lymph nodes is achieved, segmentation accuracy and model performance are improved, and the needs of in-depth analysis and precise diagnosis in clinical applications can be better met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963580A_ABST
    Figure CN119963580A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for automatically segmenting a CT (Computed Tomography) neck lymph node, and the method comprises the steps: S1, processing image data, extracting basic features through a deep convolutional neural network, and replacing an original DCNN (Digital Convolutional Neural Network) with a U-Net Efficent network; s2, multi-scale feature fusion: inputting the features extracted in the S1 into a void space pyramid pooling module, and performing multi-scale feature fusion; s3, performing feature fusion and prediction, performing up-sampling on high-level features processed by the spatial pyramid pooling module, performing cascade fusion on sampled data and low-level features, and performing segmentation output through two branches; s4, training the model, carrying out regression on the two branches at the same time, weighting loss functions of the two branches, and optimizing the model; and S5, processing the data by using the model. The method has the advantages that the lymph node segmentation task in the neck CT image is regarded as a typical image segmentation problem, detail information in the image can be effectively captured, and therefore an accurate image segmentation model structure is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition, and more specifically to a method and system for automatic segmentation of CT cervical lymph nodes. Background Art

[0002] The existing automatic segmentation algorithm for cervical lymph nodes, based on medical image analysis technology, usually divides cervical lymph nodes into 7 main anatomical groups, and can realize automatic identification and segmentation of lymph nodes to a certain extent. However, the anatomical structure of lymph nodes is extremely complex. In addition to the above 7 main groups, there are many more detailed subgroup divisions. These subgroups play an important role in the functional localization, pathological analysis and clinical treatment decision-making of lymph nodes. Therefore, in actual clinical applications and imaging diagnosis processes, the existing automatic segmentation algorithms often fail to fully consider the characteristics of lymph node subgroups, resulting in unclear and inaccurate segmentation results, and can only basically meet the needs of preliminary evaluation of the overall distribution and size of lymph nodes. This limitation has, to a certain extent, restricted the in-depth application of the algorithm in clinical practice and the improvement of accurate diagnostic capabilities. Summary of the invention

[0003] The present invention overcomes the deficiencies in the prior art and provides a method and system for automatic segmentation of CT cervical lymph nodes.

[0004] The purpose of the present invention is achieved through the following technical solutions.

[0005] The method for automatic segmentation of CT cervical lymph nodes includes: S1. Process the image data, extract basic features through deep convolutional neural network, and replace the original DCNN network with U-Net Efficient network; S2, multi-scale feature fusion, input the features extracted in S1 into the dilated space pyramid pooling module to perform multi-scale feature fusion; S3, feature fusion and prediction, upsampling the high-level features processed by the spatial pyramid pooling module, cascading the sampled data with the underlying features, and segmenting and outputting them through two branches; S4, training the model, regressing the two branches simultaneously, and weighting the loss functions of the two branches to optimize the model; S5. Use the model to process the data.

[0006] The specific method of image data processing in S1 includes: outlining the anatomical parts and image segmentation of the original CT image through the electronic plug-in, dividing the original CT image into 18 groups according to the anatomical parts, labeling the 18 groups with different colors, obtaining the labeled original data image, and generating a binary segmentation map according to the data of each category in the image, so that each original image corresponds to 18 segmentation maps containing 18 categories of data.

[0007] The basic features extracted from S1 include: Convolution features capture local texture information in the image through the local perception ability of the convolution kernel. Different convolution kernels focus on different feature patterns in the image. Structured features: detect the boundaries of objects in the image through convolutional networks, provide clear contour information, and obtain edge information of the image; Shape features: Through layer-by-layer convolution operations, the network gradually learns the overall structural features of the object, which is used to identify the shape patterns of different categories of targets; Multi-receptive field characteristics,The convolution operation realizes multi-receptive field processing by using convolution kernels of different sizes, using small receptive fields to focus on details and edge information, and using large receptive fields to capture overall texture and regional features.

[0008] The specific steps of S4 include: S41, standardize the input image data and load the data in certain batches; S42, the input image is forward propagated through the network, the bottom layer and high layer features are extracted, and the prediction results are generated through two branches respectively; S43, calculating the loss values ​​of the two branches in S42 according to the prediction results and the true labels; S44. Calculate the gradient of the weights and biases in the model parameters with respect to the loss function through back propagation, use the adam optimizer to update the model parameters, and adjust the learning rate.

[0009] In S42, the two branches include a first branch and a second branch; The first branch is used to comprehensively predict the overall segmentation map. The first branch merges the segmentation images from all channels and focuses on the overall contour of the target. The second branch is used to accurately segment each category, and the second branch tends to focus on the boundary information of a single category of the target.

[0010] The calculation formula for the two branch loss values ​​corresponding to the predicted results and the true labels in S43 is: Loss_total = α* L_focal + β * L_binary Among them, L_focal represents Focal Loss Loss Function, L_binary represents the Binary Cross-Entropy Loss loss function, and α and β are weight hyperparameters.

[0011] The model in S44 uses a learning rate decay strategy to adjust the learning rate parameters.

[0012] The system for automatic segmentation of CT cervical lymph nodes includes: Basic network, using DeepLabV3+ model as the basic network to support the whole system; Encoder: The feature extractor module in the encoder uses the xception network as the backbone network to extract image features. The spatial pyramid pooling module in the encoder uses dilated convolution and Inception structure to extract multi-scale features. The decoder is used to upsample high-level features and fuse the upsampled data with the underlying features.

[0013] The spatial pyramid pooling module in the encoder includes: 1x1 convolution is used to extract local information of input features, which plays a role in feature dimensionality reduction and refinement; Dilated convolution generates multiple feature subspaces by using dilated convolutions with different dilation rates. Different dilation rates correspond to different receptive fields, so as to capture multi-scale contextual information. Global average pooling is used to aggregate the global information of the input and obtain global features.

[0014] The atrous spatial pyramid pooling module concatenates the outputs of the 1x1 convolution branch, multiple atrous convolution branches, and the global average pooling branch in the channel dimension, and further fuses and compresses the concatenated features through an additional 1x1 convolution to generate the final output features.

[0015] The beneficial effects of the present invention are as follows: This solution uses deep learning technology and the advanced Deeplab V3+ model as the basic framework to regard the lymph node segmentation task in the cervical CT image as a typical image segmentation problem, which can effectively capture the detailed information in the image, thereby realizing an accurate image segmentation model structure. At the same time, targeted optimization is also carried out on this basis to improve the segmentation accuracy and model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is the flow chart of model training in this scheme; Figure 2 It is the flow chart of the model simulation in this scheme; Figure 3 is the image of the sample before processing in this embodiment; Figure 4 is the image after sample processing in this embodiment. DETAILED DESCRIPTION

[0017] The technical solution of the present invention is further described below through specific embodiments.

[0018] Example like Figure 1 and Figure 2 As shown in FIG. 1 , the method for automatic segmentation of CT cervical lymph nodes includes: S1. Process the image data, extract basic features through deep convolutional neural network, and replace the original DCNN network with U-Net Efficient network; S2, multi-scale feature fusion, input the features extracted in S1 into the dilated space pyramid pooling module to perform multi-scale feature fusion; S3, feature fusion and prediction, upsampling the high-level features processed by the spatial pyramid pooling module, cascading the sampled data with the underlying features, and segmenting and outputting them through two branches; S4, training the model, regressing the two branches simultaneously, and weighting the loss functions of the two branches to optimize the model; S5. Use the model to process the data.

[0019] The U-Net Efficient network in S1 optimizes the network's computational efficiency and segmentation accuracy by integrating the efficient architecture design of EfficientNet and the contextual connection method of U-Net. At the same time, feature extraction is not limited to simple convolution features, but also includes in-depth analysis of structured information such as texture, edge, and shape in the image.

[0020] The specific method of image data processing in S1 includes: in this embodiment, 3D Slicer is used to delineate anatomical parts and segment the original CT image, and the image is divided into 18 groups according to the anatomical parts, and the 18 groups are marked with different colors, and finally the marked original data image is obtained. For each category in the input data, the original image is converted into a binary image, that is, two pixel values ​​of 0 and 1.

[0021] Binarization can clearly distinguish the target and background in the image, which is convenient for the subsequent segmentation task. A binary segmentation map is generated for each category of data. It finally generates data containing 18 categories, and each category of data corresponds to the 18 segmentation maps of each original image. The above processing method not only maintains the integrity of the original information, but also provides the model with refined category features in a hierarchical manner.

[0022] Furthermore, the basic features extracted from S1 include: Convolution features capture local texture information in the image through the local perception ability of the convolution kernel. Different convolution kernels focus on different feature patterns in the image. Structured features: detect the boundaries of objects in the image through convolutional networks, provide clear contour information, and obtain edge information of the image; Shape features: Through layer-by-layer convolution operations, the network gradually learns the overall structural features of the object, which is used to identify the shape patterns of different categories of targets; Multi-receptive field characteristics,The convolution operation realizes multi-receptive field processing by using convolution kernels of different sizes, using small receptive fields to focus on details and edge information, and using large receptive fields to capture overall texture and regional features.

[0023] Convolution features use the local perception ability of convolution kernels to capture local texture information in images, such as color changes, line shapes, etc. Different convolution kernels focus on different feature patterns in images. For example, some convolution kernels are better at extracting edge features, while others focus on textures within the region.

[0024] Structured features are used to characterize edge information, detect the boundaries of objects in the image through convolutional networks, and provide clear contour information.

[0025] Shape features: Through layer-by-layer convolution operations, the network gradually learns the overall structural features of the object and recognizes the shape patterns of targets of different categories.

[0026] Multi-receptive field characteristics, the convolution operation realizes multi-receptive field processing by using convolution kernels of different sizes. Among them, the small receptive field focuses on details and edge information. The large receptive field captures the overall texture and regional characteristics.

[0027] S2 performs multi-scale fusion of the features extracted in S1 to improve segmentation performance. Its specific functions are as follows: Large-scale features are used to focus on the overall outline and macroscopic structure of the target. Large-scale features provide contextual information to the network and help identify large-area targets.

[0028] Small-scale features are used to focus on detail information, such as target edges, local textures, etc., to ensure that the model can accurately segment small objects or complex areas.

[0029] Fusion effect: The complementary fusion of large and small scale features makes the model more robust when dealing with complex scenes. For example, when the segmentation edge has blurred edges or complex textures, the large-scale features can provide global constraints, while the small-scale features make up for the lack of details in the local information.

[0030] The upsampling method in S3 is the bilinear interpolation sampling method. The cascade fusion method is to splice the channels of two features of the same size.

[0031] The specific steps of S4 include: S41, standardize the input image data and load the data in certain batches; S42, the input image is forward propagated through the network, the bottom layer and high layer features are extracted, and the prediction results are generated through two branches respectively; S43, calculating the loss values ​​of the two branches in S42 according to the prediction results and the true labels; S44. Calculate the gradient of the weights and biases in the model parameters with respect to the loss function through back propagation, use the adam optimizer to update the model parameters, and adjust the learning rate.

[0032] In S41, the input image data is standardized during the model training stage by subtracting the mean and dividing by the standard deviation.

[0033] Furthermore, the data is divided into training set, validation set and test set in a ratio of 8:1:1 to ensure that the data distribution of training data and test data is consistent.

[0034] At the same time, data enhancement techniques such as image scaling, flipping, scale transformation, and elastic deformation are used to increase data diversity, improve model robustness, and make the model better adaptable to different practical scenarios.

[0035] During the training process, the batch size is set to 16, and a single batch is used for processing during the inference phase to maintain the efficiency of inference. Data is loaded in batches and batched to ensure the stability and efficiency of gradient calculation during training. In this embodiment, the training data set contains 3002 images and the test set contains 600 images. The number of segmentation categories is 18, and all images are uniformly adjusted to a size of 512 x 512 pixels.

[0036] Further, in S42, the two branches include a first branch and a second branch; The first branch is used to comprehensively predict the overall segmentation map. The first branch merges the segmentation images from all channels and focuses on the overall contour of the target. The second branch is used to accurately segment each category, and the second branch tends to focus on the boundary information of a single category of the target.

[0037] Furthermore, the calculation formula for the two branch loss values ​​corresponding to the prediction result and the true label in S43 is: Loss_total = α* L_focal + β * L_binary Among them, L_focal represents Focal Loss Loss Function , L_binary represents the Binary Cross-Entropy Loss loss function, and α and β are weight hyperparameters.

[0038] Furthermore, in this embodiment, a relatively large learning rate is used for rapid convergence during the initial stage of model training, and a learning rate decay strategy, namely, a cosine annealing algorithm, is used in the later stage to fine-tune the parameters.

[0039] Preferably, the optimizer may also use existing electronic plug-ins such as SGD or RMSProp for better results.

[0040] The system for automatic segmentation of CT cervical lymph nodes includes: Basic network, using DeepLabV3+ model as the basic network to support the whole system; Encoder: The feature extractor module in the encoder uses the xception network as the backbone network to extract image features. The spatial pyramid pooling module in the encoder uses dilated convolution and Inception structure to extract multi-scale features. The decoder is used to upsample high-level features and fuse the upsampled data with the underlying features.

[0041] The spatial pyramid pooling module in the encoder includes: 1x1 convolution is used to extract local information of input features, which plays a role in feature dimensionality reduction and refinement. 1x1 convolution can effectively process small-scale feature information.

[0042] Dilated convolution generates multiple feature subspaces by using dilated convolutions with different dilation rates. Different dilation rates correspond to different receptive fields, so as to capture multi-scale context information. The dilation rate can be 6, 12, 18, etc.

[0043] Global average pooling (GAP) is used to aggregate the global information of the input and obtain global features. The output of global average pooling further compresses the feature dimension through a 1x1 convolution and uses bilinear interpolation to restore it to the same feature size as other branches. This branch is used to supplement the network's ability to perceive global context.

[0044] The atrous spatial pyramid pooling module concatenates the outputs of the 1x1 convolution branch, multiple atrous convolution branches, and the global average pooling branch in the channel dimension, and further fuses and compresses the concatenated features through an additional 1x1 convolution to generate the final output features.

[0045] Preferably, the system is deployed on a single CPU, using a Brower-Server architecture, with the front end accessed through a modern browser, and the back-end service efficiently deployed using Gradio. This deployment method not only ensures ease of use for users, but also ensures the scalability and stability of the system to meet the needs of different users.

[0046] In this embodiment, the method-level system provided in this solution is used to process the data sample by taking the parameters in the following table as an example. The comparison results are as follows: Figure 3 and Figure 4 shown.

[0047]

[0048] The embodiments of the present invention are described in detail above, but the contents described are only preferred embodiments of the present invention and cannot be considered to limit the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.

Claims

1. A method for automatic segmentation of CT cervical lymph nodes, characterized in that: Specific methods include: S1. Process the image data, extract basic features through deep convolutional neural network, and replace the original DCNN network with U-Net Efficient network; S2, multi-scale feature fusion, input the features extracted in S1 into the dilated space pyramid pooling module to perform multi-scale feature fusion; S3, feature fusion and prediction, upsampling the high-level features processed by the spatial pyramid pooling module, cascading the sampled data with the underlying features, and segmenting and outputting them through two branches; S4, training the model, regressing the two branches simultaneously, and weighting the loss functions of the two branches to optimize the model; S5. Use the model to process the data.

2. The method for automatic segmentation of CT cervical lymph nodes according to claim 1, characterized in that: The specific method of image data processing in S1 includes: outlining the anatomical parts and image segmentation of the original CT image through the electronic plug-in, dividing the original CT image into 18 groups according to the anatomical parts, labeling the 18 groups with different colors, obtaining the labeled original data image, and generating a binary segmentation map according to the data of each category in the image, so that each original image corresponds to 18 segmentation maps containing 18 categories of data.

3. The method for automatic segmentation of CT cervical lymph nodes according to claim 1, characterized in that: The basic features extracted from S1 include: Convolution features capture local texture information in the image through the local perception ability of the convolution kernel. Different convolution kernels focus on different feature patterns in the image. Structured features: detect the boundaries of objects in the image through convolutional networks, provide clear contour information, and obtain edge information of the image; Shape features: Through layer-by-layer convolution operations, the network gradually learns the overall structural features of the object, which is used to identify the shape patterns of different categories of targets; Multi-receptive field characteristics,The convolution operation realizes multi-receptive field processing by using convolution kernels of different sizes,using small receptive fields to focus on details and edge information,and using large receptive fields to capture overall texture and regional features.

4. The method for automatic segmentation of CT cervical lymph nodes according to claim 1, characterized in that: The specific steps of S4 include: S41, standardize the input image data and load the data in certain batches; S42, the input image is forward propagated through the network, the bottom layer and high layer features are extracted, and the prediction results are generated through two branches respectively; S43, calculating the loss values ​​of the two branches in S42 according to the prediction results and the true labels; S44. Calculate the gradient of the weights and biases in the model parameters with respect to the loss function through back propagation, use the adam optimizer to update the model parameters, and adjust the learning rate.

5. The method for automatic segmentation of CT cervical lymph nodes according to claim 4, characterized in that: In S42, the two branches include a first branch and a second branch; The first branch is used to comprehensively predict the overall segmentation map. The first branch merges the segmentation images from all channels and focuses on the overall contour of the target. The second branch is used to accurately segment each category, and the second branch tends to focus on the boundary information of a single target category.

6. The method for automatic segmentation of CT cervical lymph nodes according to claim 4, characterized in that: The calculation formula for the two branch loss values ​​corresponding to the predicted results and the true labels in S43 is: Loss_total = α* L_focal + β * L_binary Among them, L_focal represents Focal Loss Loss Function , L_binary represents the Binary Cross-EntropyLoss loss function, and α and β are weight hyperparameters.

7. The method for automatic segmentation of CT cervical lymph nodes according to claim 2, characterized in that: The model in S44 uses a learning rate decay strategy to adjust the learning rate parameters.

8. A system for automatic segmentation of CT cervical lymph nodes, used to implement the method for automatic segmentation of CT cervical lymph nodes according to any one of claims 1 to 7, characterized in that: include: Basic network, using DeepLabV3+ model as the basic network to support the whole system; Encoder: The feature extractor module in the encoder uses the xception network as the backbone network to extract image features. The spatial pyramid pooling module in the encoder uses dilated convolution and Inception structure to extract multi-scale features. The decoder is used to upsample high-level features and fuse the upsampled data with the underlying features.

9. The system for automatic segmentation of CT cervical lymph nodes according to claim 8, characterized in that: The spatial pyramid pooling module in the encoder includes: 1x1 convolution is used to extract local information of input features, which plays a role in feature dimensionality reduction and refinement; Dilated convolution generates multiple feature subspaces by using dilated convolutions with different dilation rates. Different dilation rates correspond to different receptive fields, so as to capture multi-scale contextual information. Global average pooling is used to aggregate the global information of the input and obtain global features.

10. The system for automatic segmentation of CT cervical lymph nodes according to claim 9, characterized in that: The atrous spatial pyramid pooling module concatenates the outputs of the 1x1 convolution branch, multiple atrous convolution branches, and the global average pooling branch in the channel dimension, and further fuses and compresses the concatenated features through an additional 1x1 convolution to generate the final output features.

Citation Information

Patent Citations

  • A dermatoscope image segmentation method based on a multi-branch convolutional neural network

    CN109886986A

  • Transform-based semi-supervised urban street view segmentation method

    CN117953216A

  • CT (Computed Tomography) image tumor segmentation method and device combined with convolutional network and Transform, and medium

    CN118229981A