CT image whole body multi-tumor segmentation method

By combining the pre-trained organ segmentation model and 3D convolutional neural network, text prompting and organ guidance mechanism of large language models are introduced, and the balance problem of global and local information in the whole body multi-tumor segmentation of CT images is solved, improving the segmentation accuracy and model adaptability.

CN120298433AActive Publication Date: 2025-07-11MACAO POLYTECHNIC INST

Patent Information

Application Number
CN202510414656.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing systemic multitumor segmentation methods of CT images are difficult to balance when dealing with global anatomical structure and local tumor details, lack high-level semantic information utilization of tumor type and location, and the scarce training data, resulting in insufficient generalization capabilities of the model and difficulty in adapting to diverse patient groups and complex situations.

Method used

The pre-trained organ segmentation model is used to combine 3D convolutional neural networks to integrate features through lightweight global feature encoder and complex local feature encoder, and the text prompt and organ guidance mechanism of large language models are introduced to improve the model's ability to capture tumor-specific features.

Benefits of technology

The accuracy of systemic multi-tumor segmentation is improved, and the ability to accurately locate and segment multiple tumors in complex situations is achieved, improving the robustness and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298433A_ABST
    Figure CN120298433A_ABST
Patent Text Reader

Abstract

The invention discloses a CT image whole-body multi-tumor segmentation method, which comprises the following steps: firstly, preprocessing collected CT images of multiple parts and multiple types of tumors, then inputting the preprocessed images into a pre-trained organ segmentation model, and obtaining segmentation results of different organs and local tumor images of different parts; thirdly, extracting and fusing features by utilizing global and local tumor segmentation models, and enhancing extraction of tumor specific features by taking an organ segmentation result as guide information and combining tumor text prompts generated by a large language model; and finally, carrying out step-by-step up-sampling through a decoder to obtain an accurate tumor segmentation image. According to the invention, global and local features are integrated, and an organ segmentation result is used as guidance to form a position attention module, so that the adaptability of the model to different organ tumors is enhanced; through a tumor text prompt driving mechanism of a large language model, specific hyper-parameters for each tumor are generated, and the accuracy and robustness of the model in a multi-part and multi-type tumor segmentation task are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical computer diagnostic image processing, and particularly relates to a method for segmenting multiple tumors in the whole body from CT images. Background Art

[0002] In recent years, with the rapid development of deep learning technology, significant progress has been made in the field of medical image analysis. However, the automatic segmentation of multiple tumors in whole-body CT images still faces many challenges. First, whole-body CT scans usually contain a large number of anatomical structures, and the complexity of different organs and tissues makes it difficult to accurately identify and segment tumors. Second, tumors may exhibit high heterogeneity among different patients and at different locations in the same patient, which increases the difficulty of developing a general segmentation algorithm.

[0003] Most of the existing methods for segmenting tumors from CT images are based on single-organ and single-task designs, and the models are designed and optimized for specific organs or local regions. Although these methods have achieved certain results in specific scenarios, their application scope is limited and they cannot adapt to the complex situation of multiple organs and multiple tumors in the whole body. Although some studies have attempted to use multi-stage or cascaded models to process whole-body scans, these methods are often computationally complex and difficult to apply in real time in clinical practice. Therefore, in the face of different types of tumors and diverse patient populations, existing methods often need to redesign and train new models, which increases the development cost and time.

[0004] In addition, the existing methods also face the following challenges when processing whole-body CT images: First, it is difficult to effectively balance the global anatomical structure information and local tumor details; second, there is a lack of full utilization of high-level semantic information such as tumor type and location; finally, they perform poorly when dealing with rare tumor types and imbalanced datasets. These limitations seriously affect the wide application of the whole-body multiple tumor segmentation technology for CT images in clinical practice.

[0005] On the other hand, the acquisition and annotation costs of medical image data are high, and high-quality whole-body multiple tumor datasets are relatively scarce, which limits the training effect of deep learning models, especially for rare types of tumors. How to effectively utilize limited annotated data while improving the generalization ability of the model for different types of tumors has become an urgent problem to be solved. Summary of the Invention

[0006] The purpose of the present invention is to develop a method for segmenting multiple tumors in the whole body from CT images that can simultaneously process global and local information, fuse semantic knowledge, and effectively utilize limited training data.

[0007] To achieve the above object, the present invention proposes a method for segmenting multiple tumors in the whole body from CT images and adopts the following technical solutions:

[0008] A method for whole-body multi-tumor segmentation of CT images, the method comprising the following steps:

[0009] Step 1, collect a global calibration dataset of CT images of tumors in different parts and of different types, and preprocess each of the collected CT images to form a sample set;

[0010] Step 2, construct a tumor segmentation model structure, including a pre-trained organ segmentation model, a global and local tumor segmentation model structure; the pre-trained organ segmentation model uses a 3D nnU-Net model; the global and local tumor segmentation model structure uses a 3D convolutional neural network structure, including a lightweight global feature encoder, a complex local feature encoder, a feature fusion module, a position attention module and a decoder. The organ guidance mechanism extracts the position information of the target organ through the position attention module, and introduces tumor text prompts generated by a large language model to guide the extraction of specific features of the local tumor. The feature fusion module fuses the lightweight global features with the complex local features, and finally outputs through step-by-step upsampling by the decoder;

[0011] Step 3, input each preprocessed CT image into the pre-trained organ segmentation model to obtain a target organ segmentation image and a target tumor local image; use each target organ segmentation image as the input of the complex local feature encoder, use the corresponding preprocessed CT image as the input of the lightweight global feature encoder, use the position information extracted by the organ guidance mechanism through the position attention module to guide the model to focus on the tumor features in the target organ, use the text prompts generated by the large language model based on the target tumor local image to guide the model to extract the specific features of the local tumor, and use the corresponding tumor precise segmentation image as the output to train the constructed global and local tumor segmentation model structure to obtain a trained global and local tumor segmentation model; the pre-trained organ segmentation model and the global and local tumor segmentation model constitute a tumor segmentation model;

[0012] Step 4, use the preprocessed CT image to be segmented as the input, apply the tumor segmentation model, and obtain the corresponding tumor precise segmentation image.

[0013] Furthermore, the global calibration dataset of CT images collected in Step 1 includes tumor data of different types in different organs such as bone, pancreas, kidney, colon, lung, liver, esophagus. Each CT image contains a CT scan of the whole organ, showing the positional relationship of the tumor in the overall anatomical structure, and the sizes and shapes of the tumors in the dataset are different; the preprocessing includes standardizing the physical distance and direction to ensure the consistency of the scales and orientations of different CT images, converting CT images in different formats into a unified input format, and integrating the local and global images of the same tumor in the same organ so that they can be used as the input of the pre-trained organ segmentation model at the same time.

[0014] Furthermore, in step 2, the pre-trained organ segmentation model selects the 3D nnU-Net model. Both its encoder and decoder use a single grayscale channel. The number of channels of the convolutional kernels used in the downsampling process of the encoder is successively 16, 32, 64, 128, and 256, and the corresponding image sizes after processing are successively 96×96×96, 48×48×48, 24×24×24, 12×12×12, and 6×6×6. The features extracted by the encoder are fused to the feature maps of the corresponding sizes in the decoder through skip connections. The pre-trained organ segmentation model can accurately locate the target organ area, segment out the target organ where the tumor exists, and at the same time obtain the local image of the target tumor.

[0015] Furthermore, the global and local tumor segmentation model structure extracts the global features of the CT image through a lightweight global feature encoder and extracts the detailed features of the target organ segmentation image through a complex local feature encoder. During the encoding and decoding processes, text prompts of different tumor types and anatomical locations generated by the large language model based on the local image of the target tumor are introduced to enhance the model structure's capture of tumor-specific features. The lightweight global features and complex local features are fused through a feature fusion module. During the decoding process, an organ guidance mechanism is introduced, and feature extraction is performed through a position attention module to guide the model structure to focus on the tumor features of the target organ. Finally, the corresponding accurate tumor segmentation image is output through the decoder's gradual upsampling.

[0016] Furthermore, the large language model takes the local image of the target tumor as input, extracts the target tumor type and spatial position features through convolutional encoding, and generates the target tumor text vector. During the training process, the target tumor text vector is introduced into the encoding and decoding processes of the global and local tumor segmentation model structure to guide the global and local tumor segmentation model structure to accurately capture the specific features of different tumors during the training process.

[0017] Furthermore, the proposed organ guidance mechanism takes the segmentation image of the target organ as input and performs feature extraction through a position attention module. According to the shape, size, and position of the target organ in the image, different attention weights are assigned to each position to focus on the key features related to the tumor within the target organ, and the extracted features are fused into the decoder structure, thereby enhancing the model structure's attention to the tumor features inside the target organ, suppressing interference from non-target organs, and improving the segmentation accuracy.

[0018] The present invention also protects a computer device, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete mutual communication through the communication bus. The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above-mentioned CT image whole-body multi-tumor segmentation method.

[0019] The present invention also protects a computer-readable storage medium, in which at least one executable instruction is stored, and the executable instruction causes a processor to perform the operations corresponding to the above CT image whole-body multi-tumor segmentation method.

[0020] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0021] The CT image whole-body multi-tumor segmentation method proposed by the present invention first uses a pre-trained organ segmentation model to accurately locate the organ where the tumor is located and extract the corresponding local tumor region, and then performs global and local feature extraction through a 3D convolutional neural network, realizing the effective fusion of global and local features; using the target organ segmentation result as guiding information not only helps the model more accurately identify the boundary and position of the target tumor, but also enhances the adaptability of the model to tumors in different organs; in addition, the model also introduces a text prompt-driven mechanism based on a large language model to generate specific hyperparameters for different tumor sites, further improving the accuracy and robustness of the segmentation task in complex situations; the CT image whole-body multi-tumor segmentation method proposed by the present invention significantly improves the segmentation accuracy of multi-site and multi-type tumors by combining the organ segmentation result and the tumor local region information. Description of the Drawings

[0022] Figure 1 is a flowchart of the CT image whole-body multi-tumor segmentation method provided by the present invention;

[0023] Figure 2 is a network framework diagram of the tumor segmentation model provided by the present invention;

[0024] Figure 3 is a structural diagram of the pre-trained organ segmentation model provided by the present invention;

[0025] Figure 4 is a framework diagram of the organ guidance mechanism provided by the present invention. Detailed Embodiments

[0026] The technical solutions of the present invention will be further described in detail below in conjunction with the drawings and specific embodiments.

[0027] As Figure 1 shown, a CT image whole-body multi-tumor segmentation method provided by the present invention includes the following steps:

[0028] Step 1: Collect a global calibration data set of CT images of different parts and different types of tumors, and preprocess each collected CT image to form a sample set;

[0029] Step 2: Construct the tumor segmentation model structure, including a pre-trained organ segmentation model, and global and local tumor segmentation model structures. The pre-trained organ segmentation model uses a 3D nnU-Net model. The global and local tumor segmentation model structures use a 3D convolutional neural network structure, including a lightweight global feature encoder, a complex local feature encoder, a feature fusion module, a position attention module, and a decoder. The organ guidance mechanism extracts the position information of the target organ through the position attention module and introduces tumor text prompts generated by the large language model to guide the extraction of specific features of the local tumor. The feature fusion module fuses the lightweight global features and the complex local features, and finally outputs through step-by-step upsampling by the decoder.

[0030] Step 3: Input the pre-processed CT images into the pre-trained organ segmentation model to obtain the target organ segmentation images and target tumor local images. Use the target organ segmentation images as the input of the complex local feature encoder, the corresponding pre-processed CT images as the input of the lightweight global feature encoder, use the position information extracted by the organ guidance mechanism through the position attention module to guide the model to focus on the tumor features in the target organ, use the text prompts generated by the large language model based on the target tumor local images to guide the model to extract specific features of the local tumor, and take the corresponding tumor precise segmentation images as the output to train the constructed global and local tumor segmentation model structure to obtain the trained global and local tumor segmentation model. The pre-trained organ segmentation model and the global and local tumor segmentation model constitute the tumor segmentation model.

[0031] Step 4: Use the pre-processed CT image to be segmented as the input, apply the tumor segmentation model, and obtain the corresponding tumor precise segmentation image.

[0032] The global calibration dataset of CT images collected in Step 1 includes different types of tumor data of different organs such as bone, pancreas, kidney, colon, lung, liver, and esophagus. Each CT image contains the CT scan of the entire organ, showing the positional relationship of the tumor in the overall anatomical structure, and the size and shape of the tumors in the dataset are diverse. The preprocessing includes standardizing the physical distance and direction to ensure the consistency of the scale and orientation of different CT images, converting CT images in different formats into a unified input format, and integrating the local and global images of the same tumor in the same organ so that they can be used as the input of the pre-trained organ segmentation model at the same time.

[0033] Such as Figure 3As shown, in step 2, the pre-trained organ segmentation model selects the 3D nnU-Net model. Both its encoder and decoder use a single grayscale channel. The number of channels of the convolution kernels used in the downsampling process of the encoder is 16, 32, 64, 128, and 256 in sequence, and the corresponding processed image sizes are 96×96×96, 48×48×48, 24×24×24, 12×12×12, and 6×6×6. The features extracted by the encoder are fused into the feature maps of the corresponding sizes in the decoder through skip connections. The pre-trained organ segmentation model can accurately locate the target organ area, segment the target organ where the tumor exists, and at the same time obtain the local image of the target tumor.

[0034] As Figure 2 shown, the pre-processed CT images are input into the pre-trained organ segmentation model to obtain the target organ segmentation image and the local image of the target tumor. The global and local tumor segmentation model structure extracts the global features of each CT image through a lightweight global feature encoder, extracts the detailed features of the target organ segmentation image through a complex local feature encoder, and at the same time introduces text prompts about different tumor types and anatomical locations generated by the large language model during the encoding and decoding processes. The lightweight global feature encoding and complex local feature encoding fuse the lightweight global features and complex local features through a feature fusion module. An organ guidance mechanism is introduced during the decoding process, and feature extraction is performed through a position attention module. Finally, the decoder gradually upsamples to output the accurate segmentation image of the target tumor.

[0035] As Figure 4 shown, the large prediction model takes the local image of the target tumor as input, extracts the type and spatial position information of the target tumor through convolutional encoding, and constructs a position attention mechanism in combination with the organ segmentation result. The organ segmentation result is used to guide the attention weights, enabling the model to more accurately focus on the tumor area inside the target organ and reducing interference from non-target organ areas.

[0036] During the training process, the position attention mechanism is introduced into the decoding process of the global and local tumor segmentation model, enabling the model to effectively capture the specific features of different tumor types. At the same time, the organ guidance mechanism uses the segmentation image of the target organ to extract the spatial position information of the target organ area through convolutional layer encoding, providing position guidance, enabling the model structure to more accurately focus on the tumor area inside the target organ and reducing interference from non-target organ areas.

[0037] The present invention also protects a computer device, including: a processor, a memory, a communication interface, and a communication bus. The processor, memory, and communication interface complete mutual communication through the communication bus. The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above CT image whole-body multi-tumor segmentation method.

[0038] The present invention also protects a computer-readable storage medium. At least one executable instruction is stored in the computer storage medium, and the executable instruction causes the processor to perform the operations corresponding to the above CT image whole-body multi-tumor segmentation method.

[0039] As described above, it is only the preferred specific implementation manner of the present invention and is not used to limit the present invention. Any modification, equivalent replacement, improvement, etc. made by any person skilled in the art within the technical scope disclosed by the present invention according to the technical solution and inventive concept of the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for whole-body multi-tumor segmentation of CT images, characterized in that, The method includes the following steps: Step 1: Collect the global calibration datasets of CT images of tumors in different parts and of different types, and preprocess each of the collected CT images to form a sample set; Step 2: Construct the tumor segmentation model structure, including a pre-trained organ segmentation model, and global and local tumor segmentation model structures; the pre-trained organ segmentation model uses a 3D nnU-Net model; the global and local tumor segmentation model structures use a 3D convolutional neural network structure, including a lightweight global feature encoder, a complex local feature encoder, a feature fusion module, a position attention module, and a decoder. The organ guidance mechanism extracts the position information of the target organ through the position attention module, and introduces the tumor text prompts generated by the large language model to guide the extraction of the specific features of the local tumor. The feature fusion module fuses the lightweight global features and the complex local features, and finally outputs through step-by-step upsampling by the decoder; Step 3: Input each preprocessed CT image into the pre-trained organ segmentation model to obtain the target organ segmentation image and the target local tumor image; use each target organ segmentation image as the input of the complex local feature encoder, use the corresponding preprocessed CT image as the input of the lightweight global feature encoder, use the position information extracted by the organ guidance mechanism through the position attention module to guide the model to focus on the tumor features in the target organ, use the text prompts generated by the large language model based on the target local tumor image to guide the model to extract the specific features of the local tumor, and use the corresponding accurate tumor segmentation image as the output to train the constructed global and local tumor segmentation model structure to obtain the trained global and local tumor segmentation model; the pre-trained organ segmentation model and the global and local tumor segmentation models constitute the tumor segmentation model; Step 4: Use the preprocessed CT image to be segmented as the input, apply the tumor segmentation model, and obtain the corresponding accurate tumor segmentation image.

2. The CT image whole-body multi-tumor segmentation method according to claim 1, wherein: The global calibration datasets of CT images collected in Step 1 include different types of tumor data of different organs such as bone, pancreas, kidney, colon, lung, liver, and esophagus. Each CT image contains the CT scan of the entire organ, showing the positional relationship of the tumor in the overall anatomical structure, and the sizes and shapes of the tumors in the dataset are different; the preprocessing includes standardizing the physical distance and direction to ensure the consistency of the scales and orientations of different CT images, converting CT images in different formats into a unified input format, and integrating the local and global images of the same tumor in the same organ so that they can be used as the input of the pre-trained organ segmentation model at the same time.

3. A method for whole-body multi-tumor segmentation of CT images according to claim 1, characterized in that: In step 2, the pre-trained organ segmentation model selects the 3D nnU-Net model. Both its encoder and decoder use single-channel grayscale. The number of channels of the convolutional kernels used in the downsampling process of the encoder is successively 16, 32, 64, 128, and 256, and the corresponding image sizes after image processing are successively 96×96×96, 48×48×48, 24×24×24, 12×12×12, and 6×6×6. The features extracted by the encoder are fused into the feature maps of the corresponding sizes in the decoder through skip connections. The pre-trained organ segmentation model can accurately locate the target organ area, segment the target organ where the tumor exists, and simultaneously obtain the local image of the target tumor.

4. A CT image whole-body multi-tumor segmentation method according to claim 3, characterized in that: The global and local tumor segmentation model structure extracts the global features of CT images through a lightweight global feature encoder and the detailed features of the target organ segmentation image through a complex local feature encoder. During the encoding and decoding processes, text prompts of different tumor types and anatomical locations generated by the large language model based on the local image of the target tumor are introduced to enhance the model structure's capture of tumor-specific features. The lightweight global features and complex local features are fused through a feature fusion module. During the decoding process, an organ guidance mechanism is introduced, and feature extraction is performed through a position attention module to guide the model structure to focus on the tumor features of the target organ. Finally, the decoder gradually upsamples to output the corresponding accurate tumor segmentation image.

5. A method for whole-body multi-tumor segmentation of CT images according to claim 4, characterized in that: The large language model takes the local image of the target tumor as input, extracts the target tumor type and spatial position features through convolutional encoding to generate the target tumor text vector. Subsequently, the target tumor text vector is introduced into the encoding and decoding processes of the global and local tumor segmentation model structure to guide the extraction of specific features of different tumors by the global and local tumor segmentation model structure during the training process.

6. A method for whole-body multi-tumor segmentation of CT images according to claim 5, characterized in that: The proposed organ guidance mechanism takes the segmentation image of the target organ as input and performs feature extraction through a position attention module. According to the shape, size, and position in the image of the target organ, different attention weights are assigned to each position, and the extracted features are fused into the decoder structure, thereby enhancing the model structure's attention to the tumor features inside the target organ and suppressing interference from non-target organs.

7. A computer device, characterized in that: It includes a processor, a memory, a communication interface, and a communication bus. The processor, memory, and communication interface complete mutual communication through the communication bus. The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the CT image whole-body multi-tumor segmentation method described in any one of claims 1-6.

8. A computer-readable storage medium, characterized in that: The computer storage medium stores at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the CT image whole-body multi-tumor segmentation method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Three-stage liver tumor image segmentation method based on adaptive preprocessing

    CN115018864A

  • Liver tumor segmentation method based on multi-temporal fusion and double attention mechanism

    CN115272357A

  • Lung CT image segmentation method based on global and local attention mechanisms

    CN117649385A

  • Liver tumor CT image segmentation method based on APA-UNet

    CN117746042A

  • Contrast-agent-free medical diagnostic imaging

    US20220208355A1

Cited By

  • CT image processing method, system and device, storage medium and program product

    CN121686174A