PET / CT Imaging-Based Lesion Segmentation, Fusion, and Calibration Method, Device, Medium, and Product
By adopting a convolutional neural network with a dual-branch encoding-decoding structure in PET-CT image processing, combined with a fusion downsampling module and a multimodal calibration module, the problems of low accuracy and poor robustness of automatic tumor segmentation in the PET-CT image in the prior art are solved, and a more accurate and robust tumor segmentation effect is achieved.
Patent Information
- Application Number
- CN202210625216.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-02
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-06-02
AI Technical Summary
In the prior art, the automatic tumor segmentation method based on PET-CT images has problems of low accuracy and poor robustness, especially when dealing with the accumulation of high metabolic healthy tissues and kidneys and bladder in PET images, it is difficult to accurately segment the tumor.
The deep segmentation technology based on convolutional neural network is adopted to design a dual-branch encoding-decoding structure, extract the semantic features of CT and PET images through a multimodal fusion segmentation network, and use the fusion downsampling module and the multimodal calibration module for information fusion and calibration to generate more accurate tumor segmentation results.
It improves the accuracy and robustness of automatic segmentation of a single CNN network, can more accurately segment tumors in PET-CT images, reduces interference with error messages in single-modal semantic features, and improves the performance of the model.
Smart Images

Figure CN115019041B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and particularly to a method, device, medium and product for lesion segmentation, fusion and calibration based on PET / CT imaging. Background Art
[0002] F-FDG positron emission tomography and computed tomography (PET-CT) are currently widely used in oncology. Due to the high glucose metabolism of diseased tissues, they usually show higher F-FDG uptake in PET images. And CT images can provide the anatomical structure of the abnormal uptake site. Therefore, PET-CT images integrate the advantages of both and are widely used in clinical practice (diagnosis, staging). At the same time, quantitative research based on PET-CT images shows great potential in clinical practice, such as efficacy prediction, radiomics analysis, treatment planning and prognosis evaluation, etc. Semantic segmentation is crucial in the quantitative analysis of medical images because it is a pioneer in these tasks. However, most current studies are based on the lesion regions manually outlined by radiologists. Manually segmenting lesions is very laborious and time-consuming, and is vulnerable to inter-observer and intra-observer variations. Therefore, designing an automatic and accurate tumor segmentation method has important clinical application value.
[0003] The automatic tumor segmentation method based on PET-CT images has received extensive attention. Due to the high sensitivity of PET images themselves, the initial research focused on tumor segmentation of single-modal images. The most widely used segmentation method in clinical practice is the SUV-based threshold method and related improved algorithms. In addition, researchers have also applied region growing algorithms and graph theory-based segmentation algorithms to PET tumor image segmentation. However, although PET images show high sensitivity to tumors, due to the non-specificity of 18F-FDG, high uptake also occurs in many healthy tissues with high metabolism (brain). And the accumulation of 18F-FDG in the kidneys and bladder leads to a high uptake state in PET images. In addition, due to the heterogeneity of tumors, the tumors of some patients will show a low uptake state in PET images. This is also a great challenge for the above-mentioned segmentation algorithms. And CT images belong to low-dose plain scan images, and it is almost impossible to locate the lesion region based on CT images.
[0004] Research shows that combining multi-modal images in medical image processing can provide more information, so designing an effective multi-modal fusion segmentation network is the key to accurate segmentation. Summary of the Invention
[0005] To achieve the above objects and other advantages according to the present invention, the first object of the present invention is to provide a method for lesion segmentation, fusion and calibration based on PET / CT imaging, including the following steps:
[0006] Extract semantic features. Extract the semantic features of the CT image through the first encoder in the multi-modal fusion segmentation network, and extract the semantic features of the PET image through the second encoder in the multi-modal fusion segmentation network;
[0007] Fuse semantic features. Extract the fused semantic features of the semantic features of the CT image and the PET image through the fusion downsampling module in the multi-modal fusion segmentation network to obtain the fused semantic features of the CT image and the fused semantic features of the PET image;
[0008] Decode semantic features. Upsample the fused semantic features of the CT image through the first decoder in the multi-modal fusion segmentation network, connect the upsampling result of the fused semantic features of the CT image with the semantic features of the CT image, and decode the connected features. Upsample the fused semantic features of the PET image through the second decoder in the multi-modal fusion segmentation network, connect the upsampling result of the fused semantic features of the PET image with the fused semantic features of the PET image, and decode the connected features;
[0009] Calibrate semantic features. Calibrate the semantic features of the CT image and the PET image obtained by decoding through the multi-modal calibration block in the multi-modal fusion segmentation network, and fuse the calibrated semantic features of the CT image and the PET image with the original semantic features of the CT image and the PET image to obtain updated semantic features of the CT image and the PET image;
[0010] Generate a probability map. Add the lesion images reconstructed by the first decoder and the second decoder branches, and generate an output probability map of the lesion through an activation function;
[0011] Generate pixel classification. Generate the classification of each pixel by setting a probability map threshold.
[0012] Furthermore, both the first encoder and the second encoder include a number of convolutional layers. The fusion downsampling module is arranged between adjacent convolutional layers of the first encoder and between adjacent convolutional layers of the second encoder. The outputs of the current convolutional layer of the first encoder and the corresponding convolutional layer of the second encoder are used as the inputs of the fusion downsampling module, and the output of the fusion downsampling module is used as the input of the next convolutional layer of the first encoder and the corresponding convolutional layer of the second encoder.
[0013] Furthermore, the step of fusing the semantic features includes the following steps:
[0014] Fuse the output of the current convolutional layer of the first encoder and the output of the corresponding convolutional layer of the second encoder through the fusion downsampling module;
[0015] Reduce the number of channels of the fused features through the first convolutional layer of the fusion downsampling module, then reduce the feature maps through the second convolutional layer of the fusion downsampling module, and then improve the non-linear representation of the model through a leaky rectified linear unit and several convolutional layers to obtain output features;
[0016] Downsample the output of the current convolutional layer of the first encoder through the first max pooling layer of the fusion downsampling module, add the downsampling result to the output features to obtain the first fused feature, downsample the output of the corresponding convolutional layer of the second encoder through the second max pooling layer of the fusion downsampling module, and add the downsampling result to the output features to obtain the second fused feature;
[0017] Input the first fused feature and the second fused feature into the next convolutional layer of the first encoder and the corresponding convolutional layer of the second encoder respectively.
[0018] Furthermore, both the first decoder and the second decoder include an upsampling layer and several convolutional layers. The multi-modal calibration module is arranged between adjacent convolutional layers of the first decoder and between adjacent convolutional layers of the second decoder. Take the outputs of the current convolutional layer of the first decoder and the corresponding convolutional layer of the second decoder as the inputs of the multi-modal calibration module, and take the output of the multi-modal calibration module as the inputs of the next convolutional layer of the first decoder and the corresponding convolutional layer of the second decoder.
[0019] Furthermore, the step of calibrating semantic features includes the following steps:
[0020] Activate the output of the current convolutional layer of the first decoder through the first activation function of the multi-modal calibration module to obtain the activated semantic features of the CT image, and activate the output of the corresponding convolutional layer of the second decoder through the second activation function of the multi-modal calibration module to obtain the activated semantic features of the PET image;
[0021] Feed the activated semantic features of the CT image and the PET image to the multi-scale feature extraction block of the multi-modal calibration module to calibrate the semantic features of the current PET image;
[0022] Feed the activated semantic features of the PET image and the CT image to the multi-scale feature extraction block of the multi-modal calibration module to calibrate the semantic features of the current CT image;
[0023] Fuse the calibrated semantic features of the PET image with the original semantic features of the PET image to obtain updated semantic features of the PET image;
[0024] Fuse the semantic features of the calibrated CT image with the semantic features of the original CT image to obtain the updated semantic features of the CT image.
[0025] Furthermore, the multi-scale feature extraction block includes a number of parallel convolutional layers. The semantic features of the activated PET image and the semantic features of the CT image are fused after being processed by the leaky rectified linear unit, the parallel convolutional layers, and group normalization of the multi-scale feature extraction block; the semantic features of the activated CT image and the semantic features of the PET image are fused after being processed by the leaky rectified linear unit, the parallel convolutional layers, and group normalization of the multi-scale feature extraction block.
[0026] Furthermore, in the step of generating pixel classification, set the probability map threshold to 0.5 to generate a binary classification for each pixel.
[0027] The second object of the present invention is to provide an electronic device, including: a memory on which program code is stored; a processor coupled to the memory, and when the program code is executed by the processor, a lesion segmentation fusion calibration method based on PET / CT imaging is implemented.
[0028] The third object of the present invention is to provide a computer-readable storage medium on which program instructions are stored, and when the program instructions are executed, a lesion segmentation fusion calibration method based on PET / CT imaging is implemented.
[0029] The fourth object of the present invention is to provide a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, a lesion segmentation fusion calibration method based on PET / CT imaging is implemented.
[0030] Compared with the prior art, the beneficial effects of the present invention are:
[0031] The present invention adopts a deep segmentation technology based on a convolutional neural network and a double-branch encoding-decoding structure, which not only retains the anatomical structure features of the CT image but also retains the physiological information contained in the PET image, thereby improving the accuracy and robustness of automatic segmentation of a single CNN network.
[0032] When integrating multi-modal semantic information without screening, it may bring some irrelevant information or even incorrect information, thus affecting the performance of the model. Therefore, the present invention proposes a new network structure for segmenting lesions in PET-CT images, which can achieve the fusion and calibration of multi-modal information. In the feature encoding stage, multiple encoders are used to extract the semantic features of different modal images. At the same time, a fusion downsampling module is used to extract the fused semantic features, reducing the interference of incorrect information in the single-modal semantic features. In the decoding stage, semantic features of different scales are fused through skip connections. Then, the semantic features of CT and PET images are fused through a multi-modal calibration module to achieve the mutual calibration between different modal information.
[0033] The above description is only an overview of the technical solution of the present invention. In order to understand the technical means of the present invention more clearly and implement it in accordance with the content of the specification, the following takes the preferred embodiments of the present invention and combines with the drawings to elaborate in detail as follows. The specific implementation manners of the present invention are given in detail by the following embodiments and their accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0035] Figure 1 It is a flowchart of the lesion segmentation fusion calibration method based on PET / CT imaging in Embodiment 1;
[0036] Figure 2 It is a schematic diagram of a 3D double-branch segmentation network based on the U-Net architecture in Embodiment 1;
[0037] Figure 3 It is a schematic diagram of the first two convolutional layers of the encoder in Embodiment 1;
[0038] Figure 4 It is a schematic diagram of the last three convolutional layers of the encoder in Embodiment 1;
[0039] Figure 5 It is a schematic diagram of the decoder module in Embodiment 1;
[0040] Figure 6 It is a schematic diagram of the convolutional block in Embodiment 1;
[0041] Figure 7 It is a schematic diagram of the fusion downsampling module in Embodiment 1;
[0042] Figure 8 It is a schematic diagram of the multi-modal calibration module in Embodiment 1;
[0043] Figure 9Schematic diagram of the multi-mode interaction module in Embodiment 1;
[0044] Figure 10 Schematic diagram of the electronic device in Embodiment 2;
[0045] Figure 11 Schematic diagram of the computer-readable storage medium in Embodiment 3. Detailed implementation manners
[0046] Next, in combination with the accompanying drawings and specific implementation manners, the present invention will be further described. It should be noted that, on the premise of non-conflict, any combination can be formed among the following-described embodiments or technical features to form a new embodiment.
[0047] Embodiment 1
[0048] A lesion segmentation and fusion calibration method based on PET / CT imaging uses a multi-modal fusion segmentation network, as Figure 1 shown, and includes the following steps:
[0049] Extract semantic features. Extract the semantic features of the CT image through the first encoder in the multi-modal fusion segmentation network, and extract the semantic features of the PET image through the second encoder in the multi-modal fusion segmentation network. In this embodiment, the multi-modal fusion segmentation network is a 3D double-branch segmentation network based on the U-Net architecture. The skip connections in the UNet network can achieve the fusion of high-level semantic features and low-level semantic features, thereby improving the accuracy of the model. Just as a doctor makes a clinical decision, both the global information of the lesion and the local position need to be considered to provide detailed information. The U-Net structure conforms to the clinical thinking. Therefore, it is selected as the basic network for this embodiment.
[0050] As Figure 2 shown, the 3D double-branch segmentation network based on the U-Net architecture consists of four parts: an encoder Encoder, a decoder Decoder, a fusion downsampling module FDSB, and a multi-modal calibration block MMCB. Two parallel encoders are used to extract the semantic feature images of PET and the semantic features of the CT image respectively in the encoder part.
[0051] Both the first encoder and the second encoder include several convolutional layers. The fusion downsampling module is arranged between adjacent convolutional layers of the first encoder and between adjacent convolutional layers of the second encoder. The outputs of the current convolutional layer of the first encoder and the corresponding convolutional layer of the second encoder are used as the input of the fusion downsampling module, and the output of the fusion downsampling module is used as the input of the next convolutional layer of the first encoder and the corresponding convolutional layer of the second encoder.
[0052] As Figure 2 shown, the encoder contains five convolutional layers, as Figure 3As shown, the first two convolutional layers contain two convolutions with a kernel size of 3*3*3 and a stride of 1. As Figure 4 shown, each of the last three convolutional layers is designed as a residual block, which contains three convolutions with a kernel size of 3*3*3 and a stride of 1.
[0053] For multi-encoder image segmentation networks, the currently common method is to use multiple encoder branches to independently extract the information of each modality image. However, it is difficult for this method to make the PET image and the CT image benefit from each other during the feature extraction process. In addition, the false positives of the PET image and the low contrast of the CT image (lesion and non-lesion areas in the tissue) may face the challenge of incorrect lesion localization, resulting in the entire network being unable to accurately segment the lesion. Therefore, in this embodiment, a fusion downsampling module is designed. Compared with the downsampling of single-modal image features, FDSB can reduce the influence of incorrect information by fusing the information of other modalities. In the extraction of the semantic features of the PET image, the fusion of multi-modal information helps to quickly and accurately locate the lesion area of the PET image. For the extraction of the semantic features of the CT image, the fusion of multi-modal information helps to distinguish the lesion area in the CT feature map that has similar anatomical features to the surrounding tissues.
[0054] Fuse the semantic features, and extract the fused semantic features of the semantic features of the CT image and the semantic features of the PET image through the fusion downsampling module in the multi-modal fusion segmentation network, to obtain the fused semantic features of the CT image and the fused semantic features of the PET image; specifically, it includes the following steps:
[0055] Fuse the output of the current convolutional layer of the first encoder and the output of the corresponding convolutional layer of the second encoder through the fusion downsampling module; Figure 2 In, the image on the left is the CT image, and the image on the right is the PET image. The CT image is input into the fusion downsampling module FDSB through the encoder module Encoder1, and the PET image is input into the fusion downsampling module FDSB through the decoder module Decoder1. The fusion downsampling module FDSB outputs to the encoder module Encoder2 and the decoder module Decoder2, and so on. As Figure 7 shown, the fusion downsampling module FDSB first concatenates the semantic features of the CT image and the semantic features of the PET image.
[0056] Reduce the number of channels of the fused features through the first convolutional layer of the fusion downsampling module for the fused features, then reduce the feature map through the second convolutional layer of the fusion downsampling module, and then improve the non-linear representation of the model through the leaky rectified linear unit LeakyReLU and several convolutional layers to obtain the output features; as Figure 7As shown, first, a 1×1×1 convolution is used to reduce the number of feature channels after concatenation, facilitating cross-modal information interaction. Then, a 2×2×2 convolution is used to reduce the feature map, followed by a Leaky Rectified Linear Unit (LeakyReLU), and finally, two 1×1×1 convolutions are used to improve the non-linear representation of the model.
[0057] The output of the current convolutional layer of the first encoder is downsampled by the first max pooling layer of the fusion downsampling module, and the downsampling result is added to the output feature to obtain the first fused feature. The output of the corresponding convolutional layer of the second encoder is downsampled by the second max pooling layer of the fusion downsampling module, and the downsampling result is added to the output feature to obtain the second fused feature; as Figure 7 shown, the MaxPool layer is used to downsample the encoded features of the PET image and the CT image respectively (with non-shared parameters), and finally, the result is added to the result of convolutional downsampling to obtain the final fused feature map.
[0058] The first fused feature and the second fused feature are respectively input into the next convolutional layer of the first encoder and the corresponding convolutional layer of the second encoder.
[0059] Decode the semantic features. The fused semantic features of the CT image are upsampled by the first decoder in the multi-modal fusion segmentation network. The upsampling result of the fused semantic features of the CT image is concatenated with the semantic features of the CT image, and the concatenated features are decoded. The fused semantic features of the PET image are upsampled by the second decoder in the multi-modal fusion segmentation network. The upsampling result of the semantic features of the PET image is concatenated with the fused semantic features of the PET image, and the concatenated features are decoded;
[0060] Both the first decoder and the second decoder include an upsampling layer and several convolutional layers. The multi-modal calibration module is set between adjacent convolutional layers of the first decoder and between adjacent convolutional layers of the second decoder. The outputs of the current convolutional layer of the first decoder and the corresponding convolutional layer of the second decoder are used as the inputs of the multi-modal calibration module, and the output of the multi-modal calibration module is used as the input of the next convolutional layer of the first decoder and the corresponding convolutional layer of the second decoder.
[0061] As Figure 2 shown, the 3D dual-branch segmentation network based on the U-Net architecture contains two decoder branches. As Figure 5 shown, each decoder module contains an upsampling layer and two convolutional layers. The encoded features are upsampled using a transposed convolution with a kernel size of 2 in the upsampling layer. Then, the encoded features and the deconvolution features are connected through a skip connection module. Finally, two convolutions with a kernel size of 3*3*3 are used for decoding. The specific structure of the convolutional block is as Figure 6As shown. After each convolution operation, LeakyReLU is used as the activation function, and its slope at the negative component is set to 0.01, which helps prevent the model from overfitting and vanishing gradients during backpropagation. Since 3D convolution is used and the batch size is 1, GroupNorm is used to maintain the consistency of the input data distribution.
[0062] Given that the PET-CT images input to the network are registered, these lesions should be located at the same positions in the images. However, due to the partial volume effect of the PET images, the obtained lesion areas are often not completely accurate. Therefore, in the traditional manual segmentation process, radiologists first determine the lesion range based on the high-uptake regions in the PET images and adjust the edge details of the lesions according to the anatomy of the CT images. Based on this, a plug-and-play multi-modal calibration module is designed in this embodiment.
[0063] Calibrate the semantic features, calibrate the semantic features of the decoded CT image and the semantic features of the PET image through the multi-modal calibration block in the multi-modal fusion segmentation network, and fuse the calibrated semantic features of the CT image and the semantic features of the PET image with the original semantic features of the CT image and the semantic features of the PET image to obtain updated semantic features of the CT image and the semantic features of the PET image; specifically, it includes the following steps:
[0064] Activate the output of the current convolutional layer of the first decoder through the first activation function of the multi-modal calibration module to obtain the activated semantic features of the CT image, and activate the output of the corresponding convolutional layer of the second decoder through the second activation function of the multi-modal calibration module to obtain the activated semantic features of the PET image;
[0065] As Figure 8 shown, in the multi-modal calibration module, the Sigmoid-activated semantic features of the CT image (Feature-CT) and the semantic features from the PET image (Features-PET) are fed together into the multi-scale feature extraction block MSFB. The multi-scale feature extraction block calibrates the semantic features of the current PET image to improve the segmentation accuracy.
[0066] Similarly, Feature-CT also performs a similar updated feature process; the activated semantic features of the PET image and the semantic features of the CT image are fed to the multi-scale feature extraction block of the multi-modal calibration module for calibrating the semantic features of the current CT image.
[0067] Fuse the calibrated semantic features of the PET image with the original semantic features of the PET image to obtain updated semantic features of the PET image;
[0068] Fuse the semantic features of the calibrated CT image with the semantic features of the original CT image to obtain the updated semantic features of the CT image.
[0069] The multi-scale feature extraction block includes several parallel convolutional layers. The semantic features of the activated PET image and the semantic features of the CT image are fused after passing through the LeakyReLU, parallel convolutional layers, and group normalization of the multi-scale feature extraction block; the semantic features of the activated CT image and the semantic features of the PET image are fused after passing through the LeakyReLU, parallel convolutional layers, and group normalization of the multi-scale feature extraction block. As Figure 9 shown, the MSFB consists of a series of parallel convolutions, which are composed of convolutions with kernel sizes of 1*1*1, 3*3*3, and 5*5*5.
[0070] The size of the probability map output by the network is consistent with the width, height, number of layers, and number of channels of the input PET-CT image.
[0071] Generate a probability map by adding the lesion images reconstructed by the first decoder and the second decoder branches, and generate the output probability map of the lesion through an activation function such as the Sigmoid function;
[0072] Generate pixel classification by setting the probability map threshold to generate the classification of each pixel. For example, set the probability map threshold to 0.5 to generate the binary classification (lesion or non-lesion) of each pixel.
[0073] The present invention adopts a deep segmentation technology based on a convolutional neural network, and adopts a double-branch encoding-decoding structure, which not only retains the anatomical structure features of the CT image, but also retains the physiological information contained in the PET image, thereby improving the accuracy and robustness of the automatic segmentation of a single CNN network.
[0074] When fusing multi-modal semantic information without screening, it may bring some irrelevant information or even incorrect information, thus affecting the performance of the model. Therefore, the present invention proposes a new network structure for segmenting lesions in PET-CT images, which can realize the fusion and calibration of multi-modal information. In the feature encoding stage, multiple encoders are used to extract the semantic features of different modal images. At the same time, a fusion downsampling module is used to extract the fused semantic features to reduce the interference of incorrect information in the single-modal semantic features. In the decoding stage, the semantic features of different scales are fused through skip connections. Then, the semantic features of the CT and PET images are fused through a multi-modal calibration module to achieve the mutual calibration between different modal information.
[0075] Embodiment 2
[0076] An electronic device 200, such as Figure 10As shown, including but not limited to: a memory 201, on which program code is stored; a processor 202, which is coupled to the memory, and when the program code is executed by the processor, a lesion segmentation fusion calibration method based on PET / CT imaging is implemented. For a detailed description of the method, reference can be made to the corresponding description in the above method embodiments, which will not be elaborated here.
[0077] Embodiment 3
[0078] A computer-readable storage medium, as Figure 11 shown, on which program instructions are stored, and when the program instructions are executed, a lesion segmentation fusion calibration method based on PET / CT imaging is implemented. For a detailed description of the method, reference can be made to the corresponding description in the above method embodiments, which will not be elaborated here.
[0079] Embodiment 4
[0080] A computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, a lesion segmentation fusion calibration method based on PET / CT imaging is implemented. For a detailed description of the method, reference can be made to the corresponding description in the above method embodiments, which will not be elaborated here.
[0081] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, commodity or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the element.
[0082] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments.
[0083] The above is only for the embodiments of this specification and is not used to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and transformations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of one or more embodiments of this specification. One or more embodiments of this specification, one or more embodiments of this specification, one or more embodiments of this specification, one or more embodiments of this specification.
Claims
1. A lesion segmentation, fusion and calibration method based on PET / CT imaging, characterized in that, Including the following steps: Extracting semantic features, extracting the semantic features of the CT image through the first encoder in the multi-modal fusion segmentation network, and extracting the semantic features of the PET image through the second encoder in the multi-modal fusion segmentation network; Fusing semantic features, extracting the fused semantic features of the semantic features of the CT image and the semantic features of the PET image through the fusion downsampling module in the multi-modal fusion segmentation network, to obtain the fused semantic features of the CT image and the fused semantic features of the PET image; Decoding semantic features, upsampling the fused semantic features of the CT image through the first decoder in the multi-modal fusion segmentation network, connecting the upsampling result of the fused semantic features of the CT image with the semantic features of the CT image, and decoding the connected features, upsampling the fused semantic features of the PET image through the second decoder in the multi-modal fusion segmentation network, connecting the upsampling result of the semantic features of the PET image with the fused semantic features of the PET image, and decoding the connected features; Calibrating semantic features, calibrating the semantic features of the CT image and the semantic features of the PET image obtained by decoding through the multi-modal calibration block in the multi-modal fusion segmentation network, and fusing the calibrated semantic features of the CT image and the semantic features of the PET image with the original semantic features of the CT image and the semantic features of the PET image, to obtain the updated semantic features of the CT image and the semantic features of the PET image; Generating a probability map, adding the lesion images reconstructed by the first decoder and the second decoder branches, and generating an output probability map of the lesion through an activation function; Generating pixel classification, generating the classification of each pixel by setting a probability map threshold.
2. The method for lesion segmentation, fusion and calibration based on PET / CT imaging according to claim 1, wherein: Both the first encoder and the second encoder include a plurality of convolutional layers, the fusion downsampling module is arranged between adjacent convolutional layers of the first encoder and between adjacent convolutional layers of the second encoder, taking the outputs of the current convolutional layer of the first encoder and the corresponding convolutional layer of the second encoder as the inputs of the fusion downsampling module, and taking the output of the fusion downsampling module as the input of the next convolutional layer of the first encoder and the corresponding convolutional layer of the second encoder.
3. The method for lesion segmentation, fusion and calibration based on PET / CT imaging according to claim 2, wherein The step of fusing semantic features includes the following steps: Fusing the output of the current convolutional layer of the first encoder and the output of the corresponding convolutional layer of the second encoder through the fusion downsampling module; Reducing the number of channels of the fused features through the first convolutional layer of the fusion downsampling module for the fused features, then reducing the feature map through the second convolutional layer of the fusion downsampling module, and then improving the non-linear representation of the model through a leaky rectified linear unit and a plurality of convolutional layers to obtain output features; The output of the current convolutional layer of the first encoder is downsampled by the first max pooling layer of the fusion downsampling module, and the downsampling result is added to the output feature to obtain a first fused feature. The output of the convolutional layer corresponding to the second encoder is downsampled by the second max pooling layer of the fusion downsampling module, and the downsampling result is added to the output feature to obtain a second fused feature; The first fused feature and the second fused feature are respectively input into the next convolutional layer of the first encoder and the convolutional layer corresponding to the second encoder.
4. The lesion segmentation, fusion and calibration method based on PET / CT imaging according to claim 1, wherein: Both the first decoder and the second decoder include an upsampling layer and several convolutional layers. The multi-modal calibration module is arranged between adjacent convolutional layers of the first decoder and between adjacent convolutional layers of the second decoder. The outputs of the current convolutional layer of the first decoder and the convolutional layer corresponding to the second decoder are used as the inputs of the multi-modal calibration module, and the output of the multi-modal calibration module is used as the input of the next convolutional layer of the first decoder and the convolutional layer corresponding to the second decoder.
5. The method for lesion segmentation, fusion and calibration based on PET / CT imaging according to claim 4, wherein The step of calibrating semantic features includes the following steps: The output of the current convolutional layer of the first decoder is activated by the first activation function of the multi-modal calibration module to obtain the semantic features of the activated CT image. The output of the convolutional layer corresponding to the second decoder is activated by the second activation function of the multi-modal calibration module to obtain the semantic features of the activated PET image; The semantic features of the activated CT image and the PET image are fed to the multi-scale feature extraction block of the multi-modal calibration module for calibrating the semantic features of the current PET image; The semantic features of the activated PET image and the CT image are fed to the multi-scale feature extraction block of the multi-modal calibration module for calibrating the semantic features of the current CT image; The calibrated semantic features of the PET image are fused with the original semantic features of the PET image to obtain updated semantic features of the PET image; The calibrated semantic features of the CT image are fused with the original semantic features of the CT image to obtain updated semantic features of the CT image.
6. The method for lesion segmentation, fusion and calibration based on PET / CT imaging according to claim 5, characterized in that: The multi-scale feature extraction block includes several parallel convolutional layers. The semantic features of the activated PET image and the CT image are fused after being processed by the leaky rectified linear unit, the parallel convolutional layers and group normalization of the multi-scale feature extraction block; The semantic features of the activated CT image and the PET image are fused after being processed by the leaky rectified linear unit, the parallel convolutional layers and group normalization of the multi-scale feature extraction block.
7. The method for lesion segmentation, fusion and calibration based on PET / CT imaging according to claim 1, wherein: In the step of generating pixel classification, a probability map threshold is set to 0.5 to generate a binary classification for each pixel.
8. An electronic device, characterized in that, Comprising: A memory on which program code is stored; A processor, which is connected to the memory, and when the program code is executed by the processor, implements the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, On which program instructions are stored, and when the program instructions are executed, the method according to any one of claims 1 to 7 is implemented.
10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Nasopharyngeal-carcinoma (NPC) lesion automatic-segmentation method and nasopharyngeal-carcinoma lesion automatic-segmentation systems based on deep learning
CN108257134A
Tumor prediction method of PET-CT image based on neural network and computer readable storage medium
CN112686875A