Nasopharynx cancer image segmentation method and system, terminal and storage medium
By using lightweight mixed models and inverted residual blocks in the nasopharyngeal carcinoma image segmentation model for feature learning and fusion, the problems of large amount of model parameters and high computational complexity in the prior art are solved, and efficient and accurate nasopharyngeal carcinoma tumor segmentation are achieved.
Patent Information
- Application Number
- CN202411992991.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the nasopharyngeal carcinoma image segmentation model based on deep learning cannot accurately segment the nasopharyngeal carcinoma tumor area from the image due to the large number of parameters and high computational complexity.
The lightweight hybrid model and inverted residual block are used for global feature learning and upsampling. The fusion features are fused by inverted residual block, upsampling layer and inverted residual upsampling block to reduce the number of channels of multiple fusion features to generate prediction maps.
A lightweight image segmentation model is realized, which can accurately capture complex image structures and improve the performance of nasopharyngeal carcinoma tumor segmentation.
Smart Images

Figure CN119992084A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and in particular to a nasopharyngeal carcinoma image segmentation method, system, terminal and storage medium. Background Art
[0002] Nasopharyngeal carcinoma (NPC) is a common malignant tumor of the head and neck. NPC image segmentation based on deep learning can assist clinical diagnosis. However, these models are often accompanied by a large number of parameters and high computational complexity, and cannot accurately segment the NPC tumor area from the image.
[0003] Therefore, the prior art still needs to be improved and developed. Summary of the invention
[0004] The main purpose of the present invention is to provide a nasopharyngeal carcinoma image segmentation method, system, terminal and computer-readable storage medium, aiming to solve the problem in the prior art that when nasopharyngeal carcinoma image segmentation is based on deep learning, the model is often accompanied by a large number of parameters and has high computational complexity, and the nasopharyngeal carcinoma tumor area cannot be accurately segmented from the image.
[0005] To achieve the above object, the present invention provides a nasopharyngeal carcinoma image segmentation method, the nasopharyngeal carcinoma image segmentation method comprising the following steps:
[0006] Obtain a target nasopharyngeal carcinoma image and process it through convolution to obtain a target feature map;
[0007] Perform global feature learning and upsampling on the target feature map according to a preset number of lightweight hybrid models and inverted residual blocks to obtain a preset number of fusion features;
[0008] According to the inverted residual block, the upsampling layer and the inverted residual upsampling block, a preset number of the fusion features are fused to obtain multiple fusion features;
[0009] The number of channels of the multi-fusion features is reduced according to convolution to obtain a target prediction map.
[0010] Optionally, the preset number of fusion features includes a first fusion feature, a second fusion feature, a third fusion feature, a fourth fusion feature and a fifth fusion feature;
[0011] The global feature learning and upsampling of the target feature map are performed according to a preset number of lightweight hybrid models and inverted residual blocks to obtain a preset number of fusion features, specifically including:
[0012] Inputting the target feature map into a first lightweight hybrid model, outputting a first feature, and adding the first feature to the target feature map to obtain a first fusion feature;
[0013] Inputting the first fused feature into a first inverse residual block for processing to obtain a first inverse residual feature map, inputting the first inverse residual feature map into a second lightweight hybrid model, outputting a second feature, and adding the second feature to the first inverse residual feature map to obtain a second fused feature;
[0014] Inputting the second fused feature into a second inverse residual block for processing to obtain a second inverse residual feature map, inputting the second inverse residual feature map into a third lightweight hybrid model, outputting a third feature, and adding the third feature to the second inverse residual feature map to obtain a third fused feature;
[0015] Inputting the third fused feature into a third inverted residual block for processing to obtain a third inverted residual feature map, inputting the third inverted residual feature map into a fourth lightweight hybrid model, outputting a fourth feature, and adding the fourth feature to the third inverted residual feature map to obtain a fourth fused feature;
[0016] The fourth fused feature is input into the fourth inverse residual block for processing to obtain a fourth inverse residual feature map, the fourth inverse residual feature map is input into the fifth lightweight hybrid model, the fifth feature is output, and the fifth feature is added to the fourth inverse residual feature map to obtain a fifth fused feature.
[0017] Optionally, the lightweight hybrid model includes a local representation module, a global representation module and a convolution;
[0018] The lightweight hybrid model processes the input features using point-by-point convolution and depth-by-depth convolution in the local representation module to obtain the output of the local representation module;
[0019] The features output by the local representation module are input into the global representation module for reconstruction, and separable self-attention is used for mapping, and the mapping results are folded to obtain the output of the global representation module;
[0020] The output of the global representation module is transformed through point-by-point convolution to obtain the corresponding output;
[0021] The inverted residual block performs dilated convolution on the input feature map, then performs depth-wise separable convolution processing, and uses point convolution to reduce the number of channels to obtain the output corresponding to the point convolution, and adds the output corresponding to the point convolution to the input feature map to obtain the output of the inverted residual block.
[0022] Optionally, fusing a preset number of fusion features according to the inverted residual block, the upsampling layer and the inverted residual upsampling block to obtain multiple fusion features specifically includes:
[0023] Input the fifth fusion feature into a 3×3 convolutional layer for processing, process the fourth fusion feature through a fifth inverted residual block to obtain a fourth inverted residual fusion feature, and concatenate the fourth inverted residual fusion feature with the fourth fusion feature to obtain a first multi-fusion feature;
[0024] Processing the third fusion feature through the sixth inverse residual block to obtain a third inverse residual fusion feature, processing the first multi-fusion feature through the first upsampling layer and the first inverse residual upsampling, and splicing the first multi-fusion feature with the third inverse residual fusion feature to obtain a second multi-fusion feature;
[0025] Processing the second fusion feature through the seventh inverse residual block to obtain a second inverse residual fusion feature, processing the second multi-fusion feature through a second upsampling layer and a second inverse residual upsampling, and concatenating the second inverse residual fusion feature with the second inverse residual fusion feature to obtain a third multi-fusion feature;
[0026] Processing the first fusion feature through an eighth inverse residual block to obtain a first inverse residual fusion feature, processing the third multi-fusion feature through a third upsampling layer and a third inverse residual upsampling, and concatenating the third inverse residual fusion feature with the third inverse residual fusion feature to obtain a fourth multi-fusion feature;
[0027] The fourth multi-fusion feature is processed by a fourth upsampling layer and a fourth inverse residual upsampling to obtain a multi-fusion feature.
[0028] Optionally, during the training process of the overall network corresponding to the nasopharyngeal carcinoma image segmentation method, in each training process, the outputs of the first inverse residual upsampling, the second inverse residual upsampling, the third inverse residual upsampling and the fourth inverse residual upsampling are respectively convolved to reduce the number of channels to obtain the corresponding prediction graph, and the loss function is calculated using the true label to update the corresponding parameters of the overall network.
[0029] In addition, to achieve the above-mentioned purpose, the present invention further provides a nasopharyngeal carcinoma image segmentation system, wherein the nasopharyngeal carcinoma image segmentation system comprises:
[0030] The target feature map acquisition module is used to acquire the target image map and process it through convolution to obtain the target feature map;
[0031] A fusion feature acquisition module is used to perform global feature learning and upsampling on the target feature map according to a preset number of lightweight hybrid models and inverted residual blocks to obtain a preset number of fusion features;
[0032] A multi-fusion feature acquisition module, used for fusing a preset number of fusion features according to the inverted residual block, the upsampling layer and the inverted residual upsampling block to obtain multi-fusion features;
[0033] The result output module is used to reduce the number of channels of the multi-fusion features according to convolution to obtain a target prediction map.
[0034] Optionally, performing global feature learning and upsampling on the target feature map according to a preset number of lightweight hybrid models and inverted residual blocks to obtain a preset number of fusion features specifically includes:
[0035] Inputting the target feature map into a first lightweight hybrid model, outputting a first feature, and adding the first feature to the target feature map to obtain a first fusion feature;
[0036] Inputting the first fused feature into a first inverse residual block for processing to obtain a first inverse residual feature map, inputting the first inverse residual feature map into a second lightweight hybrid model, outputting a second feature, and adding the second feature to the first inverse residual feature map to obtain a second fused feature;
[0037] Inputting the second fused feature into a second inverse residual block for processing to obtain a second inverse residual feature map, inputting the second inverse residual feature map into a third lightweight hybrid model, outputting a third feature, and adding the third feature to the second inverse residual feature map to obtain a third fused feature;
[0038] Inputting the third fused feature into a third inverted residual block for processing to obtain a third inverted residual feature map, inputting the third inverted residual feature map into a fourth lightweight hybrid model, outputting a fourth feature, and adding the fourth feature to the third inverted residual feature map to obtain a fourth fused feature;
[0039] The fourth fused feature is input into the fourth inverse residual block for processing to obtain a fourth inverse residual feature map, the fourth inverse residual feature map is input into the fifth lightweight hybrid model, the fifth feature is output, and the fifth feature is added to the fourth inverse residual feature map to obtain a fifth fused feature.
[0040] Optionally, fusing a preset number of fusion features according to the inverted residual block, the upsampling layer and the inverted residual upsampling block to obtain multiple fusion features specifically includes:
[0041] Input the fifth fusion feature into a 3×3 convolutional layer for processing, process the fourth fusion feature through a fifth inverted residual block to obtain a fourth inverted residual fusion feature, and concatenate the fourth inverted residual fusion feature with the fourth fusion feature to obtain a first multi-fusion feature;
[0042] Processing the third fusion feature through the sixth inverse residual block to obtain a third inverse residual fusion feature, processing the first multi-fusion feature through the first upsampling layer and the first inverse residual upsampling, and splicing the first multi-fusion feature with the third inverse residual fusion feature to obtain a second multi-fusion feature;
[0043] Processing the second fusion feature through the seventh inverse residual block to obtain a second inverse residual fusion feature, processing the second multi-fusion feature through a second upsampling layer and a second inverse residual upsampling, and concatenating the second inverse residual fusion feature with the second inverse residual fusion feature to obtain a third multi-fusion feature;
[0044] Processing the first fusion feature through an eighth inverse residual block to obtain a first inverse residual fusion feature, processing the third multi-fusion feature through a third upsampling layer and a third inverse residual upsampling, and concatenating the third inverse residual fusion feature with the third inverse residual fusion feature to obtain a fourth multi-fusion feature;
[0045] The fourth multi-fusion feature is processed by a fourth upsampling layer and a fourth inverse residual upsampling to obtain a multi-fusion feature.
[0046] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a nasopharyngeal carcinoma image segmentation program stored in the memory and executable on the processor, wherein the nasopharyngeal carcinoma image segmentation program implements the steps of the nasopharyngeal carcinoma image segmentation method described above when executed by the processor.
[0047] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a nasopharyngeal carcinoma image segmentation program, and when the nasopharyngeal carcinoma image segmentation program is executed by a processor, the steps of the nasopharyngeal carcinoma image segmentation method as described above are implemented.
[0048] In the present invention, a target nasopharyngeal carcinoma image is obtained and processed by convolution to obtain a target feature map; global feature learning and upsampling are performed on the target feature map according to a preset number of lightweight hybrid models and inverted residual blocks to obtain a preset number of fused features; a preset number of fused features are fused according to an inverted residual block, an upsampling layer and an inverted residual upsampling block to obtain multiple fused features; the number of channels of the multiple fused features is reduced according to convolution to obtain a target prediction map. The present invention achieves lightweighting by replacing the encoder with an inverted residual block and a lightweight hybrid model, and the lightweight hybrid model can perform global feature modeling, capture complex image structures, and improve tumor segmentation performance; in addition, the present invention uses an inverted residual block in a jump connection to learn information from different feature maps, so that an accurate segmentation map can be generated through these settings. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a flow chart of a preferred embodiment of the nasopharyngeal carcinoma image segmentation method of the present invention;
[0050] Figure 2 It is a schematic diagram of the structure of the overall network in the nasopharyngeal carcinoma image segmentation method of the present invention;
[0051] Figure 3 It is a structural schematic diagram of a lightweight hybrid model in the nasopharyngeal carcinoma image segmentation method of the present invention;
[0052] Figure 4 It is a schematic diagram of pixel information in the nasopharyngeal carcinoma image segmentation method of the present invention;
[0053] Figure 5 It is a structural diagram of a preferred embodiment of the nasopharyngeal carcinoma image segmentation system of the present invention;
[0054] Figure 6 It is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0056] Nasopharyngeal carcinoma (NPC) is a common malignant tumor of the head and neck. In the past, the corresponding images of NPC were usually manually segmented into NPC tumor areas for subsequent treatment. This process is unstable and labor-intensive due to manual processing. With the development of deep learning, a deep learning-based image segmentation model has been developed to automatically, efficiently, and accurately complete the segmentation task of NPC images, thereby assisting clinicians in making faster diagnoses. However, deep learning models often require a large amount of computing resources and storage space. These models are often accompanied by a large number of parameters and high computational complexity, and cannot accurately segment NPC tumor areas from images.
[0057] In response to one or more of the above problems, the present invention obtains a target nasopharyngeal carcinoma image and processes it through convolution to obtain a target feature map; performs global feature learning and upsampling on the target feature map according to a preset number of lightweight hybrid models and inverse residual blocks to obtain a preset number of fused features; fuses the preset number of fused features according to the inverse residual block, upsampling layer and inverse residual upsampling block to obtain multiple fused features; reduces the number of channels of the multiple fused features according to convolution to obtain a target prediction map.
[0058] The nasopharyngeal carcinoma image segmentation method described in the preferred embodiment of the present invention is as follows: Figure 1As shown, the nasopharyngeal carcinoma image segmentation method comprises the following steps:
[0059] Step S10: Obtain a target nasopharyngeal carcinoma image, and process it through convolution to obtain a target feature map.
[0060] It should be noted that the overall network for implementing the nasopharyngeal carcinoma image segmentation method in the present invention is called a lightweight visual Transformer UNet (LWV-UNet) model, which combines the advantages of convolutional neural networks and Transformer, and achieves lightweight by replacing the encoder with an inverted residual (IR) block and a MobileViTv2 block. MobileViTv2 is a lightweight hybrid model that combines the advantages of a convolutional neural network (CNN) and a visual Transformer (ViT).
[0061] At the same time, the UNet architecture is integrated into the overall network of the nasopharyngeal carcinoma image segmentation method implemented in the present invention. The UNet model is a convolutional neural network architecture for medical image segmentation, which aims to achieve accurate image segmentation by combining high resolution and context information. The network includes an encoder and a decoder, wherein the encoder part is composed of a convolution layer and a pooling layer, which is responsible for extracting the features of the input image, while gradually reducing the spatial dimension of the feature map. The number of channels of the feature map will gradually increase with the depth of the network to enhance the ability of feature extraction. The decoder part is responsible for gradually restoring the downsampled feature map to the size of the original image through an upsampling operation, that is, using an upsampling layer, to generate a segmentation result with the same size as the input image. UNet also uses a jump connection to splice the features of the corresponding layer of the encoder with the corresponding upsampled features in the decoder, thereby retaining high-resolution features and helping the decoder to restore boundary information. Finally, a 1×1 convolution is used at the top of the decoder to reduce the number of channels to the number of categories for generating a segmentation map. In the present invention, the IR block and the MobileViTv2 block are respectively used to form a decoder, and the corresponding encoder is composed of an IRUS block, IRUS, namely an inverted residual upsampling block (Inverted ResidualUpsampling).
[0062] Specifically, in the present invention, Figure 2 As shown, the target nasopharyngeal carcinoma image is obtained, that is, the input, and the original image is converted into a target feature map through a 3×3 convolution.
[0063] Step S20: Perform global feature learning and upsampling on the target feature map according to a preset number of lightweight hybrid models and inverted residual blocks to obtain a preset number of fusion features.
[0064] Specifically, after obtaining the target feature map, global feature learning and upsampling are performed on the target feature map to obtain 5 fusion features.
[0065] Furthermore, the preset number of fusion features includes a first fusion feature, a second fusion feature, a third fusion feature, a fourth fusion feature and a fifth fusion feature;
[0066] The global feature learning and upsampling of the target feature map are performed according to a preset number of lightweight hybrid models and inverted residual blocks to obtain a preset number of fusion features, specifically including:
[0067] Inputting the target feature map into a first lightweight hybrid model, outputting a first feature, and adding the first feature to the target feature map to obtain a first fusion feature;
[0068] Inputting the first fused feature into a first inverse residual block for processing to obtain a first inverse residual feature map, inputting the first inverse residual feature map into a second lightweight hybrid model, outputting a second feature, and adding the second feature to the first inverse residual feature map to obtain a second fused feature;
[0069] Inputting the second fused feature into a second inverse residual block for processing to obtain a second inverse residual feature map, inputting the second inverse residual feature map into a third lightweight hybrid model, outputting a third feature, and adding the third feature to the second inverse residual feature map to obtain a third fused feature;
[0070] Inputting the third fused feature into a third inverted residual block for processing to obtain a third inverted residual feature map, inputting the third inverted residual feature map into a fourth lightweight hybrid model, outputting a fourth feature, and adding the fourth feature to the third inverted residual feature map to obtain a fourth fused feature;
[0071] The fourth fused feature is input into the fourth inverse residual block for processing to obtain a fourth inverse residual feature map, the fourth inverse residual feature map is input into the fifth lightweight hybrid model, the fifth feature is output, and the fifth feature is added to the fourth inverse residual feature map to obtain a fifth fused feature.
[0072] Specifically, Figure 2As shown in the figure, the encoder of the network consists of an IR block and a MobileViTv2 block, and the skip connection uses the IR block to learn deep features. The decoder consists of an IRUS block. After obtaining the target feature map, the feature map is globally learned through the MobileViTv2 block. The result obtained by the MobileViTv2 block is passed through the encoder layer of the IR block and the MobileViTv2 block with a step size of 2 for 4 consecutive times to obtain the encoder result, that is, the corresponding first fusion feature, second fusion feature, third fusion feature, fourth fusion feature and fifth fusion feature.
[0073] It should be noted that the different lightweight hybrid models, inverted residual blocks and inverted residual upsampling blocks in the present invention have the same internal structure, and the step sizes of different inverted residual blocks are different. Among them, the step size of the first inverted residual block, the second inverted residual block, the third inverted residual block and the fourth inverted residual block is 2, and the step size of the fifth inverted residual block, the sixth inverted residual block, the seventh inverted residual block and the eighth inverted residual block is 1.
[0074] Further, the lightweight hybrid model includes a local representation module, a global representation module and a convolution;
[0075] The lightweight hybrid model processes the input features using point-by-point convolution and depth-by-depth convolution in the local representation module to obtain the output of the local representation module;
[0076] The features output by the local representation module are input into the global representation module for reconstruction, and separable self-attention is used for mapping, and the mapping results are folded to obtain the output of the global representation module;
[0077] The output of the global representation module is transformed through point-by-point convolution to obtain the corresponding output.
[0078] Specifically, the Mobilevitv2 module is a module designed in the MobileViT network, which aims to remodel the input features, learn global features using a separable self-attention mechanism, and finally restore the feature map to its original size. This module integrates the advantages of CNN and ViT models in representation learning.
[0079] Specific as Figure 3 As shown, assuming that the input feature map size is Where c represents the number of channels, h and w represent the height and width respectively. First, we obtain Where d, H, and W represent depth, height, and width respectively. After remodeling X1, we get Where P represents the number of pixels in each block, Represents the number of blocks, that is, the input feature map is divided into blocks, each image is divided into multiple small blocks of fixed size, and each small block is flattened. Using separable self-attention, X2 is mapped to X'2, maintaining the same shape as X2. Specifically, The three linear branches I, K, and V are projected onto Three branches, finally Among them, the processing process of separable self-attention is expressed as:
[0080] X′2=Linear(∑(SoftMax(X I )×X K )×ReLU(X V ));
[0081] where ∑, ×, and Linear represent element-wise summation, element-wise multiplication, and a linear layer with output channel d, respectively. I )×X K represents the context score of each token to the potential token and needs to be calculated between any two tokens in the multi-head attention.
[0082] Before the expansion operation, a convolution operation is used to ensure that each pixel in the convolution is enriched with information from neighboring pixels, i.e. Figure 3 The local representation part in Figure 4 As shown, each cell in the grid corresponds to a patch and pixel. When the red and blue pixels within the patch interact in the Transformer, the blue pixel has pre-encoded information from neighboring pixels, allowing comprehensive image encoding. In addition, the feedforward network in the self-attention mechanism includes two linear layers with hidden unit counts c, and LayerNorm is used instead of GoupNorm in the Transformer module. Finally, 'X'2 is folded back to the shape of X1. This process is the inverse operation of remodeling. Specifically, the flattened result is restored to small blocks, and the small blocks are re-stitched back to the original feature map. Finally, X'1 is obtained and converted into the output through point-by-point convolution. This is Figure 3 The global representation part in . Compared with the traditional self-attention method, the time complexity of the separable attention mechanism is reduced from O(k 2 ) is reduced to O(k).
[0083] Furthermore, the inverted residual block performs a dilated convolution on the input feature map, performs a depth-wise separable convolution process, and uses point convolution to reduce the number of channels to obtain an output corresponding to the point convolution, and adds the output corresponding to the point convolution to the input feature map to obtain the output of the inverted residual block.
[0084] Specifically, the IR block is a commonly used module in lightweight CNNs, which aims to improve the efficiency and performance of the model. The core idea of the IR block is to mutate the traditional residual structure to adapt to the characteristics of lightweight networks. In the traditional residual architecture, the input feature map is added to the original input after a series of convolution operations to obtain the output. In the IR block, the input feature map is added to the original input after a series of convolution operations to obtain the output. In the IR block, a lightweight dilated convolution is first performed on the input feature map, and then a depth-separable convolution is performed. Finally, the number of channels is reduced back through a point convolution, and finally the output is added to the original input feature map. In the present invention, the IR block in the encoder and decoder first uses a 1×1 convolution to double the number of input channels, and then uses a depth-separable convolution to perform training equivalent to a 3×3 convolution layer. Finally, a 1×1 convolution is used to restore the number of channels to the original number of channels, and the final output is obtained by adding it to the original input feature map.
[0085] Step S30: fusing a preset number of fusion features according to the inverse residual block, the upsampling layer and the inverse residual upsampling block to obtain multiple fusion features.
[0086] Specifically, in the present invention, after obtaining a plurality of fusion features, they are fused accordingly to obtain multiple fusion features.
[0087] Further, the fusing of a preset number of fusion features according to the inverted residual block, the upsampling layer and the inverted residual upsampling block to obtain multiple fusion features specifically includes:
[0088] Input the fifth fusion feature into a 3×3 convolutional layer for processing, process the fourth fusion feature through a fifth inverted residual block to obtain a fourth inverted residual fusion feature, and concatenate the fourth inverted residual fusion feature with the fourth fusion feature to obtain a first multi-fusion feature;
[0089] Processing the third fusion feature through the sixth inverse residual block to obtain a third inverse residual fusion feature, processing the first multi-fusion feature through the first upsampling layer and the first inverse residual upsampling, and splicing the first multi-fusion feature with the third inverse residual fusion feature to obtain a second multi-fusion feature;
[0090] Processing the second fusion feature through the seventh inverse residual block to obtain a second inverse residual fusion feature, processing the second multi-fusion feature through a second upsampling layer and a second inverse residual upsampling, and concatenating the second inverse residual fusion feature with the second inverse residual fusion feature to obtain a third multi-fusion feature;
[0091] Processing the first fusion feature through an eighth inverse residual block to obtain a first inverse residual fusion feature, processing the third multi-fusion feature through a third upsampling layer and a third inverse residual upsampling, and concatenating the third inverse residual fusion feature with the third inverse residual fusion feature to obtain a fourth multi-fusion feature;
[0092] The fourth multi-fusion feature is processed by a fourth upsampling layer and a fourth inverse residual upsampling to obtain a multi-fusion feature.
[0093] Specifically, the encoder result is restored to its original size through the decoder. Each decoder layer consists of an IRUS block. This module restores the feature map to twice its original size through an upsampling layer, reduces the number of channels to 1 / 4 of the original through a 1×1 convolutional layer, and finally uses the IR block to learn the feature map features. In the skip connection part, the IR block in each decoder layer and the feature map after the MobileViTv2 module are added. After learning the feature map of each layer through the IR block, the result is concatenated with the result of the IRUS block of the corresponding previous decoder layer. The number of IR blocks in each layer is 2, 4, 6, and 8 respectively.
[0094] Furthermore, the inverse residual upsampling includes an upsampling layer, a 1×1 convolution and an inverse residual block.
[0095] Specifically, Figure 2 As shown in Figure 1, the inverted residual upsampling processes the feature map through an upsampling layer, a 1×1 convolution, and an inverted residual block in sequence to restore the feature map to its original size.
[0096] Step S40, reducing the number of channels of the multi-fusion features according to convolution to obtain a target prediction map.
[0097] Specifically, in the present invention, after obtaining multiple fusion features, the number of channels is reduced to the number of categories through 1×1 convolution to obtain a target prediction map.
[0098] Furthermore, in the training process of the overall network corresponding to the nasopharyngeal carcinoma image segmentation method, in each training process, the outputs of the first inverse residual upsampling, the second inverse residual upsampling, the third inverse residual upsampling and the fourth inverse residual upsampling are respectively convolved to reduce the number of channels to obtain the corresponding prediction graph, and the loss function is calculated through the true label to update the corresponding parameters of the overall network.
[0099] like Figure 2 As shown in Figure 2, during training, in each training process, the outputs of the first inverse residual upsampling, the second inverse residual upsampling, the third inverse residual upsampling, and the fourth inverse residual upsampling are respectively convolved to reduce the number of channels to obtain the corresponding prediction graph, i.e., the prediction Figure 1 ,predict Figure 2 ,predict Figure 3 and prediction Figure 4 , according to these prediction graphs, the corresponding parameters of the overall network are updated.
[0100] Corresponding to the training process, the present invention uses deep supervision to calculate the loss function at different stages, thereby enhancing the network's ability to learn feature information. Specifically, multi-scale mask predictions are generated at each stage, and bilinear interpolation (BI) is used to enlarge the prediction map of each stage to restore it to the same size as the true label map, and the mask is used in the loss function.
[0101] Among them, Binary Cross-Entropy (BCE) is a commonly used loss function, mainly used for binary classification tasks. It measures the difference between the probability distribution predicted by the model and the true label, and its definition is as follows:
[0102]
[0103] Where y is the true label, the tumor area is 1, and the background area is 0. is the probability predicted by the model.
[0104] Dice loss is a loss function widely used in image segmentation tasks, especially for binary or multi-classification tasks at the pixel level, such as medical image segmentation. The Dice coefficient was originally an indicator to measure the similarity between two sets, and its value range is between 0 and 1. The closer the value is to 1, the higher the overlap, and vice versa. In deep learning image segmentation tasks, Dice loss maximizes the negative of the Dice coefficient so that it can better reflect the similarity between the model prediction results and the true label. Its definition is as follows:
[0105]
[0106] in and yi represent the model prediction result and the value of the true label at the i-th pixel, respectively. To avoid division by zero, a very small smoothing term ∈ is added. Combining the two losses, through deep supervision, the loss function of the present invention can be expressed by the following formula:
[0107]
[0108]
[0109] Among them, Bce and Dice represent binary cross entropy and Dice loss respectively. i Represents the prediction results at different stages. represents the true label, λ i are the weights of different stages. In this paper, we default to i From i=0 to i=4, they are set to 1, 0.4, 0.3, 0.2, and 0.1 respectively.
[0110] Furthermore, the present invention uses a personal data set to evaluate the performance of the method of the present invention, and the data set contains 14,863 300×300 pixel nasopharyngeal carcinoma images and the same number of mask images. The data set is divided into a training set, a validation set, and a test set in a ratio of 8:1:1, that is, 1,486 of the data sets are taken as the test set to test the generalization of the model, and then 11,891 of the data sets are taken as the training set for model training, and the remaining 1,486 data sets are used as the validation set for model training data and data for evaluating the model. The corresponding method of the present invention is implemented by PyTorch. In this experiment, a single NVIDIA GeForce RTX 3090 is used for training. The LWVit_Unet model network of this project uses the AdamW optimizer, and the initial value of the learning rate is 1e-4. The learning rate dynamically adjusts the learning rate according to the number of training steps.
[0111] At the same time, the following performance evaluation indicators are used: Precision (Precision, Pre), Sensitivity (Sen), Dice coefficient (Dice), Intersection Over Union (IoU), 95% Hausdorff distance (Hausdorff Distance 95%, HD95), and average symmetric surface distance (Average Symmetric Surface Distance, ASD). The standard definition is as follows: Assume that category 1 is a positive example, that is, the tumor area in the nasopharyngeal carcinoma image, and category 0 is a negative example, that is, the non-tumor area in the nasopharyngeal carcinoma image, which is the background area of the image. The samples whose true value is 1 and judged as 1 after deep learning are defined as true positives (True Positives, TP), the samples whose true value is 1 and judged as 0 after deep learning are defined as false negatives (False Negatives, FN), the samples whose true value is 0 and judged as 0 after deep learning are defined as true negatives (True Negatives, TN), and the samples whose true value is 0 and judged as 1 after deep learning are positioned as false positives (False Positives, FP). Therefore, the segmentation performance indicators are as follows:
[0112] Pre refers to the proportion of true positive samples among the samples predicted to be positive. Its definition is shown in the following formula. The closer its value is to 1, the better the segmentation performance is. Its representation is as follows:
[0113]
[0114] Sen refers to the probability of the positive sample in the original sample being accurately predicted as a positive sample after passing through the model. Its definition is shown in the following formula. The closer its value is to 1, the better the segmentation performance is. It is expressed as follows:
[0115]
[0116] Dice is a set similarity metric, usually used to calculate the similarity between two samples. Its definition is shown in the following formula. The closer its value is to 1, the better the segmentation performance is. It is expressed as follows:
[0117]
[0118] IoU is a set overlap metric, usually used to calculate the degree of overlap between two sets. It is defined as shown in the following formula. The closer its value is to 1, the higher the overlap and the better the segmentation performance. It is expressed as follows:
[0119]
[0120] Among them, A and B represent two point sets, namely the segmentation result and the true annotation, respectively.
[0121] HD95 is a statistical indicator used to measure the maximum distance between two point sets. It is usually used to evaluate the boundary difference between the segmentation result and the true annotation. Unlike the traditional Hausdorff distance, HD95 calculates the maximum distance of 95% in the distribution, which can effectively reduce the impact of outliers. The smaller the value, the smaller the difference between the predicted result and the true boundary, and the better the segmentation performance. Its definition is as follows:
[0122] HD95(A,B)=Percentile 95 ({d(a,B)|a∈A}∪{d(b,A)|b∈B});
[0123] Among them, A and B represent two point sets, namely the segmentation result and the true annotation, d(a,B) represents the distance from point a to the nearest point in set B, and d(b,A) represents the distance from point b to the nearest point in set A.
[0124] ASD is a symmetric metric that evaluates the average distance between two point sets. It is usually used to measure the average deviation between the segmentation boundary and the true annotation boundary. The smaller the value, the closer the segmentation result is to the true boundary and the better the segmentation performance. Its definition is as follows:
[0125]
[0126] Among them, A and B represent two point sets, namely the segmentation result and the true annotation, respectively, |A| and |B| represent the number of points in the two point sets, d(a,B) represents the distance from point a to the nearest point in set B, and d(b,A) represents the distance from point b to the nearest point in set A.
[0127] Experimental results show that the model proposed in this paper achieves excellent segmentation performance while reducing resource consumption.
[0128] The present invention obtains a target nasopharyngeal carcinoma image and processes it through convolution to obtain a target feature map; performs global feature learning and upsampling on the target feature map according to a preset number of lightweight hybrid models and inverted residual blocks to obtain a preset number of fused features; fuses the preset number of fused features according to the inverted residual block, upsampling layer and inverted residual upsampling block to obtain multiple fused features; reduces the number of channels of the multiple fused features according to convolution to obtain a target prediction map. The present invention achieves lightweighting by replacing the encoder with an inverted residual block and a lightweight hybrid model, and the lightweight hybrid model can perform global feature modeling, capture complex image structures, and improve tumor segmentation performance; in addition, the present invention uses an inverted residual block in a jump connection to learn information from different feature maps, so that an accurate segmentation map can be generated through these settings.
[0129] Furthermore, if Figure 5 As shown, based on the above nasopharyngeal carcinoma image segmentation method, the present invention also provides a nasopharyngeal carcinoma image segmentation system, wherein the nasopharyngeal carcinoma image segmentation system comprises:
[0130] The target feature map acquisition module 51 is used to acquire the target image map and process it through convolution to obtain the target feature map;
[0131] A fusion feature acquisition module 52 is used to perform global feature learning and upsampling on the target feature map according to a preset number of lightweight hybrid models and inverted residual blocks to obtain a preset number of fusion features;
[0132] A multi-fusion feature acquisition module 53 is used to fuse a preset number of fusion features according to the inverse residual block, the upsampling layer and the inverse residual upsampling block to obtain a multi-fusion feature;
[0133] The result output module 54 is used to reduce the number of channels of the multi-fusion feature according to convolution to obtain a target prediction map.
[0134] Furthermore, the global feature learning and upsampling of the target feature map are performed according to a preset number of lightweight hybrid models and inverted residual blocks to obtain a preset number of fusion features, specifically including:
[0135] Inputting the target feature map into a first lightweight hybrid model, outputting a first feature, and adding the first feature to the target feature map to obtain a first fusion feature;
[0136] Inputting the first fused feature into a first inverse residual block for processing to obtain a first inverse residual feature map, inputting the first inverse residual feature map into a second lightweight hybrid model, outputting a second feature, and adding the second feature to the first inverse residual feature map to obtain a second fused feature;
[0137] Inputting the second fused feature into a second inverse residual block for processing to obtain a second inverse residual feature map, inputting the second inverse residual feature map into a third lightweight hybrid model, outputting a third feature, and adding the third feature to the second inverse residual feature map to obtain a third fused feature;
[0138] Inputting the third fused feature into a third inverted residual block for processing to obtain a third inverted residual feature map, inputting the third inverted residual feature map into a fourth lightweight hybrid model, outputting a fourth feature, and adding the fourth feature to the third inverted residual feature map to obtain a fourth fused feature;
[0139] The fourth fused feature is input into the fourth inverse residual block for processing to obtain a fourth inverse residual feature map, the fourth inverse residual feature map is input into the fifth lightweight hybrid model, the fifth feature is output, and the fifth feature is added to the fourth inverse residual feature map to obtain a fifth fused feature.
[0140] The step of fusing a preset number of fusion features according to the inverted residual block, the upsampling layer and the inverted residual upsampling block to obtain multiple fusion features specifically includes:
[0141] Input the fifth fusion feature into a 3×3 convolutional layer for processing, process the fourth fusion feature through a fifth inverted residual block to obtain a fourth inverted residual fusion feature, and concatenate the fourth inverted residual fusion feature with the fourth fusion feature to obtain a first multi-fusion feature;
[0142] Processing the third fusion feature through the sixth inverse residual block to obtain a third inverse residual fusion feature, processing the first multi-fusion feature through the first upsampling layer and the first inverse residual upsampling, and splicing the first multi-fusion feature with the third inverse residual fusion feature to obtain a second multi-fusion feature;
[0143] Processing the second fusion feature through the seventh inverse residual block to obtain a second inverse residual fusion feature, processing the second multi-fusion feature through a second upsampling layer and a second inverse residual upsampling, and concatenating the second inverse residual fusion feature with the second inverse residual fusion feature to obtain a third multi-fusion feature;
[0144] Processing the first fusion feature through an eighth inverse residual block to obtain a first inverse residual fusion feature, processing the third multi-fusion feature through a third upsampling layer and a third inverse residual upsampling, and concatenating the third inverse residual fusion feature with the third inverse residual fusion feature to obtain a fourth multi-fusion feature;
[0145] The fourth multi-fusion feature is processed by a fourth upsampling layer and a fourth inverse residual upsampling to obtain a multi-fusion feature.
[0146] Furthermore, if Figure 6 As shown, based on the above-mentioned nasopharyngeal carcinoma image segmentation method and system, the present invention also provides a terminal accordingly, and the terminal includes a processor 10, a memory 20 and a display 30. Figure 6 Only some components of the terminal are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0147] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Further, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed in the terminal, such as the program code of the installation terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a nasopharyngeal carcinoma image segmentation program 40 is stored on the memory 20, and the nasopharyngeal carcinoma image segmentation program 40 can be executed by the processor 10, thereby realizing the nasopharyngeal carcinoma image segmentation method of the present invention.
[0148] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chip, and is used to run the program code or process data stored in the memory 20, such as executing the nasopharyngeal carcinoma image segmentation method.
[0149] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 30 is used to display information on the terminal and to display a visual user interface. The components 10-30 of the terminal communicate with each other via a system bus.
[0150] In one embodiment, when the processor 10 executes the nasopharyngeal carcinoma image segmentation program 40 in the memory 20, the following steps are implemented:
[0151] Obtain a target nasopharyngeal carcinoma image and process it through convolution to obtain a target feature map;
[0152] Perform global feature learning and upsampling on the target feature map according to a preset number of lightweight hybrid models and inverted residual blocks to obtain a preset number of fusion features;
[0153] According to the inverted residual block, the upsampling layer and the inverted residual upsampling block, a preset number of the fusion features are fused to obtain multiple fusion features;
[0154] The number of channels of the multi-fusion features is reduced according to convolution to obtain a target prediction map.
[0155] Wherein, the preset number of fusion features includes a first fusion feature, a second fusion feature, a third fusion feature, a fourth fusion feature and a fifth fusion feature;
[0156] The global feature learning and upsampling of the target feature map are performed according to a preset number of lightweight hybrid models and inverted residual blocks to obtain a preset number of fusion features, specifically including:
[0157] Inputting the target feature map into a first lightweight hybrid model, outputting a first feature, and adding the first feature to the target feature map to obtain a first fusion feature;
[0158] Inputting the first fused feature into a first inverse residual block for processing to obtain a first inverse residual feature map, inputting the first inverse residual feature map into a second lightweight hybrid model, outputting a second feature, and adding the second feature to the first inverse residual feature map to obtain a second fused feature;
[0159] Inputting the second fused feature into a second inverse residual block for processing to obtain a second inverse residual feature map, inputting the second inverse residual feature map into a third lightweight hybrid model, outputting a third feature, and adding the third feature to the second inverse residual feature map to obtain a third fused feature;
[0160] Inputting the third fused feature into a third inverted residual block for processing to obtain a third inverted residual feature map, inputting the third inverted residual feature map into a fourth lightweight hybrid model, outputting a fourth feature, and adding the fourth feature to the third inverted residual feature map to obtain a fourth fused feature;
[0161] The fourth fused feature is input into the fourth inverse residual block for processing to obtain a fourth inverse residual feature map, the fourth inverse residual feature map is input into the fifth lightweight hybrid model, the fifth feature is output, and the fifth feature is added to the fourth inverse residual feature map to obtain a fifth fused feature.
[0162] Wherein, the lightweight hybrid model includes a local representation module, a global representation module and a convolution;
[0163] The lightweight hybrid model processes the input features using point-by-point convolution and depth-by-depth convolution in the local representation module to obtain the output of the local representation module;
[0164] The features output by the local representation module are input into the global representation module for reconstruction, and separable self-attention is used for mapping, and the mapping results are folded to obtain the output of the global representation module;
[0165] The output of the global representation module is transformed through point-by-point convolution to obtain the corresponding output;
[0166] The inverted residual block performs dilated convolution on the input feature map, then performs depth-wise separable convolution processing, and uses point convolution to reduce the number of channels to obtain the output corresponding to the point convolution, and adds the output corresponding to the point convolution to the input feature map to obtain the output of the inverted residual block.
[0167] The step of fusing a preset number of fusion features according to the inverted residual block, the upsampling layer and the inverted residual upsampling block to obtain multiple fusion features specifically includes:
[0168] Input the fifth fusion feature into a 3×3 convolutional layer for processing, process the fourth fusion feature through a fifth inverted residual block to obtain a fourth inverted residual fusion feature, and concatenate the fourth inverted residual fusion feature with the fourth fusion feature to obtain a first multi-fusion feature;
[0169] Processing the third fusion feature through the sixth inverse residual block to obtain a third inverse residual fusion feature, processing the first multi-fusion feature through the first upsampling layer and the first inverse residual upsampling, and splicing the first multi-fusion feature with the third inverse residual fusion feature to obtain a second multi-fusion feature;
[0170] Processing the second fusion feature through the seventh inverse residual block to obtain a second inverse residual fusion feature, processing the second multi-fusion feature through a second upsampling layer and a second inverse residual upsampling, and concatenating the second inverse residual fusion feature with the second inverse residual fusion feature to obtain a third multi-fusion feature;
[0171] Processing the first fusion feature through an eighth inverse residual block to obtain a first inverse residual fusion feature, processing the third multi-fusion feature through a third upsampling layer and a third inverse residual upsampling, and concatenating the third inverse residual fusion feature with the third inverse residual fusion feature to obtain a fourth multi-fusion feature;
[0172] The fourth multi-fusion feature is processed by a fourth upsampling layer and a fourth inverse residual upsampling to obtain a multi-fusion feature.
[0173] Among them, in the training process of the overall network corresponding to the nasopharyngeal carcinoma image segmentation method, in each training process, the outputs of the first inverse residual upsampling, the second inverse residual upsampling, the third inverse residual upsampling and the fourth inverse residual upsampling are respectively convolved to reduce the number of channels to obtain the corresponding prediction graph, and the loss function is calculated through the true label to update the corresponding parameters of the overall network.
[0174] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a nasopharyngeal carcinoma image segmentation program, and when the nasopharyngeal carcinoma image segmentation program is executed by a processor, the steps of the nasopharyngeal carcinoma image segmentation method described above are implemented.
[0175] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the existence of other identical elements in the process, method, article or terminal including the element.
[0176] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer-readable storage medium that can be read by a computer, and the program can include the processes of the above-mentioned method embodiments when executed. The computer-readable storage medium can be a memory, a disk, an optical disk, etc.
[0177] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A nasopharyngeal carcinoma image segmentation method, characterized in that: The nasopharyngeal carcinoma image segmentation method comprises: Obtain a target nasopharyngeal carcinoma image and process it through convolution to obtain a target feature map; Perform global feature learning and upsampling on the target feature map according to a preset number of lightweight hybrid models and inverted residual blocks to obtain a preset number of fusion features; According to the inverted residual block, the upsampling layer and the inverted residual upsampling block, a preset number of the fusion features are fused to obtain multiple fusion features; The number of channels of the multi-fusion features is reduced according to convolution to obtain a target prediction map.
2. The nasopharyngeal carcinoma image segmentation method according to claim 1, characterized in that: The preset number of fusion features includes a first fusion feature, a second fusion feature, a third fusion feature, a fourth fusion feature and a fifth fusion feature; The global feature learning and upsampling of the target feature map are performed according to a preset number of lightweight hybrid models and inverted residual blocks to obtain a preset number of fusion features, specifically including: Inputting the target feature map into a first lightweight hybrid model, outputting a first feature, and adding the first feature to the target feature map to obtain a first fusion feature; Inputting the first fused feature into a first inverse residual block for processing to obtain a first inverse residual feature map, inputting the first inverse residual feature map into a second lightweight hybrid model, outputting a second feature, and adding the second feature to the first inverse residual feature map to obtain a second fused feature; Inputting the second fused feature into a second inverse residual block for processing to obtain a second inverse residual feature map, inputting the second inverse residual feature map into a third lightweight hybrid model, outputting a third feature, and adding the third feature to the second inverse residual feature map to obtain a third fused feature; Inputting the third fused feature into a third inverted residual block for processing to obtain a third inverted residual feature map, inputting the third inverted residual feature map into a fourth lightweight hybrid model, outputting a fourth feature, and adding the fourth feature to the third inverted residual feature map to obtain a fourth fused feature; The fourth fused feature is input into the fourth inverse residual block for processing to obtain a fourth inverse residual feature map, the fourth inverse residual feature map is input into the fifth lightweight hybrid model, the fifth feature is output, and the fifth feature is added to the fourth inverse residual feature map to obtain a fifth fused feature.
3. The nasopharyngeal carcinoma image segmentation method according to claim 1, characterized in that: The lightweight hybrid model includes a local representation module, a global representation module and a convolution; The lightweight hybrid model processes the input features using point-by-point convolution and depth-by-depth convolution in the local representation module to obtain the output of the local representation module; The features output by the local representation module are input into the global representation module for reconstruction, and separable self-attention is used for mapping, and the mapping results are folded to obtain the output of the global representation module; The output of the global representation module is transformed through point-by-point convolution to obtain the corresponding output; The inverted residual block performs dilated convolution on the input feature map, then performs depth-wise separable convolution processing, and uses point convolution to reduce the number of channels to obtain the output corresponding to the point convolution, and adds the output corresponding to the point convolution to the input feature map to obtain the output of the inverted residual block.
4. The nasopharyngeal carcinoma image segmentation method according to claim 2, characterized in that: The step of fusing a preset number of fusion features according to the inverted residual block, the upsampling layer and the inverted residual upsampling block to obtain multiple fusion features specifically includes: Input the fifth fusion feature into a 3×3 convolutional layer for processing, process the fourth fusion feature through a fifth inverted residual block to obtain a fourth inverted residual fusion feature, and concatenate the fourth inverted residual fusion feature with the fourth fusion feature to obtain a first multi-fusion feature; Processing the third fusion feature through the sixth inverse residual block to obtain a third inverse residual fusion feature, processing the first multi-fusion feature through the first upsampling layer and the first inverse residual upsampling, and splicing the first multi-fusion feature with the third inverse residual fusion feature to obtain a second multi-fusion feature; Processing the second fusion feature through the seventh inverse residual block to obtain a second inverse residual fusion feature, processing the second multi-fusion feature through a second upsampling layer and a second inverse residual upsampling, and concatenating the second inverse residual fusion feature with the second inverse residual fusion feature to obtain a third multi-fusion feature; Processing the first fusion feature through an eighth inverse residual block to obtain a first inverse residual fusion feature, processing the third multi-fusion feature through a third upsampling layer and a third inverse residual upsampling, and concatenating the third inverse residual fusion feature with the third inverse residual fusion feature to obtain a fourth multi-fusion feature; The fourth multi-fusion feature is processed by a fourth upsampling layer and a fourth inverse residual upsampling to obtain a multi-fusion feature.
5. The nasopharyngeal carcinoma image segmentation method according to claim 4, characterized in that: During the training process of the overall network corresponding to the nasopharyngeal carcinoma image segmentation method, in each training process, the outputs of the first inverse residual upsampling, the second inverse residual upsampling, the third inverse residual upsampling and the fourth inverse residual upsampling are respectively reduced by convolution to obtain the corresponding prediction graph, and the loss function is calculated through the true label to update the corresponding parameters of the overall network.
6. A nasopharyngeal carcinoma image segmentation system, characterized in that: The nasopharyngeal carcinoma image segmentation system comprises: The target feature map acquisition module is used to acquire the target image map and process it through convolution to obtain the target feature map; A fusion feature acquisition module is used to perform global feature learning and upsampling on the target feature map according to a preset number of lightweight hybrid models and inverted residual blocks to obtain a preset number of fusion features; A multi-fusion feature acquisition module, used for fusing a preset number of fusion features according to the inverted residual block, the upsampling layer and the inverted residual upsampling block to obtain multi-fusion features; The result output module is used to reduce the number of channels of the multi-fusion features according to convolution to obtain a target prediction map.
7. The nasopharyngeal carcinoma image segmentation system according to claim 6, characterized in that: The global feature learning and upsampling of the target feature map are performed according to a preset number of lightweight hybrid models and inverted residual blocks to obtain a preset number of fusion features, specifically including: Inputting the target feature map into a first lightweight hybrid model, outputting a first feature, and adding the first feature to the target feature map to obtain a first fusion feature; Inputting the first fused feature into a first inverse residual block for processing to obtain a first inverse residual feature map, inputting the first inverse residual feature map into a second lightweight hybrid model, outputting a second feature, and adding the second feature to the first inverse residual feature map to obtain a second fused feature; Inputting the second fused feature into a second inverse residual block for processing to obtain a second inverse residual feature map, inputting the second inverse residual feature map into a third lightweight hybrid model, outputting a third feature, and adding the third feature to the second inverse residual feature map to obtain a third fused feature; Inputting the third fused feature into a third inverted residual block for processing to obtain a third inverted residual feature map, inputting the third inverted residual feature map into a fourth lightweight hybrid model, outputting a fourth feature, and adding the fourth feature to the third inverted residual feature map to obtain a fourth fused feature; The fourth fused feature is input into the fourth inverse residual block for processing to obtain a fourth inverse residual feature map, the fourth inverse residual feature map is input into the fifth lightweight hybrid model, the fifth feature is output, and the fifth feature is added to the fourth inverse residual feature map to obtain a fifth fused feature.
8. The nasopharyngeal carcinoma image segmentation system according to claim 7, characterized in that: The step of fusing a preset number of fusion features according to the inverted residual block, the upsampling layer and the inverted residual upsampling block to obtain multiple fusion features specifically includes: Input the fifth fusion feature into a 3×3 convolutional layer for processing, process the fourth fusion feature through a fifth inverted residual block to obtain a fourth inverted residual fusion feature, and concatenate the fourth inverted residual fusion feature with the fourth fusion feature to obtain a first multi-fusion feature; Processing the third fusion feature through the sixth inverse residual block to obtain a third inverse residual fusion feature, processing the first multi-fusion feature through the first upsampling layer and the first inverse residual upsampling, and splicing the first multi-fusion feature with the third inverse residual fusion feature to obtain a second multi-fusion feature; Processing the second fusion feature through the seventh inverse residual block to obtain a second inverse residual fusion feature, processing the second multi-fusion feature through a second upsampling layer and a second inverse residual upsampling, and concatenating the second inverse residual fusion feature with the second inverse residual fusion feature to obtain a third multi-fusion feature; Processing the first fusion feature through an eighth inverse residual block to obtain a first inverse residual fusion feature, processing the third multi-fusion feature through a third upsampling layer and a third inverse residual upsampling, and concatenating the third inverse residual fusion feature with the third inverse residual fusion feature to obtain a fourth multi-fusion feature; The fourth multi-fusion feature is processed by a fourth upsampling layer and a fourth inverse residual upsampling to obtain a multi-fusion feature.
9. A terminal, characterized in that: The terminal comprises: a memory, a processor, and a nasopharyngeal carcinoma image segmentation program stored in the memory and executable on the processor, wherein the nasopharyngeal carcinoma image segmentation program implements the steps of the nasopharyngeal carcinoma image segmentation method according to any one of claims 1 to 5 when executed by the processor.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a nasopharyngeal carcinoma image segmentation program, and when the nasopharyngeal carcinoma image segmentation program is executed by a processor, the steps of the nasopharyngeal carcinoma image segmentation method according to any one of claims 1 to 5 are implemented.