An Image Semantic Segmentation Method and System Based on Differentiated Context
Through the fusion feature of MSNS and MCCM modules, the problem of weakening of redundant information and layer relationships in semantic segmentation is solved, and more accurate object positioning and edge segmentation effects are achieved.
Patent Information
- Application Number
- CN202311037168.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-17
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-08-17
AI Technical Summary
The existing semantic segmentation methods have problems such as a lot of redundant information when fusion features and weakening the relationship between layers, resulting in blurred segmentation edges and inaccurate positioning in complex situations.
Multi-scale interleaved subtraction module (MSNS) and multi-key context conversion module (MCCM) are used to fuse features, pay attention to feature maps through different convolution methods, generate differentiated information, make up for the insufficient long-distance dependence of the self-attention mechanism, and use multi-scale cascading method to fuse features at different levels.
It realizes more accurate object positioning and edge segmentation, reduces redundant information, improves the accuracy of semantic segmentation and edge enhancement effect.
Smart Images

Figure CN117253034B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and machine vision in artificial intelligence, and specifically to an image semantic segmentation method and system based on differential context. Background Art
[0002] Image segmentation is an important task in image processing and an important part of computer vision. It divides an image into several visually meaningful regions and is a pixel-level parsing process. Currently, image segmentation can be divided into semantic segmentation, instance segmentation, and panoptic segmentation according to whether pixels are divided into categories, objects, or both.
[0003] Semantic segmentation is to classify pixels using semantic labels and perform pixel-level annotation on the categories of a group of objects. It has been widely used in fields such as scene understanding, medical image analysis, robot perception, and video surveillance. Most early methods were based on image features, such as threshold-based segmentation, region-based segmentation, shape-based segmentation, etc. However, these methods all have their own applicable fields and set parameters, and their generalization ability is limited on different datasets. Moreover, the segmentation performance may also be affected by uneven lighting during photographing, a large background, and a large segmentation target.
[0004] With the development of deep learning, more and more deep learning-based methods are applied to semantic segmentation. Relying on the strong robustness and generalization of deep learning models, excellent results can be achieved in different segmentation tasks. The classic semantic segmentation network, Fully Convolutional Networks (FCN), was proposed by Long. It is the first end-to-end network for pixel-level prediction, and most subsequent networks are improved based on this network. Badrinarayanan et al. proposed SegNet, where they recorded the pooling indices in the encoder to perform non-linear upsampling. Olaf Ronneberger believed that shallow features are equally important. They used skip connection operations to merge the features of the encoder and the decoder at the channel level and proposed UNet, which achieved good results in semantic segmentation. Although pooling can expand the receptive field, it also loses some useful information. Chen proposed the Deeplabv3+ network, which introduced dilated convolution to increase the receptive field without losing information. Despite the good performance of CNN, due to the characteristics of convolution itself, local extraction cannot meet the requirements of long-distance features. Transformer has excellent characteristics for capturing long dependencies and has achieved success in natural language processing (NLP). SETR divides the image into several paths and uses Transformer for long-term dependency modeling. Since vit outputs single-scale features, SegFormer designed a hierarchical Transformer encoder to output multi-scale features and designed a lightweight multi-layer perceptron (MLP) to fuse multi-scale features to generate segmentation masks.
[0005] Most semantic segmentation methods fuse features by element-wise addition or concatenation on the channel, which makes the differences between features not significant, generates a large amount of redundant features, weakens the relationship between layers, and causes problems such as blurred edges and inaccurate localization in complex segmentation situations. To solve the above problems, we attempt to increase the difference in information and highlight more effective information for segmentation. Therefore, this patent proposes a segmentation model MSMCNet based on differential context. To construct differential information, a multi-scale staggered subtraction module (MSNS) is proposed. In addition, a multi-key context conversion module (MCCM) is proposed, which provides a good premise for MSNS to generate differential information. Therefore, a new solution needs to be proposed for the above problems. Summary of the Invention
[0006] The purpose of the present invention is to provide an image semantic segmentation method and system based on differential context.
[0007] To achieve the above object, the present invention provides the following technical solutions: An image semantic segmentation system based on differential context includes an MCCM module and an MSNS module, and further includes a feature encoding structure, a differential information fusion structure, and a full-scale feature decoding structure;
[0008] The MCCM module is used to fuse the information of isolated, local, and different-scale remote feature maps, mine context information in a self-attention manner, and provide a good premise for generating differential information;
[0009] The differential information fusion structure fuses the differential information generated in cooperation with the MSNS module, and is used to avoid the small feature information gap and a lot of redundant information brought by fusing the upsampled feature maps when restoring the image resolution;
[0010] The full-scale decoding structure is used to avoid the problem of low utilization of deep features.
[0011] Preferably, the setting of the MCCM module at least includes: First, map the input feature map to a semantic channel as the input, and through the proposed multi-scale key fusion module (MSKF), fuse the context information of isolated keys, local keys, and remote keys at different levels to obtain a rich static context key K1, while capturing long-range dependencies; merge the static context key K1 and the query Q on the channel to jointly query the context; the weight matrix is obtained through two consecutive 1×1 convolutions; secondly, filter the channels of the value V through a 1×1 convolution, and participate in the calculation of the dynamic context K2 by aggregating the value V; finally, for the dynamic context, add the static context key K1 and the dynamic context K2 element by element and then output;
[0012] The MCCM module is expressed as follows:
[0013] K / Q / V = Head(χ)
[0014] K1 = MSKF(k)
[0015] Attention = TWConv 1x1 (Concat(K1,Q))
[0016]
[0017] F out = K1 + K2
[0018] Among them, Head(χ) represents extracting multiple head representations from the input feature χ; Concat(.) represents concatenating the input feature maps on the channel, and TWConv 1x1(.) represents two consecutive convolutions with a kernel size of 1, where the result of the first convolution is activated by the Relu function and the second one is not.
[0019] Preferably, the multi-scale key fusion module (MSKF) is expressed as follows:
[0020] F f = Conv 1×1 (F in )
[0021] F i = PWC(DWC k×k, r(DWC 3×3 (F f )))
[0022] F u = abs(F1 - F2 - F3)
[0023]
[0024] where F in represents the K after mapping the input two-dimensional features, DWC k×k,r represents Depth-wise Convolution with a kernel size of k and a dilation coefficient of r;
[0025] PWC(.) represents Point-wise Convolution;
[0026] represents Element-wise Mutiplication, represents Element-wise Addition;
[0027] abs represents taking the absolute value.
[0028] The settings of the multi-scale key fusion module (MSKF) at least include: First, filter out unimportant channel information through 1×1 convolution; use DWC with a kernel size of 3 to extract local context information, and use DWC with a kernel size of 5 and a dilation coefficient of 2, DWC with a kernel size of 7 and a dilation coefficient of 3, and DWC with a kernel size of 9 and a dilation coefficient of 4 to extract different degrees of long-range context information respectively, while making up for the deficiencies of traditional self-attention and the lack of capturing long dependencies; to prevent information redundancy caused by feature fusion, choose to subtract the extracted long-range context information element-wise; after selecting important information on the channel through PWC, obtain the weight matrix for fusing multi-scale information, multiply it with the feature map after channel screening, so that the network focuses on important regions of interest and suppresses unimportant information; finally, add the input element-wise in a residual manner to prevent the network from losing information due to being too deep.
[0029] Preferably, the setting of the MSNS module at least includes: according to the inconsistent spatial sizes of the feature maps, first restoring the deep feature maps to the image resolution of the shallow layer through bilinear interpolation, and the features of the high layer Figure 1 serve as the input; two feature maps extract features through DWCs with different kernel sizes, generating feature maps with different domain ranges. Using DWCs can maintain a certain independent relationship in space and channels, ensuring the independence of information construction; using GN for group normalization, after Gelu activation, the feature maps obtained by symmetric convolutions are subtracted element-wise, generating differential information different from the input feature maps; adding the three subtracted feature maps element-wise, fusing different differential information, extracting valuable information in the fused feature map through local convolution with a kernel size of 3, establishing the position and edge features of this layer module, through BN normalization, using the ReLu function to set some invalid features to zero to reduce the impact on the value of subsequent pixels, and finally outputting a matrix containing positioning information and edge information;
[0030] The differential information fusion structure cooperates with the differential information generated by the MSNS module to avoid small feature information gaps and a large amount of redundant information when fusing the upsampled feature maps during image resolution restoration;
[0031] The MSNS module is expressed as follows:
[0032] F i = abs(Gelu(GN(DWC 1×k (F l )))-Gelu(GN(DWC k×1 (F d ))), k = 1, 3, 5
[0033]
[0034] F out = RELU(BN(Conv 3×3 (F f )))
[0035] where F l , F d represent the sum of the shallow feature maps and the upsampled deep feature maps. DWC represents Depth-wise Convolution, which reduces a certain number of parameters while ensuring the independence of information architecture;
[0036] - represents element-wise subtraction, which together with convolution generates differential information and reduces redundant information;
[0037] To prevent the obtained differential information from being negative, abs is used to take the absolute value;
[0038] GN and BN represent Group Normalization and Batch Normalization respectively;
[0039] Gelu and RELU represent the Gelu activation function and the RELU activation function respectively;
[0040] The full-scale feature decoding structure fuses the differential feature maps of different depths and outputs a prediction mask.
[0041] An image semantic segmentation method based on differential context is used for the above content: at least including the following steps:
[0042] S1: Batch preprocess the collected computed tomography data;
[0043] S2: Divide the data into a non-overlapping dataset;
[0044] S3: Augment the data in the training set;
[0045] S4: Design a loss function for accurate segmentation;
[0046] S5: Design a segmentation network based on differential context;
[0047] S6: Train the segmentation model and perform segmentation verification, and the model training is completed.
[0048] Preferably, the specific method for data preprocessing in S1 is: Set the corresponding window width and window level according to the segmentation target to increase the contrast presentation of the target area; then convert the DICOM format CT data into PNG format image data, and during the conversion process, it is required that the naming of the patient slices and masks should correspond to each other.
[0049] Preferably, the specific method for dataset division in S2 is: Randomly shuffle and divide according to the patients, and divide them into training, validation, and test sets in a ratio of 8:1:1, and the data of any one patient cannot cross between these sets.
[0050] Preferably, the data augmentation means adopted in S3 specifically include: Random rotation, random horizontal and vertical flipping to increase data diversity and prevent model overfitting, and due to limited computing resources, the image resolution size needs to be adjusted to 224×224.
[0051] Preferably, the idea and specific content of designing the segmentation loss function in S4 are as follows: To cope with the characteristics that there are segmentation targets with large scale variations in different tasks and there are small targets in some tasks, and the target areas are large, variable and complex, a mixed loss of DiceLoss and BCELoss is selected, where W1 and W2 are two loss weight hyperparameters, both equal to 0.5;
[0052] The expression of DiceLoss is as follows:
[0053]
[0054] The expression of BCELoss is as follows:
[0055]
[0056] The mixed loss of DiceLoss and BCELoss is expressed as follows:
[0057] Loss = W1 × L Dice + W2 × L BCE
[0058] where n represents the number of pixels in the image, x i represents the true value of the i-th pixel, and y i is the predicted value of the i-th pixel;
[0059] To prevent the denominator from being 0, we set the constant t = 1 × 10 -5 .
[0060] Preferably, the specific content of training and validating the model in S6 is as follows: Training is carried out through the joint loss function formulated in S4, and the model weights are updated by using the gradient backpropagation during the training of the neural network. Whether the model is better is judged according to the segmentation effect of the model on the validation set after training, and the saved model weights are updated; Finally, after the model training is completed, the segmentation effect is evaluated on the test set;
[0061] Among them, the calculation principle of the segmentation evaluation index involved in S6 is as follows: Dice and IoU measure the coincidence degree of the predicted mask and the label, and the larger the value, the better the segmentation effect; Recall is the recall rate, indicating the recall ratio, and the larger the value, the higher the recall ratio; Precision is the precision rate, indicating the number of correct ones among the samples of the recalled positive examples;
[0062] The calculation formulas are as follows:
[0063]
[0064] where TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative.
[0065] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0066] 1. The present invention proposes a multi-scale staggered subtraction module (MSNS) to focus on feature maps in different convolution manners, generating perceptual features at different levels. While retaining useful differential information, redundant information is reduced. Additionally, a multi-key context conversion module (MCCM) is proposed, which effectively fuses isolated keys, local keys, and remote keys at different scales, dynamically encoding context with weights, making up for the deficiency of the long-range dependence of the self-attention mechanism, and providing a good premise for generating differential information by MSNS. By fusing differential features at different levels in a multi-scale cascading manner, since the receptive fields of different layers are different, the object sizes and edge information contained in each layer are also different. Making full use of multi-scale features can obtain a more accurate position perception and a boundary-enhanced prediction mask;
[0067] 2. Compared with traditional methods, since the deep learning model of the present invention has good robustness, better segmentation results can be achieved in different segmentation tasks. And compared with the deep learning methods based on CNN and Transformer models, due to the use of a multi-scale staggered subtraction method to fuse features, the difference degree of features is larger and the redundant information is less, the object positioning is accurate and the edge segmentation is smooth, realizing more accurate semantic segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0069] Figure 1 It is the overall image segmentation algorithm flow chart provided by the present invention.
[0070] Figure 2 It is the overall architecture diagram of the model provided by the present invention.
[0071] Figure 3 It is the schematic diagram of the MCCM module provided by the present invention.
[0072] Figure 4 It is the schematic diagram of the MSKF module provided by the present invention.
[0073] Figure 5 It is the schematic diagram of the MSNS module provided by the present invention.
[0074] Figure 6 It is the visualization of the segmentation results of different models provided by the present invention. Detailed implementation manners
[0075] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0076] Embodiment 1:
[0077] Refer to Figures 2 - 5 , an image semantic segmentation system based on differential context, including an MCCM module and an MSNS module, and also including a feature encoding structure, a differential information fusion structure, and a full-scale feature decoding structure;
[0078] The MCCM module is used to fuse the information of isolated, local, and different-scale remote feature maps, mine the context information in a self-attention manner, and provide a good premise for generating differential information;
[0079] The differential information fusion structure fuses the differential information generated in cooperation with the MSNS module, and is used to avoid the small feature information gap and a lot of redundant information brought by fusing the upsampled feature maps when restoring the image resolution;
[0080] The full-scale decoding structure is used to avoid the problem of low utilization of deep features, fuse the differential feature maps of different depths, make full use of the position semantic information of deep features and the edge, texture and other semantic information of shallow features, and output a prediction mask.
[0081] The multi-key context conversion module (MCCM), as Figure 3 shown, the setting of the MCCM module at least includes: first, map the input feature map to a semantic channel as the input, and through the proposed multi-scale key fusion module (MSKF), fuse the context information of isolated keys, local keys, and remote keys of different degrees to obtain a rich static context key K1, while capturing long-distance dependencies; merge the static context key K1 and the query Q on the channel to jointly query the context; the weight matrix is obtained through two consecutive 1×1 convolutions; secondly, filter the channel of the value V through a 1×1 convolution, and participate in the calculation of the dynamic context K2 by aggregating the value V; finally, for the dynamic context, add the static context key K1 and the dynamic context K2 element by element and then output;
[0082] The MCCM module is expressed as follows:
[0083] K / Q / V = Head(χ)
[0084] K1 = MSKF(k)
[0085] Attention = TWConv 1x1(Concat(K1,Q))
[0086]
[0087] F out = K1 + K2
[0088] Among them, Head(χ) represents the representation of extracting multiple heads from the input feature χ; Concat(.) represents the concatenation of the input feature maps on the channels, and TWConv 1x1 (.) represents two consecutive convolutions with a kernel size of 1. The result of the first convolution is activated by the Relu function, and the second one is not.
[0089] The proposed multi-scale key fusion module (MSKF) is shown as follows: Figure 4 As shown, the multi-scale key fusion module (MSKF) is represented as follows:
[0090] F f = Conv 1×1 (F in )
[0091] F i = PWC(DWC k×k,r (DWC 3×3 (F f )))
[0092] F u = abs(F1 - F2 - F3)
[0093]
[0094] Among them, F in represents the mapped K of the input two-dimensional feature. DWC k×k,r represents the Depth-wise Convolution with a kernel size of k and a dilation coefficient of r;
[0095] PWC(.) represents the Point-wise Convolution;
[0096] represents the Element-wise Mutiplication, represents the Element-wise Addition;
[0097] abs represents taking the absolute value.
[0098] The settings of the multi-scale key fusion module (MSKF) at least include: First, filter out unimportant channel information through 1×1 convolution; use DWC with a convolution kernel size of 3 to extract local context information, and use DWC with a 5×5 kernel size and a dilation coefficient of 2, DWC with a 7×7 kernel size and a dilation coefficient of 3, and DWC with a 9×9 kernel size and a dilation coefficient of 4 to extract remote context information at different levels, while making up for the deficiencies of traditional self-attention and the lack of capturing long dependencies; to prevent information redundancy caused by feature fusion, select to subtract the extracted remote context information element by element; after selecting important information on the channel through PWC, obtain the weight matrix for fusing multi-scale information, and multiply it with the feature map after channel screening, so that the network focuses on important regions of interest and suppresses unimportant information; finally, add the input element by element in a residual manner to prevent the network from losing information due to being too deep.
[0099] The proposed multi-scale staggered subtraction module (MSNS) is as Figure 5 shown. The settings of the MSNS module at least include: According to the inconsistent spatial sizes of the feature maps, first restore the deep feature map to the image resolution of the shallow layer through bilinear interpolation, and use it as the input together with the high-level feature Figure 1 map; extract features from the two feature maps through DWC with different-sized convolution kernels to generate feature maps in different domain ranges. Using DWC can maintain a certain independent relationship in space and channels, ensuring the independence of information construction; perform group normalization using GN, and after Gelu activation, subtract the feature maps obtained by symmetric convolutions element by element to generate differential information different from the input feature maps; add the three subtracted feature maps element by element to fuse different differential information, extract valuable information from the fused feature map through local convolution with a convolution kernel size of 3, establish the position and edge features of this layer module, perform normalization through BN, and use the ReLu function to set some invalid features to zero to reduce the impact on the value of subsequent pixels. Finally, output a matrix containing positioning information and edge information;
[0100] The MSNS module is expressed as follows:
[0101] F i = abs(Gelu(GN(DWC 1×k (F l )))·-Gelu(GN(DWC k×1 (F d ))), k = 1, 3, 5
[0102]
[0103] F out = RELU(BN(Conv 3×3 (F f )))
[0104] Among them, F l , F d represents the sum of the shallow feature maps and the upsampled deep feature maps. DWC stands for Depth-wise Convolution, which reduces the number of parameters to a certain extent while ensuring the independence of the information architecture;
[0105] - represents element-wise subtraction, which together with convolution generates differential information and reduces redundant information;
[0106] To prevent the obtained differential information from being negative, the absolute value is taken using abs;
[0107] GN and BN represent Group Normalization and Batch Normalization respectively;
[0108] Gelu and RELU represent the Gelu activation function and the RELU activation function respectively;
[0109] The full-scale feature decoding structure fuses the differential feature maps of different depths and outputs a prediction mask.
[0110] The model proposed in this patent is modified from the UNet++ network structure, as Figure 2 shown. It consists of a feature encoding structure, a differential information fusion structure, and a full-scale feature decoding structure. The downsampling encoding module is used to capture the features of the segmentation target. However, most encoding modules based on semantic segmentation networks are difficult to capture and represent the features of the target in the case of complex segmentation targets or with a lot of noise. Therefore, the MCCM module is proposed to fuse the information of isolated, local, and different-scale remote feature maps, and mine the context information in a self-attention manner, providing a good premise for generating differential information.
[0111] In addition, in order to reduce the number of parameters brought by the number of channels, according to the general configuration of channel conversion, the five different-scale feature maps obtained by downsampling are uniformly converted into 64 channels, reducing the memory overhead. The differential information fusion structure fuses the differential information generated by the proposed MSNS module to solve the problem of small feature information gap and a lot of redundant information when fusing the upsampled feature maps for restoring the image resolution. At the same time, aiming at the problem of low utilization of deep features, a full-scale decoding structure is used to fuse the differential feature maps of different depths, making full use of the position semantic information of deep features and the edge, texture and other semantic information of shallow features, and outputting a prediction mask.
[0112] The input image will first be encoded by five encoding modules to encode the context, and downsampled with max pooling to extract semantic features with different receptive fields.
[0113] The features generate differential information through the MSNS module, fuse features at different levels in a multi-scale cascading manner, make full use of the multi-scale feature maps, and obtain more accurate position information and edge information. Then, the features extracted by each layer of MSNS are aggregated. Since the feature information after channel conversion has little difference and contains redundant information, the present invention does not perform feature aggregation on the converted channels. The features aggregated by the four layers are added element-wise, and after convolution, batch normalization, and the Relu activation function, the predicted mask is output.
[0114] Embodiment 2:
[0115] Refer to Figure 1 , a method for image semantic segmentation based on differential context, for the content of an image semantic segmentation system based on differential context disclosed in the above embodiment: at least including the following steps:
[0116] S1: Batch preprocess the collected computed tomography data;
[0117] S2: Divide the data into a non-crossing data set;
[0118] S3: Augment the data in the training set;
[0119] S4: Design a loss function for accurate segmentation;
[0120] S5: Design a segmentation network based on differential context;
[0121] S6: Train the segmentation model and perform segmentation verification, and the model training is completed.
[0122] The specific method for data preprocessing in S1 is: Set the corresponding window width and window level according to the segmentation target to increase the contrast presentation of the target area; then convert the DICOM format CT data into PNG format picture data, and it is required that the naming of the patient slices and masks should correspond to each other during the conversion process.
[0123] The specific method for data set division in S2 is: Randomly shuffle and divide according to patients, and divide them into training, validation, and test sets in a ratio of 8:1:1, and the data of any one patient cannot cross between these sets.
[0124] The data augmentation means adopted in S3 specifically include: random rotation, random horizontal and vertical flipping to increase data diversity and prevent model overfitting, and due to limited computing resources, the image resolution size needs to be adjusted to 224×224.
[0125] The idea and specific content of designing the segmentation loss function in S4 are as follows: To deal with the characteristics of segmentation targets with large scale variations in different tasks (such as houses and pedestrians in remote sensing images), and the existence of small targets in some tasks (such as biomedical lesions), where the target areas vary greatly and are complex, a mixed loss of DiceLoss and BCELoss is selected. Among them, W1 and W2 are two loss weight hyperparameters, both equal to 0.5;
[0126] The expression of DiceLoss is as follows:
[0127]
[0128] The expression of BCELoss is as follows:
[0129]
[0130] The mixed loss of DiceLoss and BCELoss is expressed as follows:
[0131] Loss = W1 × L Dice + W2 × L BCE
[0132] where n represents the number of pixels in the image, x i represents the true value of the i-th pixel, and y i is the predicted value of the i-th pixel;
[0133] To prevent the denominator from being 0, we set the constant t = 1 × 10 -5 .
[0134] In S6, the model is trained and validated. The specific content is as follows: Training is carried out through the joint loss function formulated in S4. The model weights are updated by the gradient backpropagation during the training process of the neural network. Whether the model is better is judged according to the segmentation effect of the model on the validation set after training, and the saved model weights are updated; finally, after the model training is completed, the segmentation effect is evaluated on the test set;
[0135] Among them, the calculation principle of the segmentation evaluation indexes involved in S6 is as follows: Dice and IoU measure the coincidence degree of the predicted mask and the label. The larger the value, the better the segmentation effect; Recall is the recall rate, indicating the recall ratio. The larger the value, the higher the recall ratio; Precision is the precision rate, indicating the number of correct ones among the samples of the recalled positive examples;
[0136] The calculation formulas are as follows:
[0137]
[0138] Among them, TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative.
[0139] Example 3:
[0140] Refer to Figure 6 , this example is used to further disclose an experimental result on the premise of the above example.
[0141] Taking the CT lesion segmentation of traumatic brain injury (TBI) as an example, after data preprocessing, 1479 CT images are obtained and divided into a training set, a validation set, and a test set according to the ratio of 8:1:1. The experiments of this patent are implemented based on the environment of Pytorch 1.12.1. The proposed model and the comparative model are trained for 200 epochs using the Adam optimizer. If the validation set loss does not decrease within 20 epochs, the model will automatically stop training to prevent overfitting. The hyperparameter settings during the experiment are as follows: the learning rate is set to 0.0005, and the batch size is set to 4. The comparative model selects a model with better performance in semantic segmentation. In the evaluation, Dice, IoU, Recall, and Precision are used as evaluation metrics.
[0142] The experimental results are shown in Table 1. The experimental results show that the proposed model has the best segmentation effect in this example. The visualization of the experimental results is as Figure 6 shown.
[0143] Table 1 Segmentation Results of Traumatic Brain Injury Dataset
[0144]
[0145] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
Claims
1. An image semantic segmentation system based on differential context, characterized in that: It includes an MCCM module and an MSNS module, and also includes a feature encoding structure, a differential information fusion structure, and a full-scale feature decoding structure; The MCCM module is used to fuse the information of remote feature maps that are isolated, local, and of different scales, mine context information in a self-attention manner, and provide a good premise for generating differential information; The setting of the MCCM module at least includes: First, map the input feature map to a semantic channel as the input. Through the proposed multi-scale key fusion module (MSKF), fuse the context information of isolated keys, local keys, and remote keys at different levels to obtain a rich static context key K1, while capturing long-range dependencies; Merge the static context key K1 and the query Q on the channel to jointly query the context; The weight matrix is obtained through two consecutive 1×1 convolutions; Secondly, filter the channels of the value V through a 1×1 convolution, and participate in the calculation of the dynamic context K2 by aggregating the value V; Finally, for the dynamic context, add the static context key K1 and the dynamic context K2 element by element and then output; The MCCM module is expressed as follows: K / Q / V = Head(χ) K1 = MSKF(k) Attention=TWConv 1x1 (Concat(K1,Q)) F out = K1 + K2 Among them, Head(χ) represents the representation of extracting multiple heads from the input feature χ; Concat(.) represents the concatenation of the input feature maps on the channels, and TWConv 1x1 (.) represents passing through two consecutive convolutions with a kernel size of 1. The result of the first convolution is activated by the Relu function, and the second convolution is not; The multi-scale key fusion module (MSKF) is expressed as follows: F f = Conv 1×1 (F in ) F i = PWC(DWC k×k,r (DWC 3×3 (F f ))) F u = abs(F1 - F2 - F3) Among them, F in represents K after mapping the input two-dimensional features, and DWC k×k,r represents Depth-wise Convolution with a convolution kernel size of k and a dilation coefficient of r; PWC(.) represents point-wise convolution; Represents Element-wise Mutiplication, Represents Element-wise Addition; abs represents taking the absolute value; The differential information fusion structure cooperates with the differential information generated by the MSNS module to avoid the small feature information gap and a lot of redundant information brought by fusing the upsampled feature map when restoring the image resolution; The MSNS module is expressed as follows: F i = abs(Gelu(GN(DWC 1×k (F l )))-Gelu(GN(DWC k×1 (F d ))), k = 1, 3, 5 F out = RELU(BN(Conv 3×3 (F f ))) Among them, F l , F d represents the sum of the shallow feature map and the upsampled deep feature map. DWC stands for Depth-wise Convolution, which reduces a certain number of parameters while ensuring the independence of the information architecture; - represents element-wise subtraction, which together with convolution generates differential information and reduces redundant information; To prevent the obtained differential information from being negative, use abs to take the absolute value; GN and BN represent Group Normalization and Batch Normalization respectively; Gelu and RELU represent the Gelu activation function and the RELU activation function respectively; The full-scale feature decoding structure fuses differential feature maps of different depths and outputs a prediction mask.
2. The image semantic segmentation system based on differential context according to claim 1, wherein: The setting of the multi-scale key fusion module at least includes: First, filter out unimportant channel information through a 1×1 convolution; Use DWC with a convolution kernel size of 3 to extract local context information, and use DWC with a 5×5 and a dilation coefficient of 2, a 7×7 and a dilation coefficient of 3, and a 9×9 and a dilation coefficient of 4 to extract remote context information at different levels, while making up for the deficiencies of traditional self-attention and the lack of capturing long dependencies; To prevent information redundancy caused by feature fusion, choose to subtract the extracted remote context information element by element; After selecting important information on the channel through PWC, obtain the weight matrix for fusing multi-scale information and multiply it with the feature map after channel screening, so that the network focuses on important regions of interest and suppresses unimportant information; Finally, add the input element by element in a residual manner to prevent the network from losing information due to excessive depth.
3. The image semantic segmentation system based on differential context according to claim 1, wherein: The settings of the MSNS module at least include: according to the inconsistent spatial sizes of the feature maps, first restore the deep feature maps to the image resolution of the shallow layer through bilinear interpolation, and use them as inputs together with the high-level feature maps; extract features from the two feature maps through DWCs with different kernel sizes to generate feature maps with different domain ranges. Using DWCs can maintain a certain independent relationship in both space and channels to ensure the independence of information construction; perform group normalization using GN, and after Gelu activation, subtract the feature maps obtained by the symmetric convolutions element by element to generate differential information different from the input feature maps; add the three subtracted feature maps element by element to fuse different differential information, and extract valuable information from the fused feature maps through local convolutions with a kernel size of 3 to determine the position and edge features of this layer module. Through BN normalization, use the RELU activation function to set some invalid features to zero to reduce the impact on the value of subsequent pixels, and finally output a matrix containing positioning information and edge information.
4. A method for image semantic segmentation based on differential context, which is used for an image semantic segmentation system based on differential context according to any one of the above claims 1-3, and is characterized in that: At least include the following steps: S1: Batch preprocess the collected computed tomography data; S2: Divide the data into a non-crossing data set; S3: Augment the data in the training set; S4: Design a loss function for precise segmentation; S5: Design a system for image semantic segmentation based on differential context; S6: Train the segmentation model and conduct segmentation verification, and the model training is completed.
5. A method for image semantic segmentation based on differentiated context according to claim 4, characterized in that: The specific method for data preprocessing in S1 is: set the corresponding window width and window level according to the segmentation target to increase the contrast presentation of the target area; then convert the DICOM format CT data into PNG format picture data, and require that the naming of patient slices and masks should correspond to each other during the conversion process.
6. The method for image semantic segmentation based on differential context according to claim 5, characterized in that: The specific method for data set division in S2 is: randomly shuffle and divide according to patients, and divide them into training, validation, and test sets at a ratio of 8:1:1, and the data of any one patient cannot cross between these sets.
7. A method for image semantic segmentation based on differential context according to claim 6, characterized in that: The data augmentation means adopted in S3 specifically include: randomly rotate, randomly horizontally and vertically flip to increase data diversity and prevent model overfitting. And due to the limitation of computing resources, the image resolution size needs to be adjusted to 224×224.
8. A method for image semantic segmentation based on differential context according to claim 7, characterized in that: The idea and specific content of designing the segmentation loss function in S4 are: in order to cope with the segmentation targets with large scale changes in different tasks and the characteristics of small targets in some tasks, where the target area changes greatly and is complex, select the mixed loss of DiceLoss and BCELoss, where W1 and W2 are two loss weight hyperparameters, both equal to 0.5; The expression of DiceLoss is as follows: The expression of BCELoss is as follows: The mixed loss of DiceLoss and BCELoss is expressed as follows: Loss=W1×L Dice +W2×L BCE Among them, n represents the number of pixels of the image, and x i represents the true value of the i-th pixel, and y i is the predicted value of the i-th pixel; To prevent the denominator from being 0, we set the constant t = 1×10 -5 .
9. A method for image semantic segmentation based on differential context according to claim 8, characterized in that: In S6, the model is trained and validated. The specific content is as follows: Training is carried out through the joint loss function formulated in S4. The weights of the model are updated by the gradient backpropagation during the training of the neural network. Whether the model is better is judged according to the segmentation effect of the model on the validation set, and the saved model weights are updated. Finally, after the model training is completed, the segmentation effect is evaluated on the test set. Among them, the calculation principle of the segmentation evaluation index involved in S6 is as follows: Dice and IoU measure the coincidence degree between the predicted mask and the label. The larger the value, the better the segmentation effect. Recall is the recall rate, indicating the recall ratio. The larger the value, the higher the recall ratio. Precision is the precision rate, indicating the number of correct samples among the recalled positive examples. The calculation formulas are as follows: Among them, TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative.