A Laparoscopic Image Segmentation Method, System, and Computer Storage Medium

Through dual-path feature extraction and context information guidance technology, the problem of low segmentation accuracy of laparoscopic images is solved, effective segmentation of discontinuous and small targets is achieved, and segmentation accuracy is improved.

CN115984289BActive Publication Date: 2025-07-11XIAN UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211574823.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-08
Publication Date
2025-07-11
Estimated Expiration
2042-12-08

AI Technical Summary

Technical Problem

The existing laparoscopic image segmentation method has poor segmentation accuracy in multi-object segmentation, especially ignoring the ambiguity, discontinuity and imbalance of the target edge, resulting in poor segmentation accuracy.

Method used

The dual-path feature extraction technology is adopted, combining the features of semantic paths and spatial paths, and the deep fusion of hierarchical semantic features and spatial detail features is used to optimize the features using the correlation prediction matrix, and the category correlation is modeled through context information guidance technology to optimize the segmentation results.

Benefits of technology

The segmentation accuracy of laparoscopic images is improved, and is especially suitable for the segmentation of discontinuous targets and small targets, providing a better segmentation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115984289B_ABST
    Figure CN115984289B_ABST
Patent Text Reader

Abstract

The present application discloses a laparoscopic image segmentation method, system and computer storage medium. The method includes: acquiring a laparoscopic image; inputting the laparoscopic image into a segmentation model to obtain a corresponding segmentation map. The segmentation method of the segmentation model for the laparoscopic image includes: extracting hierarchical semantic features and spatial detail features of the laparoscopic image; fusing the hierarchical semantic features and the spatial detail features to obtain a correlation prediction matrix; using the correlation prediction matrix to optimize the fused features after the fusion of the hierarchical semantic features and the spatial detail features to obtain an optimized feature map; performing semantic segmentation on the optimized feature map to obtain a segmentation map. The present application uses a dual-path feature extraction technology to solve the problem of information loss in a single-path network, and uses a context information guidance technology to model category correlations, and uses this information to optimize the features extracted by the dual-path, improving the segmentation accuracy of laparoscopic images and being particularly suitable for the segmentation of laparoscopic images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to a laparoscopic image segmentation method, system and computer storage medium. Background Art

[0002] Laparoscopic surgery has been increasingly favored clinically due to its small trauma and fast recovery. The rapid increase in the number of surgical videos provides a large amount of resources for medical research and surgical evaluation. Among them, the semantic segmentation of laparoscopic images is the basis for realizing computer-aided minimally invasive surgery or automating surgical process analysis. Automatically segmenting laparoscopic images is a challenging task and has important clinical significance.

[0003] In recent years, deep learning-based models have powerful feature extraction capabilities and have achieved excellent performance in many medical image segmentation tasks. However, this method usually requires a large number of labeled samples, and it is a time-consuming and laborious task to mark different targets in laparoscopic images. In the case of multi-target segmentation with a small number of samples, the segmentation accuracy of the deep network is still not good. Through research, it is found that the mainstream laparoscopic image semantic segmentation methods mainly focus on mature algorithms in the field of medical image segmentation, such as single-path networks like UNet and SegNet. These methods are not designed for laparoscopic images, and they ignore the ambiguity, discontinuity and imbalance of the object edges in laparoscopic images. Therefore, the segmentation accuracy of laparoscopic images is relatively poor. Summary of the Invention

[0004] Embodiments of this application provide a laparoscopic image segmentation method, system and computer storage medium to solve the problem that the existing image segmentation method has relatively poor segmentation accuracy for laparoscopic images.

[0005] On the one hand, embodiments of this application provide a laparoscopic image segmentation method, including:

[0006] Obtain a laparoscopic image;

[0007] Input the laparoscopic image into a segmentation model to obtain a corresponding segmentation map;

[0008] The segmentation method of the segmentation model for the laparoscopic image includes:

[0009] Extract the hierarchical semantic features and spatial detail features of the laparoscopic image;

[0010] Fuse the hierarchical semantic features and spatial detail features to obtain a correlation prediction matrix;

[0011] Optimize the fused features after fusing the hierarchical semantic features and spatial detail features by using the correlation prediction matrix to obtain an optimized feature map;

[0012] Semantically segment the optimized feature map to obtain a segmentation map.

[0013] On the other hand, an embodiment of the present application also provides a laparoscopic image segmentation system, including:

[0014] An image acquisition module for acquiring laparoscopic images;

[0015] A segmentation model for semantically segmenting the laparoscopic images to obtain corresponding segmentation maps;

[0016] The segmentation model includes:

[0017] A feature extraction sub-module for extracting hierarchical semantic features and spatial detail features of the laparoscopic images;

[0018] A correlation prediction sub-module for fusing the hierarchical semantic features and spatial detail features to obtain a correlation prediction matrix;

[0019] A feature optimization sub-module for optimizing the fused features after fusing the hierarchical semantic features and spatial detail features by using the correlation prediction matrix to obtain an optimized feature map;

[0020] An image segmentation sub-module for semantically segmenting the optimized feature map to obtain a segmentation map.

[0021] On the other hand, an embodiment of the present application also provides a computer storage medium, in which multiple computer instructions are stored, and these multiple computer instructions are used to cause a computer to execute the above method.

[0022] A laparoscopic image segmentation method, system and computer storage medium in the present application have the following advantages:

[0023] 1. The present application combines the features of the semantic path and the spatial path, and proposes a dual-path feature extraction technology, which solves the problem of information loss in single-path networks. Different from the existing methods, these features are hierarchically and deeply fused to ensure the feature learning ability of targets at different scales and improve the segmentation accuracy of laparoscopic images.

[0024] 2. The present application proposes a context information guidance technology based on the comparative learning of similar and heterogeneous targets. It models the category correlation and uses this information to optimize the features extracted by the dual-path network, and has a good segmentation effect on discontinuous targets and small targets, and is especially suitable for the segmentation of laparoscopic images.

[0025] 3. The segmentation method proposed in the present application is a dual-path segmentation framework that first considers the heterogeneity of different tissues and organs. It assists in the segmentation of laparoscopic images by obtaining the correlation of feature categories, which provides a new way for the segmentation of laparoscopic images. Description of the Drawings

[0026] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0027] Figure 1 It is a flowchart of a laparoscopic image segmentation method provided by an embodiment of the present application;

[0028] Figure 2 It is a schematic diagram of the depth fusion process of the P1 path and the P2 path provided by an embodiment of the present application;

[0029] Figure 3 It is a flowchart of context-guided feature extraction provided by an embodiment of the present application;

[0030] Figure 4 It is a flowchart of category-related matrix calculation provided by an embodiment of the present application;

[0031] Figure 5 It is a schematic diagram of the segmentation results of three dual-path segmentation methods provided by an embodiment of the present application;

[0032] Figure 6 It is a schematic diagram of the segmentation results of four images provided by an embodiment of the present application;

[0033] Figure 7 It is a comparison schematic diagram of two dual-path image segmentation methods, BiSeNetV2 and DANet, provided by an embodiment of the present application;

[0034] Figure 8 It is a comparison schematic diagram of UNet, UNet++, VAN, OCNet and the segmentation method of the present application provided by an embodiment of the present application. Detailed implementation manners

[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0036] Figure 1 It is a schematic diagram of the process of a laparoscopic image segmentation method provided by an embodiment of the present application. The embodiment of the present application provides a laparoscopic image segmentation method, including:

[0037] Step 1, obtain laparoscopic images.

[0038] Exemplarily, laparoscopic images can be obtained by a camera during laparoscopic surgery.

[0039] Step 2, input the laparoscopic images into a segmentation model to obtain corresponding segmentation maps.

[0040] Exemplarily, the method for the segmentation model to segment laparoscopic images includes:

[0041] Step 20, extract the hierarchical semantic features and spatial detail features of the laparoscopic images.

[0042] Exemplarily, the hierarchical semantic features and spatial detail features of the laparoscopic images can be extracted through the main path P1 and the auxiliary path P2 respectively. Before inputting the laparoscopic images into the main path P1 and the auxiliary path P2, it is necessary to first perform convolution processing on the laparoscopic images using a convolution module to obtain corresponding feature maps, and then input the feature maps into the main path P1 and the auxiliary path P2 respectively.

[0043] In the embodiments of the present application, two sets of convolution operations need to be performed in the convolution module. Each set of convolution operations includes four layers: a convolution layer (convolution kernel 3*3, stride = 2), a batch normalization (BN) layer, a convolution layer (convolution kernel 3*3, stride = 1), and a BN layer. After being processed by the convolution module, the resolution of the output feature map becomes 1 / 4 of the original laparoscopic image. Because there are two convolution layers with a stride of 2 in the convolution module, the number of channels of the output feature map is 64. After the convolution module completes the processing of the laparoscopic images, the feature maps are respectively input into the main path P1 and the auxiliary path P2.

[0044] When extracting features, the hierarchical semantic features of the laparoscopic images are extracted through downsampling operations on the main path P1; while the spatial detail features of the laparoscopic images are extracted through convolution operations and batch normalization (BN) operations on the auxiliary path P2. Specifically, the main path P1 has four layers, namely L1, L2, L3, and L4 in sequence, and the auxiliary path P2 also has four layers, namely A1, A2, A3, and A4 in sequence. L1-L3 perform convolution processing with a stride of 2 on the feature maps in sequence, that is, a downsampling effect is generated. The resolution of the feature map output by each layer is 1 / 8, 1 / 16, and 1 / 32 of the original laparoscopic image in sequence. As the resolution of the feature map decreases, the number of output channels gradually increases from 64 to 128, 256, and 512. In order to make the feature maps of the two paths have the same resolution, L4 performs interpolation processing on the feature map output by L3 to obtain a feature map with the same resolution as the feature map output by the convolution module, which is 1 / 4 of the original laparoscopic image. So far, the main path P1 extracts the hierarchical semantic features of the laparoscopic images at multiple scales by gradually reducing the resolution of the feature maps layer by layer.

[0045] The purpose of the auxiliary path is to make up for the loss of spatial detail features to effectively assist the learning of hierarchical semantic features. In the auxiliary path P2, A1 - A3 sequentially perform convolutional processing on the feature map with a stride of 1 and a convolution kernel of 1*1. After passing through the A1 - A3 layers, the resolution of the feature map remains unchanged, that is, the resolution of the feature map is 1 / 4 of the original image, and the number of channels is 64. To synthesize the feature maps of the two paths, A4 performs convolutional processing on the feature map output by A3 to increase the number of output channels. The resolution and the number of output channels of the feature map output by A4 and the feature map output by L4 are the same, with the resolution being 1 / 4 of the original laparoscopic image and the number of output channels being 128.

[0046] Furthermore, when extracting the hierarchical semantic features and spatial detail features of the laparoscopic image, the feature map output by L1 is fused with the feature map output by A2, the feature map output by L2 is fused with the feature map output by A3, the feature map output by A1 is fused with the feature map output by L2, and the feature map output by A2 is fused with the feature map output by L3.

[0047] Specifically, as Figure 2 shown, since the main path P1 is the main path for feature extraction, the resolution of the feature maps output by L1 - L3 in the main path P1 is smaller than the resolution of the feature maps output by A1 - A3 in the auxiliary path P2, and the number of output channels is more than that of the auxiliary path P2. Therefore, the feature map output by the main path P1 first obtains a feature map with the same resolution as the auxiliary path P2 through an upsampling operation, and then adds the upsampled output feature map in the main path P1 and the feature map of the next layer corresponding to the main path P1 in the auxiliary path P2 to obtain the corresponding feature fusion map, which is used as the actually output feature map of the corresponding layer of the auxiliary path P2. Similarly, the auxiliary path P2 obtains a feature map with the same resolution as the main path P1 through a downsampling operation, and then adds the downsampled output feature map in the auxiliary path P2 and the feature map of the next layer corresponding to the main path P1 to obtain the corresponding feature fusion map, which is used as the actually output feature map of the corresponding layer of the main path P1. During the fusion process, the upsampling operation mainly includes convolutional processing (convolution kernel 1*1, stride = 1), BN processing, and interpolation operation, while the downsampling operation includes convolutional processing (convolution kernel 1*1, stride = 2) and BN processing. Through the dual - path depth fusion strategy of the present application, each path can obtain the feature assistance of the other path. This feature fusion method not only ensures that the hierarchical semantic features of the main path P1 do not lose the detail features of the image, but also enables the spatial detail features of the auxiliary path P2 to fuse multi - scale features, effectively improving the feature learning ability of the network.

[0048] Step 21: Fuse the hierarchical semantic features and spatial detail features to obtain a correlation prediction matrix.

[0049] Exemplarily, step 21 includes: performing convolution processing on the hierarchical semantic features and the spatial detail features to respectively obtain a main feature map and an auxiliary feature map; unfolding the main feature map and the auxiliary feature map by channels respectively to obtain two feature vectors; multiplying the two feature vectors to obtain a correlation prediction matrix.

[0050] Traditional CNN (Convolutional Neural Network)-based methods are difficult to distinguish the relationships between different features, and the performance of these methods is limited because the feature learning process only depends on the final segmentation labels. To improve the feature learning ability of the network, the present application proposes a context information guidance technique. Under this technique, the features of the dual-path network are optimized by modeling the intra-class and inter-class correlations. As Figure 3 shown, the context information guidance technique consists of two parts: class correlation learning and feature optimization. In the first part, the feature maps output by the main path P1 and the auxiliary path P2 are respectively input into the convolution module. The number of output channels of the main path P1 and the auxiliary path P2 is the same, so it is represented by R, and the resolution of the feature maps is also the same, represented by M*N. The convolution module includes a convolutional layer (convolution kernel 1*1, stride = 1) and a BN layer, and then the feature map output by the convolution module is adjusted after the convolution module. Specifically, the feature map of each channel of the output feature map is flattened into a column, and the dimension of the feature map corresponding to the main path P1 is converted to L*R (L = M*N), while the dimension of the feature map corresponding to the auxiliary path P2 is converted to R*L. Since the main path P1 and the auxiliary path P2 extract different types of features, the class correlation between different positions can be obtained by multiplying the two dimension-converted feature vectors, and the resulting result is the correlation prediction matrix, whose size is L*L.

[0051] Step 22, using the correlation prediction matrix to optimize the fused feature after fusing the hierarchical semantic features and the spatial detail features to obtain an optimized feature map.

[0052] Exemplarily, step 22 includes: fusing the hierarchical semantic features and the spatial detail features to obtain a fused feature; adjusting the dimension of the fused feature; multiplying the adjusted fused feature by the correlation prediction matrix to obtain an optimized feature map.

[0053] In the second part of the context information guiding technology, the fused features of two paths are optimized through correlation prediction. Specifically, after adding and fusing the feature maps output by L4 and A4 in the two paths, the size of the fused features obtained is R*M*N. After reshaping the size of the fused features, the size of the fused features obtained is R*L. Multiply the reshaped fused features by the correlation prediction matrix to incorporate class correlation into the feature learning process. Finally, reshape the multiplied features to obtain features of size R*M*N, and then concatenate the reshaped features and the fused features to obtain features of size 2R*M*N. This feature is the optimized feature map, which includes both the fused features of the dual paths and the features optimized using class correlation.

[0054] Step 23: Perform semantic segmentation on the optimized feature map to obtain a segmentation map.

[0055] Exemplarily, after refining the optimized feature map through the context information guiding technology, semantic segmentation can be performed on the optimized feature map. Semantic segmentation mainly includes: adjusting the number of output channels using convolution operations according to the number of segmentation targets C, and the convolution kernel size is 1*1; using interpolation operations to restore the optimized feature map from 1 / 4 resolution of the original laparoscopic image to the resolution required for the segmentation map, which is called the resolution of the segmentation mask in this application.

[0056] The above segmentation model needs to be trained before actual application. The specific training method includes:

[0057] Input the training images into the segmentation network to obtain training segmentation results; determine the main loss function according to the difference between the training segmentation results and the labeled results; encode the training segmentation results to obtain a segmentation mask; perform binary conversion on the segmentation mask, and perform one-dimensional expansion on the converted segmentation mask according to the channel numbers to obtain two sub-matrices; perform exclusive OR operation on the two sub-matrices to obtain a class correlation matrix; determine the auxiliary loss function according to the difference between the class correlation matrix and the correlation prediction matrix; establish a comprehensive loss function according to the main loss function and the auxiliary loss function; adjust the parameters of the segmentation network according to the comprehensive loss function to obtain the segmentation model.

[0058] Specifically, in the training process of the segmentation model of this application, two loss functions are used, namely the main loss function and the auxiliary loss function. The main loss function F m is used to measure the difference between the training segmentation results output by the segmentation network and the segmentation mask, that is, the label corresponding to the training image. The auxiliary loss function F a is used to evaluate the difference between the correlation prediction matrix of the segmentation network and the class correlation matrix. Therefore, this application combines the two loss functions F m and F a to obtain the comprehensive loss function expressed as:

[0059] F = F m + γF a

[0060] where γ is used to control the weight of the auxiliary loss function F a The value range of γ is [0, 1]. The larger the value of γ, the greater the contribution of the auxiliary loss function. When γ equals 1, the ratio of the two types of loss functions is the same. In this application, its value is 0.05. Since the target sizes in the training segmentation results vary greatly, this application uses the Dice loss as the main loss function, which has a good effect on small targets. The Dice loss can be expressed as:

[0061]

[0062] where p is the total number of pixels in the training segmentation result, y* is the gold standard of the segmentation mask, is the training segmentation result. The auxiliary loss function F a uses the mean squared error, expressed as:

[0063]

[0064] where W * is the class correlation matrix calculated according to Figure 4 , is the correlation prediction matrix in the context information guidance, and the sizes of both W * and are L * L. The Adam optimization algorithm can be used to optimize and adjust the parameters of the segmentation network to obtain a segmentation model that meets the actual segmentation accuracy requirements, and the other parameter configurations of the network are as follows: the training cycle is 100, the initial learning rate is 0.001, the learning rate is reduced to 0.9 every 5 cycles, and the batch size is set to 2.

[0065] To solve the problem of insufficient supervision information for feature learning in traditional CNN methods, this application models the intra-class and inter-class dependencies between different pixels and uses class correlation to supervise the feature learning process. According to Figure 4The process of calculating the class correlation matrix is as follows: To control the size of the class correlation matrix, first resize the training segmentation result to M*N, and then perform one-hot encoding on the resized training segmentation result according to the number of classes C of the segmentation task to obtain the corresponding segmentation mask. Then convert the size of the segmentation mask to C*M*N, with the number of output channels being C, and each channel representing the binary segmentation mask of a target. For each channel, set the pixel value at the location of the segmented target to 1, and 0 otherwise. Then, expand the segmentation mask into a one-dimensional vector along the channel number. Therefore, two sub-matrices S and T are obtained after the adjustment operation, and the sub-matrices S and T are transposed to each other, with sizes of L*C and C*L respectively. Perform matrix multiplication on the two sub-matrices S and T according to the exclusive OR operation, that is, if two pixels have similar targets, set the value at the corresponding position of the class correlation matrix W to 1, and set the value of dissimilar targets to 0.

[0066] The class correlation matrix W contains the correlation between each pixel in the training segmentation result and all pixels. When similar targets are far apart in the image or the targets are discontinuous due to occlusion, the class correlation matrix can effectively integrate them together to assist the feature learning of similar targets in the context information guidance. On the contrary, dissimilar targets are marked as 0 in the class correlation matrix W, and the targets of different classes can be compared. This method of comparative learning can greatly promote the feature extraction of small targets that are easily confused.

[0067] The embodiment of the present application also provides a laparoscopic image segmentation system, including:

[0068] An image acquisition module, configured to acquire a laparoscopic image;

[0069] A segmentation model, configured to perform semantic segmentation on the laparoscopic image to obtain a corresponding segmentation map;

[0070] The segmentation model includes:

[0071] A feature extraction sub-module, configured to extract the hierarchical semantic features and spatial detail features of the laparoscopic image;

[0072] A correlation prediction sub-module, configured to fuse the hierarchical semantic features and spatial detail features to obtain a correlation prediction matrix;

[0073] A feature optimization sub-module, configured to optimize the fused features after fusing the hierarchical semantic features and spatial detail features by using the correlation prediction matrix to obtain an optimized feature map;

[0074] An image segmentation sub-module, configured to perform semantic segmentation on the optimized feature map to obtain a segmentation map.

[0075] The embodiments of the present application also provide a computer storage medium, in which multiple computer instructions are stored, and the multiple computer instructions are used to cause a computer to execute the above method.

[0076] Experimental comparison

[0077] Nine categories are labeled in the laparoscopic images, which are miscellaneous, tool, liver, gallbladder, top wall, intestine, fat, artery, and black, and these categories are marked with different colors. As shown in Table 1:

[0078] Table 1 Category definitions, description information, and color labels of data

[0079]

[0080]

[0081] Table 2 Performance of the dual-path network of the present application

[0082]

[0083] Table 2 quantifies the contributions of each path, named P1 and P2. These dual-path networks use different fusion strategies, and the fusion network is named the P1P2 path. BiSeNetV2 is structurally similar to the dual-path network in the present application, and both include a spatial detail branch and a hierarchical semantic branch. The spatial detail branch of BiseNetV2 is the main path, and the hierarchical semantic branch is the auxiliary path. The IoU values of the main paths of both methods are greater than 0.43, while the segmentation results of the auxiliary branch are slightly lower than those of the main branch. The proposed dual-path network adopts a deep fusion strategy, that is, multiple fusions are performed at different feature layers, while BiSeNetV2 only performs bilateral guidance fusion at the last layer. For the fusion results, the method proposed in the present application is superior to the BiSeNetV2 method in all four metrics.

[0084] Figure 5 The segmentation results of a single tool and multiple tools are visualized respectively. The results show that BiSeNetV2 has the worst effect on maintaining details, and there are serious sawtooth phenomena at the boundaries of the segmented objects. The main reason is that the P1 path and P2 path of BiSeNetV2 are downsampled 8 times and 32 times respectively, and the fusion method is relatively simple. In contrast, the P1 and P2 branches of the dual-path network of the present application are downsampled 32 times and 4 times respectively. Importantly, a deep fusion strategy is adopted. The two paths (P1 and P2) complement each other hierarchically to better retain the details of the segmented objects. The segmentation results of DANet depend on the feature extraction quality of ResNet50, and there is little difference between the two branches of DANet because they have common features. As Figure 5As shown in (a), the segmentation effect of DANet is similar to that of the dual-path network. DANet has a good segmentation effect on the target edges, but there are certain mis-segmentations in different regions. For Figure 7 In the complex scene with multiple tools shown in (b), there are significant differences in the segmentation results of the P1 path and the P2 path of the dual-path network of the present application. When the results of the two paths are fused together, the dual-path network will utilize these two paths. Compared with DANet, it has significantly fewer false positive regions and better segmentation results. Therefore, two conclusions can be drawn: 1) The dual-path method of the present application can effectively integrate the characteristics of the single-path method, and it can obtain better results than the single-path method. 2) The fusion method is important. Compared with other methods, the hierarchical depth fusion method proposed in the present application can achieve better segmentation results.

[0085] The present application adopts the context information guidance technology (CG) to model the intra-class and inter-class correlations to further optimize the image features. Ablation experiments were carried out based on the dual-path network proposed in the present application, and the segmentation of nine types of targets was performed respectively without context information guidance (W / O) and with context information guidance (W / ). The experimental results are shown in Table 3 and Figure 6 as follows.

[0086] Table 3 Comparison of the segmentation results of nine types of targets of the present application with and without CG

[0087]

[0088] As shown in Table 3, when guided by context information, all four evaluation metrics are effectively improved. In the segmentation tasks of nine types of targets, the average values of IoU, precision, recall, and F1 score are increased to 0.475, 0.577, 0.547, and 0.539 respectively. The average IoU value is increased by approximately 2%. Among them, the segmentation effect of the gallbladder is improved the most, exceeding 6%. The improvement effects for tools and fat are obvious, with the average IoU increase exceeding 4%. The average IoU values of the liver and the top wall are increased by about 2%, and the black background is increased by about 1%, reaching over 0.95. Excluding miscellaneous targets, only the segmentation effect of the intestine is negatively affected. According to the analysis of the target distribution map, the context information guidance has a good segmentation effect on targets accounting for more than 98% of the area in laparoscopic images. Due to the existence of a large number of discontinuous objects, reflections, occlusions, and blurs in laparoscopic images, the proposed context information guidance technology can effectively solve the segmentation difficulties caused by the above phenomena through category association guidance. The reason for the slightly poor segmentation effect of the intestine is as follows: 1) The intestine occupies the least area in the image (about 1.1%); 2) There is some noise in the segmentation mask of the intestine. Therefore, the annotation quality of the intestine is poor. When the target area is small and there is noise, it is difficult to learn context information from the image, and it is difficult to fully exert the comparative learning effect of the context information guidance technology. As Figure 6 shown, the segmentation results of the four images are visualized to compare the overall segmentation effect and the segmentation results of three local targets with and without context information guidance (W / O and W / ). In Figure 6Among them, A, B, C, and D represent the original images of four images and the corresponding segmentation GT. After introducing the context guidance technology, the overall average IoU value of the four images has increased. The average IoU of images B and C has increased by more than 10%. When adding context information guidance, the segmentation continuity of the gallbladder in image A has been greatly improved. In addition, the context information guidance greatly suppresses the false positive regions of the tool and fat segmentation masks. In image C, without context information guidance, the gallbladder is misclassified as a tool category. The false positive regions of the fat areas in images B and D have also been greatly reduced. Due to the small area of the intestine and the relatively isolated target, the addition of context information guidance has little impact on the IoU value. The phenomenon of scattered segmentation targets is common in laparoscopic images. For example, the fat in images A and D, the tool in image B, and the top wall in image C. The reasons for this phenomenon are: 1) there are multiple identical targets in an image, and 2) the same target is separated due to occlusion. When using context information guidance, the similarity of discrete targets effectively improves the description ability of the extracted features. The segmentation accuracy of the scattered tissues in images A and D has increased by 4%, solving the hole problem of the instrument in image B. In image C, the segmentation accuracy of the top wall has increased by 1%. It can be found that the category-related learning method adopted by the context information guidance technology proposed in this application has good performance in improving the continuity of segmentation targets, reducing errors, and the segmentation accuracy of discrete targets.

[0089] Figure 7 It shows that BiSeNetV2 has a poor segmentation effect on the target edge, while DANet and the method proposed in this application have a better segmentation effect on the edge. However, the segmentation template of DANet has more false positive regions than the method proposed in this application. For example, in Figure 7 the 1st, 2nd, and 5th rows of, DANet tends to recognize the liver as the top wall. Compared with the dual-path method, it is not difficult to find that the effective combination of semantic features and spatial features plays an important role in refining the segmentation result. The architecture of the method proposed in this application not only integrates the lightweight design idea of BiSeNetV2, but also achieves better segmentation performance than DANet by means of category relevance learning guided by context information.

[0090] Such as Figure 8As shown, UNet++ adds a series of nested and dense skip paths based on UNet, and the complex network leads to over-segmentation in the segmentation map. The VAN network mainly utilizes the large kernel attention module to combine the advantages of convolution and self-attention mechanisms, and its segmentation effect on large targets is better than that on small targets. For example, for tool classes, the segmentation effect is very good, but it is prone to errors when the scales of multiple targets in the image vary greatly. Compared with other non-dual path methods, OCNet achieves the best segmentation performance, with an IoU value reaching 0.446. The framework of OCNet is similar to the method proposed in this application. However, the differences are as follows: 1) The feature extraction methods are different. In OCNet, ResNet50 is used for feature extraction, while the method proposed in this application uses a dual-path depth fusion network for feature extraction. 2) They have different loss functions and network training methods. Since OCNet designs complex global relationship modules and local relationship modules, its network parameters (about 52M) are much larger than those of the method proposed in this application. For Figure 8 For the gallbladder in the second line, the tool in the third line, and the fat in the fifth line in

[0091] Although the preferred embodiments of this application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of this application.

[0092] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application also intends to include these changes and modifications.

Claims

1. A laparoscopic image segmentation method, characterized in that, Including: Obtain laparoscopic images; Input the laparoscopic images into a segmentation model to obtain corresponding segmentation maps; The segmentation method of the laparoscopic images by the segmentation model includes: Extract the hierarchical semantic features and spatial detail features of the laparoscopic images; Fuse the hierarchical semantic features and spatial detail features to obtain a correlation prediction matrix; Use the correlation prediction matrix to optimize the fused features after fusing the hierarchical semantic features and spatial detail features to obtain an optimized feature map; Perform semantic segmentation on the optimized feature map to obtain the segmentation map; The extraction of the hierarchical semantic features and spatial detail features of the laparoscopic images includes: Extract the hierarchical semantic features of the laparoscopic images through downsampling operations on the main path P1; Extract the spatial detail features of the laparoscopic images through convolution operations and batch normalization operations on the auxiliary path P2; Before respectively inputting the laparoscopic images into the main path P1 and the auxiliary path P2, it further includes: Perform convolution processing on the laparoscopic images using a convolution module to obtain a feature map, and input the feature map into the main path P1 and the auxiliary path P2 respectively; The main path P1 has four layers, namely L1, L2, L3, and L4 in sequence, and the auxiliary path P2 also has four layers, namely A1, A2, A3, and A4 in sequence. L1-L3 perform convolution processing with a stride of 2 on the feature map in sequence, and L4 performs interpolation processing on the feature map output by L3 to obtain a feature map with the same resolution as the feature map output by the convolution module; A1-A3 perform convolution processing with a stride of 1 on the feature map in sequence, and A4 performs convolution processing on the feature map output by A3 to increase the number of output channels. The feature map output by A4 and the feature map output by L4 have the same resolution and the same number of output channels; When extracting the hierarchical semantic features and spatial detail features of the laparoscopic images, the feature map output by L1 is fused with the feature map output by A2, the feature map output by L2 is fused with the feature map output by A3, the feature map output by A1 is fused with the feature map output by L2, and the feature map output by A2 is fused with the feature map output by L3.

2. The laparoscopic image segmentation method according to claim 1, wherein, The fusion of the hierarchical semantic features and spatial detail features to obtain a correlation prediction matrix includes: Perform convolution processing on the hierarchical semantic features and spatial detail features to respectively obtain a main feature map and an auxiliary feature map; Unfold the main feature map and the auxiliary feature map by channels respectively to obtain two feature vectors; Multiply the two feature vectors to obtain the correlation prediction matrix.

3. The laparoscopic image segmentation method according to claim 1, characterized in that The use of the correlation prediction matrix to optimize the fused features after fusing the hierarchical semantic features and spatial detail features to obtain an optimized feature map includes: Fuse the hierarchical semantic features and spatial detail features to obtain fused features; Adjust the dimension of the fused features; Multiply the adjusted fused features by the correlation prediction matrix to obtain the optimized feature map.

4. A laparoscopic image segmentation method according to claim 1, characterized in that, The training method of the segmentation model includes: Input the training images into the segmentation network to obtain the training segmentation results; Determine the main loss function according to the difference between the training segmentation results and the labeled results; Encode the training segmentation results to obtain a segmentation mask; Perform binarization conversion on the segmentation mask, and perform one-dimensional expansion on the converted segmentation mask according to the channel numbers to obtain two sub-matrices; Perform exclusive OR operation on the two sub-matrices to obtain a category-related matrix; Determine the auxiliary loss function according to the difference between the category-related matrix and the correlation prediction matrix; Establish a comprehensive loss function according to the main loss function and the auxiliary loss function; Adjust the parameters of the segmentation network according to the comprehensive loss function to obtain the segmentation model.

5. A system applying the laparoscopic image segmentation method according to any one of claims 1-4, characterized in that, It includes: An image acquisition module for acquiring laparoscopic images; A segmentation model for performing semantic segmentation on the laparoscopic images to obtain corresponding segmentation maps; The segmentation model includes: A feature extraction sub-module for extracting hierarchical semantic features and spatial detail features of the laparoscopic images; A correlation prediction sub-module for fusing the hierarchical semantic features and spatial detail features to obtain a correlation prediction matrix; A feature optimization sub-module for optimizing the fused features after fusing the hierarchical semantic features and spatial detail features by using the correlation prediction matrix to obtain an optimized feature map; An image segmentation sub-module for performing semantic segmentation on the optimized feature map to obtain the segmentation map; The method for the feature extraction sub-module to extract the hierarchical semantic features and spatial detail features of the laparoscopic images includes: Extract the hierarchical semantic features of the laparoscopic images through downsampling operations on the main path P1; Extract the spatial detail features of the laparoscopic images through convolution operations and batch normalization operations on the auxiliary path P2; Before inputting the laparoscopic images into the main path P1 and the auxiliary path P2 respectively, it further includes: Perform convolution processing on the laparoscopic images by using a convolution module to obtain a feature map, and input the feature map into the main path P1 and the auxiliary path P2 respectively; The main path P1 has four layers, namely L1, L2, L3, and L4 in sequence, and the auxiliary path P2 also has four layers, namely A1, A2, A3, and A4 in sequence. L1-L3 perform convolution processing on the feature map with a stride of 2 in sequence, and L4 performs interpolation processing on the feature map output by L3 to obtain a feature map with the same resolution as the feature map output by the convolution module; A1-A3 perform convolution processing on the feature map with a stride of 1 in sequence, and A4 performs convolution processing on the feature map output by A3 to increase the number of output channels. The feature map output by A4 and the feature map output by L4 have the same resolution and the same number of output channels; When extracting the hierarchical semantic features and spatial detail features of the laparoscopic images, the feature map output by L1 is fused with the feature map output by A2, the feature map output by L2 is fused with the feature map output by A3, the feature map output by A1 is fused with the feature map output by L2, and the feature map output by A2 is fused with the feature map output by L3.

6. A computer storage medium, characterized in that, The computer storage medium stores multiple computer instructions, and the multiple computer instructions are used to cause a computer to execute the method described in any one of claims 1-4.

Citation Information

Patent Citations

  • Image translation method based on context-aware attention

    CN111222515A