A method and system for segmenting surgical instruments in spinal endoscope images
By using convolutional neural networks in spinal endoscopic images for multi-scale convolution and multi-gradient synchronous extraction, combined with recognition of attention and spot interference characteristics for convolutional fusion, the problem of insufficient accuracy of surgical instrument contour segmentation in the prior art is solved, and a more robust and accurate surgical instrument segmentation effect is achieved.
Patent Information
- Application Number
- CN202510016484.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-06
AI Technical Summary
The prior art is difficult to effectively segment the surgical instrument profile in spinal endoscopic images, especially in the face of light changes, tissue occlusion and highlight reflection, segmentation accuracy is limited.
Through multi-scale convolution operation based on convolution neural network, the spot interference characteristics of surgical instruments in contour edge detection are extracted, and multi-gradient synchronous extraction is performed based on texture information variance to obtain the contour feature maps under each gradient branch. Then, convolutional fusion is performed based on the recognition attention of the gradient branch channel and the spot interference characteristics to generate the fusion profile of the surgical instrument, and finally image segmentation is performed.
It significantly improves the robustness and accuracy of surgical instrument segmentation, can effectively process complex spinal endoscopic images, enhances the recognition ability of the edge features of surgical instruments, and reduces the contour detection error caused by a single gradient branch.
Smart Images

Figure CN119399227B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image segmentation, and more specifically, to a method and system for segmenting spinal endoscope images of surgical instruments. Background Art
[0002] Image segmentation plays a key role in spinal endoscopic images, especially in the precise segmentation of surgical instruments. Endoscopic images are usually complex and contain multiple interference factors, such as light spots, instrument occlusions and the complex structure of the spine. These factors make it difficult to identify and locate the contours of surgical instruments. Image segmentation technology can provide clear instrument contours by distinguishing surgical instruments from the background in the image, thereby helping doctors achieve more precise navigation and operation during surgery.
[0003] In the existing technology, traditional spinal endoscopic image segmentation methods (such as algorithms based on threshold, region or edge detection) have limited performance in percutaneous spinal detection scenarios, and it is difficult to effectively deal with the interference of illumination changes, tissue occlusion and high light reflection. At the same time, although algorithms based on specific tools (such as genetic algorithms and active contour models) have certain optimization capabilities in contour detection, they are still insufficient in real-time and robustness. In addition, although image segmentation methods based on deep learning have made major breakthroughs, when processing spinal endoscopic images, the retention of surgical instrument edges and gradient information is still insufficient, resulting in limited accuracy of surgical instrument contour segmentation. In contrast, by performing multi-gradient branch feature extraction on spinal endoscopic images to obtain contour features of surgical instruments, different levels of image semantic information can be effectively retained, thereby significantly improving the accuracy of image segmentation. Therefore, how to achieve multi-gradient synchronous extraction of surgical instrument contour features in spinal endoscopic images to improve the robustness of surgical instrument segmentation has become a difficult problem faced by the industry. Summary of the invention
[0004] The present application provides a method and system for segmenting surgical instruments in spinal endoscope images, which can realize multi-gradient synchronous extraction of surgical instrument contour features in spinal endoscope images, thereby improving the robustness of surgical instrument segmentation.
[0005] In a first aspect, the present application provides a method for segmenting surgical instruments in a spinal endoscope image, comprising the following steps:
[0006] Acquire spinal endoscopic images with spinal detail information obscured by surgical instruments;
[0007] Performing a multi-scale convolution operation on the spinal endoscope image based on a convolutional neural network to obtain the light spot interference characteristics of the surgical instrument in the contour edge detection;
[0008] Determine the information variance of the texture in the spinal endoscopic image, perform multi-gradient synchronous extraction on the contour information of the surgical instrument in the spinal endoscopic image based on the information variance of the texture, obtain the contour feature map of the surgical instrument under each gradient branch, and then determine the recognition attention of the channel corresponding to each gradient branch according to the texture information difference between each contour feature map;
[0009] All contour feature maps are convolutionally fused according to the recognition attention of the corresponding channel of each gradient branch and the light spot interference feature to obtain a fused contour of the surgical instrument, and then the spinal endoscope image is segmented based on the fused contour to obtain a segmentation feature map of the surgical instrument.
[0010] Preferably, performing a multi-scale convolution operation on the spinal endoscope image based on a convolutional neural network to obtain the spot interference features of the surgical instrument in the contour edge detection specifically includes:
[0011] Build a recognition model for spot recognition based on convolutional neural network;
[0012] Designing convolution kernels of multiple scales in the recognition model, and then capturing the light spot area in the spinal endoscope image through the convolution kernels of each scale, so as to obtain the light spot features at each scale;
[0013] The light spot interference characteristics of the surgical instrument in the contour edge detection are determined based on all the light spot characteristics.
[0014] Preferably, determining the information variance of the texture in the spinal endoscopic image specifically includes:
[0015] dividing the spinal endoscopy image into a plurality of local image regions;
[0016] The grayscale features of pixels in each local image area are extracted through the grayscale co-occurrence matrix;
[0017] The information variance of the texture in the spinal endoscopy image is determined based on all the grayscale features.
[0018] Preferably, determining the recognition attention of the channel corresponding to each gradient branch according to the texture information difference between each contour feature map specifically includes:
[0019] Determine the texture entropy of each contour feature map;
[0020] Performing global average pooling on each contour feature map, extracting global texture features of each contour feature map, and then determining the texture difference between each contour feature map and other contour feature maps;
[0021] The recognition attention of each gradient branch corresponding to the channel is determined by all texture differences and texture entropy of each contour feature map.
[0022] Preferably, performing convolution fusion on all contour feature maps according to the recognition attention of the channel corresponding to each gradient branch and the light spot interference feature to obtain the fused contour of the surgical instrument specifically includes:
[0023] Determine the fusion coefficient of the contour feature map under each gradient branch according to the recognition attention of the channel corresponding to each gradient branch;
[0024] All contour feature maps are weightedly fused based on each fusion coefficient to obtain a feature fusion map;
[0025] Convolution interference compensation is performed on the feature fusion image according to the light spot interference feature, and a surgical instrument contour feature extraction operation is further performed to obtain a fusion contour of the surgical instrument.
[0026] Preferably, performing image segmentation on the spinal endoscope image based on the fusion contour to obtain a segmentation feature map of the surgical instrument specifically includes:
[0027] determining a pixel segmentation region of a surgical instrument in the spinal endoscope image according to the fusion contour;
[0028] A segmentation feature map of the surgical instrument is generated through the pixel segmentation area.
[0029] Preferably, the spinal endoscope image includes surgical instruments in an in vivo environment.
[0030] In a second aspect, the present application provides a spinal endoscope image surgical instrument segmentation system, comprising:
[0031] An acquisition module, used for acquiring an endoscopic image of the spine in which detailed information of the spine is obscured by surgical instruments;
[0032] A processing module, used for performing a multi-scale convolution operation on the spinal endoscope image based on a convolutional neural network to obtain a light spot interference feature of the surgical instrument in contour edge detection;
[0033] The processing module is further used to determine the information variance of the texture in the spinal endoscope image, perform multi-gradient synchronous extraction on the contour information of the surgical instrument in the spinal endoscope image based on the texture information variance, obtain the contour feature map of the surgical instrument under each gradient branch, and then determine the recognition attention of the channel corresponding to each gradient branch according to the texture information difference between each contour feature map;
[0034] The execution module is used to perform convolution fusion on all contour feature maps according to the recognition attention of the channel corresponding to each gradient branch and the light spot interference feature to obtain the fused contour of the surgical instrument, and then perform image segmentation on the spinal endoscope image based on the fused contour to obtain the segmentation feature map of the surgical instrument.
[0035] In a third aspect, the present application provides a computer device, comprising a memory and a processor, wherein the memory stores codes, and the processor is configured to obtain the codes and execute the above-mentioned spinal endoscopic image surgical instrument segmentation method.
[0036] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, which implements the above-mentioned spinal endoscopic image surgical instrument segmentation method when executed by a processor.
[0037] The technical solution provided by the embodiments disclosed in this application has the following beneficial effects:
[0038] In an embodiment of the present application, a spinal endoscopic image in which surgical instruments obstruct detail information of the spine is obtained; a multi-scale convolution operation is performed on the spinal endoscopic image based on a convolutional neural network to obtain a spot interference feature of the surgical instrument in contour edge detection; the information variance of the texture in the spinal endoscopic image is determined, and based on the information variance of the texture, multi-gradient synchronous extraction is performed on the contour information of the surgical instrument in the spinal endoscopic image to obtain a contour feature map of the surgical instrument under each gradient branch, and then the recognition attention of the corresponding channel of each gradient branch is determined based on the texture information difference between each contour feature map; all the contour feature maps are convoluted and fused based on the recognition attention of the corresponding channel of each gradient branch and the spot interference feature to obtain a fused contour of the surgical instrument, and then the spinal endoscopic image is segmented based on the fused contour to obtain a segmentation feature map of the surgical instrument.
[0039] It can be seen that the present application performs convolution fusion on all contour feature maps through the recognition attention and spot interference features of the corresponding channel of each gradient branch to obtain the fused contour of the surgical instrument, and then performs image segmentation on the spinal endoscopic image based on the fused contour to obtain the segmentation feature map of the surgical instrument; first, a multi-scale convolution operation is performed on the spinal endoscopic image, which can effectively identify and suppress the spot interference caused by instrument occlusion, thereby ensuring that subsequent processing can focus on the extraction of spinal tissue and surgical instrument contours; then, multi-gradient synchronous extraction is performed on the contour information of the surgical instrument in the spinal endoscopic image to obtain the contour feature map of the surgical instrument under each gradient branch, and the local contour of the surgical instrument in the spinal endoscopic image is collaboratively captured through gradient branches at different levels. Details, regional continuity and global semantic information can enhance the accuracy of recognition of edge features of surgical instruments; finally, all contour feature maps are convolutionally fused according to the recognition attention and spot interference features of the corresponding channel of the gradient branch to obtain the fused contour of the surgical instrument. Through the convolution fusion of multi-gradient contour features, the semantic information between different levels can be integrated, thereby enhancing the recognition ability of surgical instruments in complex backgrounds, enabling the segmentation model to effectively distinguish between instruments and surrounding tissue structures, avoiding the contour detection error caused by a single gradient branch, thereby improving the segmentation accuracy of surgical instruments in complex scenes; in summary, the present application scheme can realize multi-gradient synchronous extraction of contour features of surgical instruments in spinal endoscopy images, thereby improving the robustness of surgical instrument segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is an exemplary flow chart of a method for segmenting surgical instruments in spinal endoscope images according to some embodiments of the present application;
[0041] Figure 2 are spinal endoscopic images of surgical instruments before segmentation according to some embodiments of the present application, wherein (a) is a spinal endoscopic image of a rongeur before segmentation, (b) is a spinal endoscopic image of a burr before segmentation, and (c) is a spinal endoscopic image of a probe before segmentation;
[0042] Figure 3 are spinal endoscopy images after surgical instrument segmentation according to some embodiments of the present application, wherein (a) is a spinal endoscopy segmentation feature image after rongeur segmentation, (b) is a spinal endoscopy segmentation feature image after bur segmentation, and (c) is a spinal endoscopy segmentation feature image after probe segmentation;
[0043] Figure 4 is a schematic diagram of a process for determining information variance according to some embodiments of the present application;
[0044] Figure 5is a schematic diagram of the structure of a spinal endoscope image surgical instrument segmentation system according to some embodiments of the present application;
[0045] Figure 6 It is a structural schematic diagram of a computer device for implementing a spinal endoscope image surgical instrument segmentation method according to some embodiments of the present application. DETAILED DESCRIPTION
[0046] In order to better understand the technical solution of the present application, the technical solution of the present application will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0047] refer to Figure 1 , which is an exemplary flow chart of a spinal endoscope image surgical instrument segmentation method according to some embodiments of the present application, and the spinal endoscope image surgical instrument segmentation method 100 mainly includes the following steps:
[0048] In step 101, an endoscopic image of the spine containing detailed information of the spine obstructed by surgical instruments is acquired.
[0049] In a specific implementation, a high-definition endoscope device is first used to percutaneously acquire real-time video of the spinal region, and then frame extraction is performed on the acquired video to select key image frames blocked by surgical instruments, and the key image frames are used as spinal endoscope images in the embodiments of the present application. The spinal endoscope image includes surgical instruments in an in vivo environment. As a specific embodiment, the surgical instruments may be, for example, bone-holding forceps, a grinding drill, or a probe. Figure 2 is an endoscopic image of the spine before the surgical instrument is segmented according to some embodiments of the present application, wherein Figure 2 (a) is a schematic diagram of the spinal endoscopy image before the rongeur is divided. Figure 2 (b) is an endoscopic image of the spine before the drill is segmented. Figure 2 (c) is the spinal endoscopy image before the probe is segmented. These images are some of the key images captured during percutaneous spinal endoscopic surgery, in which two main objects can be observed: one is the anatomical structure of spinal surgery, including bony tissue and soft tissue; the other is the characteristics of surgical instruments; the anatomical structure appears relatively blurred in the image. Due to the local overexposure caused by the endoscopic light source, some bony areas show higher brightness. At the same time, the surgical instruments have a significant gloss on their surface due to the high reflectivity of metal materials, which forms a certain visual contrast with the background tissue. However, these contrasts are limited by the light changes in the surgical environment and the texture similarity between tissues, making the boundary between the instrument and the background unclear and easy to confuse. In addition, since the instrument may be at different angles and depths during the operation, its shape and position have large dynamic changes in the image, which increases the difficulty of the image segmentation algorithm in subsequent processing.
[0050] In step 102, a multi-scale convolution operation is performed on the spinal endoscope image based on a convolutional neural network to obtain the spot interference characteristics of the surgical instrument in the contour edge detection.
[0051] In some embodiments, performing a multi-scale convolution operation on the spinal endoscope image based on a convolutional neural network to obtain the spot interference feature of the surgical instrument in the contour edge detection can be achieved by the following steps:
[0052] Build a recognition model for spot recognition based on convolutional neural network;
[0053] Designing convolution kernels of multiple scales in the recognition model, and then capturing the light spot area in the spinal endoscope image through the convolution kernels of each scale, so as to obtain the light spot features at each scale;
[0054] The light spot interference characteristics of the surgical instrument in the contour edge detection are determined based on all the light spot characteristics.
[0055] When implementing it, first, you can choose Faster The R-CNN deep learning model is used as the basic framework of the recognition model, and the recognition model is pre-trained using historical spinal endoscopy image data, and a multi-scale convolution module is designed in the recognition model. The multi-scale convolution module contains at least two or more convolution kernels. In the embodiment of the present application, three convolution kernels of different sizes (such as 3×3, 5×5 and 7×7) are respectively configured in the multi-scale convolution module to capture the spot features of different scales in the image; secondly, the spinal endoscopy image is used as the input of the recognition model, and the recognition model gradually extracts the feature map of the spot area in the convolution layer through forward propagation. The shallow convolution kernel (i.e., 3×3) captures the local edge details of the spot, the middle convolution kernel (i.e., 5×5) captures the global spot morphology, and the deep convolution kernel (i.e., 7×7) captures the global spot distribution. The shallow convolution kernel captures the local edge details of the spot as the spot feature at this scale, the middle convolution kernel captures the global spot morphology as the spot feature at this scale, and the deep convolution kernel captures the global spot distribution as the spot feature at this scale;Then, a feature pyramid network can be used to fuse the light spot features at all scales to obtain the light spot fusion features, and then the light spot fusion features are input into the pre-constructed saliency detection network, and the light spot and the surgical instrument contour are extracted through the convolution layer in the saliency detection network, and the overlapping part of the light spot and the surgical instrument contour is used as the light spot interference feature of the surgical instrument in the contour edge detection. In specific implementation, the saliency detection network in the present application can be trained based on the existing U-Net structural framework, that is, first, a specified number of spinal endoscope images are prepared, and the salient areas (light spot areas and surgical instrument contour areas) of all spinal endoscope images are annotated, and then all the annotated spinal endoscope images are divided into a test set and a training set according to a ratio of 2:8. Secondly, the existing U-Net structural framework is selected as the structural framework of the saliency detection network, and the existing cross entropy loss function is set as the loss function of the saliency detection network. Then, the training set is input into the saliency detection network with the above-mentioned U-Net structural framework as the structural framework and the cross entropy loss function as the loss function, and the light spot and the surgical instrument are extracted through forward propagation. The significant features of the machine contour are obtained by using the Adam algorithm, and then the weights of the significant detection network (such as the convolution kernel size, the weight of the fully connected layer, and the normalized weight) are adjusted according to the loss function. The generalization performance of the network model is gradually improved through multiple rounds of training. Finally, the detection performance of the significant detection network is evaluated using the test set, and the Dice coefficient and IoU evaluation index are used to quantify the recognition ability of the significant detection network for the significant area, and finally the training of the significant detection network is completed. Among them, the U-Net structure framework has a symmetrical encoder part and a decoder part, which can effectively perform high-resolution contour extraction. Among them, the encoder part is responsible for extracting the high-level features of the spinal endoscope image layer by layer. Using convolution and pooling operations, the encoder part will gradually reduce the spatial resolution and enhance the abstract features. The decoder part performs upsampling operations based on the features extracted by the encoder, and gradually restores the spatial resolution of the spinal endoscope image, so that the network can perform high-precision segmentation output, and then obtain the light spot and surgical instrument contour. It should be noted that the light spot interference feature in this application refers to the artifact feature generated by the light spot in the spinal endoscope image. ;
[0056] In step 103, the information variance of the texture in the spinal endoscopic image is determined, and based on the information variance of the texture, multi-gradient synchronous extraction is performed on the contour information of the surgical instrument in the spinal endoscopic image to obtain the contour feature map of the surgical instrument under each gradient branch, and then the recognition attention of the corresponding channel of each gradient branch is determined according to the texture information difference between each contour feature map.
[0057] In some embodiments, reference Figure 4As shown, this figure is a schematic diagram of the process of determining information variance in some embodiments of the present application. In this embodiment, determining the information variance of the texture in the spinal endoscopy image can be achieved by using the following steps:
[0058] In step 1031, the spinal endoscopy image is divided into a plurality of local image regions;
[0059] In step 1032, the grayscale features of the pixels in each local image region are extracted by using a grayscale co-occurrence matrix;
[0060] In step 1033, the information variance of the texture in the spinal endoscopy image is determined based on all the grayscale features.
[0061] In the specific implementation, first, the spinal endoscope image is divided into multiple local image areas of fixed size (such as 8×8 or 16×16 pixel blocks); then, a grayscale co-occurrence matrix is generated for each local image area, and the grayscale differences of all adjacent pixels are counted through the grayscale co-occurrence matrix, and then all the grayscale differences are combined into a set to describe the grayscale features of the pixels in the corresponding local image area; finally, the grayscale differences in all grayscale features are substituted into the variance formula to calculate the grayscale variance, and then the grayscale variance is used as the information variance of the texture in the spinal endoscope image.
[0062] It should be noted that the information variance in the present application is an indicator for measuring the complexity of texture in spinal endoscopic images. The larger the value of the information variance, the more complex the texture changes in the spinal endoscopic images and the richer the information contained. The smaller the value of the information variance, the smoother the texture changes in the image texture of the spinal endoscopic images and the less information contained.
[0063] In some embodiments, based on the information variance of the texture, multi-gradient synchronous extraction of the contour information of the surgical instrument in the spinal endoscope image is performed to obtain the contour feature map of the surgical instrument under each gradient branch, which can be achieved by the following steps:
[0064] Build a feature extraction model for surgical instrument contour information extraction based on YOLOv8n;
[0065] According to the information variance of the texture, the fine granularity of the shallow gradient branch, the middle gradient branch and the deep gradient branch in the feature extraction model in extracting the contour feature map of the surgical instrument is set;
[0066] The spinal endoscope image is used as the initialization parameter of the feature extraction model, and then the contour feature map of the surgical instrument under each gradient branch is synchronously extracted through the feature extraction model.
[0067] In the specific implementation, first, the lightweight structure of YOLOv8n can be used as the basic framework of the feature extraction model, and the multi-scale feature extraction capability of its backbone network can be retained; then, the feature extraction module in the original YOLOv8n network framework is improved. The improved feature extraction module has three branches (i.e., shallow gradient branch, middle gradient branch and deep gradient branch) to perform multi-gradient synchronous processing on the image. Specifically, the fine-grainedness of each gradient branch of the feature extraction model can be dynamically adjusted according to the information variance of the texture in the spinal endoscopic image. For example, a small convolution kernel (such as 3×3) is used in the shallow branch to capture local edge details, and a feature pyramid network is combined in the middle branch to perform multi-scale feature aggregation. In the deep branch, the receptive field is increased (such as the convolution kernel is 5×5) to extract global contour features; finally, the spinal endoscopic image is used as the input of the feature extraction model, and the contour feature map of each gradient branch is obtained through forward propagation calculation. In the specific implementation, the spinal The endoscopic image is used as the input of the feature extraction model. It is passed layer by layer through the feature extraction model, and after the operations of the convolution layer, activation function, and pooling layer in sequence, a gradually deepening feature representation can be extracted. In each convolution layer, the feature extraction model uses the convolution kernel to perform a convolution operation on the input image to extract local features. The activation function (such as ReLU) introduces a nonlinear relationship, and the pooling layer helps to reduce the size of the feature map and retain important information. For each gradient branch (that is, a convolution kernel branch of different scales), after the above processing, a feature map corresponding to each gradient branch will be obtained. These feature maps are gradually merged through forward propagation, and finally the corresponding contour feature maps are output. Among them, the shallow gradient branch is used to extract the local texture edge information of the surgical instrument contour, the middle gradient branch is used to extract the continuity information of the edge area of the surgical instrument contour, and the deep gradient branch is used to extract the global contour shape and semantic features, and synchronously output the contour feature maps of each gradient branch.
[0068] It should be noted that the fine-grainedness in this application refers to the feature extraction model's ability to resolve and characterize surgical instrument contour information and its accuracy, which is mainly reflected in the levels of capture of local details and global features by different gradient branches; it should also be noted that the shallow gradient branches in this application focus on the fine extraction of local edge details (such as texture details or local boundaries), the middle gradient branches focus on capturing the continuity of regional contours and the correlation between local features, and the deep gradient branches focus on global contour shape and semantic information. By setting the fine-grainedness of the shallow, middle and deep gradient branches, the contour features of surgical instruments can be decomposed and extracted synchronously at different scales and semantic levels, thereby ensuring the balanced expression of global and local features.
[0069] In some embodiments, determining the recognition attention of the channel corresponding to each gradient branch according to the texture information difference between each contour feature map can be implemented by the following steps:
[0070] Determine the texture entropy of each contour feature map;
[0071] Performing global average pooling on each contour feature map, extracting global texture features of each contour feature map, and then determining the texture difference between each contour feature map and other contour feature maps;
[0072] The recognition attention of each gradient branch corresponding to the channel is determined by all texture differences and texture entropy of each contour feature map.
[0073] In the specific implementation, first, for each contour feature map, all pixel values in the contour feature map are normalized to obtain the normalized pixel values of all pixel values, and then all the normalized pixel values are substituted into the information entropy calculation formula, and the calculated information entropy is used as the texture entropy of the contour feature map, and then the texture entropy of all contour feature maps can be obtained. The texture entropy is an indicator for quantifying texture complexity. The higher the texture entropy, the more texture detail information the contour feature map contains; then, for each contour feature map, the built-in pooling function in the deep learning framework (such as PyTorch) can be used to perform global average pooling on the contour feature maps respectively. For example, torch.nn.AdaptiveAvgPool2d((1,1)) is called in PyTorch to perform a global average operation on each channel of the input contour feature map, compress the H×W two-dimensional contour feature map into a single value, generate a global feature value for each channel, and then all channels are averaged. The global eigenvalues of the contour feature map form a global feature vector, and the global feature vector is used as the global texture feature of the contour feature map, and then the global texture feature of each contour feature map is obtained, and then the global texture feature of each two contour feature maps is measured by cosine similarity to obtain the texture difference between each two contour feature maps, and then the average texture difference between the contour feature map and all other contour feature maps is used as the texture difference between the contour feature map and other contour feature maps, and the texture information difference between the contour feature maps can be described by the texture difference, and then the texture difference between each contour feature map and other contour feature maps is obtained; finally, a contour feature map is selected as the selected contour feature map, and the product of the texture difference between the selected contour feature map and other contour feature maps and the texture entropy of the selected contour feature map is used as the recognition attention of the corresponding channel of the gradient branch where the selected contour feature map is located. The recognition attention of the corresponding channel of each gradient branch can be obtained by the above method.
[0074] It should be noted that the recognition attention in this application refers to the contribution of the corresponding channel in the surgical instrument contour detection. The recognition attention can enhance the segmentation model's perception of the surgical instrument contour features, while suppressing redundant features, thereby improving the robustness of the surgical instrument contour feature extraction.
[0075] In step 104, all contour feature maps are convolutionally fused according to the recognition attention of the corresponding channel of each gradient branch and the light spot interference feature to obtain a fused contour of the surgical instrument, and then the spinal endoscope image is segmented based on the fused contour to obtain a segmentation feature map of the surgical instrument.
[0076] In some embodiments, convolution fusion is performed on all contour feature maps according to the recognition attention of the channel corresponding to each gradient branch and the light spot interference feature to obtain the fused contour of the surgical instrument, which can be achieved by the following steps:
[0077] Determine the fusion coefficient of the contour feature map under each gradient branch according to the recognition attention of the channel corresponding to each gradient branch;
[0078] All contour feature maps are weightedly fused based on each fusion coefficient to obtain a feature fusion map;
[0079] Convolution interference compensation is performed on the feature fusion image according to the light spot interference feature, and a surgical instrument contour feature extraction operation is further performed to obtain a fusion contour of the surgical instrument.
[0080] In the specific implementation, first, a gradient branch is selected as the selected gradient branch, and the ratio between the recognition attention of the channel corresponding to the selected gradient branch and the sum of all recognition attentions is used as the fusion coefficient of the contour feature map under the selected gradient branch, and the fusion coefficient of the contour feature map under the remaining gradient branches is further determined; then, each fusion coefficient is used as the weight of each contour feature map in the weighted fusion process, and then all the weights are used to weightedly fuse each contour feature map to obtain a feature fusion map, which integrates the texture, edge and semantic information of multiple gradient branches; finally, a compensation convolution module is designed, and the spot interference feature is used as the compensation guidance parameter of the compensation convolution module. The compensation guidance parameters can guide the contour segmentation model to perform spot compensation on the contour of the surgical instrument, thereby reducing the degree of interference of the spot on the segmentation of surgical instruments in the spinal endoscope image. It should also be noted that the compensation convolution module in the present application can suppress the intensity of the spot interference feature area and enhance the edge features of the non-interference area. The feature fusion map is further convolved through the compensation convolution module, and the spot area in the feature fusion map is weakened through the convolution layer in the compensation convolution module. The compensated feature fusion map is used as the compensated contour map, and the compensated contour map is input into the pre-trained contour feature extraction model, and the fusion contour of the surgical instrument is output through the contour feature extraction model.
[0081] It should be noted that the role of convolution fusion in the present application is to integrate contour feature maps from different gradient branches into a unified feature expression through dynamic weight allocation and weighted fusion mechanism, thereby enhancing the comprehensive perception ability of the contour of surgical instruments. Specifically, convolution fusion uses the recognition attention of the channel to determine the fusion coefficient of each feature map, highlighting the contribution of key gradient branches, and at the same time combines the compensation mechanism of the spot interference feature to optimize the texture and edge details of the fused feature map through convolution operation, and suppress the influence of the interference area on the feature expression. This convolution fusion method can effectively improve the clarity and accuracy of the contour of surgical instruments, and can provide more reliable input for subsequent segmentation and detection tasks.
[0082] In some embodiments, performing image segmentation on the spinal endoscope image based on the fusion contour to obtain a segmentation feature map of the surgical instrument can be achieved by using the following steps:
[0083] determining a pixel segmentation region of a surgical instrument in the spinal endoscope image according to the fusion contour;
[0084] A segmentation feature map of the surgical instrument is generated through the pixel segmentation area.
[0085] In the specific implementation, first, the fusion contour is aligned with the spinal endoscope image at the pixel level, and the highlighted area of the fusion contour is used as the initial mask. A segmentation model is further pre-trained based on Mask R-CNN, and the initial mask is used as the guiding parameter of the segmentation model to guide the contour segmentation process through the fusion contour. Then, the shape and boundary of the surgical instrument area are captured through the multi-scale feature extraction module of the segmentation model, and the pixel segmentation area of the surgical instrument is obtained by generating a preliminary segmentation through forward propagation; then, the pixel segmentation area obtained by the preliminary segmentation is post-processed, and the boundary of the pixel area is optimized through the conditional random field (CRF), and the small noise area in the pixel area is removed and the segmentation holes are filled in combination with morphological operations (such as expansion and erosion) to obtain the segmentation feature map of the surgical instrument. Figure 3 is a spinal endoscope image segmented by surgical instruments according to some embodiments of the present application, and Figure 2 Compared with the images (a)-(c) of Fig. 1, the spinal endoscope segmentation feature map after surgical instrument segmentation can better reflect its pixel range and boundary information. Figure 3 (a) is the spinal endoscope segmentation feature image after rongeur segmentation. Figure 3 (b) is the spinal endoscope segmentation feature image after grinding segmentation. Figure 3 (c) is the spinal endoscope segmentation feature image after probe segmentation. Figure 3The segmented instruments in (a)-(c) are marked with boxes and will not be described in detail here. The above method can effectively improve the accuracy and robustness of surgical instrument area segmentation.
[0086] It should be noted that the present application scheme uses multi-gradient synchronous extraction, and gradient branches at different levels collaborate to capture local details, regional continuity and global semantic information, thereby enhancing the precise modeling of edge features of surgical instruments; secondly, multi-gradient fusion integrates cross-level semantic information, enhances the ability to recognize surgical instruments in complex backgrounds, and enables the model to effectively distinguish between instruments and surrounding tissues; finally, the fusion feature and spot interference feature compensation mechanism further improves the robustness of the model to interference areas, ensuring the boundary clarity and regional integrity of the segmentation results, thereby improving the detection accuracy and segmentation performance of the algorithm as a whole.
[0087] On the other hand, in some embodiments, the present application provides a spinal endoscope image surgical instrument segmentation system, referring to Figure 5 , which is a schematic diagram of the structure of a spinal endoscope image surgical instrument segmentation system according to some embodiments of the present application, the spinal endoscope image surgical instrument segmentation system 400 includes: an acquisition module 401, a processing module 402 and an execution module 403, which are respectively described as follows:
[0088] Acquisition module 401, in this application, acquisition module 401 is mainly used to acquire spinal endoscopic images with spinal detail information obstructed by surgical instruments;
[0089] Processing module 402, in the present application, the processing module 402 is used to perform a multi-scale convolution operation on the spinal endoscope image based on a convolutional neural network to obtain a light spot interference feature of the surgical instrument in contour edge detection;
[0090] The processing module 402 in the present application is also used to determine the information variance of the texture in the spinal endoscope image, perform multi-gradient synchronous extraction on the contour information of the surgical instrument in the spinal endoscope image based on the information variance of the texture, obtain the contour feature map of the surgical instrument under each gradient branch, and then determine the recognition attention of the channel corresponding to each gradient branch according to the texture information difference between each contour feature map;
[0091] Execution module 403. In the present application, execution module 403 is mainly used to perform convolution fusion on all contour feature maps according to the recognition attention of the channel corresponding to each gradient branch and the light spot interference feature to obtain the fused contour of the surgical instrument, and then perform image segmentation on the spinal endoscope image based on the fused contour to obtain the segmentation feature map of the surgical instrument.
[0092] In addition, the present application also provides a computer device, which includes a memory and a processor, wherein the memory stores codes, and the processor is configured to obtain the codes and execute the above-mentioned spinal endoscopic image surgical instrument segmentation method.
[0093] In some embodiments, reference Figure 6 , which is a schematic diagram of the structure of a computer device for implementing a spinal endoscope image surgical instrument segmentation method according to some embodiments of the present application. The spinal endoscope image surgical instrument segmentation method in the above embodiment can be Figure 6 The computer device 500 shown in the figure is implemented, and the computer device 500 includes at least one processor 501, a communication bus 502, a memory 503 and at least one communication interface 504.
[0094] The processor 501 may be a general-purpose central processing unit (CPU) or an application-specific integrated circuit (ASIC).
[0095] The communication bus 502 may be used to transmit information between the above-mentioned components.
[0096] The memory 503 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including a compressed optical disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 503 may exist independently and be connected to the processor 501 via the communication bus 502. The memory 503 may also be integrated with the processor 501.
[0097] The memory 503 is used to store the program code for executing the solution of the present application, and the execution is controlled by the processor 501. The processor 501 is used to execute the program code stored in the memory 503. The program code may include one or more software modules. The spinal endoscope image surgical instrument segmentation method in the above embodiment can be implemented by the processor 501 and one or more software modules in the program code in the memory 503.
[0098] The communication interface 504 uses any transceiver or other device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0099] In a specific implementation, as an embodiment, a computer device may include multiple processors, each of which may be a single-CPU processor or a multi-CPU processor. The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0100] The above-mentioned computer device may be a general-purpose computer device or a special-purpose computer device. In a specific implementation, the computer device may be a desktop computer, a portable computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device or an embedded device. The embodiment of the present application does not limit the type of computer device.
[0101] In addition, the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned spinal endoscope image surgical instrument segmentation method is implemented.
[0102] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0103] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A method for segmenting surgical instruments in spinal endoscope images, characterized in that: The steps include: Acquire spinal endoscopic images with spinal detail information obscured by surgical instruments; Performing a multi-scale convolution operation on the spinal endoscope image based on a convolutional neural network to obtain a light spot interference feature of the surgical instrument in contour edge detection, wherein the light spot interference feature refers to an artifact feature generated by a light spot in the spinal endoscope image; Determine the information variance of the texture in the spinal endoscopic image, perform multi-gradient synchronous extraction on the contour information of the surgical instrument in the spinal endoscopic image based on the information variance of the texture, obtain the contour feature map of the surgical instrument under each gradient branch, and then determine the recognition attention of the channel corresponding to each gradient branch according to the texture information difference between each contour feature map; Performing convolution fusion on all contour feature maps according to the recognition attention of the channel corresponding to each gradient branch and the light spot interference feature to obtain a fused contour of the surgical instrument, and then performing image segmentation on the spinal endoscope image based on the fused contour to obtain a segmentation feature map of the surgical instrument; Among them, the multi-scale convolution operation is performed on the spinal endoscope image based on the convolutional neural network to obtain the spot interference features of the surgical instrument in the contour edge detection, which specifically include: Build a recognition model for spot recognition based on convolutional neural network; Designing convolution kernels of multiple scales in the recognition model, and then capturing the light spot area in the spinal endoscope image through the convolution kernels of each scale, so as to obtain the light spot features at each scale; Determine the light spot interference characteristics of the surgical instrument in the contour edge detection according to all the light spot characteristics; Among them, determining the recognition attention of each gradient branch corresponding to the channel according to the texture information difference between each contour feature map specifically includes: Determine the texture entropy of each contour feature map; Performing global average pooling on each contour feature map, extracting global texture features of each contour feature map, and then determining the texture difference between each contour feature map and other contour feature maps, wherein the texture difference between the contour feature maps is described by the texture difference; Determine the recognition attention of the corresponding channel of each gradient branch through all texture differences and texture entropy of each contour feature map; Among them, all contour feature maps are convoluted and fused according to the recognition attention of the corresponding channel of each gradient branch and the light spot interference feature to obtain the fused contour of the surgical instrument, which specifically includes: Determine the fusion coefficient of the contour feature map under each gradient branch according to the recognition attention of the channel corresponding to each gradient branch; All contour feature maps are weightedly fused based on each fusion coefficient to obtain a feature fusion map; Convolution interference compensation is performed on the feature fusion image according to the light spot interference feature, and a surgical instrument contour feature extraction operation is further performed to obtain a fusion contour of the surgical instrument.
2. The method according to claim 1, characterized in that Determining the information variance of the texture in the spinal endoscopy image specifically includes: dividing the spinal endoscopy image into a plurality of local image regions; The grayscale features of pixels in each local image area are extracted through the grayscale co-occurrence matrix; The information variance of the texture in the spinal endoscopy image is determined based on all the grayscale features.
3. The method according to claim 1, characterized in that The image segmentation of the spinal endoscope image based on the fusion contour to obtain the segmentation feature map of the surgical instrument specifically includes: determining a pixel segmentation region of a surgical instrument in the spinal endoscope image according to the fusion contour; A segmentation feature map of the surgical instrument is generated through the pixel segmentation area.
4. The method according to claim 1, characterized in that The spinal endoscope image contains surgical instruments in an in vivo environment.
5. A spinal endoscope image surgical instrument segmentation system, which uses the method according to any one of claims 1 to 4 to segment spinal endoscope image surgical instruments, characterized in that: The system includes: An acquisition module, used for acquiring an endoscopic image of the spine in which detailed information of the spine is obscured by surgical instruments; A processing module, used for performing a multi-scale convolution operation on the spinal endoscope image based on a convolutional neural network to obtain a light spot interference feature of the surgical instrument in contour edge detection; The processing module is further used to determine the information variance of the texture in the spinal endoscope image, perform multi-gradient synchronous extraction on the contour information of the surgical instrument in the spinal endoscope image based on the texture information variance, obtain the contour feature map of the surgical instrument under each gradient branch, and then determine the recognition attention of the channel corresponding to each gradient branch according to the texture information difference between each contour feature map; The execution module is used to perform convolution fusion on all contour feature maps according to the recognition attention of the channel corresponding to each gradient branch and the light spot interference feature to obtain the fused contour of the surgical instrument, and then perform image segmentation on the spinal endoscope image based on the fused contour to obtain the segmentation feature map of the surgical instrument.
6. A computer device, comprising a memory and a processor, wherein the memory stores a code, characterized in that: The processor is configured to obtain the code and execute the spinal endoscope image surgical instrument segmentation method as described in any one of claims 1 to 4.
7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the spinal endoscope image surgical instrument segmentation method as described in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Real-time neurosurgical operating instrument segmentation method and device based on an endoscope image, and storage medium
CN112396601A
Endoscope image processing method and endoscope system thereof
CN119168994A