Robot medical image segmentation and feature extraction method for precise operation
By using a dynamic segmentation network model of a bidirectional attention architecture in robot-assisted surgery, combined with Transformer and convolutional coding branches, the precise segmentation of anatomical structures and instruments in surgical videos is achieved, solving the problem of poor segmentation performance in the prior art, and improving the accuracy and reliability of the surgery.
Patent Information
- Application Number
- CN202510706042.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The prior art is difficult to effectively segment anatomical structures and instruments in static images and dynamic videos in robot-assisted surgery, especially in the case of rapid movement of the instrument and interactive occlusion, resulting in poor segmentation performance.
A dynamic segmentation network model with bidirectional attention architecture is proposed, combining Transformer coded branches and convolution coded branches, and through multi-stage feature fusion units and decoding output modules, the precise segmentation and feature extraction of anatomical tissues and dynamic instruments in surgical scenarios is realized.
The pixel-level segmentation of the anatomical structure and instruments in the surgical video is realized, the problems of instrument motion blur and occlusion are solved, segmentation performance and reliability are improved, and reliable support is provided for precise surgery.
Smart Images

Figure CN120236083A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of real-time medical image segmentation for robot-assisted surgery, and more specifically, to a method for robot-assisted medical image segmentation and feature extraction for precise surgery. Background Art
[0002] With the development of robot-assisted minimally invasive laparoscopic surgery, accurate surgical medical image segmentation has become crucial. However, existing methods mainly face two challenges: one is the limitation of static images, including the ambiguity of local features between different anatomical structures and the fine-grained structural complexity of deformable instruments; the other is the complexity of dynamic videos, such as blurring caused by rapid movement of instruments and inevitable interactive occlusions. Although existing methods mainly focus on spatial feature extraction, they ignore the temporal dependence in the surgical video stream, resulting in poor performance in surgical scene segmentation tasks. Summary of the Invention
[0003] The object of the present invention is to propose a method for robot-assisted medical image segmentation and feature extraction for precise surgery, which realizes precise segmentation and feature extraction of anatomical tissues and dynamic instruments in the surgical scene by combining temporal dynamics and structural asymmetry, so as to provide reliable support for precise surgery.
[0004] To achieve the above object, the present invention proposes a method for robot-assisted medical image segmentation and feature extraction for precise surgery, including: Constructing a dynamic segmentation network model with a bidirectional attention architecture, the dynamic segmentation network model including an input preprocessing module, a Transformer encoding branch, a convolutional encoding branch, a multi-stage feature fusion unit, and a decoding output module; wherein, the input preprocessing module is used to extract features from video frames and generate a multi-scale feature pyramid; the Transformer encoding branch is used to perform temporal feature modeling based on the input feature pyramid and output temporally enhanced features; the convolutional encoding branch is used to enhance anatomical features and surgical instrument features through convolutional operations based on the input feature pyramid and output spatially enhanced features; the multi-stage feature fusion unit is used to perform multi-stage iterative fusion on the outputs of the Transformer encoding branch and the convolutional encoding branch and output the final fused features; the decoding output module is used to input the fused features into a Transformer decoder for decoding to generate pixel-level segmentation masks of anatomical structures and surgical instruments; Training the dynamic segmentation network model using a training data set; During robot-assisted surgery, using the trained dynamic segmentation network model, medical image segmentation and feature extraction are performed based on the real-time input medical image video sequence, and the segmentation mask output by the model is transmitted to the surgical robot control system for surgical safety warning and navigation assistance.
[0005] Optionally, the input preprocessing module includes: A ResNet-50 backbone network for extracting features from the input consecutive video frames to generate feature maps of different resolutions for each frame, where the feature maps contain the spatial feature information of each frame; A temporal dimension expansion unit for splicing the feature maps of consecutive multiple frames along the time dimension to form a spatio-temporal feature tensor; A 3D convolutional layer for applying 3D convolution operations to the spatio-temporal feature tensor for progressive downsampling to generate a multi-scale feature pyramid, where the multi-scale feature pyramid includes multiple feature layers of different resolutions, and each feature layer contains time dimension information.
[0006] Optionally, the Transformer encoding branch includes a time query propagator, a multi-head attention layer, and a dimension expansion layer connected in sequence; The time query propagator includes: A feature projection unit for splicing the consecutive multiple frames of features at a certain level in the feature pyramid along the time dimension, projecting the spliced features through a learnable key-value weight matrix to generate a key-value matrix; A time query generation unit for performing convolution operations on the spliced features to extract high-activation regions, using Top-K to select key content features to generate content queries, and generating position queries based on the position embeddings of the previous frame, middle frame, and next frame of the input through forward propagation and backward propagation; The multi-head attention layer is used to perform multi-head attention calculations based on the content queries, the position queries, and the key-value matrix, and refine the features; The dimension expansion layer is used to expand the channels of the refined features and output enhanced features with temporal consistency.
[0007] Optionally, the generation formula for the content query is:
[0008] Where is the query content vector for capturing key region features; is the feature map of the t-th frame; is the consecutive T frames of features spliced along the time dimension; Conv represents a 2D convolution with a convolution kernel size of 3×3 and an output channel number of D ; Indicates selecting the top K high response regions based on activation values to generate a sparse content query vector.
[0009] Optionally, the generation formula for the position query is:
[0010] where The position query vector of the t th frame, used to encode spatio-temporal position information; is the consecutive T frame features concatenated along the time dimension; is a learnable weight matrix; Conv represents 2D convolution; represents a fully connected layer that keeps the dimension of the position query as D ; and respectively represent the position query vectors of the t+1 th frame and the t-1 th frame.
[0011] Optionally, the multi-head attention calculation formula is:
[0012] where is the query content vector; The position query vector of the t th frame; K is the key matrix generated by feature projection; is the transpose matrix of K; V is the value matrix generated by feature projection; represents the temporal aggregation of the position query; ⊙ represents element-wise multiplication; is the scaling factor for stabilizing the gradient; represents normalization along the key dimension, and the output is the attention weight matrix; Attention(Q, K, V) represents multi-head attention calculation, and the output is the attention-weighted feature matrix, where each row corresponds to the enhanced feature of a key region.
[0013] Optionally, the convolutional encoding branch includes an aggregated asymmetric feature pyramid module and a feature compression layer; The aggregated asymmetric feature pyramid module includes: A spatio-temporal aggregation layer for performing convolution operations on consecutive multi-frame features at a certain level in the feature pyramid to obtain spatio-temporal aggregation features; Anatomical perception enhancement branch for performing symmetric convolution operations on the spatio-temporal aggregation feature map to enhance the anatomical structure features of irregular polygons and output multi-scale anatomical features; The instrument perception enhancement branch is used to perform a convolution operation on the spatio-temporal aggregation feature map using an asymmetric strip convolution pair to enhance the instrument features of regular strips and output multi-scale instrument features; The attention weighted fusion layer is used to perform attention fusion calculation on the enhanced anatomical features, instrument features, and original aggregation features to generate spatially enhanced features; The feature compression layer is used to compress the spatially enhanced features to the same channel dimension as the Transformer encoding branch.
[0014] Optionally, the calculation formula of the spatio-temporal aggregation layer is:
[0015] Where, is the spatio-temporal aggregation feature; is the continuous T frame features concatenated along the time dimension; Conv 5×5×T represents a convolution kernel of size 5×5×T; The calculation formula of the anatomical perception enhancement branch is:
[0016] The calculation formula of the instrument perception enhancement branch is:
[0017] Where, k m is the spatial size of the symmetric convolution kernel; t m is the depth of the time dimension.
[0018] Optionally, the attention fusion calculation formula is:
[0019] Where, ⨂ represents element-wise multiplication; Conv 1×1×T represents spatio-temporal compression convolution to generate an attention map; is the spatially enhanced feature of attention weighted fusion.
[0020] Optionally, the multi-stage feature fusion unit adds the outputs of the Transformer encoding branch and the convolution encoding branch element-wise, and through multiple residual iterations with parameter sharing, generates the final fusion feature.
[0021] The beneficial effects of the present invention are as follows: The present invention constructs a dynamic segmentation network model with a bidirectional attention architecture. In the Transformer encoding branch, by integrating the global temporal modeling of the Transformer and the local spatial perception of convolution, spatio-temporal feature co-optimization is achieved (modeling cross-frame dependencies through a dynamic query mechanism based on a temporal query propagator, enhancing frame feature representations using temporal consistency constraints, and solving the problems of instrument motion blur and occlusion); in the convolution encoding branch, anatomical features and surgical instrument features are enhanced through convolution operations to extract spatially enhanced features (separating and enhancing the feature expressions of anatomical tissues (irregular shapes) and surgical instruments (long strip structures) by aggregating an asymmetric feature pyramid module using symmetric and asymmetric convolutions). Then, through a multi-stage feature fusion unit, the features extracted from the Transformer encoding branch and the convolution encoding branch are iteratively fused in multiple stages to achieve progressive optimization of spatial accuracy and temporal consistency, and the final fused features are output. Finally, the fused features are decoded through a decoding output module to generate pixel-level segmentation masks of anatomical structures and surgical instruments, thereby achieving precise segmentation of anatomical structures and instruments in surgical videos.
[0022] The system of the present invention has other characteristics and advantages that will be apparent from or will be described in detail in the accompanying drawings and the subsequent detailed description incorporated herein. These accompanying drawings and the detailed description together are used to explain the specific principles of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] By describing the exemplary embodiments of the present invention in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present invention will become more apparent. In the exemplary embodiments of the present invention, the same reference numerals generally represent the same components.
[0024] Figure 1 FIG. shows a schematic diagram of the dynamic segmentation network model in an embodiment of the present invention.
[0025] Figure 2 FIG. shows a schematic diagram of the temporal query propagator (TQP) in an embodiment of the present invention.
[0026] Figure 3 FIG. shows a schematic diagram of the aggregating asymmetric feature pyramid module (AAFP) in an embodiment of the present invention. DETAILED DESCRIPTION
[0027] The present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art. Embodiment
[0028] This embodiment provides a method for robot-assisted medical image segmentation and feature extraction for precise surgery, including: S1: Construct a dynamic segmentation network model with a bidirectional attention architecture. The dynamic segmentation network model includes an input preprocessing module, a Transformer encoding branch, a convolutional encoding branch, a multi-stage feature fusion unit, and a decoding output module. Among them, the input preprocessing module is used to extract features from video frames and generate a multi-scale feature pyramid; the Transformer encoding branch is used to perform temporal feature modeling based on the input feature pyramid and output temporally enhanced features; the convolutional encoding branch is used to enhance anatomical features and surgical instrument features through convolutional operations based on the input feature pyramid and output spatially enhanced features; the multi-stage feature fusion unit is used to perform multi-stage iterative fusion on the outputs of the Transformer encoding branch and the convolutional encoding branch and output the final fused features; the decoding output module is used to input the fused features into the Transformer decoder for decoding to generate pixel-level segmentation masks for anatomical structures and surgical instruments. Specifically, as Figure 1 shown, a dynamic segmentation network model with a bidirectional attention architecture is constructed, including a Transformer encoding branch and a convolutional encoding branch, which iteratively fuse global temporal dependencies and local spatial features through multiple stages. Generate dynamic spatio-temporal queries through a Time Query Propagator (TQP), combine content queries and position queries, calculate multi-head attention weights, and generate occlusion-resistant spatio-temporal feature tubes; through an Aggregated Asymmetric Feature Pyramid module (AAFP), perform symmetric convolution and asymmetric strip convolution operations on anatomical structures and surgical instruments respectively to fuse spatio-temporal aggregation features. Then, add the outputs of the Transformer branch and the convolutional branch element by element, and generate the final fused features after M times of parameter-sharing residual iterations; finally, input the fused features into the Transformer decoder to generate pixel-level segmentation masks.
[0029] Furthermore, in this embodiment, the input preprocessing module includes: ResNet-50 backbone (fully connected layer removed), which is used to extract features from the input consecutive video frames, generate feature maps of different resolutions for each frame, and the feature maps contain the spatial feature information of each frame. For example: input single-frame images in sequence → ResNet-50 → output 2D feature maps of different levels (resolutions gradually decreasing, such as 1 / 4, 1 / 8 of the original image resolution, etc.) (such as C1: 1 / 4, C2: 1 / 8, C3: 1 / 16, C4: 1 / 32).
[0030] Temporal dimension expansion unit, which is used to splice the feature maps of consecutive multiple frames along the time dimension to form a spatio-temporal feature tensor. For example, splice the ResNet output features of consecutive T frames along the time dimension to form a spatio-temporal feature tensor; 3D convolutional layer, which is used to gradually downsample the spatio-temporal feature tensor by applying 3D convolution operations to generate a multi-scale feature pyramid. The multi-scale feature pyramid includes multiple feature layers of different resolutions, and each feature layer contains time dimension information. For example: apply multiple groups of 3D convolutions (kernel size 3×3×3) to the spatio-temporal feature tensor to generate a multi-scale spatio-temporal feature pyramid (resolutions: 1 / 4, 1 / 8, 1 / 16, 1 / 32).
[0031] In one example, the input preprocessing module performs multi-scale feature extraction on the input video frames based on the ResNet-50 network combined with 3D convolution, and outputs a multi-scale feature pyramid with resolutions of {1 / 4, 1 / 8, 1 / 16, 1 / 32} and feature dimensions of {256, 512, 1024, 2048}. Preprocessing reduces the computational burden and retains multi-scale spatial information, providing basic features for the subsequent encoding stage.
[0032] Furthermore, in this embodiment, the Transformer encoding branch includes a time query propagator (TQP), a multi-head attention layer, and a dimension expansion layer connected in sequence; the Transformer encoding branch dynamically generates spatio-temporal feature tubes through the time query propagator (TQP) to model the temporal dependencies between video frames and solve the problems of motion blur and occlusion; uses the multi-head attention layer to capture global context semantic information. By extracting anti-occlusion temporal consistency features, the tracking ability of the model for dynamic objects (such as surgical instruments) is enhanced.
[0033] As Figure 2 shown, the time query propagator (TQP) includes: Feature projection unit: used to splice the consecutive multi-frame features of a certain level (such as 1 / 4 level features) in the feature pyramid along the time dimension, and project the spliced features through a learnable key-value weight matrix to generate a key-value matrix. For example, input consecutive T frame features (such as T=5), extract the multi-scale feature pyramid through the input preprocessing module , and then splice the multi-frame features along the time dimension to obtain the spliced features ; Then project the spliced features through the learnable key weight matrix and the value weight matrix to generate the cross-frame key-value matrix K and W . Through multi-frame key-value projection, even if the target is occluded in a certain frame (such as the surgical forceps being covered by tissue), its position and semantics can still be restored through the projection features of adjacent frames.
[0034] Temporal query generation unit: used to perform convolution operations on the spliced features to extract high-activation regions, and adopt Top-K to select key content features to generate content queries, and generate position queries through forward and backward propagation based on the position embeddings of the previous frame, middle frame, and subsequent frame of the input; Temporal queries include content queries and position queries, and each selected temporal query (Q Temp ) corresponds to a spatio-temporal feature tube; Among them, the generation formula of the content query is:
[0035] Among them, is the query content vector, used to capture the key region features; is the feature map of the t-th frame; is the continuous T frame features spliced along the time dimension; Conv represents 2D convolution, the convolution kernel size is 3×3, and the number of output channels is D ; represents selecting the top K high-response regions based on the activation values to generate a sparse content query vector.
[0036] The generation formula of the position query is:
[0037] Among them, The position query vector of the t -th frame, used to encode spatio-temporal position information; is the continuous T frame features spliced along the time dimension; is the learnable weight matrix; Conv represents 2D convolution; represents the fully connected layer, keeping the dimension of the position query as D ; and respectively represent the position query vectors of the t+1 -th frame and the t-1 -th frame.
[0038] The multi - head attention layer is used to perform multi - head attention calculation and feature refinement based on content queries, position queries, and key - value matrices; the multi - head attention calculation formula is:
[0039] Among them, is the query content vector; The t position query vector of the K frame; is the key matrix generated by feature projection; V is the transpose matrix of K; is the value matrix generated by feature projection;
[0040] The dimension expansion layer is used to expand the channels of the refined features and output enhanced features with temporal consistency.
[0041] Furthermore, in this embodiment, the convolutional encoding branch includes an Aggregated Asymmetric Feature Pyramid module (AAFP) and a feature compression layer; the convolutional encoding branch uses symmetric convolutions through the Aggregated Asymmetric Feature Pyramid module (AAFP) to enhance the local features of irregular anatomical structures (such as the liver and gallbladder), and at the same time extracts the details of long - strip instruments (such as surgical forceps) through asymmetric strip convolutions. Thus, the spatial detail features are optimized to distinguish target categories with significantly different morphologies.
[0042] As Figure 3 shown, the Aggregated Asymmetric Feature Pyramid module includes: The spatio - temporal aggregation layer: used to perform convolution operations on consecutive multi - frame features at a certain level in the feature pyramid (such as the 1 / 4 - level features with the same input as the Transformer encoding branch) to obtain spatio - temporal aggregation features; for example, fusing spatio - temporal information through a 5×5×T convolution; the calculation formula of the spatio - temporal aggregation layer is:
[0043] Among them, is the spatio - temporal aggregation feature; is the consecutive T frame features concatenated along the time dimension; Conv 5×5×T represents a convolution kernel of size 5×5×T; Anatomical perception enhancement branch: It is used to perform symmetric convolution operations on the spatio-temporal aggregation feature map to enhance the anatomical structure features of irregular polygons and output multi-scale anatomical features. The calculation formula of the anatomical perception enhancement branch is: ; Instrument perception enhancement branch: It is used to perform convolution operations on the spatio-temporal aggregation feature map using asymmetric strip convolution pairs to enhance the instrument features of regular strips and output multi-scale instrument features; The calculation formula of the instrument perception enhancement branch is:
[0044] where, k m is the spatial size of the symmetric convolution kernel; t m is the depth of the time dimension.
[0045] Attention weighted fusion layer: It is used to perform attention fusion calculation on the enhanced anatomical features, instrument features and original aggregation features to generate spatially enhanced features. The attention fusion calculation formula is:
[0046] where, ⨂ represents element-wise multiplication; Conv 1×1×T represents spatio-temporal compression convolution to generate an attention map; is the spatially enhanced feature of attention weighted fusion.
[0047] The feature compression layer is used to compress the spatially enhanced features to the same channel dimension as the Transformer encoding branch.
[0048] Furthermore, in this embodiment, the multi-stage feature fusion unit adds the outputs of the Transformer encoding branch and the convolutional encoding branch element-wise, and through multiple residual iterations with parameter sharing, generates the final fusion features. Through multi-stage (such as 4-stage) residual connections, the global temporal features of the Transformer branch and the local spatial features of the convolutional branch are fused layer by layer to achieve feature complementarity and progressive optimization. By combining global semantics and local details, more robust spatio-temporal fusion features are generated.
[0049] Specifically, the number of stages of multi-stage feature fusion is consistent with the levels of the feature pyramid (e.g., 4 levels: 1 / 4, 1 / 8, 1 / 16, 1 / 32), and the processing order is to refine step by step from the deep layer (low resolution) to the shallow layer (high resolution). For example, the current layer feature of the input is the multi-scale feature (spatiotemporal feature splicing) of the 1 / 8 layer. The Transformer encoding branch generates dynamic queries, models cross-frame dependencies, and calculates global spatiotemporal weights through the multi-head attention mechanism to output temporally enhanced features; the convolutional encoding branch optimizes the anatomical and instrument features through symmetric / non-symmetric convolutions, compresses the enhanced features to the same dimension as the temporally enhanced features, and outputs spatially enhanced features, and then fuses the features of the two branches by element-wise addition. Then, cross-stage transmission is completed through residual connections, that is, the fused features output by the current layer are transmitted to the next stage (such as the 1 / 4 layer) through parameter-sharing residual connections, layer by layer fusing the global temporal features of the Transformer branch and the local spatial features of the convolutional branch, and outputting a multi-level feature pyramid to achieve feature complementarity and progressive optimization.
[0050] Furthermore, in this embodiment, the decoding output module upsamples the fused features to the original resolution and generates pixel-level segmentation masks through the segmentation head; the joint cross-entropy loss and Dice loss are used to optimize the model training.
[0051] Specifically, the input of the decoding output module is the fused features (1 / 4, 1 / 8, 1 / 16, 1 / 32 layers) output by the multi-stage feature fusion unit. Each layer of feature contains time dimension information (the spatiotemporal interaction results of consecutive T frames). The decoder adopts a progressive upsampling + cross-layer skip connection structure to gradually restore spatial details and integrate global semantics. For example, first, deep feature initialization is performed. The deepest layer feature (1 / 32 layer, with the lowest resolution but the richest semantic information) is input, and the number of channels is compressed through a 3×3 convolution to reduce the computational amount. Then, layer-by-layer upsampling and feature fusion are performed. Upsampling is performed step by step from the deep layer to the shallow layer and concatenated with the corresponding hierarchical bidirectional attention features; after fusion, asymmetric convolution blocks 1×3 and 3×1 convolutions are used to enhance the edges of slender instruments, and residual connections are used: the original features are retained to avoid gradient disappearance; attention weights in the time dimension are introduced during fusion to suppress sudden noise to ensure smooth transitions in the segmentation results of adjacent frames (such as continuous instrument movement trajectories). Finally, bilinear interpolation or transposed convolution is used to upsample the highest-resolution decoded features (1 / 4 layer) to the size of the input image. The loss calculation can use a hybrid loss function:
[0052] Among them, and are the cross-entropy loss and Dice loss respectively; is the cross-entropy loss weight, which is used to control the contribution degree of the cross-entropy loss to the total loss; It is the Dice loss weight, which is used to control the contribution degree of the Dice loss to the total loss.
[0053] S2: Use the training data set to train the dynamic segmentation network model; Specifically, the dynamic segmentation network model in this embodiment is a fully supervised model, and its training depends on the temporal segmentation mask with pixel-level annotation. The training data can be constructed by oneself or directly use existing public data sets, such as the EndoVis2018 data set (including 15 laparoscopic video sequences, resolution 1280×1024, 12 types of annotations), etc. The model training process focuses on spatio-temporal feature learning and surgical scene characteristic optimization, and is completed by combining traditional methods such as data preprocessing, hybrid loss function, and multi-stage training strategy. The training process will not be elaborated here.
[0054] S3: During the robot-assisted surgery process, use the trained dynamic segmentation network model to perform medical image segmentation and feature extraction based on the real-time input medical image video sequence, and transmit the segmentation mask output by the model to the surgical robot control system for surgical safety warning and navigation assistance.
[0055] Specifically, during the robot-assisted surgery (such as cholecystectomy), the endoscopic video stream is collected in real time and scaled to the model input size in real time, and continuous T frames (such as T = 5 frames) are cached as the input window. Use the trained dynamic segmentation network model to perform image segmentation and feature extraction in real time, and output the segmentation mask to the surgical robot control system for surgical safety warning, such as superimposing the segmentation result on the endoscopic image and highlighting the key structures (such as marking the cystic triangle area in red and the instrument contour in blue), and navigation assistance, such as inputting the segmentation mask into the safety evaluation module to calculate the distance between the instrument and the vulnerable tissue in real time and trigger a warning (such as an alarm when the distance < 2mm).
[0056] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments.
Claims
1. A method for robot-assisted medical image segmentation and feature extraction for precise surgery, characterized in that Including: Construct a dynamic segmentation network model with a bidirectional attention architecture. The dynamic segmentation network model includes an input preprocessing module, a Transformer encoding branch, a convolutional encoding branch, a multi-stage feature fusion unit, and a decoding output module. Among them, the input preprocessing module is used to extract features from video frames and generate a multi-scale feature pyramid. The Transformer encoding branch is used to perform temporal feature modeling based on the input feature pyramid and output temporally enhanced features. The convolutional encoding branch is used to enhance anatomical features and surgical instrument features through convolutional operations based on the input feature pyramid and output spatially enhanced features. The multi-stage feature fusion unit is used to perform multi-stage iterative fusion on the outputs of the Transformer encoding branch and the convolutional encoding branch and output the final fused features. The decoding output module is used to input the fused features into a Transformer decoder for decoding and generate pixel-level segmentation masks of anatomical structures and surgical instruments. Use a training dataset to train the dynamic segmentation network model. During the robot-assisted surgery process, use the trained dynamic segmentation network model to perform medical image segmentation and feature extraction based on the real-time input medical image video sequence, and transmit the segmentation mask output by the model to the surgical robot control system for surgical safety warning and navigation assistance.
2. The method according to claim 1, characterized in that The input preprocessing module includes: A ResNet-50 backbone network, which is used to extract features from the input continuous video frames and generate feature maps with different resolutions for each frame. The feature maps contain the spatial feature information of each frame. A temporal dimension expansion unit, which is used to splice the feature maps of multiple consecutive frames along the time dimension to form a spatio-temporal feature tensor. A 3D convolutional layer, which is used to gradually downsample the spatio-temporal feature tensor by applying 3D convolutional operations to generate a multi-scale feature pyramid. The multi-scale feature pyramid includes multiple feature layers with different resolutions, and each feature layer contains time dimension information.
3. The method according to claim 1, wherein The Transformer encoding branch includes a time query propagator, a multi-head attention layer, and a dimension expansion layer connected in sequence. The time query propagator includes: A feature projection unit, which is used to splice the continuous multi-frame features of a certain level in the feature pyramid along the time dimension, project the spliced features through a learnable key-value weight matrix, and generate a key-value matrix. A time query generation unit, which is used to perform convolutional operations on the spliced features to extract high-activation regions, use Top-K to select key content features to generate content queries, and generate position queries through forward propagation and backward propagation based on the position embeddings of the previous frame, middle frame, and next frame of the input. The multi-head attention layer is used to perform multi-head attention calculations based on the content query, the position query, and the key-value matrix and perform feature refinement. The dimension expansion layer is used to expand the channels of the refined features and output temporally consistent enhanced features.
4. The method according to claim 3, wherein The generation formula for the content query is: , Among them, is the query content vector, which is used to capture the key region features; is the feature map of the t-th frame; is the continuous T frame features concatenated along the time dimension; Conv represents 2D convolution with a kernel size of 3×3 and an output channel number of D ; represents selecting the top K high-response regions based on the activation values to generate a sparse content query vector.
5. The method according to claim 4, wherein The generation formula for the position query is: , Among them, The t position query vector of the th frame is used to encode spatio-temporal position information; T is the continuous frame features concatenated along the time dimension; is a learnable weight matrix; Conv represents 2D convolution; D denotes a fully connected layer that keeps the dimension of the position query as ; and t+1 respectively represent the position query vectors of the t-1 th frame and the t-1 th frame.
6. The method according to claim 5, wherein The multi-head attention calculation formula is: , in, is the query content vector; No. t The position query vector of the frame; K is the bond matrix generated by feature projection; is the transposed matrix of K; V is the value matrix generated by feature projection; represents the time aggregation of location query; ⊙ represents element-by-element multiplication; is the scaling factor used to stabilize the gradient; It represents normalization along the key dimension, and the output is the attention weight matrix; Attention(Q,K,V) represents multi-head attention calculation, and outputs the attention weighted feature matrix, where each row corresponds to the enhanced feature of a key area.
7. The method according to claim 1, wherein The convolutional coding branch includes an aggregated asymmetric feature pyramid module and a feature compression layer; The aggregated asymmetric feature pyramid module includes: A spatio-temporal aggregation layer for performing a convolutional operation on consecutive multi-frame features at a certain level in the feature pyramid to obtain spatio-temporal aggregation features; An anatomy-aware enhancement branch for performing symmetric convolutional operations on the spatio-temporal aggregation feature map to enhance the anatomical structure features of irregular polygons and output multi-scale anatomical features; An instrument-aware enhancement branch for performing convolutional operations on the spatio-temporal aggregation feature map using an asymmetric strip convolution pair to enhance the regular strip-shaped instrument features and output multi-scale instrument features; An attention-weighted fusion layer for performing attention fusion calculations on the enhanced anatomical features, instrument features, and original aggregated features to generate spatially enhanced features; The feature compression layer is used to compress the spatially enhanced features to the same channel dimension as the Transformer coding branch.
8. The method according to claim 7, wherein The calculation formula of the spatio-temporal aggregation layer is: , Among them, is the spatio-temporal aggregation feature; is the continuous T frame features concatenated along the time dimension; Conv 5×5×T represents a convolutional kernel with a size of 5×5×T; The calculation formula of the anatomy-aware enhancement branch is: ; The calculation formula of the instrument-aware enhancement branch is: , Among them, k m is the spatial size of the symmetric convolution kernel; t m is the depth of the time dimension.
9. The method according to claim 8, wherein The attention fusion calculation formula is: , Among them, ⨂ represents element-wise multiplication; Conv 1×1×T represents spatio-temporal compressive convolution to generate an attention map; is the spatially enhanced feature for attention-weighted fusion.
10. The method according to claim 1, characterized in that, The multi-stage feature fusion unit adds the outputs of the Transformer coding branch and the convolutional coding branch element by element, and through multiple residual iterations with parameter sharing, generates the final fusion feature.
Citation Information
Patent Citations
Image segmentation method based on double attention fusion
CN116012581A
Medical image segmentation method and system based on double-branch embedded attention mechanism
CN116309650A
Lightweight medical image segmentation network, method and equipment based on multi-path pyramid
CN117274607A
Weak supervision image segmentation method based on deep learning
CN117911429A
Medical image segmentation method based on dynamic multi-scale conditional diffusion model
CN118247509A
Cited By
Valve side sleeve partial discharge identification method and system based on artificial intelligence
CN120408283A
SMIL-based cholecystectomy CVS evaluation system, method and equipment
CN120894734A
Smil-based cholecystectomy cvs assessment system, method and apparatus
CN120894734B
Staged identification method and system based on surgical instrument feature fusion
CN122289620A
Regulation and control method and system of optical storage direct flexible system, computer equipment and medium
CN122456450A