Remote surgical instrument positioning and tracking method and system based on image recognition

By using convolutional neural networks and cross-frame temporal modeling structures, combined with attention mechanisms, the problems of unstable instrument recognition and trajectory interruption in remote surgery were solved, achieving high-precision and continuous instrument tracking, and improving the real-time performance and reliability of the remote surgery system.

CN120876538AInactive Publication Date: 2025-10-31JIANGSU YIMILU HEALTH TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510995717.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-31
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing remote surgical image recognition methods are prone to recognition errors or loss in scenarios with occlusion, overlap, or changes in lighting, and lack effective filtering correction mechanisms, resulting in frequent trajectory interruptions and failing to meet the requirements of high accuracy and real-time performance.

Method used

By employing a convolutional neural network and a cross-frame temporal modeling structure, combined with an attention mechanism, image features are extracted and the instrument trajectory is constructed. Trajectory interruptions caused by occlusion or frame loss are eliminated through filtering correction, thereby achieving continuous tracking.

Benefits of technology

It improves the accuracy and continuity of surgical instrument recognition, enhances the real-time performance and reliability of the system, reduces the impact of errors and occlusions on the trajectory, and improves the precision and response efficiency of remote control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876538A_ABST
    Figure CN120876538A_ABST
Patent Text Reader

Abstract

The invention discloses a remote surgical instrument positioning and tracking method and system based on image recognition. The method comprises the following steps that S1, image data are collected and preprocessed; s2, inputting the image data into a convolutional neural network, and extracting features; s3, inputting the frame-level feature map set into a cross-frame modeling structure, fusing space and time sequence features, and screening a target area in combination with an attention mechanism; s4, recoding the feature map sequence, calculating an instrument category probability and a confidence score, and outputting an identification calibration result; s5, center coordinates are extracted, a position vector sequence is constructed, and track association is carried out based on Euclidean distance; s6, filtering and correcting the trajectory, smoothing the trajectory by adopting a state prediction mechanism, and complementing an interrupt part; and S7, coding the instrument track as structured data, and transmitting the structured data to a remote terminal to realize position feedback. According to the invention, the accuracy of instrument identification and the continuity of tracking in a remote operation are improved, and high-precision real-time feedback of the position of the surgical instrument is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical devices and precision positioning technology, and in particular to a method and system for remote surgical instrument positioning and tracking based on image recognition. Background Technology

[0002] With the continuous development of telemedicine technology, remote surgical systems are gradually becoming an important component driving intelligent medical services. Especially in high-risk, high-precision surgical scenarios, the lead surgeon needs to use remote image information to control surgical instruments in real time. However, existing remote surgical image-assisted systems generally rely on manual observation or traditional visual tracking algorithms for instrument identification and position determination, resulting in problems such as low positioning accuracy, unstable identification, and frequent trajectory interruptions. These systems struggle to meet the high requirements of real-time image processing and spatial continuity in remote surgery.

[0003] Currently, most common image recognition methods employ static image feature extraction techniques, performing instrument identification on a separate basis for each image frame. This fails to fully exploit the temporal relationships between consecutive image frames, leading to frequent identification errors or loss of instruments in scenarios involving occlusion, overlap, or changes in lighting. Furthermore, some methods only fit instrument trajectories based on the positional differences of single frames, neglecting the continuity and inertia of surgical procedures. This can easily cause abrupt changes, drift, or discontinuities in the trajectory, severely impacting the accuracy of remote physicians' assessment of the surgical status.

[0004] Furthermore, existing instrument tracking solutions lack effective filtering and correction mechanisms and state prediction capabilities when faced with real-world environmental interference such as video frame rate fluctuations, image signal compression, and network transmission delays. They are unable to intelligently complete trajectory interruptions, leading to potential problems such as delayed operation response and broken path feedback in remote surgical control.

[0005] Therefore, how to provide a remote surgical instrument positioning and tracking method and system based on image recognition is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose a remote surgical instrument positioning and tracking method and system based on image recognition. This invention makes full use of convolutional neural networks, cross-frame temporal modeling structures and attention mechanisms, and describes in detail the entire process from image feature extraction, instrument identification and calibration, trajectory construction to filtering correction. It has the advantages of high positioning accuracy, strong tracking continuity and excellent anti-occlusion ability.

[0007] The remote surgical instrument positioning and tracking method based on image recognition according to an embodiment of the present invention includes the following steps:

[0008] S1. Acquire image data from remote surgical scenarios and perform preprocessing;

[0009] S2. Input the preprocessed image data into the convolutional neural network to extract the spatial texture features, edge structure features and color distribution features contained in each frame of the image, and output a set of frame-level feature maps.

[0010] S3. Input the set of frame-level feature maps into the cross-frame temporal modeling structure to construct a feature map sequence that integrates spatial features and temporal relationships, and combine the attention mechanism to filter target regions with high matching degree;

[0011] S4. Perform feature recoding on the feature map sequence, calculate the device category probability and confidence score for each target region, and output the device identification and calibration results with category labels and confidence scores.

[0012] S5. Extract the center coordinate information based on the instrument identification and calibration results, construct a sequence of instrument position vectors for consecutive frames, and perform trajectory association based on Euclidean distance to form a set of instrument trajectories;

[0013] S6. Perform filtering correction on the instrument trajectory set, use state prediction mechanism to smooth the trajectory curve, and eliminate trajectory interruptions caused by occlusion or frame loss to form a continuous instrument tracking trajectory.

[0014] S7. Encode the instrument tracking trajectory into structured position data and transmit it to the remote control terminal to achieve real-time feedback of the surgical instrument position during remote operation.

[0015] Optionally, the image data includes a sequence of consecutive frame color images and depth images acquired from a multi-angle camera device.

[0016] Optionally, the preprocessing includes size normalization, resolution unification, noise filtering, and edge enhancement.

[0017] Optionally, S2 specifically includes:

[0018] S21. Input the preprocessed consecutive frame images into a multi-branch convolutional neural network. The first branch extracts local texture patterns and grayscale change information, the second branch extracts multi-directional edge contour maps, and the third branch extracts color channel response distribution maps.

[0019] S22. In the first branch, a three-layer stacked convolutional layer and max pooling layer combination structure is set up. A small receptive field kernel is used to extract local texture of the image in a fine-grained manner while keeping the spatial resolution unchanged, and a spatial texture feature map is generated.

[0020] S23. In the second branch, set up an edge filtering layer group with orthogonal direction receptive kernels to extract the gradient response of the image in the vertical, horizontal and diagonal directions respectively, and fuse the directional gradient magnitude map and the edge position activation map to generate an edge structure feature map.

[0021] S24. Perform color channel separation operation in the third branch, build a channel attention mechanism for each color channel, enhance the response value of the relevant color area of ​​the device, and generate a color distribution feature map after fusion.

[0022] S25. The three types of feature maps are fused into a frame-level composite feature map set by channel splicing.

[0023] Optionally, S3 specifically includes:

[0024] S31. Input the set of frame-level feature maps into the bidirectional temporal modeling structure in chronological order. Extract the spatial state transformation features between the current frame and the previous frame in the forward channel, extract the dynamic context information between the current frame and the subsequent frame in the backward channel, and output the cross-frame feature sequence.

[0025] S32. Construct a temporal association graph structure based on cross-frame feature sequences. Use each frame feature map as a node, calculate the feature similarity and spatial consistency between adjacent nodes, and assign connection weights based on inter-frame continuity to complete the structured modeling of the feature map sequence.

[0026] S33. An attention mechanism is embedded in the structured feature map sequence to construct a spatial attention submodule and a temporal attention submodule. The spatial attention submodule is used to enhance the response intensity of the target region within a single frame, and the temporal attention submodule is used to enhance the trajectory continuity of the target region between frames.

[0027] S34. Based on the attention scores of the two modules, select image regions with high matching degree in consecutive frames, determine the image regions as the target regions of the candidate devices, and complete region alignment on all frames.

[0028] Optionally, S4 specifically includes:

[0029] S41. Perform region convolution operations on the selected target regions in each frame with the feature map sequence of the corresponding frame in turn, and perform hierarchical normalization and residual compression on the semantic features of each region.

[0030] S42. A multi-head transformation structure is used to calculate the matching response of each target region, and the device category probability P is generated through a normalized exponential function. c (i);

[0031] S43. After obtaining the probability of the device category, construct the confidence scoring function for the target region. The calculation formula is as follows:

[0032]

[0033] Where S(i) represents the confidence score of the i-th target region, Ω iLet A(x,y) represent the spatial coordinate region of the i-th target region in the corresponding frame, and let M represent the response value at coordinates (x,y). i (t) represents the matching score of the i-th target region between the t-th frame and the reference frame, T represents the total number of consecutive frames used for temporal matching, log2(·) represents the logarithmic function, and N represents the total number of all instrument categories;

[0034] S44. Based on the device category probability and confidence score, determine the device category and corresponding confidence level for each target area, and output the device identification and calibration results.

[0035] Optionally, the formula for calculating the probability of the device category in the target area is:

[0036]

[0037] Among them, P c (i) represents the probability that the i-th target region belongs to the c-th type of device. This represents the recoded feature value of the i-th target region in the j-th channel. This represents the weighting coefficient of the c-th type of device on the j-th channel. b represents the weighting coefficient of the k-th type of device on the j-th channel. c b represents the offset parameter of the c-th type of device. k Let represent the offset parameter of the k-th type of device, d represent the number of feature channels, N represent the total number of all device categories, exp(·) represent the natural exponential function to the base e, and ln(·) represent the logarithmic function to the base e.

[0038] Optionally, S5 specifically includes:

[0039] S51. In each frame of the image, based on the instrument recognition and calibration results, extract the minimum bounding box of the calibration area, and use the geometric center of the bounding box as the two-dimensional coordinate position of the instrument in the current frame to construct a frame-level center coordinate set.

[0040] S52. Arrange the set of frame-level center coordinates in chronological order to form a sequence of instrument position vectors. Combine the instrument category probability and confidence score of the corresponding frame to perform position removal and interpolation repair on low-confidence samples, and retain high-confidence samples in continuous frames for trajectory construction.

[0041] S53. Perform pairing calculations on the center coordinate sequences of similar instruments in all frames, and use Euclidean distance to perform trajectory association on the coordinate points in adjacent frames. The calculation formula for the trajectory matching scoring function is as follows:

[0042]

[0043] Where D(i,j) represents the matching score between the instrument position points in the i-th frame and the j-th frame, x i y i Let x and x represent the horizontal and vertical coordinates of the instrument center in the i-th frame, respectively. j y j Let S1 and S2 be the horizontal and vertical coordinates of the instrument center in the j-th frame, respectively, and S(i) and S(j) be the confidence scores of the i-th and j-th target regions, respectively.

[0044] S54. Construct the shortest matching path between instruments in consecutive frames based on the matching score value, forming a set of trajectories for each instrument in multiple frames of images.

[0045] Optionally, S6 specifically includes:

[0046] S61. Perform segmented sliding processing on the position vector sequence of each trajectory in the instrument trajectory set, extract the horizontal and vertical coordinate sequences in each sliding segment, calculate the continuous coordinate difference sequence, and construct the velocity change curve and acceleration trend curve of the sliding segment.

[0047] S62. Perform position state prediction operation in the region where there are missing points in the trajectory sequence. Based on the velocity change curve and acceleration trend curve, generate the position prediction vector for the corresponding time point and insert the prediction result into the original trajectory sequence to eliminate trajectory interruption caused by occlusion or frame loss.

[0048] S63. Perform curve filtering correction on the completed trajectory sequence, construct a time-centered weighted position fusion window at each trajectory point, and use a Gaussian weighted function to smooth the coordinate values ​​of the trajectory points to obtain a continuous and smooth instrument tracking trajectory.

[0049] A remote surgical instrument positioning and tracking system based on image recognition according to an embodiment of the present invention includes:

[0050] The image processing module is used to acquire image data in remote surgical scenarios and perform preprocessing.

[0051] The feature extraction module is used to input the preprocessed image data into the convolutional neural network, extract the spatial texture features, edge structure features and color distribution features contained in each frame of the image, and output a set of frame-level feature maps;

[0052] The temporal modeling module is used to input the set of frame-level feature maps into the cross-frame temporal modeling structure, construct a feature map sequence that integrates spatial features and temporal relationships, and combine an attention mechanism to filter target regions with high matching degree;

[0053] The instrument recognition module is used to re-encode the feature map sequence, calculate the instrument category probability and confidence score for each target region, and output the instrument recognition calibration result with category label and confidence score;

[0054] The trajectory association module is used to extract the center coordinate information based on the instrument identification and calibration results, construct a sequence of instrument position vectors for consecutive frames, and perform trajectory association based on Euclidean distance to form a set of instrument trajectories;

[0055] The trajectory optimization module is used to perform filtering and correction on the instrument trajectory set, use a state prediction mechanism to smooth the trajectory curve, and eliminate trajectory interruptions caused by occlusion or frame loss to form a continuous instrument tracking trajectory.

[0056] The transmission feedback module is used to encode the instrument tracking trajectory into structured position data and transmit it to the remote control terminal to realize real-time feedback of the position of surgical instruments during remote operation.

[0057] The beneficial effects of this invention are:

[0058] First, this invention constructs a multi-branch convolutional neural network to jointly extract spatial texture features, edge structure features, and color distribution features from surgical images. This enables accurate identification of various surgical instruments in complex surgical environments, effectively improving the depth and robustness of image understanding and laying a high-quality feature foundation for subsequent tracking tasks.

[0059] Secondly, this invention introduces a cross-frame temporal modeling structure and attention mechanism, which integrates dynamic contextual information between consecutive image frames to achieve temporal association of surgical instruments and enhancement of target regions. This significantly improves the recognition stability and tracking continuity in scenarios with occlusion, overlap, and image disturbance, and solves the problem of frequent target loss in complex operation scenarios using traditional methods.

[0060] Finally, this invention combines state prediction and filtering correction strategies to intelligently complete interruptions in the instrument trajectory caused by frame loss or occlusion, and eliminates trajectory jumps and jitters through smoothing operations, forming a stable and continuous tracking path. Simultaneously, this invention encodes the instrument position trajectory into structured data for transmission to the remote terminal, significantly improving the accuracy feedback and response efficiency of remote control, and enhancing the system's real-time performance and reliability. Attached Figure Description

[0061] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0062] Figure 1 This is a flowchart of the remote surgical instrument positioning and tracking method based on image recognition proposed in this invention;

[0063] Figure 2 This is a flowchart of the multi-branch feature extraction and temporal fusion structure of the remote surgical instrument positioning and tracking method based on image recognition proposed in this invention;

[0064] Figure 3 This is a block diagram of the remote surgical instrument positioning and tracking system based on image recognition proposed in this invention. Detailed Implementation

[0065] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0066] refer to Figure 1-2 A remote surgical instrument positioning and tracking method based on image recognition includes the following steps:

[0067] S1. Acquire image data from remote surgical scenarios and perform preprocessing;

[0068] S2. Input the preprocessed image data into the convolutional neural network to extract the spatial texture features, edge structure features and color distribution features contained in each frame of the image, and output a set of frame-level feature maps.

[0069] S3. Input the set of frame-level feature maps into the cross-frame temporal modeling structure to construct a feature map sequence that integrates spatial features and temporal relationships, and combine the attention mechanism to filter target regions with high matching degree;

[0070] S4. Perform feature recoding on the feature map sequence, calculate the device category probability and confidence score for each target region, and output the device identification and calibration results with category labels and confidence scores.

[0071] S5. Extract the center coordinate information based on the instrument identification and calibration results, construct a sequence of instrument position vectors for consecutive frames, and perform trajectory association based on Euclidean distance to form a set of instrument trajectories;

[0072] S6. Perform filtering correction on the instrument trajectory set, use state prediction mechanism to smooth the trajectory curve, and eliminate trajectory interruptions caused by occlusion or frame loss to form a continuous instrument tracking trajectory.

[0073] S7. Encode the instrument tracking trajectory into structured position data and transmit it to the remote control terminal to achieve real-time feedback of the surgical instrument position during remote operation.

[0074] This invention achieves automatic identification and continuous tracking of surgical instruments in remote surgical scenarios by constructing a complete process method that includes image acquisition, feature extraction, temporal modeling, recognition and calibration, trajectory generation and feedback output, thereby improving the accuracy and response efficiency of remote operations.

[0075] In this embodiment, the image data includes a sequence of consecutive frame color images and depth images acquired from a multi-angle camera device.

[0076] This invention introduces a multi-angle camera device to acquire color and depth image sequences, which enhances the spatial information representation capability of instrument images and improves the accuracy and robustness of target recognition and positioning.

[0077] In this embodiment, the preprocessing includes size normalization, resolution unification, noise filtering, and edge enhancement.

[0078] This invention performs size normalization, resolution unification, noise filtering, and edge enhancement during the image preprocessing stage, which effectively improves the image quality and consistency of feature representation during the feature extraction stage, and enhances the stability of network training and recognition.

[0079] In this embodiment, S2 specifically includes:

[0080] S21. Input the preprocessed consecutive frame images into a multi-branch convolutional neural network. The first branch extracts local texture patterns and grayscale change information, the second branch extracts multi-directional edge contour maps, and the third branch extracts color channel response distribution maps.

[0081] S22. In the first branch, a three-layer stacked convolutional layer and max pooling layer combination structure is set up. A small receptive field kernel is used to extract local texture of the image in a fine-grained manner while keeping the spatial resolution unchanged, and a spatial texture feature map is generated.

[0082] S23. In the second branch, set up an edge filtering layer group with orthogonal direction receptive kernels to extract the gradient response of the image in the vertical, horizontal and diagonal directions respectively, and fuse the directional gradient magnitude map and the edge position activation map to generate an edge structure feature map.

[0083] S24. Perform color channel separation operation in the third branch, build a channel attention mechanism for each color channel, enhance the response value of the relevant color area of ​​the device, and generate a color distribution feature map after fusion.

[0084] S25. The three types of feature maps are fused into a frame-level composite feature map set by channel splicing.

[0085] This invention employs a multi-branch convolutional neural network structure to extract and fuse texture, edge, and color features in images in parallel, thereby enhancing the expressive dimension and discriminative ability of the feature maps and providing stronger feature support for subsequent target screening and recognition.

[0086] In this embodiment, S3 specifically includes:

[0087] S31. Input the set of frame-level feature maps into the bidirectional temporal modeling structure in chronological order. Extract the spatial state transformation features between the current frame and the previous frame in the forward channel, extract the dynamic context information between the current frame and the subsequent frame in the backward channel, and output the cross-frame feature sequence.

[0088] S32. Construct a temporal association graph structure based on cross-frame feature sequences. Use each frame feature map as a node, calculate the feature similarity and spatial consistency between adjacent nodes, and assign connection weights based on inter-frame continuity to complete the structured modeling of the feature map sequence.

[0089] S33. An attention mechanism is embedded in the structured feature map sequence to construct a spatial attention submodule and a temporal attention submodule. The spatial attention submodule is used to enhance the response intensity of the target region within a single frame, and the temporal attention submodule is used to enhance the trajectory continuity of the target region between frames.

[0090] S34. Based on the attention scores of the two modules, select image regions with high matching degree in consecutive frames, determine the image regions as the target regions of the candidate devices, and complete region alignment on all frames.

[0091] This invention utilizes a bidirectional temporal modeling structure and a space-time attention mechanism to fuse inter-frame dynamic features, constructing a continuous feature representation of the target region, thereby improving the recognition stability and target region selection accuracy in occluded and blurred scenes.

[0092] In this embodiment, S4 specifically includes:

[0093] S41. Perform region convolution operations on the selected target regions in each frame with the feature map sequence of the corresponding frame in turn, and perform hierarchical normalization and residual compression on the semantic features of each region.

[0094] S42. A multi-head transformation structure is used to calculate the matching response of each target region, and the device category probability P is generated through a normalized exponential function. c (i);

[0095] S43. After obtaining the probability of the device category, construct the confidence scoring function for the target region. The calculation formula is as follows:

[0096]

[0097] Where S(i) represents the confidence score of the i-th target region, Ω i Let A(x,y) represent the spatial coordinate region of the i-th target region in the corresponding frame, and let M represent the response value at coordinates (x,y). i (t) represents the matching score of the i-th target region between the t-th frame and the reference frame, T represents the total number of consecutive frames used for temporal matching, log2(·) represents the logarithmic function, and N represents the total number of all instrument categories;

[0098] S44. Based on the device category probability and confidence score, determine the device category and corresponding confidence level for each target area, and output the device identification and calibration results.

[0099] This invention achieves the identification and calibration of target instruments by combining feature recoding, classification probability calculation and confidence scoring mechanism, which effectively improves classification accuracy and category identification confidence, and reduces false identification and missed identification.

[0100] In this embodiment, the formula for calculating the probability of the device category in the target area is:

[0101]

[0102] Among them, P c (i) represents the probability that the i-th target region belongs to the c-th type of device. This represents the recoded feature value of the i-th target region in the j-th channel. This represents the weighting coefficient of the c-th type of device on the j-th channel. b represents the weighting coefficient of the k-th type of device on the j-th channel. c b represents the offset parameter of the c-th type of device. k Let represent the offset parameter of the k-th type of device, d represent the number of feature channels, N represent the total number of all device categories, exp(·) represent the natural exponential function to the base e, and ln(·) represent the logarithmic function to the base e.

[0103] This invention constructs a multi-channel weighted probability formula for device categories, which improves the ability to distinguish between different device categories and provides stronger support for accurate recommendation of category labels.

[0104] In this embodiment, S5 specifically includes:

[0105] S51. In each frame of the image, based on the instrument recognition and calibration results, extract the minimum bounding box of the calibration area, and use the geometric center of the bounding box as the two-dimensional coordinate position of the instrument in the current frame to construct a frame-level center coordinate set.

[0106] S52. Arrange the set of frame-level center coordinates in chronological order to form a sequence of instrument position vectors. Combine the instrument category probability and confidence score of the corresponding frame to perform position removal and interpolation repair on low-confidence samples, and retain high-confidence samples in continuous frames for trajectory construction.

[0107] S53. Perform pairing calculations on the center coordinate sequences of similar instruments in all frames, and use Euclidean distance to perform trajectory association on the coordinate points in adjacent frames. The calculation formula for the trajectory matching scoring function is as follows:

[0108]

[0109] Where D(i,j) represents the matching score between the instrument position points in the i-th frame and the j-th frame, x i y i Let x and x represent the horizontal and vertical coordinates of the instrument center in the i-th frame, respectively. j y j Let S1 and S2 be the horizontal and vertical coordinates of the instrument center in the j-th frame, respectively, and S(i) and S(j) be the confidence scores of the i-th and j-th target regions, respectively.

[0110] S54. Construct the shortest matching path between instruments in consecutive frames based on the matching score value, forming a set of trajectories for each instrument in multiple frames of images.

[0111] This invention constructs a trajectory matching scoring function based on center point coordinates and instrument confidence, and adopts a trajectory scoring mechanism that integrates Euclidean distance and confidence weight, thereby improving the stability and trajectory continuity of cross-frame instrument matching.

[0112] In this embodiment, S6 specifically includes:

[0113] S61. Perform segmented sliding processing on the position vector sequence of each trajectory in the instrument trajectory set, extract the horizontal and vertical coordinate sequences in each sliding segment, calculate the continuous coordinate difference sequence, and construct the velocity change curve and acceleration trend curve of the sliding segment.

[0114] S62. Perform position state prediction operation in the region where there are missing points in the trajectory sequence. Based on the velocity change curve and acceleration trend curve, generate the position prediction vector for the corresponding time point and insert the prediction result into the original trajectory sequence to eliminate trajectory interruption caused by occlusion or frame loss.

[0115] S63. Perform curve filtering correction on the completed trajectory sequence, construct a time-centered weighted position fusion window at each trajectory point, and use a Gaussian weighted function to smooth the coordinate values ​​of the trajectory points to obtain a continuous and smooth instrument tracking trajectory.

[0116] This invention employs a sliding window and state prediction mechanism to intelligently complete missing trajectory segments, and introduces Gaussian smoothing filtering to handle trajectory jitter, effectively solving the trajectory interruption problem caused by occlusion and frame loss, and improving the consistency and accuracy of tracking results.

[0117] refer to Figure 3 A remote surgical instrument positioning and tracking system based on image recognition includes:

[0118] The image processing module is used to acquire image data in remote surgical scenarios and perform preprocessing.

[0119] The feature extraction module is used to input the preprocessed image data into the convolutional neural network, extract the spatial texture features, edge structure features and color distribution features contained in each frame of the image, and output a set of frame-level feature maps;

[0120] The temporal modeling module is used to input the set of frame-level feature maps into the cross-frame temporal modeling structure, construct a feature map sequence that integrates spatial features and temporal relationships, and combine an attention mechanism to filter target regions with high matching degree;

[0121] The instrument recognition module is used to re-encode the feature map sequence, calculate the instrument category probability and confidence score for each target region, and output the instrument recognition calibration result with category label and confidence score;

[0122] The trajectory association module is used to extract the center coordinate information based on the instrument identification and calibration results, construct a sequence of instrument position vectors for consecutive frames, and perform trajectory association based on Euclidean distance to form a set of instrument trajectories;

[0123] The trajectory optimization module is used to perform filtering and correction on the instrument trajectory set, use a state prediction mechanism to smooth the trajectory curve, and eliminate trajectory interruptions caused by occlusion or frame loss to form a continuous instrument tracking trajectory.

[0124] The transmission feedback module is used to encode the instrument tracking trajectory into structured position data and transmit it to the remote control terminal to realize real-time feedback of the position of surgical instruments during remote operation.

[0125] This invention constructs a modular remote surgical instrument tracking system that integrates image processing, feature extraction, modeling and analysis, trajectory construction, and feedback execution. It has the advantages of clear structure, fast response, and flexible deployment, thereby improving the intelligence level of remote surgical systems.

[0126] Example 1:

[0127] To verify the feasibility of this invention in practice, it was applied to a laparoscopic minimally invasive surgery scenario within a remote surgical assistance platform. In this scenario, the surgeon at the control end needs to determine instrument paths and execute remote operations based on surgical images. Due to the variety of instruments, the complexity of operation paths, and the presence of image occlusion, instrument crossing, and background interference in the surgical environment, traditional methods often encounter problems such as instrument confusion, positioning drift, or trajectory interruption during the identification and tracking stages, severely affecting the stability and accuracy of remote control. This invention simultaneously acquires and preprocesses color and depth image sequences, inputs them into a multi-branch convolutional neural network, extracts the spatial texture, edge structure, and color distribution features of the instrument targets, constructs a frame-level composite feature map, and introduces a bidirectional temporal modeling structure and a space-time attention mechanism to enhance the feature continuity between image frames.

[0128] In practical applications, the system is deployed at the edge processing end of a remote control platform and interconnected in real time with the remote surgical control console. During a series of laparoscopic surgeries, the system processed a total of 10,472 image frames, automatically identifying and tracking three types of instruments: grasping forceps, cutting forceps, and electrocautery hooks, each with different color textures and edge shapes. The system outputs the corresponding instrument center coordinates and category label in each frame and feeds the position data back to the control terminal in real time for path fitting. In the recognition phase, the feature recoding mechanism and confidence scoring model employed in this invention significantly improve target discrimination capabilities, achieving an instrument recognition accuracy of 93.8% even under strong occlusion conditions, a 12.4 percentage point improvement compared to the traditional single-frame convolutional recognition model. In the tracking phase, the system effectively avoids the frequency of instrument interruptions through state prediction and trajectory filtering correction, increasing the completeness of continuous frame tracking from 81.5% in the original system to 96.2%.

[0129] In terms of frame processing efficiency, after executing the algorithm of this invention at the edge, the image processing frame rate is stably maintained at over 28 frames per second, meeting the requirements of remote real-time control. Meanwhile, in comparative tests, the traditional method has an average trajectory reconstruction error of 4.5 pixels when dealing with image occlusion, while the average error of this invention is only 1.2 pixels, representing a reduction rate of 73.3%. Furthermore, even under an environment simulating a 30% reduction in network bandwidth, this invention can still maintain a remote feedback latency of less than 160 milliseconds, far superior to traditional cloud-based reconstruction methods.

[0130] This invention demonstrates excellent performance in multiple dimensions, including instrument recognition accuracy, trajectory continuity, error suppression, and system response delay, indicating its significant technical advantages in demanding remote surgical tasks.

[0131] Table 1. Comparison of performance test data of the present invention in remote laparoscopic surgical instrument tracking.

[0132]

[0133] This embodiment fully demonstrates the technical feasibility and performance advantages of the present invention in remote surgical instrument tracking tasks, and provides reliable technical support for precise control in complex surgical scenarios.

[0134] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A remote surgical instrument positioning and tracking method based on image recognition, characterized in that, Includes the following steps: S1. Acquire image data from remote surgical scenarios and perform preprocessing; S2. Input the preprocessed image data into the convolutional neural network to extract the spatial texture features, edge structure features and color distribution features contained in each frame of the image, and output a set of frame-level feature maps. S3. Input the set of frame-level feature maps into the cross-frame temporal modeling structure to construct a feature map sequence that integrates spatial features and temporal relationships, and combine the attention mechanism to filter target regions with high matching degree; S4. Perform feature recoding on the feature map sequence, calculate the device category probability and confidence score for each target region, and output the device identification and calibration results with category labels and confidence scores. S5. Extract the center coordinate information based on the instrument identification and calibration results, construct a sequence of instrument position vectors for consecutive frames, and perform trajectory association based on Euclidean distance to form a set of instrument trajectories; S6. Perform filtering correction on the instrument trajectory set, use state prediction mechanism to smooth the trajectory curve, and eliminate trajectory interruptions caused by occlusion or frame loss to form a continuous instrument tracking trajectory. S7. Encode the instrument tracking trajectory into structured position data and transmit it to the remote control terminal to achieve real-time feedback of the surgical instrument position during remote operation.

2. The remote surgical instrument positioning and tracking method based on image recognition according to claim 1, characterized in that, The image data includes a sequence of consecutive frame color images and depth images acquired from a multi-angle camera device.

3. The remote surgical instrument positioning and tracking method based on image recognition according to claim 1, characterized in that, The preprocessing includes size normalization, resolution unification, noise filtering, and edge enhancement.

4. The remote surgical instrument positioning and tracking method based on image recognition according to claim 1, characterized in that, S2 specifically includes: S21. Input the preprocessed consecutive frame images into a multi-branch convolutional neural network. The first branch extracts local texture patterns and grayscale change information, the second branch extracts multi-directional edge contour maps, and the third branch extracts color channel response distribution maps. S22. In the first branch, a three-layer stacked convolutional layer and max pooling layer combination structure is set up. A small receptive field kernel is used to extract local texture of the image in a fine-grained manner while keeping the spatial resolution unchanged, and a spatial texture feature map is generated. S23. In the second branch, set up an edge filtering layer group with orthogonal direction receptive kernels to extract the gradient response of the image in the vertical, horizontal and diagonal directions respectively, and fuse the directional gradient magnitude map and the edge position activation map to generate an edge structure feature map. S24. Perform color channel separation operation in the third branch, build a channel attention mechanism for each color channel, enhance the response value of the relevant color area of ​​the device, and generate a color distribution feature map after fusion. S25. The three types of feature maps are fused into a frame-level composite feature map set by channel splicing.

5. The remote surgical instrument positioning and tracking method based on image recognition according to claim 1, characterized in that, S3 specifically includes: S31. Input the set of frame-level feature maps into the bidirectional temporal modeling structure in chronological order. Extract the spatial state transformation features between the current frame and the previous frame in the forward channel, extract the dynamic context information between the current frame and the subsequent frame in the backward channel, and output the cross-frame feature sequence. S32. Construct a temporal association graph structure based on cross-frame feature sequences. Use each frame feature map as a node, calculate the feature similarity and spatial consistency between adjacent nodes, and assign connection weights based on inter-frame continuity to complete the structured modeling of the feature map sequence. S33. An attention mechanism is embedded in the structured feature map sequence to construct a spatial attention submodule and a temporal attention submodule. The spatial attention submodule is used to enhance the response intensity of the target region within a single frame, and the temporal attention submodule is used to enhance the trajectory continuity of the target region between frames. S34. Based on the attention scores of the two modules, select image regions with high matching degree in consecutive frames, determine the image regions as the target regions of the candidate devices, and complete region alignment on all frames.

6. The remote surgical instrument positioning and tracking method based on image recognition according to claim 1, characterized in that, S4 specifically includes: S41. Perform region convolution operations on the selected target regions in each frame with the feature map sequence of the corresponding frame in turn, and perform hierarchical normalization and residual compression on the semantic features of each region. S42. A multi-head transformation structure is used to calculate the matching response of each target region, and the device category probability P is generated through a normalized exponential function. c (i); S43. After obtaining the probability of the device category, construct the confidence scoring function for the target region. The calculation formula is as follows: Where S(i) represents the confidence score of the i-th target region, Ω i Let A(x,y) represent the spatial coordinate region of the i-th target region in the corresponding frame, and let M represent the response value at coordinates (x,y). i (t) represents the matching score of the i-th target region between the t-th frame and the reference frame, T represents the total number of consecutive frames used for temporal matching, log2() represents the logarithmic function, and N represents the total number of all instrument categories; S44. Based on the device category probability and confidence score, determine the device category and corresponding confidence level for each target area, and output the device identification and calibration results.

7. The remote surgical instrument positioning and tracking method based on image recognition according to claim 6, characterized in that, The formula for calculating the probability of the device category in the target area is: Among them, P c (i) represents the probability that the i-th target region belongs to the c-th type of device. This represents the recoded feature value of the i-th target region in the j-th channel. This represents the weighting coefficient of the c-th type of device on the j-th channel. b represents the weighting coefficient of the k-th type of device on the j-th channel. c b represents the offset parameter of the c-th type of device. k Let represent the offset parameter of the k-th type of device, d represent the number of feature channels, N represent the total number of all device categories, exp(·) represent the natural exponential function to the base e, and ln(·) represent the logarithmic function to the base e.

8. The remote surgical instrument positioning and tracking method based on image recognition according to claim 1, characterized in that, S5 specifically includes: S51. In each frame of the image, based on the instrument recognition and calibration results, extract the minimum bounding box of the calibration area, and use the geometric center of the bounding box as the two-dimensional coordinate position of the instrument in the current frame to construct a frame-level center coordinate set. S52. Arrange the set of frame-level center coordinates in chronological order to form a sequence of instrument position vectors. Combine the instrument category probability and confidence score of the corresponding frame to perform position removal and interpolation repair on low-confidence samples, and retain high-confidence samples in continuous frames for trajectory construction. S53. Perform pairing calculations on the center coordinate sequences of similar instruments in all frames, and use Euclidean distance to perform trajectory association on the coordinate points in adjacent frames. The calculation formula for the trajectory matching scoring function is as follows: Where D(i,j) represents the matching score between the instrument position points in the i-th frame and the j-th frame, x i y i Let x and x represent the horizontal and vertical coordinates of the instrument center in the i-th frame, respectively. j y j Let S1 and S2 be the horizontal and vertical coordinates of the instrument center in the j-th frame, respectively, and S(i) and S(j) be the confidence scores of the i-th and j-th target regions, respectively. S54. Construct the shortest matching path between instruments in consecutive frames based on the matching score value, forming a set of trajectories for each instrument in multiple frames of images.

9. The remote surgical instrument positioning and tracking method based on image recognition according to claim 1, characterized in that, S6 specifically includes: S61. Perform segmented sliding processing on the position vector sequence of each trajectory in the instrument trajectory set, extract the horizontal and vertical coordinate sequences in each sliding segment, calculate the continuous coordinate difference sequence, and construct the velocity change curve and acceleration trend curve of the sliding segment. S62. Perform position state prediction operation in the region where there are missing points in the trajectory sequence. Based on the velocity change curve and acceleration trend curve, generate the position prediction vector for the corresponding time point and insert the prediction result into the original trajectory sequence to eliminate trajectory interruption caused by occlusion or frame loss. S63. Perform curve filtering correction on the completed trajectory sequence, construct a time-centered weighted position fusion window at each trajectory point, and use a Gaussian weighted function to smooth the coordinate values ​​of the trajectory points to obtain a continuous and smooth instrument tracking trajectory.

10. A remote surgical instrument positioning and tracking system based on image recognition, comprising the remote surgical instrument positioning and tracking method based on image recognition as described in any one of claims 1 to 9, characterized in that, include: The image processing module is used to acquire image data in remote surgical scenarios and perform preprocessing. The feature extraction module is used to input the preprocessed image data into the convolutional neural network, extract the spatial texture features, edge structure features and color distribution features contained in each frame of the image, and output a set of frame-level feature maps; The temporal modeling module is used to input the set of frame-level feature maps into the cross-frame temporal modeling structure, construct a feature map sequence that integrates spatial features and temporal relationships, and combine an attention mechanism to filter target regions with high matching degree; The instrument recognition module is used to re-encode the feature map sequence, calculate the instrument category probability and confidence score for each target region, and output the instrument recognition calibration result with category label and confidence score; The trajectory association module is used to extract the center coordinate information based on the instrument identification and calibration results, construct a sequence of instrument position vectors for consecutive frames, and perform trajectory association based on Euclidean distance to form a set of instrument trajectories; The trajectory optimization module is used to perform filtering and correction on the instrument trajectory set, use a state prediction mechanism to smooth the trajectory curve, and eliminate trajectory interruptions caused by occlusion or frame loss to form a continuous instrument tracking trajectory. The transmission feedback module is used to encode the instrument tracking trajectory into structured position data and transmit it to the remote control terminal to realize real-time feedback of the position of surgical instruments during remote operation.