Method and system for identifying change in position information of body surface positioning member for surgery

By installing a micro camera at the end of the puncture surgery robot and using the training model to track the position of the position of the position in real time, the real-time monitoring of the position of the position in the puncture surgery is solved, and the safety and controllability of the surgery are improved.

CN119700257BActive Publication Date: 2025-07-22SUZHOU PNEMED TECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510214501.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-22
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

In the prior art, puncture surgical robots cannot monitor the position changes of the body surface positioning parts in real time during needle insertion, resulting in insufficient safety and controllability of the surgery, especially when the patient is breathing or abnormal movements, errors are easily generated.

Method used

A micro camera is embedded at the end of the robotic arm of the puncture surgery robot. Through the trained sample training model, the X, Y, and Z position information of the positioning member is tracked in real time, forming a real-time data flow, and the displacement curve is presented in the display screen to realize real-time observation and tracking of the positioning member.

Benefits of technology

Real-time observable and controllable process of the needle insertion process is achieved, reducing errors caused by breathing or abnormal movements, and improving the safety and reliability of the operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119700257B_ABST
    Figure CN119700257B_ABST
Patent Text Reader

Abstract

A method for identifying changes in the position information of a body surface positioning part for surgery includes the following steps: A micro camera is embedded at the end of the needle insertion device at the end of the robotic arm of the puncture surgery robot. The micro camera continuously captures the positioning part located on the patient's body surface during the time period of the puncture needle insertion process, obtains all the images to be processed, and after initial processing, outputs them as the first images to be processed in sequence; All the first images to be processed are sequentially sent to a pre-trained first sample training model, and the position information of the center point of the positioning part in the X and Y directions in the image plane of all the first images to be processed is inferred to form a second image to be processed; All the second images to be processed are sequentially sent to a second sample training model, and the longitudinal Z position information of the center point of the positioning part is inferred; The X, Y, and Z position information of the center point of the positioning part in all the second images to be processed is sequentially output to form a real-time output data stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of puncture surgical robots, and specifically, to a method and system for identifying changes in the position information of a body surface positioning member for surgery. Background Art

[0002] With the rapid development of medical technology, the working process of a puncture surgical robot is to plan and determine the lesion position, puncture target point, body surface needle insertion point, puncture path, etc. according to the patient's CT image, combined with the doctor's diagnosis and the auxiliary diagnosis and treatment functions provided by the surgical robot, and then send data and instructions to the needle insertion device installed at the end of the robotic arm of the puncture surgical robot. The needle insertion device drives the puncture needle to accurately move to the specified position, and then enters the puncture target point at the lesion position from the body surface needle insertion point.

[0003] Currently, the observation of the body surface surgical needle insertion point (positioning member) of a patient is mostly achieved through a camera (optical tracking system or infrared camera) installed at a distance, that is, the needle insertion point is found by positioning the center point of the positioning member with the help of the camera. However, during the needle insertion process of the puncture surgery after the positioning is completed, it is impossible to further position the positioning member (the so-called microscopic observation) using a camera (optical tracking system or infrared camera) installed at a distance. That is, the needle insertion process of the puncture needle is in a state of no monitoring. If the position of the positioning member is not further tracked and positioned, although the needle insertion process only takes 2 - 4 seconds, if the patient moves and / or has abnormal movements during these 2 - 4 seconds, it will have a significant impact on the puncture surgery, and there is no monitoring at all during the instant of needle insertion, which belongs to the surgical blind area, cannot ensure the safety of the surgery, and no records are made for the surgical process either.

[0004] Movement refers to the movement of the organs due to breathing during the deep breathing control or deep breath - holding control process, which drives the positioning member located on the patient's body surface to move.

[0005] Abnormal movement refers to any unexpected activity during the surgery, including the patient or other interfering factors, which causes unnecessary displacement and poses a surgical risk.

[0006] During the instant process, due to the extremely fast puncture speed, the naked eye needs to slow down by 10 times to observe, and it is impossible to achieve real - time tracking, positioning, and observation using a camera installed at a distance. Summary of the Invention

[0007] The purpose of the present invention is to provide a method and system for identifying changes in the position information of a body surface positioning member for surgery. By using the identification method and system of the present invention, it is possible to achieve real - time tracking, positioning, and observation of the positioning member during the puncture process, so that the puncture needle is always in a state of observable / controllable / traceable during the needle insertion process.

[0008] The method for identifying the change of the position information of the surgical body surface positioning piece in the present invention includes the following steps:

[0009] S1: A micro camera is embedded at the end of the needle insertion device at the end of the robotic arm of the puncture surgical robot, and at the same time, the puncture surgical robot system is initialized and the environment is set;

[0010] S2: The micro camera continuously takes pictures of the positioning piece located on the patient's body surface during the time period of the needle insertion process of the puncture needle. The positioning piece is used for positioning during the puncture operation of the puncture robot. All to-be-processed images are obtained, and all the obtained to-be-processed images are initialized and output as the first to-be-processed images in sequence;

[0011] S3: All the first to-be-processed images are sequentially sent to the pre-trained first sample training model, and the position information of the center point of the positioning piece in the X and Y directions in the image plane of all the first to-be-processed images is deduced according to the first sample training model, and a second to-be-processed image including the X and Y position information is formed;

[0012] S4: All the second to-be-processed images are sequentially sent to the second sample training model. According to the depth estimation method in the second sample training model, the longitudinal depth value of all the second to-be-processed images from the micro camera is first deduced, and further the longitudinal Z position information of the center point of the positioning piece is deduced;

[0013] S5: The X, Y, and Z position information of the center point of the positioning piece in all the second to-be-processed images is output in sequence to form a real-time output data stream, and further made into a displacement curve that can be presented on the display screen.

[0014] Preferably, after step S1, a step of judging whether the video transmission stream is normal is added. If it is not normal, the program is closed and the puncture surgical robot exits the operation; if it is normal, the following steps are executed.

[0015] Preferably, in step S2, the first frame of the to-be-processed image is first converted into a first array with three channels, that is, the digital processing of the image. The first array with three channels contains the content information of each pixel point, color, and size of the first frame of the to-be-processed image. At the same time, the coordinate values of the three channels in the first array are defined, and then the subsequent second frame of the to-be-processed image is converted into a second array with three channels according to the coordinate values of the three channels in the first array. On the basis of the second array, the third frame of the to-be-processed image is converted into a third array with three channels. In this way, all the obtained to-be-processed images are initialized in sequence.

[0016] Preferably, all the first to-be-processed images are sequentially and repeatedly executed in the first sample training model according to the following steps:

[0017] S3.1) Convert the first image to be processed into a grayscale image, and apply Gaussian filtering for smoothing and noise reduction;

[0018] S3.2) Set the dynamic threshold to 125 - 255, and perform dynamic threshold processing on the first image to be processed after step S3.1);

[0019] S3.3) Use the canny edge detection algorithm to scan the first image to be processed after step S3.2), enhance the shape, features and contour edges of the positioning parts in the first image to be processed, and decompose the image into smaller segments or regions according to the image features of the positioning parts to form multiple first target regions;

[0020] S3.4) Connect the contours of multiple first target regions to form multiple separate regions;

[0021] S3.5) Screen out the second target region from multiple separate regions according to the area size of the positioning part, draw the contour of the second target region, and fit the contour of the second target region according to the preset ellipse formula to form multiple ellipse frames;

[0022] S3.6) Select the target ellipse frame according to the aspect ratio of the long and short axes of the ellipse, and obtain the coordinate values of the center point of the target ellipse frame, which are the X and Y position information of the center point of the positioning part in this plane.

[0023] Preferably, in step S3.6), select an ellipse with an aspect ratio of the long and short axes between 0.8 - 1.2.

[0024] Preferably, the training process of the second sample training model is as follows:

[0025] Obtain training data, collect a large number of unlabeled monocular images from a public large-scale dataset, called the pseudo-label set, labeled as Du = {ui} N i=1; at the same time, collect a large number of labeled images from the public large-scale dataset, called the labeled set, labeled as Dl = {(xi,di)}M i=1;

[0026] Construct a monocular depth recognition network, which includes two groups of encoders, the first and the second, and a group of decoders;

[0027] Use the constructed monocular depth recognition network to train a powerful monocular depth estimation model S1, specifically as follows:

[0028] 3.1) Training of the initial depth estimation model

[0029] Use the labeled images in the labeled set D1 to pre-train an initial depth estimation model T;

[0030] 3.2) Generate pseudo-labels

[0031] Use the initial depth estimation model T to predict the unlabeled images in the pseudo-labeled set Du, generating pseudo-depth labels d * i ;

[0032] 3.3) Joint training

[0033] Combine the labeled images in the labeled set D1 and the unlabeled images in the pseudo-labeled set Du to form a joint dataset D = Dl ∪ Du;

[0034] 3.4) Training of the depth estimation model

[0035] Use the joint dataset D to train the initial depth estimation model T to obtain a depth estimation model S;

[0036] 3.5) Feature alignment loss

[0037] Use a 3D acquisition device to collect the grayscale image and depth map of the positioning part, and use the pre-trained DINOv2 model to extract the features of the positioning part from them. Align the features extracted by the depth estimation model S with the features extracted in the DINOv2 encoder according to the following formula;

[0038] Where:

[0039] cos(f i , f’ i ) measures the cosine similarity between two feature vectors;

[0040] f i is the feature extracted by the depth estimation model S;

[0041] f’ i is the feature extracted from the DINOv2 model;

[0042] 3.6) Through multiple rounds of training and optimization in steps 3.1 - 3.5), a powerful monocular depth estimation model S1, that is, the second sample training model, is obtained.

[0043] The system for identifying the change of the position information of the surgical body surface positioning part in the present invention includes a preprocessing module, a startup module, an identification module, a depth inference module, and an output module,

[0044] The preprocessing module is composed of a package deployment environment warm-up module and a neural network training module, and is used for the initialization and environment configuration of the puncture surgical robot system;

[0045] The startup module consists of a strong logic judgment module and a data flow monitoring module, and is used to judge whether the video stream captured by the micro camera can be transmitted normally. If the transmission stream cannot be detected during the judgment process, the program will be closed and exited.

[0046] The recognition module is the first sample training module, and is used to recognize the X and Y position information of the center point of the positioning part in the plane of each frame of image transmitted by the startup module.

[0047] The depth inference module is the second sample training module, and applies zero-shot generalization depth estimation model inference, and is used to recognize the Z position information between the center point of the positioning part in each frame of image transmitted by the recognition module and the micro camera longitudinally.

[0048] The output module is used to synchronously output the X and Y position information obtained by the first sample training module and the Z position information obtained by the second sample training module to the display screen of the puncture surgical robot system, forming a real-time output data stream.

[0049] In the field of thoracoabdominal puncture surgery navigation, due to the large volume and the undulation with the patient's posture and breathing of soft parts such as the chest and abdomen, the method in the present invention can flexibly / without occlusion track the puncture point of the positioning part for positioning in real time. By using a high-precision and high-real-time positioning and tracking system, the movement error of the in-vivo lesion organs caused by observing the breathing undulation by fitting the body surface can be reduced. Description of the Drawings

[0050] Figure 1 is a flowchart of the method for identifying the change of the position information of the surgical surface positioning part in the present invention.

[0051] Figure 2 is a schematic diagram of the method for identifying the change of the position information of the surgical surface positioning part in the present invention.

[0052] Figure 3 is a structural block diagram of the system for identifying the change of the position information of the surgical surface positioning part in the present invention.

[0053] Figure 4 is a working principle diagram of the method for identifying the change of the position information of the surgical surface positioning part in the present invention.

[0054] Figure 5 is a three-dimensional structural schematic diagram of the positioning part in the present invention. Detailed Embodiments

[0055] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are merely exemplary and are not intended to limit the protection scope of the present invention. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention. At the same time, regarding relational terms such as "first" and "second", there is no substantial sequential relationship, and they are only used to distinguish different components.

[0056] As Figure 4 and Figure 5 shown, the positioning member 1 in the present invention is made of aluminum oxide ceramic material, and its color is white throughout. Its outer diameter is 30 mm and its thickness is 5 mm. A central circle 2 is provided in the middle, the diameter of the central circle 2 is 10 mm, and an inverted conical concave arc surface is formed between the inner wall surface of the central circle 2 and the outer circumference of the positioning member 1.

[0057] The method for identifying the change of the position information of the body surface positioning member for surgery in the present invention first embeds a micro camera at the end of the needle insertion device installed at the end of the robotic arm of the puncture surgery robot. The shooting direction of the micro camera is parallel to the installation direction of the puncture needle, and the installation position of the micro camera is within a distance of 10 - 30 mm from the vertical position of the puncture needle after installation, so that the movement process of the puncture needle is completely within the shooting range of the micro camera. Preferably, a micro camera with a focal length of 5 - 50 mm, 30 fps (Frames Per Second), and a diameter of only 1.8 mm is embedded and installed at the end of the needle insertion device. The head of this micro camera is equipped with 4 LED light beads for illuminating the local environment during the surgery. In this embodiment, the micro camera is installed at the end of the robotic arm and starts working when it moves to a position 10 - 20 mm away from the patient's body surface with the robotic arm. It has the advantages of clear observation, being able to change with the surgical pose without being blocked, etc., and can clearly capture the entire process of whether the position of the positioning member changes during the puncture needle insertion process. Using the micro camera in this embodiment, multiple frames of 2D color images with a size of 400 * 400 and a clarity of 480p can be continuously captured. Then, with the method and system in the present invention, the spatial position information of the center point of the positioning member in the image, that is, the position coordinate information of X, Y, and Z, can be obtained. By continuously outputting the spatial position information of the positioning member in each frame of the image, close-range observation can be achieved, and the captured images can be transmitted in real time, with good real-time performance, which can effectively solve the problem of being unable to observe the entire process of puncture needle insertion in real time.

[0058] Based on the installation of the micro camera, it is necessary to initialize and set the environment of the puncture surgical robot system first. Specifically, necessary libraries are imported and the environment is configured, such as: V4L2, RTSP, go2rtc, opencv or capvideo. At the same time, the pre-trained first sample training model and second sample training model (also called the monocular depth estimation algorithm model or the monocular depth estimation neural network training model) are loaded.

[0059] When the puncture surgical robot system is started, the system for identifying the change of the position information of the body surface positioning part in the present invention is preheated and necessary package libraries are deployed to ensure the stability of the subsequent processing flow.

[0060] As Figure 3 shown, the system for identifying the change of the position information of the body surface positioning part in the present invention includes a preprocessing module, a start module, an identification module, a depth inference module and an output module.

[0061] The preprocessing module is composed of a package deployment environment preheating module and a neural network training module, and is used for the initialization and environment configuration of the puncture surgical robot system, preheating the environment for the system for identifying the change of the position information of the body surface positioning part in the present invention and deploying necessary package libraries to ensure the stability of the subsequent processing flow.

[0062] The start module is composed of a strong logic judgment module and a data flow monitoring module, and is used to judge whether the video stream captured by the micro camera can be transmitted normally. If the transmission stream of the video cannot be found during the judgment process, the program is closed and exited, so as to further judge whether the identification system in the present invention can run normally.

[0063] The identification module is the first sample training module, and is used to identify the X and Y position information of the center point of the positioning part in the plane where the image is located in each frame of image transmitted by the start module. The identification module can be a pre-trained deep learning module or an image processing module for processing the image to be processed transmitted by the start module.

[0064] The depth inference module is the second sample training module, and applies the zero-shot generalization depth estimation model for inference, and is used to identify the Z position information between the center point of the positioning part and the longitudinal direction of the micro camera in each frame of image transmitted by the identification module.

[0065] The output module is used to synchronously output the X and Y position information obtained by the first sample training module and the Z position information obtained by the second sample training module to the display screen of the puncture surgery robot system. The position change of the needle insertion point on the patient's body surface can be further output through the coordinate values of the center point X, Y, and Z of the positioning part in the image. With the help of the first frame image, the second frame image, the third frame... A plurality of consecutive frames form a real-time output data stream, which is collected by the puncture surgery robot system and made into a displacement curve according to the data. If the coordinate values of the center point of the positioning part do not change, a straight line will be presented on the display screen. If they change, a changing displacement curve will be presented according to the specific values of X, Y, and Z, which is convenient for medical staff to observe in real time, so as to observe whether the patient's position changes during the puncture surgery. Embodiment 1

[0066] As Figure 1 and Figure 2 shown, the method for identifying the change of the position information of the body surface positioning part for surgery in the present invention is executed according to the following steps.

[0067] S1: A micro camera is embedded at the end of the needle insertion device at the end of the robotic arm of the puncture surgery robot, and at the same time, the puncture surgery robot system is initialized and the environment is set; the end of the needle insertion device refers to the bottom end of the needle insertion device facing the patient. The installation position of the micro camera is 20 mm away from the vertical position of the puncture needle after installation, and the shooting direction of the micro camera is parallel to the puncture needle.

[0068] S2: The micro camera continuously shoots the positioning part 1 on the patient's body surface in real time within 2-4 seconds during the needle insertion process of the puncture needle. The positioning part is used to assist the puncture robot in positioning during the puncture surgery on the patient's chest and abdomen. All the obtained 2D images to be processed are acquired. Usually, 20-60 frames of images to be processed are obtained. In this embodiment, it is 30 frames, and all the images to be processed are sequentially transmitted to the recognition module in real time. The specific process of this step is as follows.

[0069] S2.1: First, the start module judges whether the video transmission stream is normal. If it is not normal, the program is closed and the puncture robot exits the surgery; if it is normal, the next step is executed.

[0070] S2.2: The video stream is obtained through the V4L2 interface from the RTSP source or by opening the video file through the opencv library. Each frame of the image in the video stream, that is, all the images to be processed, is obtained. All the images to be processed are initialized according to the transmission order. The coordinate values of the previous frame of the image are used during the initialization process, and then each frame of the image is read in sequence for repeated processing.

[0071] The specific initialization process is as follows: First, the first frame of the captured image (the first one transmitted) is converted into a first array with three channels, that is, the digital processing of the image. The first array with three channels contains content information such as each pixel point, color, and size of the first frame of the image. At the same time, the coordinate values of the three channels in the first array are defined by the system. Then, according to the coordinate values of the three channels in the first array, the subsequent second frame of the image is converted into a second array with three channels, and on the basis of the second array, the third frame of the image is converted into a third array with three channels. In this way, all the to-be-processed images obtained are initialized in sequence.

[0072] S2.3: Screen out all the first to-be-processed images containing the positioning part from all the to-be-processed images that have completed the initialization process. Since the positioning part 1 has a specific shape, color, and size, all the first to-be-processed images containing the positioning part can be screened out according to the characteristics of the positioning part in the image (such as white, circular, and the pixel area of the image accounting for 50,000 - 500,000). Then, all the first to-be-processed images are transmitted to the recognition module in sequence, which can accelerate the processing speed of the recognition module, reduce the processing time, and achieve the purpose of real-time observation and monitoring.

[0073] S3: Send all the first to-be-processed images to the recognition module in sequence. The pre-trained first sample training model extracts the contour data of the positioning part from each frame of the first to-be-processed images and infers the X and Y position information of the center point of the positioning part in the image plane in all the first to-be-processed images.

[0074] In this step S3, since the first sample training model is pre-trained by the neural network training module and integrates the deep learning algorithm, only by inputting all the first to-be-processed images in step S2, the coordinate values of X and Y of the positioning part in each frame of the image can be directly output, that is, the second to-be-processed images containing the X and Y position information are output.

[0075] The training process of the first sample training model is as follows.

[0076] a) The pre-recorded and captured video is intercepted frame by frame using a video tool, and each frame of the image is annotated to form an annotated data set. In this embodiment, it is required to capture the positioning part 1 placed on the patient's body surface at a position 10 - 20 mm on the patient's body surface. That is, the images captured from a distance are not recognized, and only the images captured from a short distance are recognized and processed. Therefore, the images captured from a short distance and recognized are selected for annotation sample processing and annotated as positive samples.

[0077] b) Send all positive samples into the neural network for training, and set the required parameters, such as: learning rate, optimizer, loss function, activation function, training batch size, training image size, etc. In this embodiment, it is default that 25 parameters can be set. The specific training process is easy to implement for those skilled in the computer field and will not be elaborated here.

[0078] c) After training, it is found that there are many misjudgments. Then return to the video tool machine to intercept the misjudged image frames and label them as negative samples.

[0079] d) Combine the negative samples and positive samples together to form a balanced data set. In this embodiment, a data set of positive and negative samples is constructed in a ratio of 1:1, and then the labeled data set is used for model training to obtain the first sample model training. This is easy to implement for those skilled in the computer field and will not be elaborated here.

[0080] S4: Sequentially send all the second images to be processed to the depth inference module. The pre-trained second sample training model first infers the depth distance of each frame of the second image to be processed from the micro camera according to the depth estimation method, and then further infers the longitudinal Z position information between the center point of the positioning part in each frame of the second image to be processed and the micro camera.

[0081] The training process of the second sample training model is as follows.

[0082] Obtain training data. Collect 62 million unlabeled monocular images from 6 - 8 public large-scale data sets, called the pseudo-labeled set, labeled as Du = {ui} N i=1; at the same time, collect 1.5 million labeled images from 6 - 8 public large-scale data sets, called the labeled set, labeled as Dl = {(xi,di)}M i=1.

[0083] Construct a monocular depth recognition network. The monocular depth recognition network includes two groups of encoders and decoders: the first and the second. Each group of encoders includes a local encoder and a global encoder. The local encoder in the first group of encoders is used to extract the local information of each frame of image in the labeled set Dl and generate a local feature map, and the global encoder in the first group of encoders is used to extract the global information of each frame of image in the labeled set Dl and generate a global feature map; the local encoder in the second group of encoders is used to extract the local information of each frame of image in the pseudo-labeled set Du and generate a local feature map, and the global encoder in the second group of encoders is used to extract the global information of each frame of image in the pseudo-labeled set Du and generate a global feature map; the decoder is used to sample the feature maps provided by the first group of encoders and the second group of encoders to obtain the recognized depth map.

[0084] Use the constructed monocular depth recognition network to train a powerful monocular depth estimation model S1, specifically as follows.

[0085] 3.1) Training of the initial depth estimation model

[0086] Pre-train an initial depth estimation model (also known as the teacher model T) using 1.5 million labeled images in the labeled set D1.

[0087] 3.2) Generating pseudo-labels

[0088] Use the initial depth estimation model T to predict the unlabeled images in the pseudo-labeled set Du to generate pseudo-depth labels (predicted depth values) d * i .

[0089] 3.3) Joint training

[0090] Combine the labeled images in the labeled set D1 and the unlabeled images in the pseudo-labeled set Du to form a joint dataset D = Dl ∪ Du. During this training process, the depth value is converted to the disparity space through a specific formula, specifically d = f / z, where d represents the disparity value, f represents the focal length of the camera, and z represents the depth value (i.e., the actual distance of the object from the camera). Convert the depth information to the disparity information on the image plane.

[0091] 3.4) Training of the depth estimation model

[0092] Use the joint dataset D to train the initial depth estimation model T to obtain a depth estimation model (referred to as the student model S).

[0093] During the training process of the depth estimation model S, use the affine invariant mean absolute error loss ρ to calculate the loss between the predicted depth value d ∗ i in the initial depth estimation model T and the true depth value d i and then obtain the depth value through the weighted average method formula (2)....

[0094] (1)

[0095] (2) where:

[0096] d ∗ i and d i are respectively the predicted depth value and the true depth value obtained in the initial depth estimation model T;

[0097] ρ is the affine invariant mean absolute error loss;

[0098] t(d) and s(d) are respectively the predicted depth value and the true depth value obtained in the depth estimation model S.

[0099]

[0100]

[0101] During the training process of the depth estimation model S, strong perturbations (such as strong color distortion and strong spatial distortion) are injected into the unlabeled images, forcing the depth estimation model to seek additional visual knowledge and learn robust representations.

[0102] Combined with the pseudo-depth label d * i and the strong perturbation, the unlabeled loss Lu of the unlabeled image is calculated according to the following formula.

[0103]

[0104] The two losses are added by the weighted average method:

[0105]

[0106] 3.5) Feature alignment loss

[0107] Use a 3D acquisition device to collect the grayscale image and depth image of the positioning part. Use the pre-trained DINOv2 model to extract the features of the positioning part, such as the specific shape, color, and size of the positioning part. Align the features extracted by the depth estimation model S with the features extracted in the DINOv2 encoder according to the following formula.

[0108] Where:

[0109] cos(f i , f' i ) measures the cosine similarity between two feature vectors.

[0110] f i is the feature extracted by the depth estimation model S.

[0111] f' i is the feature extracted from the DINOv2 model.

[0112] 3.6) Through multiple rounds of training and optimization in steps 3.1 - 3.5), a powerful monocular depth estimation model S1, that is, the second sample training model, is obtained. It effectively utilizes the comprehensive information of labeled and unlabeled images, improves the accuracy and generalization ability of depth estimation, and further enhances the robustness and semantic understanding ability of the model by introducing strong perturbation injection and feature alignment loss, so that the depth of the positioning part can be quickly and accurately estimated in various scenarios.

[0113] S5: Synchronously output the X and Y position information of the center point of the positioning member in each frame of image within the plane of the image and the Z position information in the longitudinal direction between the center point of the positioning member and the micro camera in steps S3 and S4, and display them on the display screen of the puncture robot control system. Embodiment 2

[0114] This Embodiment 2 is basically the same as Embodiment 1, the difference lies in the processing method of step S3, but both can quickly and accurately output the X and Y position information of the center point of the positioning member in the plane of the image in all the first images to be processed, and the processing process for each frame of the first image to be processed is the same, and sequential processing is performed on all the first images to be processed. Now, the processing process of the first frame of the first image to be processed is listed as follows: Figure 2 as shown below:

[0115] 1) Convert the first image to be processed into a grayscale image, and apply Gaussian filtering for smoothing and noise reduction according to the following formula.

[0116]

[0117] f(x, y) is the function mapping value

[0118] 2) Further perform dynamic threshold processing on the first image to be processed, and the set threshold for dynamic threshold processing is 125 - 255.

[0119] 3) Use the Canny edge detection algorithm to process the contours and edges of the image

[0120] 3.1) Edge detection: Use the Canny algorithm to scan the first image to be processed, enhance the shape, features, and contour edges of the positioning member in the image by applying a Gaussian filter (to reduce noise), calculating gradients (to identify potential edges), and then thinning the feature edges to improve clarity.

[0121] 3.2) Image segmentation: Use the Canny algorithm to decompose the image into smaller segments or regions based on the color and shape features of the positioning member, and the regions of interest related to the positioning member can be found from them according to the features of the positioning member, that is, multiple first target regions are formed.

[0122] 3.3) Parameter dynamic adjustment: Use the Canny algorithm to adjust the parameters of each found first target region, such as edge sensitivity.

[0123] 3.4) Adaptive texture: Use the Canny algorithm to process each first target area separately. For example, relatively smooth first target areas can be processed quickly with lower sensitivity, while first target areas with complex patterns require higher sensitivity processing. Ensure that the necessary details of each part of the image captured by the Canny algorithm can be utilized, so that different parts of the image can be identified and adapted, making it effective and efficient for tasks such as feature extraction and segmentation.

[0124] 4) Connect the contours of multiple first target areas to form multiple separate areas;

[0125] 5) Screen the contours according to the area size, select the contours with an area occupancy ratio in the range of 50,000 - 500,000 pixel values, that is, screen out the contours of the second target areas that match the positioning parts from multiple separate areas according to the area size of the positioning parts, and draw the contours of the second target areas. Fit the contours of the second target areas according to the preset ellipse formula to form multiple ellipse frames;

[0126] 6) Screen the ellipses according to the aspect ratio of the major and minor axes, and select the ellipses with an aspect ratio of the major and minor axes between 0.8 - 1.2.

[0127] 7) According to the following ellipse formula, the coordinate values X and Y of the center point of the ellipse frame can be obtained.

[0128]

[0129] Where (x0, y0) is the coordinate of the center point of the ellipse, and a and b are the semi - axis lengths of the ellipse along the x - axis and y - axis respectively.

[0130] In summary, through steps such as data acquisition, model training, edge detection, border drawing, and Gaussian filtering, the present invention realizes the effective processing and analysis of video frames. It has the advantages of flexible and simple tracking, sensitive and accurate capture of position changes, so that the picture and position information can be fed back in real time, and the position changes of the positioning parts can be monitored and observed in real time.

Claims

1. A system for identifying changes in the position information of a surgical body surface positioning member, comprising a preprocessing module, a startup module, an identification module, a depth inference module, and an output module, characterized in that The preprocessing module consists of a package deployment environment preheating module and a neural network training module, and is used for the initialization and environment configuration of the puncture surgical robot system; The startup module consists of a strong logic judgment module and a data flow monitoring module, and is used for judging whether the video stream captured by the micro camera can be transmitted normally. If the transmission stream cannot be found during the judgment process, the program is closed and exited; The identification module is a first sample training module, and is used for identifying the X and Y position information of the center point of the positioning member in the plane of the image in each frame of image transmitted by the startup module; The depth inference module is a second sample training module, and applies a zero-sample generalization depth estimation model for inference, and is used for identifying the Z position information between the center point of the positioning member and the longitudinal direction of the micro camera in each frame of image transmitted by the identification module; The output module is used to synchronously output the X and Y position information obtained by the first sample training module and the Z position information obtained by the second sample training module to the display screen of the puncture surgical robot system, forming a real-time output data stream; A micro camera is embedded at the end of the needle insertion device at the end of the robotic arm of the puncture surgical robot system. The micro camera continuously captures the positioning member located on the patient's body surface during the time period of the puncture needle insertion. The positioning member is used for positioning during the puncture surgery of the puncture surgical robot system. The startup module acquires all the images to be processed, performs initialization processing on all the acquired images to be processed, and outputs them to the identification module in sequence as the first image to be processed; The first sample training module of the identification module infers and obtains the X and Y position information of the center point of the positioning member in the plane of the image in all the first images to be processed, forming a second image to be processed including the X and Y position information; The second sample training model first infers and obtains the longitudinal depth value of all the second images to be processed from the micro camera, and further infers and obtains the longitudinal Z position information of the center point of the positioning member.

2. The system for identifying the change in the position information of the body surface positioning member for surgery according to claim 1, wherein The X, Y, and Z position information of the center point of the positioning member is used to make a displacement curve that can be presented on the display screen.

3. The system for identifying the change in the position information of the surgical body surface positioning member according to claim 2, wherein The startup module converts the first frame of image to be processed into a first array with three channels, that is, the digital processing of the image. The first array with three channels contains the content information of each pixel point, color, and size of the first frame of image to be processed. At the same time, the coordinate values of the three channels in the first array are defined, and then the subsequent second frame of image to be processed is converted into a second array with three channels according to the coordinate values of the three channels in the first array. On the basis of the second array, the third frame of image to be processed is converted into a third array with three channels, and so on. All the acquired images to be processed are initialized in sequence.

4. The system for identifying the change of the position information of the body surface positioning piece for surgery according to claim 1, wherein The identification module converts the first image to be processed into a grayscale image, and performs smoothing and noise reduction processing using Gaussian filtering; the identification module performs dynamic threshold processing on the first image to be processed after Gaussian filtering, and sets the dynamic threshold to 125 - 255.

5. The system for identifying the change in the position information of the surgical body surface positioning member according to claim 4, wherein The recognition module uses the canny edge detection algorithm to scan and enhance the shape, features, and contour edges of the positioning part in the first image to be processed after the above processing, and decomposes the image into smaller segments or regions according to the image features of the positioning part to form multiple first target regions; and performs connectivity processing on the contours of the multiple first target regions to form multiple separate regions.

6. The system for identifying the change in the position information of the surgical body surface positioning member according to claim 5, characterized in that, The recognition module filters out the second target region from the multiple separate regions according to the area size of the positioning part, draws the contour of the second target region, fits the contour of the second target region according to the preset ellipse formula to form multiple ellipse frames; and selects the target ellipse frame according to the aspect ratio of the long and short axes of the ellipse, and obtains the coordinate values of the center point of the target ellipse frame, which are the X and Y position information of the center point of the positioning part in this plane.

7. The system for identifying the change in the position information of the body surface positioning part for surgery according to claim 6, characterized in that, The aspect ratio of the long and short axes of the ellipse is between 0.8 and 1.

2.

8. The system for identifying the change in the position information of the body surface positioning member for surgery according to any one of claims 1-7, characterized in that, The second sample training model uses a 3D acquisition device to acquire the grayscale image and depth image of the positioning part, uses the pre-trained DINOv2 model to extract the features of the positioning part from them, and aligns the features extracted by the depth estimation model with the features extracted in the DINOv2 encoder according to the following formula; Wherein: cos(f i , f' i ) measures the cosine similarity between two feature vectors; f i Features extracted by the depth estimation model; f’ i are the features extracted from the DINOv2 model.

Citation Information

Patent Citations

  • Navigation system and readable storage medium

    CN119367052A

  • Vascular puncture device and control method thereof

    WO2024070933A1