An autonomous identification and positioning method, system and medium for medical devices
Through the Arrow OBB-YOLO network structure, the arrow bounding box is generated using data preprocessing and optimization training algorithms, which solves the accuracy and real-time problems of medical device recognition and positioning, and realizes the rapid and accurate identification and real-time positioning of the device.
Patent Information
- Application Number
- CN202210093561.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-01-26
AI Technical Summary
In the prior art, medical device identification methods based on cooperative marks are susceptible to contamination and interference from blood, resulting in inaccurate identification and positioning, and image segmentation methods cannot meet clinical real-time requirements.
Arrow OBB-YOLO network structure is adopted, and the arrow bounding box is generated through data preprocessing, optimization training algorithms and output constraint models, and combined with device number judgment and focus generation area algorithm to achieve rapid and accurate identification and real-time positioning of the device.
It realizes the rapid and accurate identification of a variety of medical devices, and provides key information such as the position and angle of the device in real time, improving the accuracy and real-timeness of automatic identification of medical devices.
Smart Images

Figure CN114549640B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a method, a system and a medium for autonomous recognition and positioning of medical devices. Background Art
[0002] The application of robot-assisted control endoscopes in surgical operations has received extensive attention and rapid development. Compared with traditional minimally invasive endoscopic surgeries, it can effectively alleviate problems such as unstable and inaccurate display screen images during manual operations. Years of research have shown that using computer-assisted robots to achieve automatic adjustment of endoscopes can reduce the problem of distracted attention during doctor-patient interaction. One of the key issues in using computer-assisted robots is the recognition and positioning of surgical instruments. With the rapid development of image-related technologies, vision-based surgical instrument recognition and positioning methods have received extensive attention. In image-based recognition methods, cooperation marks are usually added to surgical instruments, and the reference points of medical devices are indirectly obtained by recognizing the cooperation marks. Although the recognition method based on cooperation marks can quickly achieve the recognition and positioning of surgical instruments, the marks on the instruments are easily contaminated and interfered by blood and the like, resulting in inaccurate recognition and positioning. Deep learning has excellent performance in image detection and recognition, and it shows excellent robustness to the interference of light and shadow, so it has gradually become the most effective method for endoscopic surgical instrument recognition. Among various expression methods of medical devices using deep learning, the instrument is labeled through a bounding box, and its center is used to approximately equivalent the position of the instrument, but the information of the instrument is not accurately expressed comprehensively. Although using the image segmentation method to achieve pixel-level classification can fully express the instrument, it cannot meet the real-time requirements in clinical practice. Summary of the Invention
[0003] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is to provide a method, a system and a medium for autonomous recognition and positioning of medical devices, which can achieve the rapid and accurate recognition of multiple medical devices, and accurately give key information such as the position and angle of the instruments in real time, and improve the accuracy and real-time performance of automatic recognition of medical devices.
[0004] To achieve the above object, the present invention provides a method for autonomous recognition and positioning of medical devices, including the following steps:
[0005] Obtain pictures, then perform data preprocessing, perform data augmentation on the labeled data, and then divide the data into a training set, a validation set and a test set; wherein, the data augmentation includes inversion, cropping and splicing;
[0006] The preprocessed data is optimized and trained by an optimized training algorithm. By optimizing the error expression and adjusting the error weights, the training convergence speed of the arrow bounding box is accelerated. The optimal weights obtained from the optimized training are used for real-time prediction, and the output of the real-time prediction is ensured to keep the arrow within the bounding box through an output constraint model.
[0007] According to the output constraint model, a corresponding arrow bounding box is generated in a coordinate transformation manner through an arrow bounding box generation algorithm.
[0008] The corresponding focus generation area algorithm is selected based on the number of instruments, and a corresponding tracking area is generated according to the arrow bounding box to achieve instrument positioning and provide guidance for instrument tracking.
[0009] Furthermore, the optimized training algorithm adopts the Arrow OBB-YOLO network structure.
[0010] Furthermore, the output constraint model restricts the arrow coordinates output by the network between [0, 1] through σ(x), and it is agreed that this is the coordinate relative to the upper left corner of the bounding box, where:
[0011]
[0012] And it is converted into the relative coordinate relative to the upper left corner of the photo through the following formula, that is:
[0013]
[0014]
[0015]
[0016]
[0017] where, b ax1 , b ay1 is the relative coordinate of the arrow tail, b ax2 , b ay2 is the relative coordinate of the arrow head, t x1 , t y1 is the output corresponding to the arrow tail in the network prediction result, t x2 , t y2 is the output corresponding to the arrow head in the network prediction result, b x , b y is the pixel coordinate of the center of the predicted bounding box, b w , b h is the width and height of the predicted bounding box, and W, H represent the width and height of the picture.
[0018] Based on the above coordinate transformation, the arrow is constrained within the bounding box, and the prediction of the bounding box with an arrow is realized using a neural network.
[0019] Furthermore, the corresponding focus generation area algorithm is selected by judging the number of instruments, and the corresponding tracking area is generated according to the arrow bounding box to achieve instrument positioning and provide guidance for instrument tracking. Specifically:
[0020] A direction vector is generated using the deviation between the focus and the center of the image, and the depth of the endoscope is adjusted using the position ratio of the target area in the image. The multiple instruments are divided into three categories: the first category is a single instrument, the second category is two instruments, and the third category is three or more instruments. Assume O i , R i represent the center and radius of the tracking area in the i-th case, c x , c y represent the horizontal and vertical coordinates of the center respectively, b j ax1 , b j ay1 represent the tail coordinates of the j-th arrow, b j ax2 , b j ay2 represent the tail coordinates of the j-th arrow
[0021] For a single instrument, the tip of the arrow is directly used as the field of view focus, and at the same time, the focus is used as the center and the length of the arrow is used as the radius to generate the target area, that is:
[0022]
[0023] For two instruments, the midpoint of the line connecting the tips of the two arrows is used as the focus, and at the same time, the focus is used as the center and the distance from the end of the two arrows farther from the center is used as the radius to generate the target area, that is:
[0024]
[0025] where, L1 = (c x - b1 ax1 ) 2 + (c y - b1 ay1 ) 2 , L2 = (c x - b2 ax1 ) 2 + (c y - b2 ay1 ) 2 , and max(L1, L2) represents the maximum value of the two.
[0026] For three or more devices, by selecting three main devices, the problem is converted into a problem of three devices. Using the center of the circumcircle of the arrow tips of the three devices as the focus, with this focus as the center and the maximum distance from the arrow end to this focus as the radius, a target area is generated, that is:
[0027]
[0028] Among them, the calculation of L1 and L2 is the same as before, and L3 = (c x -b3 ax1 ) 2 +(c y -b3 ay1 ) 2 , O3 = (c x , c y ) is the circumcircle of the three arrow head points, and their mathematical relationship can be expressed as:
[0029]
[0030] According to the formula, the coordinates of the circumcircle center can be obtained as:
[0031]
[0032] Among them,
[0033] Furthermore, the Arrow OBB-YOLO network structure adds the prediction of arrow coordinates. The arrow coordinates are composed of the vectors of two coordinates. Define the pixel coordinates of the tail and head of the arrow as (b ax1 , b ay1 ), (b ax2 , b ay2 ), and the output of the network is defined as t x1 , t y1 , t x2 , t y2 , t x1 , t y1 is the output corresponding to the arrow tail in the network prediction result, and t x2 , t y2 is the output corresponding to the arrow head in the network prediction result. The output dimension of the network is N gx ·N gy ·N a ·(N cls +5+4), where N gx , N gy represents the number of divided grids, N a represents the number of preset anchor boxes, and N cls represents the number of predicted categories.
[0034] The present invention also provides an autonomous identification and positioning system for medical devices, including an image acquisition module, a model training module, a bounding box generation module, and a tracking area generation module, wherein:
[0035] The image acquisition module is used to acquire the images required for model training and the images for real-time detection;
[0036] The model training module mainly trains the deep model according to the images and labels using an optimization algorithm to obtain the optimal weights for real-time detection;
[0037] The bounding box generation module is used to detect medical devices in the images acquired in real time and generate arrow bounding boxes according to the algorithm;
[0038] The tracking area generation module is used to generate corresponding tracking areas according to the recognition results of the devices, realize device positioning, and guide device tracking.
[0039] The present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the steps of the above method.
[0040] The beneficial effects of the present invention are as follows:
[0041] The present invention can realize the rapid and accurate recognition of a variety of medical devices, and give key information such as the position and angle of the devices in real time and accurately, improving the accuracy and real-time performance of automatic medical device recognition.
[0042] The following will further illustrate the concept, specific structure and technical effects of the present invention in conjunction with the drawings to fully understand the purpose, features and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a comparison diagram of traditional surgical procedures and computer-aided procedures.
[0044] Figure 2 is the overall flowchart of the recognition and positioning method based on the Arrow OBB network of the present invention.
[0045] Figure 3 is the model diagram of the arrow bounding box of the present invention.
[0046] Figure 4 is the schematic diagram of the output model of the bounding box of the present invention.
[0047] Figure 5 is the Arrow OBB-YOLO network structure diagram of the present invention.
[0048] Figure 6 is the schematic diagram of arrow calculation of the present invention.
[0049] Figure 7 This is the flow chart of the step-by-step training optimization algorithm of the present invention. Detailed implementation manner
[0050] As Figure 1 shown, in traditional surgical operations, doctors need to judge whether the instrument deviates from the center of the image according to the position of the instrument in the endoscopic display screen. Then, according to past experience, the endoscope is finely adjusted to achieve the tracking of the surgical instrument by the endoscope. For an automatic tracking system, the identification, positioning, judgment, and tracking of the instrument are all left to the computer to complete independently. In the substitution process from manual judgment to automatic computer completion, accurate identification and positioning of the surgical instrument are the most crucial link and also the prerequisite for realizing automatic instrument tracking.
[0051] To improve the accuracy and real-time performance of instrument identification and positioning, the present invention proposes a general framework for the method and system of autonomous identification and positioning of medical devices. It mainly consists of an image acquisition module, a model training module, an "arrow bounding box" prediction module, and a tracking area generation module. The image acquisition module is used to acquire the images required for model training and the images for real-time detection. The model training module mainly trains the deep model using the optimization algorithm based on the images and labels to obtain the optimal weights for real-time detection. The bounding box generation module is used to detect medical devices in the images acquired in real time and generate arrow bounding boxes according to the algorithm. The tracking area generation module is used to generate corresponding tracking areas according to the identification results of the instruments, realize instrument positioning, and guide instrument tracking.
[0052] The algorithms of the medical device autonomous identification and positioning system mainly include: data preprocessing, optimization training algorithm, "arrow bounding box" generation algorithm, instrument number judgment, and focus area generation algorithm. Data preprocessing is the first step after obtaining the data, mainly performing data augmentation (inversion, cropping, and splicing) on the labeled data, and then dividing the data into a training set, a validation set, and a test set. The optimization training algorithm speeds up the training convergence speed of the arrow bounding box by optimizing the error expression and adjusting the error weights. The optimal weights obtained by the optimization training are used for real-time prediction, and the output of the real-time prediction is ensured by the output constraint that the arrow falls inside the bounding box. The "arrow bounding box" generation algorithm uses the constrained model output and generates the corresponding "arrow bounding box" by coordinate transformation. Then, the corresponding focus generation area algorithm is selected through the instrument number judgment, and the corresponding tracking area is generated according to the "arrow bounding box" to realize instrument positioning and provide guidance for instrument tracking.
[0053] For the feature map of the output layer, we have preset A anchor boxes and divided the feature map into W f ×H f grids, so the feature map will predict a total of A bounding box. For the recognition of n categories, the specific definition of each bounding box is Where Represents the center coordinates of the predicted bounding box, Represents its width and height, Represents the confidence of this bounding box, Represents the probability of each category. We define the true bounding box as Where Is a one-hot matrix, the corresponding category is 1, and the values of other categories are 0.
[0054] Based on the above analysis, we can define the positioning error of the bounding box as:
[0055] L loc = θ i,j L(b i , g j ) (1)
[0056] Where, only when the predicted bounding box and the target bounding box are in the same area and the GIOU value is the largest, θ i,j = 1, and in other cases it is 0. And L(b i , g i ) = 1 - GIoU(b i , g i ).
[0057] Similarly, the confidence error and category error of the bounding box are defined as:
[0058]
[0059]
[0060] Where, L BCE (x, y) = {l i ... n}}, x, y ∈ R 1×n , l i = -w i (y i ogσ(x i ) + (1 - y i )log(1 - σ(x i ))).
[0061] Therefore, for the above bounding box model, we define the error model affected by multiple parameters as L(w) = L loc + L obj + L cls , where w is the parameter of the model.
[0062] Based on the identification of the key parts of the instrument using the bounding box, we use two points to represent the center of the key parts of the medical device and the tip, these two important pieces of information. At the same time, connecting the two points, the arrow pointing from the center to the tip represents the angular direction information of the instrument. The schematic diagram is as Figure 3 shown.
[0063] Based on the above model, the predicted i th arrow-assisted bounding box is defined as where the arrow coordinates The actual j th arrow-assisted bounding box is defined as where the arrow coordinates
[0064] Under this assumption, the other errors of the bounding box are defined in the same way as the previous bounding box. In addition, we additionally define the prediction error of the arrow and give the optimization objective, that is:
[0065]
[0066] When calculating the arrow error, consider the distance error of the corresponding point coordinates, that is:
[0067]
[0068] Furthermore, the total error of the arrow-assisted bounding box model can be written as:
[0069] L(w) = L loc + L obj + L cls + L a (6)
[0070] Therefore, our objective function is:
[0071]
[0072] The improved Arrow OBB-YOLO network structure proposed by the present invention is as Figure 5 shown. ArrowOBB-YOLO draws on the unified prediction characteristics of the YOLO series of networks and directly predicts the bounding box with an arrow at one time using the network. Similar to the YOLOv3 network, it is based on the modification of the Draknet53 network, and the key is the modification of its output.
[0073] The output size of each prediction layer of YOLO is N gx · N gy · N a · (N cls + 5), N gx · N gy represents the number of grids, Na represents the number of preset anchor boxes. Each YOLO layer predicts N gx ·N gy ·N a bounding boxes, and each predicted bounding box has N cls +5 parameters, namely the position predictions t x , t y , t w , t h , the confidence prediction t obj , and N cls numbers between 0 and 1 representing class probabilities. In YOLOv3, the network output is converted into prediction results through the following formula, c x , c y refers to the offset of the grid relative to the upper left corner, p w , p h are the pre-set width and height of the anchor box, that is:
[0074] b x = σ(t x ) + c x (8)
[0075] b y = σ(t y ) + c y (9)
[0076]
[0077]
[0078] Figure 4 is the schematic diagram of the output model of the bounding box. Based on YOLO, Arrow-YOLO adds the prediction of arrow coordinates. From the arrow bounding box modeling process in the previous text, we know that our arrow coordinates are composed of the vectors of two coordinates. We define the pixel coordinates of the starting point and the ending point of the arrow as (b ax1 , b ay1 ), (b ax2 , b ay2 ), and the network output is defined as t x1 , t y1 , t x2 , t y2 , so the output dimension of the network becomes N gx ·N gy ·N a ·(N cls +5+4).
[0079] During the prediction process of the arrow, if the pixel coordinates of the arrow are directly predicted, the arrow may fall outside the bounding box. Therefore, we use coordinate transformation to constrain the coordinates of the arrow. Thus, in the present invention, the arrow coordinates output by the network are restricted to between [0, 1] by σ(x), which we agree to be the coordinates relative to the upper left corner of the bounding box, and are converted into the relative coordinates relative to the upper left corner of the photo through the following formula, that is:
[0080]
[0081]
[0082]
[0083]
[0084] where b ax1 , b ay1 are the relative coordinates of the arrow tail, b ax2 , b ay2 are the relative coordinates of the arrow head, t x1 , t y1 are the outputs corresponding to the arrow tail in the network prediction result, t x2 , t y2 are the outputs corresponding to the arrow head in the network prediction result, b x , b y are the pixel coordinates of the center of the predicted bounding box, b w , b h are the width and height of the predicted bounding box, and W, H represent the width and height of the picture.
[0085] Based on the above coordinate transformation, we have successfully constrained the arrow within the bounding box and realized the prediction of the bounding box with an arrow using a neural network.
[0086] In the actual calculation process, the coordinate error cannot well express the difference of the arrows. As Figure 6 shown, the error between arrow a1 and arrow a2
[0087] The error between arrow a1 and arrow a3 is From the figure, we can see that the endpoints of the blue arrow and the green arrow fall on the same circle, so L d (a1, a2) = L d (a1, a3), but according to the length and angle information of the arrow, obviously the blue arrow is more similar to the red arrow, so the calculation of the error needs to consider the length and angle information of the arrow.
[0088] Define the length error between arrows as:
[0089]
[0090] Similarly, we define the angular error between arrows as:
[0091]
[0092]
[0093] Since problems such as gradient explosion may occur when solving atan() multiple times, we transform this error model and define the angular error as:
[0094]
[0095] Among them, through the transformation of trigonometric functions, the conversion of errors is made more convenient, that is:
[0096]
[0097] Combining the coordinate error, length error, and angular error, the arrow error can finally be defined as:
[0098]
[0099] Although the above model successfully realizes the prediction output of the "arrow bounding box", the convergence effect and stability of the model are not good. In the process of calculating the arrow error, if only the coordinates relative to the bounding box are used, the prediction effect cannot be reflected, while using the coordinates relative to the entire image, the problem of coupling between arrow prediction and bounding box appears. This is because we limit the arrow inside the bounding box. Before the bounding box is accurately predicted, the error of the arrow will be very large and there is no clear convergence direction.
[0100] To address the above problems, we propose a step-by-step training method, and its process is as Figure 7 shown. Based on the above analysis, we know that the reason for the instability of the arrow error is that the bounding box prediction has not converged. So we initially set the weight of the arrow error to w, so that the error calculation of the model only considers the error of the bounding box. Update and record the mAP of the model on the validation set. After judging that the prediction effect of the model on the bounding box meets the requirements and is stable, then introduce the error of the arrow by modifying the weight.
[0101] Based on the above "arrow bounding box" model, we propose a focal region model. The deviation between the focus and the center of the image is used to generate a direction vector, and the proportion of the target region in the image is used to adjust the depth of the endoscope. The generation method of the "field of view focal region" here divides multiple instruments into three categories: the first category is a single instrument, the second category is 2 instruments, and the third category is 3 or more instruments.
[0102] For a single instrument, directly use the tip of the arrow as the focus of the field of view, and at the same time use this focus as the center of the circle and the length of the arrow as the radius to generate the target area, that is:
[0103]
[0104] For two instruments, use the midpoint of the line connecting the tips of the two arrows as the focus, and at the same time use this focus as the center of the circle and the distance from the end of the two arrows farther from the center as the radius to generate the target area, that is:
[0105]
[0106] where, L1 = (c x -b1 ax1 ) 2 +(c y -b1 ay1 ) 2 , L2 = (c x -b2 ax1 ) 2 +(c y -b2 ay1 ) 2 .
[0107] For three or more instruments, by selecting three main instruments, convert the problem into a problem of three instruments. Use the center of the circumscribed circle of the arrow tips of the three instruments as the focus. Take this focus as the center of the circle and the farthest distance from the end of the arrow to this focus as the radius to generate the target area, that is:
[0108]
[0109] where, the calculation of L1 and L2 is the same as before, and L3 = (c x -b3 ax1 ) 2 +(c y -b3 ay1 ) 2 . O3 = (c x , c y ) is the circumscribed circle of the three arrow head points. Their mathematical relationship can be expressed as:
[0110]
[0111] According to Equation (25), the coordinates of the center of the circumscribed circle can be obtained as:
[0112]
[0113] where,
[0114] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0115] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art shall fall within the protection scope determined by the claims.
Claims
1. An autonomous identification and positioning method for a medical device, characterized in that, It includes the following steps: Obtain images, then perform data preprocessing, perform data augmentation on the labeled data, and then divide the data into a training set, a validation set, and a test set; among them, data augmentation includes inversion, cropping, and stitching; Optimize and train the preprocessed data through an optimized training algorithm. By optimizing the error expression and adjusting the error weights, the training convergence speed of the arrow bounding box is accelerated; the optimal weights obtained from the optimized training are used for real-time prediction, and the output of the real-time prediction is ensured by the output constraint model that the arrow falls within the bounding box; According to the output constraint model, generate a corresponding arrow bounding box in a coordinate transformation manner through the arrow bounding box generation algorithm; Select the corresponding focus generation area algorithm through the instrument number judgment, generate a corresponding tracking area according to the arrow bounding box, realize instrument positioning, and provide guidance for instrument tracking; Among them, the optimized training algorithm adopts the Arrow OBB-YOLO network structure; The output constraint model restricts the arrow coordinates output by the network between [0, 1] through σ(x), and agrees that it is the coordinate relative to the upper left corner of the bounding box, where: And convert it into the relative coordinate relative to the upper left corner of the photo through the following formula, that is: Among them, b ax1 , b ay1 are the relative coordinates of the arrow tail, b ax2 , b ay2 are the relative coordinates of the arrow head, t x1 , t y1 is the output corresponding to the arrow tail in the network prediction result, t x2 , t y2 is the output corresponding to the arrow head in the network prediction result, b x , b y are the pixel coordinates of the center of the predicted bounding box, b w , b h are the width and height of the predicted bounding box, where W and H represent the width and height of the image; Based on the above coordinate transformation, the arrow is constrained within the bounding box, and the prediction of the bounding box with an arrow is realized by using a neural network; The Arrow OBB-YOLO network structure adds the prediction of arrow coordinates, which are composed of the vectors of two coordinates. The pixel coordinates of the tail and head of the arrow are defined as (b ax1 ,b ay1 ),(b ax2 ,b ay2 ). The output of the network is defined as t x1 ,t y1 ,t x2 ,t y2 ,t x1 ,t y1 . t x2 ,t y2 is the output corresponding to the tail of the arrow in the network prediction result, and t x2 ,t y2 is the output corresponding to the head of the arrow in the network prediction result. The output dimension of the network is N gx ·N gy ·N a ·(N cls +5+4), where N gx , N gy represents the number of divided grids, N a represents the number of preset anchor boxes, and N cls represents the number of predicted categories.
2. The method for autonomous identification and positioning of a medical device according to claim 1, wherein The step of selecting the corresponding focus generation area algorithm through the instrument number judgment, generating a corresponding tracking area according to the arrow bounding box, realizing instrument positioning, and providing guidance for instrument tracking is specifically: Generate a direction vector using the deviation between the focus and the center of the image, and adjust the depth of the endoscope using the position ratio of the target area in the image. The multi-instruments are divided into three categories: the first category is a single instrument, the second category is two instruments, and the third category is three or more instruments; Assume O i ,R i represent the center and radius of the tracking area in the i-th case, c x ,c y represent the horizontal and vertical coordinates of the center respectively, b j ax1 ,b j ay1 represent the tail coordinates of the j-th arrow, b j ax2 ,b j ay2 represent the tail coordinates of the j-th arrow; For a single instrument, directly use the tip of the arrow as the field of view focus, and at the same time use this focus as the center of the circle and the length of the arrow as the radius to generate the target area, that is: For two instruments, use the midpoint of the line connecting the tips of the two arrows as the focus, and at the same time use this focus as the center of the circle and the distance from the end of the two arrows farther from the center as the radius to generate the target area, that is: where, L1 = (c x - b1 ax1 ) 2 + (c y - b1 ay1 ) 2 ; L2 = (c x - b2 ax1 ) 2 + (c y - b2 ay1 ) 2 ; max(L1, L2) represents the maximum value of the two; For three or more instruments, by selecting three main instruments, convert the problem into a problem of three instruments, use the center of the circumcircle of the tips of the arrows of the three instruments as the focus, and use this focus as the center of the circle and the farthest distance from the end of the arrow to this focus as the radius to generate the target area, that is: Among them, the calculations of L1 and L2 are the same as those in the previous text, and L3 = (c x -b3 ax1 ) 2 +(c y -b3 ay1 ) 2 , and O3 = (c x , c y ) is the circumcircle of the three arrowhead points, and their mathematical relationship can be expressed as: According to formula (9), the coordinates of the center of the circumcircle can be obtained as: Where:
3. An independent identification and positioning system for medical devices, characterized in that, It includes an image acquisition module, a model training module, a bounding box generation module, and a tracking area generation module, where: The image acquisition module is used to acquire the images required for model training and the images for real-time detection; The model training module trains the deep model according to the images and labels by using the optimized training algorithm, obtains the optimal weights for real-time detection, and the output of the real-time prediction is ensured by the output constraint model that the arrow falls within the bounding box; The bounding box generation module is used to detect medical devices in the images acquired in real time, and generate a corresponding arrow bounding box in a coordinate transformation manner according to the arrow bounding box generation algorithm; The tracking area generation module is used to generate a corresponding tracking area according to the recognition result of the instrument, realize instrument positioning, and guide instrument tracking; Among them, the optimized training algorithm adopts the Arrow OBB-YOLO network structure; The output constraint model restricts the arrow coordinates output by the network between [0, 1] through σ(x), and it is agreed that they are the coordinates relative to the upper left corner of the bounding box, where: And it is converted into the relative coordinates relative to the upper left corner of the photo through the following formula, that is: Among them, b ax1 , b ay1 are the relative coordinates of the arrow tail, b ax2 , b ay2 are the relative coordinates of the arrow head, t x1 , t y1 is the output corresponding to the arrow tail in the network prediction result, t x2 , t y2 is the output corresponding to the arrow head in the network prediction result, b x , b y are the pixel coordinates of the center of the predicted bounding box, b w , b h are the width and height of the predicted bounding box, where W and H represent the width and height of the image; Based on the above coordinate transformation, the arrow is constrained within the bounding box, and the prediction of the bounding box with an arrow is realized by using a neural network; The Arrow OBB-YOLO network structure adds the prediction of arrow coordinates, which are composed of the vectors of two coordinates. The pixel coordinates of the tail and head of the arrow are defined as (b ax1 , b ay1 ), (b ax2 , b ay2 ). The output of the network is defined as t x1 , t y1 , t x2 , t y2 , t x1 , t y1 . t x2 , t y2 is the output corresponding to the tail of the arrow in the network prediction result, and t gx , t gy is the output corresponding to the head of the arrow in the network prediction result. The output dimension of the network is N gx ·N gy ·N a ·(N cks + 5 + 4), where N gx , N gy represents the number of divided grids, N a represents the number of preset anchor boxes, and N cls represents the number of predicted categories.
4. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 2.
Citation Information
Patent Citations
Multi-people posture recognition method based on optical flow positioning and sliding window detection
CN106611157A
Gesture tracking and recognition method based on deep learning
CN111709310A