Needle tip recognition method and device and needle tip recognition model training method and device
Through the key point detection network combined with contour recognition technology, the needle tip position is automatically marked and iteratively trained, solving the instability and high cost of needle tip recognition in surgical robots, and achieving efficient and real-time needle tip recognition.
Patent Information
- Application Number
- CN202510350788.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-22
AI Technical Summary
In the prior art, in the identification of needle tips in surgical robots, there are problems of identification instability and high cost of manual labeling, especially in complex backgrounds and lighting changes, which are difficult to meet real-time needs.
The key point detection network is used to perform preliminary labeling through contour recognition technology, and the pin tip position is automatically determined using the extremely subtle characteristics of the needle tip, and the key point detection network is iteratively trained to form a self-iteration labeling mechanism, reducing manual labeling costs and improving recognition accuracy and real-timeness.
It effectively reduces the cost of manual labeling for needle tip recognition, improves the accuracy and real-time recognition, and can accurately capture needle tip information in complex backgrounds, meeting the real-time needs of ophthalmic surgery.
Smart Images

Figure CN120355890A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical image recognition, and particularly to a method and device for tip recognition and tip recognition model training. Background Art
[0002] Surgical robots have gradually become an important tool for assisting ophthalmologists in ophthalmic surgeries due to their high precision, stability, and repeatability. During the calibration process of surgical robots, after each Remote Center of Motion (RCM) movement, it is necessary to identify the tip position to calculate the error and optimize the system control parameters.
[0003] However, traditional object detection models rely on rectangular box annotations. When facing the small tip positioning scenario, not only is the manual annotation workload large, but it is also easily affected by factors such as background complexity, illumination changes, and tissue reflection, resulting in unstable recognition. Although the segmentation model can provide more detailed pixel-level information, its computational complexity is higher and it cannot meet the real-time requirements of ophthalmic surgeries.
[0004] Therefore, how to ensure the real-time performance of tip recognition while reducing the manual annotation cost required for model training is a technical problem that needs to be solved currently. Summary of the Invention
[0005] This application provides a method and device for tip recognition and tip recognition model training to solve the technical problem of how to ensure the real-time performance of tip recognition while reducing the manual annotation cost required for model training.
[0006] To solve the above technical problem, in a first aspect, an embodiment of this application provides a method for training a tip recognition model, including:
[0007] Collect a number of tip images with the same tip orientation to obtain a first data set;
[0008] Label each tip image in the first data set through contour recognition technology to obtain a second data set;
[0009] According to the currently obtained second data set, iteratively train a preset key point detection network, and obtain annotation data according to the key point detection network obtained in each iteration to update the second data set until the accuracy of the key point detection network meets the preset conditions, and output the finally iteratively obtained key point detection network as the tip recognition model.
[0010] Compared with the prior art, the embodiments of the present application have the following beneficial effects: Since the tip of the needle is an extremely small and specific point, if a target detection model is used, it will not only lead to inaccurate detection of the tip of the needle, but also the manual annotation cost required to train the target detection model for identifying the tip of the needle target is higher. Instead, through the key point detection network, only the tips of the needle images in the dataset need to be annotated, effectively reducing the cost required for annotating data. At the same time, when performing data annotation for the first time, contour recognition technology is used for annotation. When the key point detection network is trained with the data annotated for the first time, the tip recognition results output by the key point detection network updated each time can be used as annotations to update the second data, forming a self-iterative annotation mechanism, which reduces the manual annotation cost while improving the model accuracy. Finally, the key point detection network requires less computational effort compared to the semantic segmentation model. In the face of real-time tip recognition scenarios, it can improve the recognition real-time performance while ensuring the recognition accuracy.
[0011] In some embodiments of the first aspect of the present application, obtaining annotation data according to the key point detection network obtained in each iteration to update the second dataset includes:
[0012] Re-collect a number of needle tip images, and input the re-collected needle tip images into the key point detection network obtained in the current iteration to obtain the detection boxes and key points corresponding to the needle tip images;
[0013] Annotate the corresponding needle tip images according to the detection boxes and key points;
[0014] Update the second dataset used in the current iteration according to the newly annotated needle tip images.
[0015] Compared with the prior art, the above embodiments have the following beneficial effects: Through the self-iterative annotation mechanism, in each training iteration, needle tip images are re-collected, and the detection boxes and key points are automatically output by the key point detection network obtained in the current iteration. There is no need to manually annotate the new images, but instead, the network is used to automatically generate annotation data, effectively reducing the dataset annotation cost. Further, the dataset is continuously updated according to the newly annotated data, enabling the model to obtain more accurate feedback in each iteration, thereby gradually optimizing the accuracy of key point detection.
[0016] In some embodiments of the first aspect of the present application, annotating each needle tip image in the first dataset through contour recognition technology to obtain the second dataset includes:
[0017] Identify the needle tip contours in each needle tip image in the first dataset through contour recognition technology; the needle tip contours include a number of contour points;
[0018] Determine the needle tip recognition judgment condition according to the needle tip orientation, and determine the first contour point that meets the needle tip recognition judgment condition from the needle tip contour;
[0019] Mark the corresponding needle tip image according to the needle tip contour and the first contour point, and obtain the second data set.
[0020] Compared with the prior art, the above embodiments have the following beneficial effects: Utilizing the extremely fine feature of the needle tip, the needle tip must be the extreme value of the needle body contour in a certain orientation. Therefore, automatically determining the needle tip position using the image contour and specific judgment conditions can effectively improve the accuracy of the first labeled data, and at the same time avoid the cumbersome operation of needing to label the entire rectangular box in traditional object detection.
[0021] In some embodiments of the first aspect of the present application, the step of marking the corresponding needle tip image according to the needle tip contour and the first contour point to obtain the second data set includes:
[0022] Calculate the minimum bounding rectangle of the needle tip contour;
[0023] Use the minimum bounding rectangle as the detection frame and the first contour point as the key point, and mark the corresponding needle tip image according to the detection frame and the key point.
[0024] Compared with the prior art, the above embodiments have the following beneficial effects: Defining the detection area through the minimum bounding rectangle and using this detection area for data annotation can enable the subsequent needle tip recognition model to quickly determine the area where the needle tip is located when recognizing the key point, further accurately lock the needle tip position, thereby improving the stability and efficiency of key point recognition.
[0025] In some embodiments of the first aspect of the present application, the step of determining the needle tip recognition judgment condition according to the needle tip orientation includes:
[0026] Determine the axis orientation with the smallest included angle formed by the needle tip orientation and the axis orientation of the coordinate system of the needle tip image;
[0027] When the axis orientation is the positive direction of the vertical axis, determine that the needle tip recognition judgment condition is that the vertical axis coordinate value corresponding to the first contour point is the largest;
[0028] When the axis orientation is the negative direction of the vertical axis, determine that the needle tip recognition judgment condition is that the vertical axis coordinate value corresponding to the first contour point is the smallest;
[0029] When the axis orientation is the positive direction of the horizontal axis, determine that the needle tip recognition judgment condition is that the horizontal axis coordinate value corresponding to the first contour point is the largest;
[0030] When the axis orientation is the negative direction of the horizontal axis, it is determined that the needle tip recognition judgment condition is that the horizontal axis coordinate value corresponding to the first contour point is the smallest.
[0031] Compared with the prior art, the above embodiments have the following beneficial effects: corresponding judgment conditions are given for different needle tip orientations, improving the robustness and consistency of the recognition process; further, key points are determined through quantitative coordinate value comparison, effectively reducing the annotation errors caused by image angle changes.
[0032] In some embodiments of the first aspect of the present application, the key point detection network includes: a first backbone network, a first multi-scale feature fusion module, and a first detection head;
[0033] Among them, the first backbone network is used to extract the first needle body feature of the needle tip image;
[0034] The first multi-scale feature fusion module extracts and fuses features of different scales according to the first needle body feature to obtain a first multi-scale fusion feature;
[0035] The first detection head is used to identify the key points and detection frames of the needle tip image according to the first multi-scale fusion feature.
[0036] Compared with the prior art, the above embodiments have the following beneficial effects: by extracting and fusing the features of the needle body at different scales, it is ensured that the needle tip information can still be accurately captured under complex backgrounds or tiny structures; at the same time, compared with the semantic segmentation method, the key point detection network has a simple structure and less computational complexity, effectively improving the real-time performance of key point recognition.
[0037] In the second aspect, the embodiments of the present application further provide a needle tip recognition method, including:
[0038] Obtain a needle tip recognition model; wherein, the needle tip recognition model is trained and obtained by using the needle tip recognition model training method in any one of the embodiments of the present application;
[0039] Input the needle tip image to be recognized into the needle tip recognition model, identify the key points of the needle tip image to be recognized, and use the key points as the needle tip recognition result.
[0040] Compared with the prior art, the above embodiments have the following beneficial effects: by using the trained key point detection network as the needle tip recognition model, the computational complexity required is smaller than that of the semantic segmentation model. In the face of real-time needle tip recognition scenarios, the recognition real-time performance can be improved while ensuring the recognition accuracy.
[0041] In some embodiments of the second aspect of the present application, the tip recognition model includes: a second backbone network, a second multi-scale feature fusion module, and a second detection head; inputting the tip image to be recognized into the tip recognition model to recognize the key points of the tip image to be recognized includes:
[0042] Extracting the second needle body feature of the tip image to be recognized through the second backbone network;
[0043] Through the second multi-scale feature fusion module, combining the second needle body feature, extracting and fusing features of different scales, and obtaining a second multi-scale fusion feature;
[0044] Inputting the second multi-scale fusion feature into the second detection head to recognize the key points of the tip image.
[0045] Compared with the prior art, the above embodiments have the following beneficial effects: By extracting and fusing the features of the needle body at different scales, it is ensured that the tip information can still be accurately captured under complex backgrounds or tiny structures; at the same time, compared with the semantic segmentation method, the key point detection network structure is simple and the calculation amount is less, effectively improving the real-time performance of key point recognition.
[0046] In a third aspect, an embodiment of the present application further provides a tip recognition model training device, including: a data acquisition module, a second data set acquisition module, and an iterative update module;
[0047] Among them, the data acquisition module is used to acquire a number of tip images with the same tip orientation to obtain a first data set;
[0048] The second data set acquisition module is used to label each tip image in the first data set through contour recognition technology to obtain a second data set;
[0049] The iterative update module is used to iteratively train a preset key point detection network according to the currently obtained second data set, and obtain labeled data according to the key point detection network obtained in each iteration to update the second data set until the accuracy of the key point detection network meets the preset conditions, and output the finally iteratively obtained key point detection network as the tip recognition model.
[0050] In a fourth aspect, an embodiment of the present application further provides a tip recognition device, including: a tip recognition model acquisition module and a tip recognition module;
[0051] Among them, the tip recognition model acquisition module is used to acquire a tip recognition model; among them, the tip recognition model is trained and obtained through any one of the tip recognition model training devices in the embodiments of the present application;
[0052] The tip recognition module is configured to input the tip image to be recognized into the tip recognition model, recognize the key points of the tip image to be recognized, and use the key points as the tip recognition result. Description of the Drawings
[0053] Figure 1 It is a schematic flowchart of a method for training a tip recognition model provided in some embodiments of the present application;
[0054] Figure 2 It is a schematic diagram of a tip image provided in some embodiments of the present application;
[0055] Figure 3 It is a schematic flowchart of a method for extracting the first contour points provided in some embodiments of the present application;
[0056] Figure 4 It is a schematic structural diagram of a key point detection network provided in some embodiments of the present application;
[0057] Figure 5 It is another schematic flowchart of a method for training a tip recognition model provided in some embodiments of the present application;
[0058] Figure 6 It is a schematic flowchart of a tip recognition method provided in some embodiments of the present application;
[0059] Figure 7 It is a schematic structural diagram of a tip recognition model training device provided in some embodiments of the present application;
[0060] Figure 8 It is a schematic structural diagram of a tip recognition device provided in some embodiments of the present application. Detailed Description of the Embodiments
[0061] Traditional object detection models rely on rectangular box annotations. When facing the small tip positioning scenario, not only is the manual annotation workload large, but also it is easily affected by factors such as background complexity, illumination changes, and tissue reflection, resulting in unstable recognition. Although the segmentation model can provide more detailed pixel-level information, the computational complexity is greater, and it cannot meet the real-time requirements of ophthalmic surgery.
[0062] In order to solve the above technical problems, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0063] Embodiment 1
[0064] Please refer to Figure 1 , a method for training a tip recognition model provided by an embodiment of the present application, including S101 to S103, specifically including:
[0065] S101: Collect a number of tip images with the same tip orientation to obtain a first data set.
[0066] Further, in some embodiments of the present application, the tip orientation in the tip image can be any orientation, and the present application does not make specific restrictions on the orientation. However, the tip images collected in the same batch should have the same orientation to facilitate batch processing of image data. Exemplarily, as Figure 2 shown in the schematic diagram of the tip image, the tip orientation can always be downward. In the embodiments of the present application, the case where the tip is downward will be taken as an example to specifically illustrate the solution.
[0067] S102: Label each tip image in the first data set through contour recognition technology to obtain a second data set.
[0068] Further, in some embodiments of the present application, labeling each tip image in the first data set through contour recognition technology to obtain a second data set includes:
[0069] Recognize the tip contour in each tip image in the first data set through contour recognition technology; the tip contour includes a number of contour points;
[0070] Determine the tip recognition judgment condition according to the tip orientation, and determine the first contour point that meets the tip recognition judgment condition from the tip contour;
[0071] Label the corresponding tip image according to the tip contour and the first contour point to obtain a second data set.
[0072] Further, in some embodiments of the present application, recognizing the tip contour in each tip image in the first data set through contour recognition technology includes: performing binarization processing on the tip image to convert the RGB image into a binary image containing only black and white pixel values; further performing contour extraction on the binary image to obtain the tip contour of the needle body. The contour extraction method includes, but is not limited to, any algorithm that can extract edge contours such as Canny edge detection, Sobel operator, Laplacian operator, etc. The present application does not limit the algorithm used.
[0073] Further, in some embodiments of the present application, the tip contour includes a number of contour points, and each contour point corresponds to a specific coordinate value. The coordinate system corresponding to this coordinate value can be a coordinate system established with any point as the origin. Subsequently, the tip recognition judgment condition is determined based on the tip orientation and the arbitrarily established coordinate system.
[0074] Taking advantage of the extremely fine tip, the tip must be the extreme value of the needle body contour in a certain orientation. Therefore, automatically determining the tip position using the image contour and specific judgment conditions can effectively improve the accuracy of the first labeled data and avoid the cumbersome operation of labeling the entire rectangular box in traditional object detection.
[0075] Further, in some embodiments of the present application, according to the tip contour and the first contour point, labeling the corresponding tip image to obtain the second dataset, including:
[0076] Calculating the minimum bounding rectangle of the tip contour;
[0077] Taking the minimum bounding rectangle as the detection box and the first contour point as the key point, and labeling the corresponding tip image according to the detection box and the key point.
[0078] Further, in some embodiments of the present application, calculating the minimum bounding rectangle of the tip contour can be implemented by minimum bounding rectangle algorithms such as AABB bounding box and oriented bounding box OBB. The present application does not limit this algorithm.
[0079] Defining the detection area through the minimum bounding rectangle and using this detection area for data labeling can enable the subsequent tip recognition model to quickly determine the area where the tip is located when recognizing the key point, further accurately lock the tip position, thereby improving the stability and efficiency of key point recognition.
[0080] Further, in some embodiments of the present application, according to the tip orientation, determining the tip recognition judgment conditions, including:
[0081] Determining the axis orientation with the smallest included angle formed by the tip orientation and the axis orientation of the tip image coordinate system;
[0082] When the axis orientation is the positive direction of the vertical axis, then determine that the tip recognition judgment condition is that the vertical axis coordinate value corresponding to the first contour point is the largest;
[0083] When the axis orientation is the negative direction of the vertical axis, then determine that the tip recognition judgment condition is that the vertical axis coordinate value corresponding to the first contour point is the smallest;
[0084] When the axis orientation is the positive direction of the horizontal axis, then determine that the tip recognition judgment condition is that the horizontal axis coordinate value corresponding to the first contour point is the largest;
[0085] When the axis orientation is the negative direction of the horizontal axis, then determine that the tip recognition judgment condition is that the horizontal axis coordinate value corresponding to the first contour point is the smallest.
[0086] Exemplarily, such as Figure 2The shown tip orientation is downward. For example, according to the pixel coordinates of the image (taking the upper left corner of the tip image as the origin, the vertical axis as the y-axis, and the horizontal axis as the x-axis, at this time the angle formed by the tip orientation and the positive direction of the vertical axis of the tip image coordinate system is the smallest), the tip recognition and judgment condition at this time is that the contour point with the largest y coordinate value among all contour points is the first contour point. Specifically:
[0087] C = {p1, p2, …, p n}
[0088] where C is the set formed by all contour points, and each point p i has coordinates, p i = (x i , y i );
[0089] C ′ = sort(C, key = y)
[0090] where C ′ is the sorted point set; sort(C, key = y) is to sort the set C by the y value.
[0091] Finally, the contour point with the largest y coordinate value in the sorted C ′ is used as the first contour point.
[0092] Corresponding judgment conditions are given for different tip orientations to improve the robustness and consistency of the recognition process; further, key points are determined by quantitative coordinate value comparison, effectively reducing the annotation error caused by image angle changes.
[0093] Further, referring to Figure 3 , when the tip is downward and the pixel coordinates are used as the coordinate system, the process of extracting the first contour point from the tip image is as follows: binarize the tip image to obtain a binary image; extract the tip contour from the binary image and calculate the minimum bounding rectangle according to the tip contour; sort the contour points in the tip contour by the y coordinate value; use the contour point with the largest y coordinate value as the first contour point; output the first contour point and the minimum bounding rectangle.
[0094] S103: Iteratively train a preset key point detection network according to the currently obtained second data set, and update the second data set according to the key point detection network obtained in each iteration until the accuracy of the key point detection network meets the preset conditions, and output the key point detection network obtained in the final iteration as the tip recognition model.
[0095] Further, in some embodiments of the present application, obtaining annotation data according to the key point detection network obtained in each iteration to update the second data set includes:
[0096] Re - collect a number of tip images, and input the re - collected tip images into the key - point detection network obtained in the current iteration to obtain the detection boxes and key points corresponding to the tip images;
[0097] Label the corresponding tip images according to the detection boxes and key points;
[0098] Update the second data set used in the current iteration according to the newly labeled tip images.
[0099] Furthermore, in some embodiments of the present application, when the detection boxes and key points are obtained (whether the detection boxes and key points are obtained by contour extraction or by the key - point detection network), when labeling the corresponding tip images according to the detection boxes and key points, the labelme annotation format is specifically used for storage to form an annotation file. Further, the qualified annotation files are saved, and the saved annotation files are converted from the labelme format to other formats, such as coco, voc, yolo and other formats.
[0100] Through the self - iterative annotation mechanism, in each training iteration, the tip images are re - collected, and the key - point detection network obtained in the current iteration is used to automatically output the detection boxes and key points. There is no need to manually annotate the new images, but the network is used to automatically generate annotation data, effectively reducing the cost of dataset annotation; further, the dataset is continuously updated according to the newly labeled data, so that the model can obtain more accurate feedback in each iteration, thereby gradually optimizing the accuracy of key - point detection.
[0101] Furthermore, in some embodiments of the present application, the key - point detection network includes: a first backbone network, a first multi - scale feature fusion module, and a first detection head;
[0102] Among them, the first backbone network is used to extract the first needle - body feature of the tip image;
[0103] The first multi - scale feature fusion module extracts and fuses features of different scales according to the first needle - body feature to obtain the first multi - scale fusion feature;
[0104] The first detection head is used to identify the key points and detection boxes of the tip image according to the first multi - scale fusion feature.
[0105] Exemplarily, refer to Figure 4 , in some embodiments of the present application, the principle of the key - point detection network is as follows:
[0106] 1) After pre - processing the image with a tip, input it into the backbone network of the key - point network to extract the needle - body feature;
[0107] 2) After the Backbone feature extraction, a multi-scale feature fusion (MSFF) module is adopted to improve the network's ability to capture features at different scales. The MSFF module fuses features from different scales through techniques such as pyramid pooling, attention mechanism, transposed convolution, or upsampling to enhance the model's accurate recognition ability of the key point positions.
[0108] 3) After passing through the MSFF module, the fused feature map is fed into the detection head (Head), which is responsible for generating the final key point predictions, namely the detection box (BOX) and the key points (KeyPoints).
[0109] Furthermore, in some embodiments of the present application, the above key point detection network is any model that can identify keys, such as HRnet, yolo-pose, ViTPose, etc. The present application does not limit the specific model used by the key point detection network.
[0110] By extracting and fusing the features of the needle body at different scales, it is ensured that the needle tip information can still be accurately captured under complex backgrounds or tiny structures; at the same time, compared with the semantic segmentation method, the key point detection network has a simple structure and less computational complexity, effectively improving the real-time performance of key point recognition.
[0111] Furthermore, to more clearly describe the solution of the present application, refer to Figure 5 , which is another schematic flowchart of a method for training a needle tip recognition model provided in some embodiments of the present application, specifically including:
[0112] Step 1: For the construction of the model training set, it is divided into two stages. In stage 1, the contour needle tip extraction module generates annotation information based on the first input needle tip image or needle tip video.
[0113] Step 2: Based on the annotation information generated in step 1, an annotation file is formed.
[0114] Step 3: Use the data processing module to manually screen the annotation file, retain the files that meet the annotations, and convert them into a specific data format, such as coco, voc, yolo, etc., to obtain the second data set.
[0115] Step 4: Divide the second data set in step 3 into a training set, a validation set, and a test set to form a model training set.
[0116] Step 5: Based on the model training set, train a key point detection network.
[0117] Step 6: Deploy the key point detection network to the inference module for inferring and predicting the needle tip points.
[0118] Step 7: When subsequent tip images or tip videos are input, input the tip images or tip videos into the inference module, and form a new annotation file with the key point and detection box information obtained by the inference module, instead of using the annotation file obtained by the contour tip extraction module, so as to further obtain the second data set in stage 2 and be used for the training and update of the subsequent key point detection network.
[0119] In summary, a method for training a tip recognition model provided by an embodiment of the present application has the following beneficial effects: adopting deep learning technology, with automatic feature learning; since the tip is an extremely small and specific point, if a target detection model is used, it will not only lead to inaccurate tip detection, but also the manual annotation cost required for training the target detection model for identifying tip targets is higher. However, through the key point detection network, only the tips of the tip images in the data set need to be annotated, effectively reducing the cost required for annotating data; at the same time, when performing data annotation for the first annotation, annotation is performed through contour recognition technology. When the key point detection network is trained with the data obtained from the first annotation, the tip recognition results output by the key point detection network updated iteratively each time can be used as annotations to update the second data, forming a self-iterative annotation mechanism, which reduces the manual annotation cost while improving the model accuracy; finally, the key point detection network requires less computational effort than the semantic segmentation model. Facing real-time tip recognition scenarios, it can improve the recognition real-time performance while ensuring the recognition accuracy.
[0120] Embodiment 2
[0121] Reference Figure 6 , a tip recognition method provided in an embodiment of the present application, includes S104 to S105, specifically:
[0122] S104: Obtain a tip recognition model; wherein, the tip recognition model is trained and obtained by any one of the tip recognition model training methods in the embodiments of the present application.
[0123] S105: Input the tip image to be recognized into the tip recognition model, recognize the key points of the tip image to be recognized, and use the key points as the tip recognition result.
[0124] Further, in some embodiments of the present application, the tip recognition model includes: a second backbone network, a second multi-scale feature fusion module, and a second detection head; inputting the tip image to be recognized into the tip recognition model and recognizing the key points of the tip image to be recognized includes:
[0125] Extract the second needle body feature of the tip image to be recognized through the second backbone network;
[0126] Through the second multi-scale feature fusion module, combine the second needle body feature, extract and fuse features of different scales, and obtain the second multi-scale fusion feature;
[0127] Input the second multi-scale fusion feature into the second detection head to identify the key points of the needle tip image.
[0128] Understandably, in some embodiments of the present application, the needle tip recognition model is a key point detection network as shown in Figure 4 The principle of this key point detection network is as follows:
[0129] 1) After preprocessing the image with the needle tip, input it into the backbone network of the key point network to extract the needle body features;
[0130] 2) After feature extraction by the Backbone, a multi-scale feature fusion (MSFF) module is used to improve the network's ability to capture features at different scales. The MSFF module fuses features from different scales through techniques such as pyramid pooling, attention mechanism, deconvolution, or upsampling to enhance the model's accurate recognition ability of the key point positions;
[0131] 3) After passing through the MSFF module, the fused feature map is sent to the detection head (Head), which is responsible for generating the final key point prediction, that is, the detection box (BOX) and key points (KeyPoints).
[0132] By extracting and fusing the features of the needle body at different scales, it is ensured that the needle tip information can still be accurately captured under complex backgrounds or tiny structures; at the same time, compared with the semantic segmentation method, the key point detection network has a simpler structure and less computational complexity, effectively improving the real-time performance of key point recognition.
[0133] In summary, a needle tip recognition method provided by the embodiments of the present application has the following beneficial effects: By using the trained key point detection network as the needle tip recognition model, compared with the semantic segmentation model, it requires less computational complexity. In the face of real-time needle tip recognition scenarios, it can improve the recognition real-time performance while ensuring the recognition accuracy; compared with traditional image processing, it can handle more complex background changes and has self-adaptive ability and generalization ability.
[0134] Embodiment III
[0135] Reference Figure 7 , a needle tip recognition model training device provided in the embodiments of the present application, includes: a data acquisition module 201, a second data set acquisition module 202, and an iterative update module 203.
[0136] Further, in some embodiments of the present application, the data acquisition module 201 is configured to acquire a plurality of tip images with the same tip orientation to obtain a first data set; the second data set acquisition module 202 is configured to label each tip image in the first data set through contour recognition technology to obtain a second data set; the iterative update module 203 is configured to iteratively train a preset key point detection network according to the currently obtained second data set, and obtain labeled data according to the key point detection network obtained in each iteration to update the second data set until the accuracy of the key point detection network meets a preset condition, and output the finally iteratively obtained key point detection network as a tip recognition model.
[0137] Further, in some embodiments of the present application, the obtaining labeled data according to the key point detection network obtained in each iteration to update the second data set includes: re-acquiring a plurality of tip images, inputting the re-acquired tip images into the key point detection network obtained in the current iteration, obtaining the detection frame and key points corresponding to the tip images; labeling the corresponding tip images according to the detection frame and key points; and updating the second data set used in the current iteration according to the newly labeled tip images.
[0138] Further, in some embodiments of the present application, the labeling each tip image in the first data set through contour recognition technology to obtain a second data set includes: identifying the tip contours in each tip image in the first data set through contour recognition technology; the tip contours include a plurality of contour points; determining a tip recognition judgment condition according to the tip orientation, and determining a first contour point that meets the tip recognition judgment condition from the tip contours; and labeling the corresponding tip images according to the tip contours and the first contour point to obtain the second data set.
[0139] Further, in some embodiments of the present application, the labeling the corresponding tip images according to the tip contours and the first contour point to obtain the second data set includes: calculating the minimum circumscribed rectangle of the tip contours; using the minimum circumscribed rectangle as the detection frame and the first contour point as the key point, and labeling the corresponding tip images according to the detection frame and key points.
[0140] Further, in some embodiments of the present application, determining the needle tip recognition judgment condition according to the needle tip orientation includes: determining the axis orientation with the smallest included angle formed by the needle tip orientation and the coordinate system axis orientation of the needle tip image; when the axis orientation is the positive direction of the vertical axis, determining that the needle tip recognition judgment condition is that the vertical axis coordinate value corresponding to the first contour point is the largest; when the axis orientation is the negative direction of the vertical axis, determining that the needle tip recognition judgment condition is that the vertical axis coordinate value corresponding to the first contour point is the smallest; when the axis orientation is the positive direction of the horizontal axis, determining that the needle tip recognition judgment condition is that the horizontal axis coordinate value corresponding to the first contour point is the largest; when the axis orientation is the negative direction of the horizontal axis, determining that the needle tip recognition judgment condition is that the horizontal axis coordinate value corresponding to the first contour point is the smallest.
[0141] Further, in some embodiments of the present application, the key point detection network includes: a first backbone network, a first multi-scale feature fusion module, and a first detection head; wherein, the first backbone network is used to extract the first needle body feature of the needle tip image; the first multi-scale feature fusion module extracts and fuses features of different scales according to the first needle body feature to obtain a first multi-scale fusion feature; the first detection head is used to identify the key points and detection frames of the needle tip image according to the first multi-scale fusion feature.
[0142] It can be understood that the above device item embodiments correspond to the method item embodiments of the present invention. A needle tip recognition model training device provided by the embodiments of the present invention can implement any one of the method item embodiments of the present invention, that is, the needle tip recognition model training method provided in Embodiment 1.
[0143] In summary, a needle tip recognition model training device provided by the embodiments of the present application has the following beneficial effects: adopting deep learning technology, it has automatic feature learning; since the needle tip is an extremely small and specific point, if a target detection model is used, it will not only lead to inaccurate needle tip detection, but also the manual annotation cost required to train the target detection model for identifying the needle tip target is higher. However, through the key point detection network, only the needle tips of the needle tip images in the dataset need to be annotated, effectively reducing the cost required for annotating data; at the same time, when performing data annotation for the first time, contour recognition technology is used for annotation. When the key point detection network is trained with the data annotated for the first time, the needle tip recognition results output by the key point detection network updated each time can be used as annotations to update the second data, forming a self-iterative annotation mechanism, which reduces the manual annotation cost while improving the model accuracy; finally, the key point detection network requires less computational effort than the semantic segmentation model. In the face of real-time needle tip recognition scenarios, it can improve the recognition real-time performance while ensuring the recognition accuracy.
[0144] Embodiment 4
[0145] Reference Figure 8 This is a tip recognition device provided by an embodiment of the present application, including: a tip recognition model acquisition module 204 and a tip recognition module 205.
[0146] Further, in some embodiments of the present application, the tip recognition model acquisition module 204 is configured to acquire a tip recognition model; wherein, the tip recognition model is trained and acquired by any tip recognition model training device in the embodiments of the present application; the tip recognition module 205 is configured to input the tip image to be recognized into the tip recognition model, recognize the key points of the tip image to be recognized, and use the key points as the tip recognition result.
[0147] Further, in some embodiments of the present application, the tip recognition model includes: a second backbone network, a second multi-scale feature fusion module, and a second detection head; the step of inputting the tip image to be recognized into the tip recognition model and recognizing the key points of the tip image to be recognized includes: extracting second needle body features of the tip image to be recognized through the second backbone network; through the second multi-scale feature fusion module, combining the second needle body features, extracting and fusing features of different scales to obtain second multi-scale fusion features; inputting the second multi-scale fusion features into the second detection head to recognize the key points of the tip image.
[0148] It can be understood that the above device item embodiments correspond to the method item embodiments of the present invention. A tip recognition device provided by the embodiments of the present invention can implement any one of the method item embodiments of the present invention, that is, the tip recognition method provided in Embodiment 2.
[0149] In summary, a tip recognition device provided by an embodiment of the present application has the following beneficial effects: by using the trained key point detection network as the tip recognition model, the required computational complexity is smaller than that of the semantic segmentation model. In the face of real-time tip recognition scenarios, it can improve the recognition real-time performance while ensuring the recognition accuracy; compared with traditional image processing, it can handle more complex background changes and has self-adaptive ability and generalization ability.
[0150] Embodiment 5
[0151] Based on the above embodiments of the tip recognition and tip recognition model training method, another embodiment of the present application provides a tip recognition and tip recognition model training terminal device. The tip recognition and tip recognition model training terminal device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the tip recognition and tip recognition model training method of any embodiment of the present application.
[0152] Exemplarily, in this embodiment, the computer program may be divided into one or more modules. The one or more modules are stored in the memory and executed by the processor to complete the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the tip recognition and tip recognition model training device.
[0153] The tip recognition and tip recognition model training device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The tip recognition and tip recognition model training terminal device may include, but is not limited to, a processor and a memory.
[0154] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the tip recognition and tip recognition model training device, and connects various parts of the entire tip recognition and tip recognition model training device through various interfaces and circuits. The memory may be used to store the computer program and / or modules. The processor realizes various functions of the tip recognition and tip recognition model training device by running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash device, or other volatile solid-state storage devices.
[0155] Embodiment Six
[0156] Based on the embodiments of the above needle tip recognition and needle tip recognition model training method, another embodiment of the present application provides a storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute the needle tip recognition and needle tip recognition model training method of any embodiment of the present application.
[0157] In this embodiment, the above storage medium is a computer-readable storage medium. The computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0158] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above description is only for the specific embodiments of the present application and is not used to limit the protection scope of the present application. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for training a tip recognition model, characterized in that, Including: Collect a number of tip images with the same tip orientation to obtain a first data set; Label each tip image in the first data set through contour recognition technology to obtain a second data set; According to the currently obtained second data set, iteratively train a preset key point detection network, and obtain labeled data according to the key point detection network obtained in each iteration to update the second data set until the accuracy of the key point detection network meets the preset conditions, and output the finally iteratively obtained key point detection network as a tip recognition model.
2. The method for training a tip recognition model according to claim 1, wherein The step of obtaining labeled data according to the key point detection network obtained in each iteration to update the second data set includes: Re-collect a number of tip images, and input the re-collected tip images into the key point detection network obtained in the current iteration to obtain the detection frame and key points corresponding to the tip images; Label the corresponding tip images according to the detection frame and key points; Update the second data set used in the current iteration according to the newly labeled tip images.
3. The method for training a tip recognition model according to claim 1, characterized in that The step of labeling each tip image in the first data set through contour recognition technology to obtain a second data set includes: Identify the tip contours in each tip image in the first data set through contour recognition technology; the tip contours include a number of contour points; Determine the tip recognition judgment condition according to the tip orientation, and determine the first contour point that meets the tip recognition judgment condition from the tip contours; Label the corresponding tip images according to the tip contours and the first contour point to obtain the second data set.
4. The method for training a tip recognition model according to claim 3, wherein, The step of labeling the corresponding tip images according to the tip contours and the first contour point to obtain the second data set includes: Calculate the minimum bounding rectangle of the tip contour; Use the minimum bounding rectangle as the detection frame and the first contour point as the key point, and label the corresponding tip images according to the detection frame and key points.
5. The method for training a tip recognition model according to claim 3, wherein The step of determining the tip recognition judgment condition according to the tip orientation includes: Determine the axis orientation with the smallest included angle formed by the tip orientation and the coordinate axis orientation of the tip image; When the axis orientation is the positive direction of the vertical axis, it is determined that the tip recognition judgment condition is that the vertical axis coordinate value corresponding to the first contour point is the largest; When the axis orientation is the negative direction of the vertical axis, it is determined that the tip recognition judgment condition is that the vertical axis coordinate value corresponding to the first contour point is the smallest; When the axis orientation is the positive direction of the horizontal axis, it is determined that the tip recognition judgment condition is that the horizontal axis coordinate value corresponding to the first contour point is the largest; When the axis orientation is the negative direction of the horizontal axis, it is determined that the tip recognition judgment condition is that the horizontal axis coordinate value corresponding to the first contour point is the smallest.
6. A method for training a tip recognition model according to any one of claims 1 to 5, characterized in that, The key point detection network includes: a first backbone network, a first multi-scale feature fusion module, and a first detection head; Among them, the first backbone network is used to extract the first needle body feature of the tip image; The first multi-scale feature fusion module extracts and fuses features of different scales according to the first needle body feature to obtain a first multi-scale fusion feature; The first detection head is used to identify the key points and detection frames of the tip image according to the first multi-scale fusion feature.
7. A method for tip recognition, characterized in that, It includes: Obtain a tip recognition model; wherein, the tip recognition model is trained and obtained by the tip recognition model training method according to any one of claims 1 to 6. Input the tip image to be recognized into the tip recognition model, identify the key points of the tip image to be recognized, and use the key points as the tip recognition result.
8. A tip recognition method according to claim 7, characterized in that, The tip recognition model includes: a second backbone network, a second multi-scale feature fusion module, and a second detection head; inputting the tip image to be recognized into the tip recognition model and identifying the key points of the tip image to be recognized includes: Extract the second needle body feature of the tip image to be recognized through the second backbone network. Through the second multi-scale feature fusion module, combine the second needle body feature, extract and fuse features of different scales, and obtain a second multi-scale fusion feature. Input the second multi-scale fusion feature into the second detection head to identify the key points of the tip image.
9. A training device for a tip recognition model, characterized in that, It includes: A data acquisition module, a second data set acquisition module, and an iterative update module. Among them, the data acquisition module is used to collect a number of tip images with the same tip orientation to obtain a first data set. The second data set acquisition module is used to label each tip image in the first data set through contour recognition technology to obtain a second data set. The iterative update module is used to iteratively train a preset key point detection network according to the currently obtained second data set, and obtain labeled data according to the key point detection network obtained each time to update the second data set until the accuracy of the key point detection network meets the preset conditions, and output the finally iteratively obtained key point detection network as the tip recognition model.
10. A tip recognition device, characterized in that, It includes: A tip recognition model acquisition module and a tip recognition module. Among them, the tip recognition model acquisition module is used to obtain a tip recognition model; wherein, the tip recognition model is trained and obtained by the tip recognition model training device according to claim 9. The tip recognition module is used to input the tip image to be recognized into the tip recognition model, identify the key points of the tip image to be recognized, and use the key points as the tip recognition result.
Citation Information
Cited By
Pentahedral needle tip sharpness evaluation system and method based on deep learning
CN121121086A