A guide wire tip detection method based on YOLOV4Tiny network

By optimizing the feature extraction and fusion methods of the YOLOV4Tiny network, the problem of low accuracy in guidewire tip detection was solved, achieving high-precision and high-robust guidewire tip detection.

CN116051951BActive Publication Date: 2025-11-28BEIJING INFORMATION SCI & TECH UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211300801.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-24
Publication Date
2025-11-28
Estimated Expiration
2042-10-24

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in detecting the end of the guidewire, making it difficult to achieve precise tracking, especially when the end of the guidewire is a small target, which is prone to being missed.

Method used

Based on the YOLOv4Tiny network, optimizations were made to improve the feature extraction capability and detection accuracy of small targets by improving the residual structure in the feature extraction network, adding an attention mechanism module, and hybrid dilated convolution.

Benefits of technology

It achieves precise detection of the guidewire tip, improves detection accuracy and robustness, increases average precision by 21.59%, achieves average accuracy of 97.6%, and has a detection error of less than 5%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116051951B_ABST
    Figure CN116051951B_ABST
Patent Text Reader

Abstract

The application provides a guide wire end detection method based on a YOLOV4Tiny network, comprising: acquiring guide wire image data, performing feature extraction on the guide wire image data, wherein the residual structure in the feature extraction network is divided into multiple groups, a plurality of filters are used to perform feature extraction on the guide wire image data input into the residual structure, wherein one filter performs feature extraction on the guide wire image data of one group of the residual structure, and the extracted feature map and the guide wire image data of the next group of the residual structure are input into the next filter for feature extraction until the feature extraction of the guide wire image data of each group of the residual structure is completed, the feature maps extracted from the guide wire image data of each group of the residual structure are input into a 1*1 convolution filter in a tensor connection mode to obtain a fused feature map. The application improves the feature extraction capability, detection accuracy and robustness of the algorithm for the guide wire end.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of image recognition, in particular to a guide wire tip detection method based on a YOLOV4Tiny network. BACKGROUND

[0002] Percutaneous interventional surgery is to insert a guide wire through the blood vessels and send it to the lesion site in the heart cavity under the condition of contrast fluoroscopy to obtain lesion information and treat the lesion. The operation process is complex and the operation technique is difficult. At present, the traditional percutaneous interventional surgery is difficult and the treatment time is long. The radiation produced by the contrast agent used in the operation can cause potential harm to the doctors and patients. The medical application of robot technology provides accurate position information of the catheter tip for doctors, reducing the use of contrast agent. In the process of interventional surgery, the accuracy of the operation is higher and the safety is stronger. In order to realize the accurate tracking of the position of the guide wire tip in percutaneous interventional surgery, different institutions at home and abroad have done a lot of research. W. Frumkin et al. of Winchester Medical Center Research Department designed Amigo remote navigation system, which can determine whether the catheter tip successfully reaches the specified position through fluorescent image tracking and endocardial electrogram. Experiments show that it is feasible to track the catheter tip through fluorescent image tracking, but the pacemaker threshold and endocardial electrogram must be manually controlled. Justin A. Borgstadt1 et al. of the University of Wisconsin designed a double particle filter, which can real-time locate the position of the catheter in minimally invasive surgery through a fluorescent stereoscopic imaging system and an electromagnetic pose sensor. However, because the catheter tip target is too small, it is difficult to accurately position the catheter tip, and this method is highly dependent on high-precision sensors and is easily disturbed by white noise in the operation, which has poor robustness. Peng Wang et al. of Siemens Research Department do not rely on 3D information to track the motion of the guide wire in a single-view fluoroscopy sequence, with a tracking accuracy of less than 0.4 mm. This algorithm has high accuracy in guide wire positioning, but it is difficult to position the guide wire due to the large range of activity of the guide wire tip. Adrian Barbu et al. of the Princeton Siemens Research Center designed a detection and segmentation curve model to truly use medical fluorescent video to track and locate the guide wire tip, with a tracking accuracy of 1 mm and a real-time speed of 1 frame / s. The accuracy and real-time speed of the algorithm have room for optimization. Chang et al. of the University of London proposed a deformable b-spline fitting method based on image processing, which can accurately display the shape of the catheter and enhance the robustness of the algorithm by using regional probability algorithm for fitting without relying on gradient. However, it is easy to miss detection when the guide wire moves quickly. Ihsan Ullah et al. of Kyungpook National University in South Korea proposed a convolutional neural network (CNN) based image tracking method, which can real-time track the guide wire tip in X-ray video sequence through detection and segmentation network, with a real-time speed of 19 ms. However, the accuracy of the algorithm is highly dependent on the previous detection results, and the detection accuracy of the next frame will decrease significantly after the detection of the previous frame fails. In 2020, Alexey et al. proposed YOLOV4 target detection algorithm, which can alleviate the problem of gradient information disappearance in neural network, increase the input resolution of the network, and make the detection of small targets more accurate. However, YOLOV4 slows down the detection speed because it combines multiple modules, making it difficult to detect in real time.In March 2022, Yan et al. of Nanjing University of Aeronautics and Astronautics proposed a method of adding attention mechanism CA to improve the algorithm, which increased the detection accuracy of ground small target from 42.3% to 94.6%, and the speed reached 58.8 frames per second. In November 2020, Jiang et al. of Northeastern University proposed developing YOLOV4Tiny algorithm on embedded devices for target detection. The detection speed can reach 294FPS, and the real-time effect is very good, but the average detection accuracy can only reach 38%. SUMMARY

[0003] In order to solve the technical problem of low guide wire tip detection rate in the prior art, one object of the present application is to provide a guide wire tip detection method based on YOLOV4Tiny network, the method comprising the following method steps:

[0004] Step S1, acquiring guide wire image data,

[0005] Step S2, feature extraction is performed on the guide wire image data, wherein the residual structure in the feature extraction network is divided into multiple groups, and a plurality of filters are used to extract features from the guide wire image data input into the residual structure,

[0006] Wherein, one filter extracts features from the guide wire image data of one group of residual structures, and the extracted feature map and the guide wire image data of the next group of residual structures are input into the next filter for feature extraction until the guide wire image data of each group of residual structures is completely extracted,

[0007] The feature maps extracted from the guide wire image data of each group of residual structures are tensor-connected and input into a 1x1 convolution filter to obtain a fused feature map.

[0008] Step S3, input the feature map obtained in step S2 into the feature fusion network to fuse the feature maps of different sizes;

[0009] Step S4, input the fused feature map of step 3 into the feature detection network to obtain the pixel coordinate value of the guide wire tip.

[0010] Preferably, a channel attention module is provided after step S2 to perform weighting processing on the fused feature map.

[0011] Preferably, the attention module comprises a global average pooling layer, a 1D convolution with a convolution kernel K, and an activation function.

[0012] Preferably, in step S2, a plurality of filters extract features from the guide wire image data input into the residual structure by mixing dilated convolution.

[0013] The application provides a guide wire tip detection method based on a YOLOV4Tiny network, realizes accurate detection of a guide wire tip, and improves feature extraction capability, detection accuracy and robustness of an algorithm for small targets (guide wire tips) by optimizing residual structures in a feature extraction network, adding an attention mechanism module and mixed dilated convolution.

[0014] The guide wire tip detection method based on the YOLOV4Tiny network provided by the application is based on a Yolov4Tiny network architecture, first optimizes residual structures in a feature extraction network, simultaneously adds an attention mechanism module, improves feature extraction capability and detection accuracy of an algorithm for small targets on the premise of not increasing computational complexity, and expands a receptive field of an image on the premise of not reducing image resolution and not changing relative spatial positions of pixels by adding a mixed dilated convolution network. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the specific embodiments of the application or the technical solutions in the prior art, the drawings needed in the specific embodiments or the prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0016] Figure 1 A schematic diagram of a YOLOV4Tiny network structure is schematically shown.

[0017] Figure 2 A schematic diagram of a residual structure in the feature extraction network after optimization of the application is shown.

[0018] Figure 3 A structural schematic diagram of the attention module of the application is shown.

[0019] Figure 4 A schematic diagram of dilated convolution kernels and ordinary convolution kernels is shown.

[0020] Figure 5 A feature map obtained by the mixed dilated convolution under different dilated rates of the application is shown.

[0021] Figure 6 A relationship between a loss value of a training set, a loss value of a test set and an iteration number in an embodiment of the application is shown.

[0022] Figure 7 A chessboard corner point recognition schematic diagram in an embodiment of the application is shown.

[0023] Figure 8 A precision recall rate curve schematic diagram is shown.

[0024] Figure 9 A schematic diagram showing a comparison of the recall and measurement accuracy of the YOLOv4Tiny network of the present invention with those of the YOLOv4Tiny network before the improvement is presented.

[0025] Figure 10 A schematic diagram of the detection results at the end of the guidewire is shown. Detailed Implementation

[0026] To make the above and other features and advantages of the present invention clearer, the invention will be further described below with reference to the accompanying drawings. It should be understood that the specific embodiments given herein are for the purpose of explanation to those skilled in the art and are exemplary only, not restrictive.

[0027] Detection of the end of the guidewire in interventional surgery is crucial for achieving precise surgical control and is essential for ensuring the safety of the procedure. To address the low accuracy of guidewire end detection in existing technologies, and considering the small and thin tip of the interventional guidewire, as well as the algorithm's requirements for feature extraction capability and detection accuracy at small target guidewire ends, this invention proposes a guidewire end detection method based on the YOLOv4Tiny network. The algorithm is improved primarily through three aspects: optimization of the feature extraction network, addition of an attention mechanism, and hybrid dilated convolution.

[0028] This invention proposes a method for detecting the end of a guidewire based on a YOLOv4Tiny network, comprising the following steps:

[0029] Step S1: Obtain guidewire image data.

[0030] Step S2: Construct the YOLOv4Tiny network and optimize it.

[0031] like Figure 1 The diagram shown illustrates the YOLOv4Tiny network structure. According to an embodiment of the present invention, the YOLOv4Tiny network comprises three parts: a feature extraction network, a feature fusion network, and a feature detection network.

[0032] The feature extraction network extracts feature maps of different sizes from the input samples at multiple scales. The feature fusion network merges feature maps of different sizes to enhance the detection of small targets. Finally, the target coordinate information is obtained through the feature detection network.

[0033] According to an embodiment of the present invention, feature extraction is performed on guidewire image data using a feature extraction network.

[0034] Since the guide wire and the blood vessel outer contour are linear structures, the differences in some structure pictures are small, which is easy to cause false detection, the application distinguishes the two by increasing fine granularity division, enhances the extraction of guide wire features, optimizes the residual structure in the feature extraction network, such as Figure 2 As shown in the schematic diagram of the optimized residual structure in the feature extraction network of the application, the residual structure in the feature extraction network is divided into multiple groups, and the guide wire image data of the input residual structure is extracted by multiple filters.

[0035] Among them, one filter extracts the feature map of the guide wire image data of one group of residual structures, and inputs the extracted feature map and the guide wire image data of the next group of residual structures into the next filter for feature extraction, until the guide wire image data of each group of residual structures is completely extracted.

[0036] The feature map extracted from the guide wire image data of each group of residual structures is input into a 1x1 convolution filter for tensor connection to obtain a fused feature map.

[0037] In specific embodiments, in order to improve the detection of the feature map, the input features of the residual structure in the original feature extraction network are divided into four groups: x1, x2, x3, x4, a smaller group of convolution filters (three in the embodiment) is used to replace the 64-channel 3x3 convolution filter, each filter is a 3x3 filter with 16 channels.

[0038] Each 3x3, 16 filter extracts a feature map from the guide wire image data of a group of residual structures (for example, x2), and inputs the extracted feature map and the guide wire image data of the group of residual structures (for example, x3) into the next 3x3, 16 filter for feature extraction, and the process is repeated 3 times until the guide wire image data of each group of residual structures is completely extracted.

[0039] After processing the feature map, all the groups of feature maps are sent to a 1x1 convolution filter for tensor connection to obtain the feature map.

[0040] The application divides the original channel into four branches, respectively performs convolution, realizes fine-grained feature fusion through inter-block fusion and channel splicing, enhances the feature extraction of small targets, and reduces the false detection probability caused by similar feature confusion.

[0041] The optimized feature extraction network parameters are shown in Table 1, and the parameter corresponding meaning [output channel, kernel size, stride, padding].

[0042] Table 1: Optimized feature extraction network parameters

[0043]

[0044]

[0045] Due to the characteristics of the guidewire tip being small, thin and sharp, the pixels of the feature extraction part in the picture are few, and the pixels of the background part are too many. In a 2448*2048 pixel picture, the guidewire tip usually has only 10*10 pixels, or even less, which leads to the system having to process a large amount of meaningless parameters, reducing the calculation speed and affecting the real-time detection. Moreover, the small target contains few pixels, resulting in less original information provided, and it is extremely easy to miss detection, so it is necessary to try to select a small and shallow network to extract features.

[0046] Based on this, the feature fusion of the feature map extracted by the feature extraction network is carried out, in order to avoid the calculation burden caused by the feature redundancy of the extracted feature map channel, leading to the increase of detection error, an attention module (ECA) is added after the residual module, the redundant features are ignored, and the fused feature map is weighted processed, the ability of accurate positioning of the small target (guidewire tip) is improved.

[0047] As shown in the structural schematic diagram of the attention module of the application, Figure 3 The attention module includes a global average pooling layer (Global Average Pooling, GAP), a 1D convolution with a convolution kernel K and an activation function. In order to reduce the number of channels and reduce the amount of calculation, the global average pooling layer is used to compress the feature map with an original size of C*H*W to C*1*1. Then a 1*k one-dimensional convolution is used to share the learning parameters to make each channel share the weight information, realizing cross-channel information interaction. This way not only avoids dimension reduction, but also greatly reduces the complexity of the model.

[0048] The size of the convolution kernel K and the number of channels C can be obtained by the following relationship:

[0049]

[0050] According to the embodiment of the application, in the process of extracting features from the guidewire image data of the input residual structure by multiple filters in the foregoing, the guidewire image data of the input residual structure is extracted by a mixed dilated convolution.

[0051] Common neural networks usually use down-sampling to increase the receptive field (Receptive Filed), and then use up-sampling to compensate to restore the original size of the image. The essence of down-sampling is extraction, and the purpose is to reduce the dimension of features and retain effective information to obtain more feature maps of different sizes. Since the guidewire tip is a small target, the pixel information is very small, and the information is easily lost in the process of down-sampling, resulting in the loss of the target.

[0052] This invention replaces ordinary convolution with hybrid dilated convolution. By changing the dilation rate, it learns receptive maps at different scales, avoiding the use of downsampling and reducing the probability of missing small objects. It increases the receptive field while maintaining the size of the feature map. The receptive field refers to the coverage area of ​​pixels in the current-sized feature map on the original image. The larger the receptive field, the more semantic information from the original image is contained in the feature map. The formula for calculating the receptive field is as follows:

[0053] RF i+1 =(2 i+2 -1)×(2 i+2 +1),

[0054] Where i+1 is the layer number, RF i+1 It is the receptive field of the feature map of the (i+1)th layer.

[0055] like Figure 4 The diagram shows a comparison between dilated convolution kernels and ordinary convolution kernels. Figure 4 The 3x3 color block on the left side of the middle ( Figure 4 The gray blocks in the diagram represent ordinary convolution kernels. Figure 4 The right side is filled with a 0 color block ( Figure 4 The dilated convolution kernel (containing white blocks) expands the size of a regular convolution by padding with zeros, without altering the feature map units involved in the computation. Dilated convolution can acquire more feature maps of varying sizes, increasing the receptive field, while avoiding downsampling. However, a large number of holes (zero blocks) can lead to the loss of target information. To address this drawback, hybrid dilated convolution uses a zigzag, cyclic dilation rate to cover all holes (zero blocks).

[0056] like Figure 5 The image shown is a feature map obtained by hybrid dilation convolution under different dilation rates according to the present invention. Figure 5 In (a), (b), and (c), the outputs are all feature maps of the same size, but due to different dilation rates, the receptive fields of the output images are different. In (a), the dilation rate rat = 1, and the original convolution kernel is 3×3, which is a regular convolution. As shown in the figure, the receptive field is relatively small. In (b), the dilation rate rat = 2, and the convolution kernel after adding holes (0 color blocks) is 5×5. The corresponding receptive field can be calculated as: (2 (a +2) (c) -1 = 7. When the dilation rate rat = 5, the kernel size can be changed to 11×11, and the receptive field increases to (2) (a +2))-1=127. With the number of relevant parameters unchanged, a larger receptive field does not have the cost of additional calculation. The use of different dilated rate of empty (0 patch) convolution enhances the receptive field extraction to different scale features instead of the original down sampling, avoiding the precision loss in the process of down sampling. The purpose of realizing higher precision multi-scale feature extraction and fusion is achieved, and the calculation amount is not increased.

[0057] Step S3, input the feature map obtained in step S2 into the feature fusion network, and fuse the feature maps of different sizes.

[0058] Step S4, input the feature map fused in step S3 into the feature detection network, and obtain the pixel coordinate value of the guidewire tip.

[0059] Embodiment

[0060] The embodiment of the application constructs a data set to test the effectiveness of the guidewire tip detection method based on the YOLOV4Tiny network provided by the application.

[0061] Data acquisition and model training

[0062] During the experimental test, the experimental system built includes a data acquisition module, a motion module for controlling the guidewire, an acrylic blood vessel model, a Medtronic medical guide catheter EBU3.5, and a guidewire.

[0063] The data acquisition module uses a Hikvision MV-CA050-10GC industrial camera to obtain the target guidewire sample, uses a UR5 robot arm to control the guidewire motion, and uses an acrylic blood vessel model and a medical guide wire to simulate a medical environment.

[0064] The medical guide wire is transported to the specified target by the UR5 robot arm industrial control guidewire motion, and the CMOS is used to obtain the target picture sample from the overhead angle, to obtain the original data. The original data is labeled and sample set is made using the lambelImg labeling software, so that each picture obtains a corresponding xml file, which records the picture path, picture color channel, size, and size of the labeling box. The labeled xml data set is converted into a txt data required for training, and the data set is divided into a test set and a training set in a ratio of 9:1. The training set data is input into the improved YOLOv4Tiny network, and the parameters of the network are adjusted through the learning of the training set. Finally, the test set data is input into the adjusted network, and the pixel coordinate value of the target point in the picture is obtained.

[0065] To better train the proposed method, images are preprocessed to expand the database and increase the number of samples using Mosaic data augmentation. Instance segmentation is used to crop the images proportionally and then reassemble them into 416×416 images. The number of small targets in each sample is randomly increased to increase the data for matching target samples. This improves the detection accuracy of the algorithm. Figure 6 The diagram illustrates the relationship between the loss value of the training set, the loss value of the test set, and the number of iterations in one embodiment of the present invention. It shows that the loss value is 3 before the first iteration, decreases rapidly with increasing iteration count, and changes slowly after 200 iterations. After 250 iterations, the loss value gradually approaches 0, indicating that the network training has achieved the expected results.

[0066] To accurately obtain the transformation relationship between image plane coordinates and spatial coordinates, high-precision calibration and orientation of the camera's interior orientation parameters, exterior orientation parameters, and distortion coefficients are required. First, by taking multi-station photos of the control field and high-precision orientation target in space, the camera's interior orientation parameters and distortion coefficients are calibrated using a bundle adjustment algorithm. After parameter optimization, the image plane error of the control points is less than 1 / 35 pixel. Second, using a checkerboard pattern and the provided known spatial corner coordinate constraints, a corner recognition algorithm is used to extract the image plane coordinates of corresponding corner points in the image. Using the known spatial coordinates and their corresponding image plane coordinates, a resection algorithm is used to perform spatial pose orientation of the camera's exterior orientation parameters, such as... Figure 7 The diagram shown is a schematic representation of checkerboard corner point recognition in one embodiment of the present invention. Finally, under the condition of obtaining an image by camera orthogonal imaging, the image positioning coordinates of the guidewire end are mapped from the image coordinate system to the spatial coordinate system using the already obtained camera interior orientation parameters, distortion coefficients, and exterior orientation parameters, to obtain the spatial positioning coordinates of the guidewire end.

[0067] Analysis of Experimental Results

[0068] The experiment evaluated the model performance using tracking accuracy metrics such as precision (P), recall (R), average precision (AP), and detection speed (FPS).

[0069] The theoretical formulas for precision (P) and recall (R) derived from the confusion matrix are as follows:

[0070]

[0071]

[0072] TP represents the number of correct detections, FP represents the number of false detections, and FN represents the number of false negatives. Theoretically, precision and recall are inversely proportional (PR decrease curve); increasing precision will decrease recall. A balance needs to be found through repeated training, i.e., the confidence threshold. Furthermore, average precision (AP) is expressed as:

[0073]

[0074] Average precision (AP) is the prediction accuracy of a certain category at different confidence thresholds, while detection speed (FPS) represents the number of frames that the model can detect per second, and is often used to measure the speed of an algorithm.

[0075] To test whether the improved algorithm is more effective than the original YOLOV4Tiny algorithm in detecting the end of the guidewire, four methods were trained separately, and the effectiveness of the improvement was judged by the mean precision (AP). The higher the AP value, the better the detection effect. The AP value is the area enclosed by the precision-recall curve (PR) formed by recall (R) on the x-axis and precision (P) on the y-axis.

[0076] like Figure 8 The diagram shows the precision and recall curves. Figure 8 In the table, Precision4 shows the relationship between precision and recall obtained from the original algorithm test, Precision3 shows the relationship between precision and recall obtained after improving the feature extraction network, Precision2 shows the relationship between precision and recall obtained after introducing hybrid dilated convolution, and Precision1 shows the relationship between precision and recall obtained after adding an attention mechanism.

[0077] The results show that the improved methods all achieve an improvement in average accuracy compared to the original network. Adding an attention mechanism resulted in the largest improvement in average accuracy (AP) compared to the original network, increasing it by 17.47%. This was followed by a feature extraction network that replaced the original convolutions with hybrid dilated convolutions, achieving an improvement of 4.03%. By adding these methods, the detection accuracy of the original YOLOv4Tiny network was significantly improved. Figure 8 The Precision4 in the original text was upgraded to Precision1, and the average accuracy (AP) was improved from 76.01% to 97.6%. The results show that the improved solution of the present invention is effective.

[0078] Testing revealed that the present invention can improve the success rate of guidewire end position detection. To obtain the best detection performance, all the improved parts of the present invention were integrated into a YOLOv4Tiny network. The guidewire end detection results obtained after training were compared with the original YOLOv4Tiny detection results.

[0079] likeFigure 9 The comparison diagram of the recall rate and the measurement accuracy of the YOLOv4Tiny network of the application and the recall rate and the measurement accuracy of the YOLOv4Tiny network before improvement is shown, wherein (a) is the recall rate of the YOLOv4Tiny network before improvement, (b) is the recall rate of the YOLOv4Tiny network of the application, (c) is the measurement accuracy of the YOLOv4Tiny network before improvement, and (d) is the measurement accuracy of the YOLOv4Tiny network of the application.

[0080] By increasing Figure 9 As shown in the comparison, the recall rate (R) of the original network is increased from 68.09% to 89.36%. As can be seen from the obtained results, the improved algorithm reduces the number of missed detections (FN), increases the number of correct detections (TP), and improves the recall rate. The accuracy of the original YOLOv4Tiny network is increased from 80% to 97.67% by improvement. As can be seen from the obtained results, the improved algorithm reduces the number of false detections (FP), increases the number of correct detections (TP), and improves the accuracy.

[0081] Table 2 Test results of YOLOv4Tiny network before and after improvement

[0082]

[0083] As can be seen from Table 2, the detection accuracy AP of the improved algorithm can reach 97.6%, which is increased by 21.59% compared with the original YOLOv4Tiny algorithm. It is proved that the target detection algorithm is effective and feasible for accurate detection of small targets (the tip of the guide wire). It can effectively solve the problem that small target pixels are too few and easy to be missed, and can meet the demand for accurate detection of small targets.

[0084] In order to verify the effectiveness of the improved YOLOv4Tiny algorithm, the UR5 mechanical arm is used to control the guide wire to form a motion curve. As shown in Figure 10 The guide wire tip detection result diagram is shown in Figure 10 As can be seen from the figure, the detection trajectory of the guide wire tip by the improved algorithm is closer to the real trajectory, and the maximum average error percentage before and after the algorithm is improved is reduced from 15% to 5%, and the error rate is obviously reduced.

[0085] The guide wire tip detection method based on the YOLOV4Tiny network provided by the application improves the accuracy of the algorithm before improvement by 21.59%, and the average accuracy reaches 97.6%. The detection error of the guide wire tip in the simulated blood vessel trajectory is less than 5%. The application provides an effective method for detecting the tip of the guide wire in interventional surgery, and has a wide application prospect in the field of biomedical robot measurement.

[0086] Although the embodiments of the present application have been shown and described above, it is understood that the above-described embodiments are exemplary and are not to be construed as limiting the present application, and that variations, modifications, substitutions and changes can be made by those skilled in the art without departing from the scope of the present application.

Claims

1. A guide wire tip detection method based on YOLOV4Tiny network, characterized in that, The method comprises the following method steps: Step S1, acquiring guidewire image data, Step S2, performing feature extraction on the guidewire image data, wherein the residual structure in the feature extraction network is divided into multiple groups, and a plurality of filters are used to perform feature extraction on the guidewire image data input into the residual structure, Wherein, one filter performs feature extraction on the guidewire image data of one group of residual structures, and the extracted feature map and the guidewire image data of the next group of residual structures are input into the next filter for feature extraction, until the feature extraction of the guidewire image data of each group of residual structures is completed, The feature maps extracted from the guidewire image data of each group of residual structures are tensor-connected and input into a 1x1 convolution filter to obtain a fused feature map; An attention module is added after the residual module to ignore redundant features and perform weighted processing on the fused feature map; the attention module comprises a global average pooling layer, a 1D convolution with a convolution kernel K, and an activation function; The size of the convolution kernel K and the number of channels C can be obtained by the following relationship: ; In the process of feature extraction on the guidewire image data input into the residual structure, the mixed dilated convolution is used to perform feature extraction on the guidewire image data input into the residual structure; Step S3, inputting the feature map obtained in step S2 into a feature fusion network to fuse feature maps of different sizes; Step S4, inputting the fused feature map in step 3 into a feature detection network to obtain pixel coordinate values of the guidewire tip.