Improved yolov8n model training method and device for surgical knotting action video recognition
Through the improved yolov8n model, combined with dynamic convolution module and ghost network, the problem of high cost of surgical knotted motion capture hardware in the prior art is solved, and low-cost real-time surgical motion capture and real-time feedback of wrong actions is achieved.
Patent Information
- Application Number
- CN202510197082.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
In the prior art, the hardware cost of motion capture technology for surgical knotting is high, making it difficult to achieve low-cost real-time surgical motion capture and real-time feedback of wrong actions.
Through computer vision technology and deep learning technology, an improved yolov8n model was designed, combining dynamic convolution modules and ghost networks for surgical knotting action video recognition. This model trains the improved yolov8n model by obtaining the surgical action video training set, and obtains the trained model to achieve motion capture.
It realizes the capture of surgical actions at low computational cost, and provides a low-cost surgical knotting motion capture detector, which can perform motion recognition and feedback in real-time situations.
Smart Images

Figure CN120126050A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image recognition technology, and in particular, to an improved yolov8n model training method for surgical knot-tying action video recognition, an improved yolov8n model training device for surgical knot-tying action video recognition, and a surgical knot-tying action video recognition method. Background Art
[0002] In clinical practice, the knot-tying action is one of the basic skills often used by doctors. This usually involves using surgical sutures to ligate or suture tissues or blood vessels to stop bleeding or connect tissues. The correct execution of knot-tying is crucial for the success of the surgery and the recovery of the patient. Especially for the intern group, in a study, it was confirmed that there is still a large gap between the knot-tying level of interns and their mentors. The average tensile strength of knots tied by interns is considered to be greater than that of their mentors, and there is a greater probability of causing wound dehiscence. Therefore, developing a lightweight network that is convenient to deploy is of great significance for real-time surgical motion capture and real-time feedback on incorrect actions. Currently, the mainstream motion capture technologies include optical motion capture technology, inertial motion capture technology, and laser motion capture technology. However, the hardware costs of these motion capture technologies are relatively high. Taking the most commonly used optical motion capture as an example, an optitrack motion capture system with only 8 PrimeX41 infrared cameras costs more than $60,000, which is very expensive for most application scenarios. In recent years, with the development of computer vision, using deep learning methods to complete tasks such as image classification, object detection, and image segmentation has achieved good results. Based on this, the present invention designs a low-cost surgical knot-tying motion capture detector through computer vision technology and deep learning technology, and realizes the capture of surgical motions at low computational cost.
[0003] Therefore, it is desirable to have a technical solution to solve or at least mitigate the above deficiencies of the prior art. Summary of the Invention
[0004] The purpose of the present invention is to provide an improved yolov8n model training method for surgical knot-tying action video recognition to at least solve one of the above technical problems.
[0005] The present invention provides the following solutions:
[0006] According to one aspect of the present invention, there is provided an improved yolov8n model training method for surgical knot-tying action video recognition, and the improved yolov8n model training method for surgical knot-tying action video recognition includes:
[0007] Obtain a surgical motion video training set;
[0008] Obtain an improved YOLOv8n model, where the improved YOLOv8n model includes a dynamic convolution module and a GhostNet;
[0009] Train the improved YOLOv8n model with the surgical action video training set to obtain the trained YOLOv8n model.
[0010] Optionally, the obtaining of the improved YOLOv8n model includes:
[0011] Modify the bottleneck layer in the c2f layer of the YOLOv8n model, change the depthwise conv layer in the bottleneck layer to n Ghost modules, and change the cv2 layer in the bottleneck layer to a dynamic convolution layer.
[0012] Optionally, the output of the bottleneck layer in the modified c2f layer is expressed as follows:
[0013]
[0014] * represents a convolution operation, ω primary,k is the k-th standard convolution kernel, let ω k be the k-th dynamic convolution kernel, there are K dynamic convolution kernels in total, α k is the dynamically generated weight.
[0015] Optionally, the comprehensive loss function of the improved YOLOv8n model is:
[0016] ζ = λ cls ζ cls + λ loc ζ loc + λ pose ζ pose ;
[0017] where ζ cls refers to the classification loss, which measures the gap between the predicted class and the true class, and uses cross-entropy loss; ζ loc refers to the localization loss, which measures the gap between the predicted bounding box and the true bounding box; ζ pose refers to the pose estimation loss, which measures the gap between the predicted key point positions and the true key point positions; uses mean squared error, λ cls 、λ loc 、λ pose These are the weight hyperparameters corresponding to the loss terms, used to balance the contributions of different loss terms to the total loss.
[0018] Optionally, the calculation formula of the ζ cls classification loss is as follows:
[0019]
[0020] Among them, N is the number of samples, C is the number of categories, is the sample and the true label belonging to category c, and the probability of category c predicted by the model.
[0021] Optionally, the ζ loc The calculation formula of the positioning loss is as follows:
[0022]
[0023] Among them, is the sample and the true bounding box, and the bounding box predicted by the model.
[0024] Optionally, the ζ pose The calculation formula of the pose estimation loss is as follows:
[0025]
[0026] Among them, K is the number of key points, is the sample and the true position of the j-th key point, and the position of the j-th key point predicted by the model.
[0027] Optionally, the surgical action video training set includes knot-tying actions in daily scenarios and also includes knot-tying actions during actual surgical procedures.
[0028] This application also provides an improved yolov8n model training device for surgical knot-tying action video recognition. The improved yolov8n model training device for surgical knot-tying action video recognition includes:
[0029] The training set acquisition module is used to acquire the surgical action video training set;
[0030] The model acquisition module, and the model acquisition module is used to acquire the improved yolov8n model. The improved yolov8n model includes a dynamic convolution module and a ghost network;
[0031] The training module, and the training module is used to train the improved yolov8n model through the surgical action video training set to obtain the trained yolov8n model.
[0032] This application also provides a surgical knot-tying action video recognition method. The surgical knot-tying action video recognition method includes:
[0033] Obtain the improved yolov8n model for surgical knot-tying action video recognition as described above;
[0034] Obtain the video to be recognized;
[0035] Input the video to be recognized into the improved yolov8n model for surgical knot-tying action video recognition to obtain the recognition result.
[0036] The training method of the improved yolov8n model for surgical knot-tying action video recognition in this application designs a low-cost yolov8n model for surgical knot-tying action capture through computer vision technology and deep learning technology, and realizes the capture of surgical actions at low computational cost. Description of the Drawings
[0037] Figure 1 is a schematic flowchart of the training method of the improved yolov8n model for surgical knot-tying action video recognition in an embodiment of this application;
[0038] Figure 2 is a structural diagram of the lightweight network model - GD-Det for capturing and quantifying surgical actions using key point detection technology provided in an embodiment of this application;
[0039] Figure 3 is a diagram of the surgical action dataset in an embodiment of this application;
[0040] Figure 4 is a schematic diagram of data annotation in an embodiment of this application;
[0041] Figure 5 is a schematic diagram of other lightweight modules in an embodiment of this application;
[0042] Figure 6 is a schematic diagram of the ghost convolution module in an embodiment of this application;
[0043] Figure 7 is a schematic diagram of the dynamic convolution module in an embodiment of this application;
[0044] Figure 8 is a composition diagram of the C2f-GhostDynamic convolution layer in an embodiment of this application;
[0045] Figure 9 is a comparison diagram of the number of parameters and GFLOPs after the comparative experiment and ablation experiment in an embodiment of this application;
[0046] Figure 10 is a comparison diagram of precision, recall rate, and loss function after the comparative experiment and ablation experiment in an embodiment of this application;
[0047] Figure 11 It is a schematic diagram of the prediction result of the GD-Det model in an embodiment of the present application.
[0048] Figure 12 It is a schematic structural diagram of the C2f module in the prior art.
[0049] Figure 13 A comparison schematic diagram between the bottleneck module of the prior art and the bottleneck module of the present application. Detailed implementation manners
[0050] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0051] As Figure 1 shown, the improved yolov8n model training method for surgical knot-tying action video recognition includes:
[0052] Obtain a surgical action video training set;
[0053] Obtain an improved yolov8n model, where the improved yolov8n model includes a dynamic convolution module and a GhostNet;
[0054] Train the improved yolov8n model with the surgical action video training set to obtain a trained yolov8n model.
[0055] Refer to Figure 12 and Figure 13 In this embodiment, the obtaining of the improved yolov8n model includes:
[0056] Modify the bottleneck layer in the c2f layer of the yolov8n model, change the depthwise conv layer in the bottleneck layer to n Ghost modules, and change the cv2 layer in the bottleneck layer to a dynamic convolution layer.
[0057] In this embodiment, the output of the bottleneck layer in the modified c2f layer is expressed as follows:
[0058]
[0059] * represents a convolution operation, ω primary,k is the kth standard convolution kernel. Let ω k be the kth dynamic convolution kernel, and there are a total of K dynamic convolution kernels. α k is the dynamically generated weight.
[0060] In this embodiment, the comprehensive loss function of the improved YOLOv8n model is as follows:
[0061] ζ = λ cls ζ cls + λ loc ζ loc + λ pose ζ pose ;
[0062] Among them, ζ cls refers to the classification loss, which measures the gap between the predicted class and the true class, and uses the cross-entropy loss; ζ loc refers to the localization loss, which measures the gap between the predicted bounding box and the true bounding box; ζ pose refers to the pose estimation loss, which measures the gap between the predicted key point positions and the true key point positions; uses the mean squared error, λ cls 、λ loc 、λ pose These are the weight hyperparameters of the corresponding loss terms, which are used to balance the contributions of different loss terms to the total loss.
[0063] In this embodiment, the calculation formula of the ζ cls classification loss is as follows:
[0064]
[0065] Among them, N is the number of samples, C is the number of classes, is the true label that the sample belongs to class c, is the probability of class c predicted by the model.
[0066] In this embodiment, the calculation formula of the ζ loc localization loss is as follows:
[0067]
[0068] Among them, is the true bounding box of the sample , is the bounding box predicted by the model.
[0069] In this embodiment, the calculation formula of the ζ pose pose estimation loss is as follows:
[0070]
[0071] Among them, K is the number of key points, is the true position of the j-th key point of the sample , is the position of the j-th key point predicted by the model.
[0072] In this embodiment, the surgical action video training set includes knot-tying actions in daily scenarios and also includes knot-tying actions during actual surgical procedures.
[0073] The technical solution of the present application will be further elaborated below. It should be understood that this example does not constitute any limitation to the present application.
[0074] As Figure 2 shown, it is the structural diagram of the improved yolov8n model - GD-Det provided by the present invention. In this embodiment, each C2f module contains the above-mentioned ghost module. The main steps are as follows:
[0075] (1) Obtain the surgical action video training set. Specifically, extract the surgical action video into pictures. The dataset used in the present invention is generated from the knot-tying action videos recorded by doctors. To ensure the diversity of the dataset, our dataset not only includes knot-tying actions in daily scenarios but also includes knot-tying actions during actual surgical procedures.
[0076] After obtaining the pictures, perform key point annotation on the important hand joint points in the pictures. Some examples of the dataset are Figure 3 shown. An example of the image annotation interface is Figure 4 shown. The images are annotated using Labelme, and the label format is JavaScript (JSON). The entire hand area is annotated using "create rectangle", while the key points are annotated using "create point". Under the guidance of doctors, 18 important key points of the knot-tying action are annotated in each picture. After the annotation is completed, all json files are converted into txt files for the next experiment.
[0077] (2) Obtain the improved yolov8n model. Yolov8-pose integrates multi-task capabilities, including object detection, instance segmentation, and key point detection. This model contains five variants, namely n, s, m, l, and x, and all variants share a consistent network architecture.
[0078] In the prior art, the C3 module is one of the key components used in YOLOv5, Figure 5 which shows the structure of the C3 module. The design of the C3 module is based on Cross Stage Partial Network (CSPNet), aiming to enhance the feature extraction ability and improve the computational efficiency of the network. The C3 module consists of multiple bottleneck structures, which combine convolution, batch normalization, and activation functions to effectively process redundant gradients and improve information transmission; in YOLOv8, the C3 module is replaced by the C2f module to achieve more efficient feature extraction and gradient flow. In Figure 6The C2f module structure is shown. The C2f module adopts a chunk operation, evenly dividing the input feature map into multiple sub-parts. Each sub-part is independently processed through a bottleneck layer or other convolutional operations, and finally the processed features are concatenated together. The C2f module allows for richer gradient flow, enhances the feature extraction ability, and reduces the computational complexity; The FasterNet Block is an efficient module in YOLOv8. Its structure diagram is as Figure 7 shown. By replacing the Bottleneck in C2f with the FasterNet Block, the feature extraction efficiency is significantly improved. It aims to reduce the computational complexity and the number of parameters by using Partial Convolution (PConv), thereby improving the inference speed and efficiency of the model. PConv selectively performs convolutional operations on the input channels without affecting other channels, thus significantly reducing the floating-point operations (FLOPs) and improving the inference speed; Figure 8 In [4], the EMA (Efficient Multi-scale Attention) module is proposed to solve the problems that traditional Convolutional Neural Networks (CNNs) often have high computational complexity and insufficient feature fusion when dealing with multi-scale information. By introducing multi-scale feature representations and an efficient attention mechanism in the C2f layer to process the multi-scale information of images. To ensure computational efficiency, EMA reduces the dimension of the features before calculating the attention to reduce the computational complexity, and localizes the features, calculating the attention only within the local area, thereby reducing the amount of computation.
[0079] (3) The part of the improved yolov8n model in this application can be called the GD-Det model structure. The C2f layer in YOLOv8 includes three processes: two consecutive convolutional operations, a feature fusion operation, and a residual connection. The purpose is to keep the network simple and efficient while extracting features.
[0080] In the prior art, the formula for the standard C2f layer is as follows:
[0081]
[0082] Among them, represents the input feature map, γ represents the output feature map, and Bottleneck represents the bottleneck layer.
[0083] However, even when using the yolov8n-pose model with the smallest number of parameters and amount of computation for keypoint detection, its actual number of parameters is 3.31M, and the GFLOPs can reach 9.3. Therefore, we still need to further lightweight the network.
[0084] In this embodiment, GhostNet is a lightweight neural network architecture designed to reduce the computational cost and the number of parameters through an efficient feature generation mechanism, thereby improving the inference speed and energy efficiency of the network. The core idea of GhostNet is to decompose each feature map into two parts: one is the "base feature" generated through standard convolution operations, and the other is the "ghost feature" generated from the base feature through simpler linear operations. In this way, the network can significantly reduce the computational overhead while maintaining performance. The schematic diagram of the Ghost convolution module is as shown in Figure 5 and the formula for the feature map generated by the Ghost module can be expressed as:
[0085]
[0086] where ω primary is the standard convolution kernel, is the number of ghost features generated for each input feature map, is the linear transformation function.
[0087] Dynamic convolution is a convolution operation that generates convolution kernels dynamically based on the input. When X is the input and Y is the output, the mixture-of-experts model in the middle of the input and output is generated by weighted summation of each expert layer with the dynamically generated coefficient α. The dynamic weight α is generated dynamically according to different input samples. That is, for the input X, first perform global average pooling on it to fuse the information into a vector, and then use a two-layer MLP module with a softmax activation function to dynamically generate the coefficient.
[0088] The mixture of experts (MoE) is an ensemble learning method that combines multiple specialized sub-models (experts) to form an overall model, and each "expert" contributes in its area of expertise. In the MoE model in GD-Det, the input data first passes through the gating network, which dynamically selects a part of the expert models for activation according to the characteristics of the input data, and the softmax activation function is used here. The activated expert models will process the input data and generate corresponding outputs. Finally, these outputs will be combined to form the final prediction result of the overall model.
[0089] The formula for generating the dynamic weight α is as follows:
[0090] α = softmax(MLP(pool((X))) (3)
[0091] k is the number of dynamic convolution kernels, ω k is the k-th dynamic convolution kernel. α k is the dynamically generated weight, and * represents the convolution operation.
[0092] Among them, the MLP (Multilayer Perceptron) is a basic artificial neural network model, which consists of an input layer, at least one or more hidden layers, and an output layer. The input layer receives the original data or features and passes them to the next layer. Each input node represents a feature, and the number of nodes in the input layer is determined by the dimension of the features. The hidden layer is the intermediate layer connecting the input layer and the output layer. Each hidden layer contains multiple neurons (nodes). Each neuron is connected to all nodes in the previous layer and outputs a value obtained by processing the weighted sum through an activation function. The output layer receives the signals from the last hidden layer and outputs the prediction results of the model. Each connection has a corresponding weight, indicating the strength of the connection, which is used to adjust the influence of the input. Each neuron has a bias term, which is used to adjust the activation threshold of the neuron.
[0093] The dynamic convolution formula can be expressed as:
[0094]
[0095] k is the number of dynamic convolution kernels, ω k is the k-th dynamic convolution kernel. α k is the dynamically generated weight, usually generated through an attention mechanism, and * represents the convolution operation.
[0096] GD-Det (Ghost-Dynamic-Detection) combines dynamic convolution and ghost network, and changes the traditional C2f layer to the C2f-GhostDynamic-Conv layer for experiments. The model structure is as Figure 7 shown.
[0097] Assume the input feature map is the output feature map is γ, and we have K dynamic convolution kernels. Each dynamic convolution kernel generates several ghost features.
[0098] The GD-Det formula can be expressed as:
[0099]
[0100] * represents the convolution operation, ω primary,k is the k-th standard convolution kernel. Let ω k be the k-th dynamic convolution kernel. There are a total of K dynamic convolution kernels, and α k is the dynamically generated weight.
[0101] When calculating the loss function, the comprehensive loss function usually includes classification loss, localization loss, and key point loss. The definitions are as follows:
[0102] ζ = λ cls ζ cls+λ loc ζ loc +λ pose ζ pose (6)
[0103] Among them, ζ cls refers to the classification loss, which measures the gap between the predicted class and the true class. Cross-Entropy Loss is used. ζ loc refers to the localization loss, which measures the gap between the predicted bounding box and the true bounding box. ζ pose refers to the pose estimation loss, which measures the gap between the predicted key point positions and the true key point positions. Mean Squared Error (MSE) is used. λ cls 、λ loc 、λ pose These are the weight hyperparameters corresponding to the loss terms, used to balance the contributions of different loss terms to the total loss.
[0104] ζ cls The calculation formula for the classification loss is as follows:
[0105]
[0106] Among them, N is the number of samples, C is the number of classes, is the true label of the sample belonging to class c, is the probability of class c predicted by the model.
[0107] ζ loc The calculation formula for the localization loss is as follows:
[0108]
[0109] Among them, is the true bounding box of the sample , is the bounding box predicted by the model.
[0110] ζ pose The calculation formula for the pose estimation loss is as follows:
[0111]
[0112] Among them, K is the number of key points, is the true position of the j-th key point of the sample , is the position of the j-th key point predicted by the model.
[0113] In this embodiment, this application further includes:
[0114] The yolov8n model of this application is verified through comparative experiments and ablation experiments.
[0115] Specifically as follows:
[0116] Comparative experiments and ablation experiments. The lightweight evaluation metrics of the improved yolov8n model include the number of parameters (Params) and the number of floating-point operations per second (GFLOPs), and these two metrics are used to measure the lightweight degree of the model; the accuracy of knotting action detection uses precision (P), recall (R), and mean average precision (mAP) as performance metrics. Their calculation formulas can be seen in Eq(9)-(11).
[0117]
[0118] Among them, true positive (TP) represents the instance where the model accurately predicts the number of target object samples, while false negative (FN) corresponds to the situation where the model fails to correctly detect the number of target object samples. On the contrary, false positive (FP) represents the instance where the model wrongly identifies a non-target object as a target object. The precision-recall (PR) curve is constructed by calculating the precision and recall values at different confidence thresholds. Then, the average precision (AP) is calculated by integrating the precision values for each recall rate along the PR curve.
[0119] In this part, a total of 8 groups of comparative experiments were conducted, including five baseline models of yolov8-pose and three lightweight improved models (yolov8n-pose+F: replacing Bottleneck with FasterNet block; yolov8n-pose+FE: adding EMA attention mechanism in C2f or its derivative structure; GD-Det: replacing the ordinary C2f layer with C2f-GhostDynamic-Conv layer).
[0120] In the experiment, first, five models of the yolov8-pose series: yolov8n-pose, yolov8s-pose, yolov8m-pose, yolov8l-pose, and yolov8x-pose were compared. The basic architectures of these models are the same, but their parameter numbers and computational amounts gradually increase. Among them, the smallest yolov8n-pose model has 3.3M parameters and 9.3 GFLOPs; the largest yolov8x-pose model has 69.5M parameters and 263.2 GFLOPs. The former has 95% fewer parameters and 96.5% fewer GFLOPs than the latter. The running results of these five models on the dataset are shown in Table 1.
[0121] Table 1 Comparison results of yolo series model tests
[0122]
[0123] See Figure 9 , it can be seen from the experimental result graph that the maximum difference in the precision results of each model in the yolov8-pose series running on the dataset is 3.1%, the maximum difference in the recall precision is 3.2%, and the maximum difference in the mean average precision is 4.2%. On this basis, the parameter quantity of their smallest model and largest model differs by 21 times, and the GFLOPs differ by 28 times. Therefore, using the yolov8n-pose model as the baseline model is the most in line with the lightweight requirements on the premise that the results such as precision and recall do not differ much.
[0124] In the second stage of the experiment, a comparative experiment was conducted between the baseline model yolov8n-pose and the lightweight models yolov8n-pose+F, yolov8n-pose+FE, and our GD-Det. The experimental results are shown in Table 2.
[0125] Table 2 Test comparison results of lightweight models
[0126]
[0127]
[0128] It can be seen from the experimental results that compared with the baseline model yolov8n-pose, the parameter quantity of the GD-Det model is reduced to 75% of the baseline model, and the GFLOPs are reduced to 71.8% of the baseline. On this basis, there is only a 0.8% precision loss and a 0.6% mean average precision loss; in addition, the parameter quantity and GFLOPs of our GD-Det model are lower than those of the other two lightweight models yolov8n-pose+F and yolov8n-pose+FE. However, on this basis, the precision, recall, and mean average precision are all higher than those of these two models.
[0129] In the previous content, dynamic convolution, ghost convolution, and GD-Det have been introduced in detail. Therefore, in this section, in order to further demonstrate the superiority of GD-Det, an ablation experiment was conducted. The yolov8n-pose network alone, the yolov8n-pose network with dynamic convolution added (yolov8n-pose+D), the yolov8n-pose network with ghost convolution added (yolov8n-pose+G), and the GD-Det network were compared respectively. The experimental results are shown in Table 3.
[0130] Table 3 Ablation experiment results
[0131]
[0132] When only ghost convolution is added to the baseline model, due to the fact that ghost convolution generates ghost features and thus reduces the number of parameters, the accuracy of the corresponding yolov8n-pose+G model also decreases as the number of parameters is reduced; while when dynamic convolution is added, the number of parameters of the model will increase significantly. Therefore, by combining ghost convolution and dynamic convolution and replacing the C2f layer in yolov8n-pose with the C2f-GhostDynamic layer, a lightweight model with significantly reduced number of parameters and computational volume and basically unchanged accuracy can be obtained. Figure 9 Figure 4 shows the comparison chart of the number of parameters and GFLOPs experimental results of these models. Figure 10 Figure 6 shows the comparison chart of the experimental results of the precision, recall rate and loss function of these models.
[0133] As can be seen from the result chart, in terms of the number of parameters and GFLOPs, GD-Det is the lowest, and in terms of precision, it is only 0.8% lower than the baseline model yolov8n-pose; even in the comparison of the recall rate, GD-Det can be comparable to the baseline model. Except for the yolov8n-pose baseline model and yolov8n-pose+D with increased number of parameters by adding Dynamic-Conv, GD-Det shows a faster convergence speed compared to other models. This observation indicates that our GD-Det effectively reduces the loss value and enhances the detection performance by integrating Ghost-Conv and Dynamic-Conv.
[0134] Comparison chart of GD-Det prediction results. The comparison chart of the prediction results of the baseline model, lightweight model, other models for ablation experiments and the GD-Det model on the same chart is as Figure 11 shown.
[0135] The present application also provides an improved yolov8n model training device for surgical knot-tying action video recognition. The improved yolov8n model training device for surgical knot-tying action video recognition includes a training set acquisition module, a model acquisition module and a training module, wherein,
[0136] The training set acquisition module is used to acquire a surgical action video training set;
[0137] The model acquisition module is used to acquire an improved yolov8n model, and the improved yolov8n model includes a dynamic convolution module and a ghost network;
[0138] The training module is used to train the improved yolov8n model through the surgical action video training set to obtain a trained yolov8n model.
[0139] The present application also provides a method for recognizing surgical knot-tying action videos, and the method for recognizing surgical knot-tying action videos includes:
[0140] Obtain the improved yolov8n model for recognizing surgical knot-tying action videos as described above;
[0141] Obtain the video to be recognized;
[0142] Input the video to be recognized into the improved yolov8n model for recognizing surgical knot-tying action videos, so as to obtain the recognition result. Specifically, the recognition result is to recognize the finger key points, such as the joint points of the hand.
[0143] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An improved yolov8n model training method for surgical knotting action video recognition, characterized in that: The improved yolov8n model training method for surgical knotting action video recognition includes: Obtain a surgical action video training set; Obtain an improved yolov8n model, wherein the improved yolov8n model includes a dynamic convolution module and a ghost network; The improved yolov8n model is trained by using the surgical action video training set to obtain the trained yolov8n model.
2. The improved yolov8n model training method for surgical knotting action video recognition as claimed in claim 1, characterized in that: The method of obtaining the improved yolov8n model includes: The bottleneck layer in the c2f layer of the yolov8n model is modified by changing the depthwise conv layer in the bottleneck layer to n ghost modules and changing the cv2 layer in the bottleneck layer to a dynamic convolution layer.
3. The improved yolov8n model training method for surgical knotting action video recognition as claimed in claim 2, characterized in that: The bottleneck layer output in the modified c2f layer is expressed as follows: * represents the convolution operation, ω primary,k is the kth standard convolution kernel, let ω k is the kth dynamic convolution kernel, there are K dynamic convolution kernels in total, α k is a dynamically generated weight.
4. The improved yolov8n model training method for surgical knotting action video recognition as claimed in claim 3, characterized in that: The comprehensive loss function of the improved yolov8n model is: g=l cls g cls +λ loc g loc +λ pose g pose ; Among them, cls Refers to classification loss, which measures the gap between the predicted category and the true category, using cross entropy loss; ζ loc The position loss measures the gap between the predicted bounding box and the true bounding box; ζ pose Refers to the pose estimation loss, which measures the gap between the predicted key point position and the true key point position; using the mean square error, λ cls , loc , pose These are the weight hyperparameters for the corresponding loss terms, which are used to balance the contribution of different loss terms to the total loss.
5. The improved yolov8n model training method for surgical knotting action video recognition as claimed in claim 4, characterized in that: The ζ cls The classification loss calculation formula is as follows: Where N is the number of samples, C is the number of categories, and y i,c is the true label of sample i belonging to category c, is the probability of class c predicted by the model.
6. The improved yolov8n model training method for surgical knotting action video recognition as claimed in claim 4, characterized in that: The ζ loc The positioning loss calculation formula is as follows: in, It is a sample The ground-truth bounding box, is the bounding box predicted by the model.
7. The improved yolov8n model training method for surgical knotting action video recognition as claimed in claim 4, characterized in that: The ζ pose The posture estimation loss calculation formula is as follows: Among them, K is the number of key points, p i,j is the true position of the jth key point of sample i, is the position of the jth keypoint predicted by the model.
8. The improved yolov8n model training method for surgical knotting action video recognition as claimed in claim 1, characterized in that: The surgical action video training set includes knotting actions in daily scenes and also includes knotting actions in actual surgical procedures.
9. An improved yolov8n model training device for surgical knotting action video recognition, characterized in that: The improved yolov8n model training device for surgical knotting action video recognition comprises: A training set acquisition module, wherein the training set acquisition module is used to acquire a surgical action video training set; A model acquisition module, wherein the model acquisition module is used to acquire an improved yolov8n model, wherein the improved yolov8n model includes a dynamic convolution module and a ghost network; A training module, wherein the training module is used to train the improved yolov8n model through the surgical action video training set, thereby obtaining the trained yolov8n model.
10. A method for identifying surgical knotting action videos, characterized in that: The surgical knotting action video recognition method comprises: Obtaining the improved yolov8n model for surgical knotting action video recognition as described in any one of claims 1 to 9; Get the video to be identified; The video to be identified is input into the improved yolov8n model for surgical knotting action video recognition to obtain the recognition result.
Citation Information
Patent Citations
Training method of surgical action recognition model, medium and equipment
CN113705320A
Automatic generation and detection method for whole PCB cutting path based on YOLOv8 algorithm
CN117593319A
Double-cavity tracheal intubation assisting method and system based on YOLOv5
CN118969233A
Ship fire detection method, device and equipment and storage medium
CN119418284A