Medical action training method and device based on machine vision model and storage medium
Through deep learning technology based on machine vision models, the dispensing movements of medical staff are automatically identified and evaluated, and the problems of low efficiency and inaccurate evaluation in the existing technology are solved, efficient and accurate motion analysis and personalized improvement suggestions are achieved, and the professional skills and work safety of medical staff are improved.
Patent Information
- Application Number
- CN202510512633.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-27
AI Technical Summary
The existing methods of medical staff's movement training rely on manual monitoring, have low efficiency, inaccurate evaluation, cannot correct irregular operations in real time, and have high labor and time costs.
The medical action training method based on machine vision model is adopted, and the improved YOLOv10-pose deep learning model is used to identify the dispensing movements of medical staff, locate the key points of the hand, analyze the spatial distribution and motion trajectory, judge the movement normativeness through template matching, and generate improvement suggestions.
It realizes automatic identification and evaluation of the movements of medical staff, improves the efficiency and accuracy of training, and can generate personalized improvement suggestions based on specific movements, helps medical staff improve their professional skills, reduce occupational fatigue, and improve work efficiency and safety.
Smart Images

Figure CN120048004A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a medical action training method, device and storage medium based on a machine vision model. Background Art
[0002] In the medical industry, training of medical staff's movements is an important part of ensuring medical quality and safety. Especially in the process of dispensing medicine, the standardization of medical staff's movements directly affects the quality of medicine preparation and the probability of medical staff's own injury. If medical staff make irregular movements during the process of dispensing medicine, it may lead to dispensing errors, affect the treatment effect, and even endanger the patient's life. Therefore, real-time and accurate monitoring of medical staff's movements is crucial.
[0003] At present, the main method of training medical staff's movements is manual monitoring. Many medical institutions still mainly use manual monitoring during training, and supervisors observe the medical staff's operations to judge the standardization of the movements. This detection method has complicated operation procedures and the test results are easily affected by subjective factors. It is difficult to accurately and consistently evaluate the movements of medical staff. In addition, manual monitoring cannot be carried out in real time, and it is often evaluated after the fact. It is difficult to correct irregular operations in time, which is inefficient and has high labor and time costs. Summary of the invention
[0004] In order to solve the above-mentioned technical problems of low efficiency and inaccurate evaluation of medical staff action training, the present invention provides a medical action training method, device and storage medium based on a machine vision model.
[0005] According to one aspect of the present invention, a medical action training method based on a machine vision model is provided, comprising: According to another aspect of the present invention, a medical action training device based on a machine vision model is also provided, including: collecting standard action videos of dispensing medicine and building an expert library as a detection standard; filming the actual dispensing process of medical staff and collecting actual dispensing action videos; using an improved YOLOv10-pose deep learning model to perform hand action recognition on the actual dispensing action video, locate hand key points, and crop and save the model output frame, filter out the background, and retain the hand area image; analyze the spatial distribution and motion trajectory of each key point based on the located hand key points; template match the cropped actual dispensing action with the standard action in the expert library, compare the similarity of the actual dispensing action with the standard action in spatial distribution and motion trajectory, and judge whether the medical staff's action meets the standard; if the medical staff's action does not meet the standard, generate improvement suggestions based on the standard action in the expert library, and guide the medical staff to standardize the non-standard action according to the standard action in the expert library.
[0006] Optionally, the improved YOLOv10-pose deep learning model is used to perform hand motion recognition on an actual medication dispensing action video. Before locating the hand key points, the method includes: collecting images or videos of hand motions during medication dispensing, including standard medication dispensing actions and non-standard medication dispensing actions, and annotating the collected images or videos, wherein the annotation content includes category labels and key points; preprocessing the hand motion images or videos using data enhancement technology; segmenting the preprocessed images or videos into a training set, a validation set, and a test set; constructing an improved YOLOv10-pose deep learning model, loading the training set, the validation set, and the test set to train and evaluate the model to obtain an optimal model; and using the optimal model to perform hand motion recognition during actual medication dispensing.
[0007] Optionally, the preprocessing of the hand motion image or video using data enhancement technology includes: adjusting the collected hand motion image or video to a uniform preset size to ensure that the input data meets the input requirements of the YOLOv10-pose model; and performing standardization and data enhancement processing on the size-adjusted medicine dispensing hand motion image, wherein the data enhancement processing includes denoising and clarity improvement.
[0008] Optionally, the improved YOLOv10-pose deep learning model is an RD-YOLOv10 model, and the RD-YOLOv10 model uses an improved RD-Pose detection head to replace the original Pose detection head, and uses Focal Loss to replace the original classification loss function.
[0009] Optionally, the spatial distribution and motion trajectory of each key point are analyzed based on the located hand key points; the cropped actual dispensing action is template matched with the standard action in the expert library, and the similarity in spatial distribution and motion trajectory between the actual dispensing action and the standard action is compared, and whether the action of the medical staff meets the standard is judged, including: calculating the first spatial relative position relationship between two adjacent key points in the actual dispensing action based on the located hand key points; traversing each standard action in the expert library one by one, comparing the actual dispensing action with each standard action in the expert library in turn, and finding the standard action that best matches the actual dispensing action based on the hand action characteristics of the actual dispensing action; obtaining the second spatial relative position relationship between corresponding key points in the best matching standard action; calculating the relative error between the actual action and the standard action based on the first spatial relative position relationship and the second spatial relative position relationship; judging whether the relative error is less than or equal to a preset error threshold; if the relative error is less than or equal to the preset error threshold, it is considered that the action of the medical staff meets the standard; otherwise, it is considered that the action of the medical staff does not meet the standard.
[0010] Optionally, the improvement suggestions include graphical interfaces, voice prompts or report forms. If the medical staff's actions do not meet the standards, correction prompts are given through a graphical interface or voice, or standard action reference videos or action decomposition diagrams are provided for medical staff to learn.
[0011] Optionally, after generating improvement suggestions based on standard actions in the expert database, the method also includes: detecting whether the medical staff has completed the medication dispensing, and generating an operation evaluation report after the medication dispensing is completed, the operation evaluation report including the proportion of standard actions of the medical staff in medication dispensing, the types of non-standard actions and improvement suggestions, and a historical trend chart of action optimization.
[0012] Optionally, the method uses a structured light camera to capture the medication dispensing action, the resolution of the structured light camera is at least 1080p, the frame rate is at least 30fps, and the collected image or video data is published to the ROS2 topic in the sensor_msgs / Image format.
[0013] According to another aspect of the present invention, a medical action training device based on a machine vision model is also provided, including: an expert library construction module, which is used to collect standard action videos of dispensing medicine and build an expert library as a detection standard; a data acquisition module, which is used to shoot the actual dispensing process of medical staff and collect actual dispensing action videos; an action recognition module, which is used to use the improved YOLOv10-pose deep learning model to perform hand action recognition on the actual dispensing action video and locate the hand key points; an image cropping module, which is used to crop and save the model output frame, filter out the background, and retain the hand area image; a motion comparison module, which is used to analyze the spatial distribution and motion trajectory of each key point based on the located hand key points; template matching the cropped actual dispensing action with the standard action in the expert library, comparing the similarity of the actual dispensing action with the standard action in spatial distribution and motion trajectory, and judging whether the action of the medical staff meets the standard; a feedback module, which is used to generate improvement suggestions based on the standard actions in the expert library if the action of the medical staff does not meet the standard, and guide the medical staff to standardize the non-standard actions based on the standard actions in the expert library.
[0014] According to another aspect of the present invention, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method are implemented.
[0015] The present invention can automatically identify and evaluate the medication dispensing actions of medical staff, improve the efficiency and accuracy of training, and can generate personalized improvement suggestions based on the specific actions of medical staff, which helps to improve the professional skills of medical staff in a targeted manner. Real-time detection and evaluation of the actions of medical staff in the medication dispensing process can help medical staff develop correct action habits. By combining advanced machine vision technology and deep learning algorithms, the actions of medical staff can be efficiently and accurately analyzed, irregular operations in the actions can be identified, and targeted improvement suggestions can be provided. The present invention has the characteristics of high efficiency, real-time and accuracy, and can be widely used in medical scenarios to help medical staff optimize medication dispensing postures and actions, reduce occupational fatigue, and improve work efficiency and safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is a flow chart of a medical action training method based on a machine vision model according to an embodiment of the present invention; Figure 2 is a schematic diagram of equipment layout in an embodiment of the present invention; Figure 3 This is a schematic diagram of the RD-YOLOv10 network structure; Figure 4 It is a schematic diagram of the RD-Pose structure; Figure 5 This is the Focal Loss effect diagram corresponding to different focus factors; Figure 6 Compare the results before and after model improvement; Figure 7 Schematic diagram of the relative position relationship of the key points of the hand; Figure 8 The present invention is a structural block diagram of a medical action training device based on a machine vision model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0017] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only embodiments of a part of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present application. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0018] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0019] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0020] Embodiment 1 of the present invention:
[0021] Reference Figure 1 , Figure 1 is a flow chart of a medical action training method based on a machine vision model according to an embodiment of the present invention, such as Figure 1 As shown, the method includes: S1, collect standard action videos of dispensing medicine and build an expert database as the detection standard; This method aims to train the standardization of hand movements when dispensing medicine. If there is a large deviation between the medical staff's actual dispensing movements and the standard dispensing movements in the expert database, it means that the medical staff's movements are not standardized and need to be corrected.
[0022] Before real-time detection, it is necessary to collect standard action videos of dispensing medicine to build an expert library. These videos should cover all key steps and correct hand movements. The movements must be standard, smooth, safe and comfortable, and avoid movement errors, sudden pauses or unnecessary movement amplitudes. The movements must be ergonomic as much as possible, and the most labor-saving posture and method should be used to achieve the desired effect. Using movements that meet these requirements as templates, an expert library is built for feature comparison with the action images of the actual dispensing process of the medical staff to be tested.
[0023] S2, filming the actual medication dispensing process of medical staff and collecting actual medication dispensing action videos; The actual hand operation of medical staff in the process of dispensing medicine is captured by a high-definition camera or a depth camera. Figure 2 , Figure 2Schematic diagram of equipment layout in an embodiment of the present invention. Figure 2 As shown in the figure, the camera is installed at an appropriate position on the workbench or operating area to ensure that it can cover the key movement range of medical staff, especially the hand operation area. The collected image or video data will be transmitted to the data processing module through the standard ROS2 (Robot Operating System 2, the second generation robot operating system) communication interface.
[0024] In one implementation, a structured light camera is integrated into the ROS2 environment, and the structured light camera is used to capture the dispensing action. The resolution of the structured light camera is required to be at least 1080p, the frame rate is at least 30fps, and accurate depth information is obtained. The collected image or video data is published to the ROS2 topic in the sensor_msgs / Image format. Each node in ROS2 is responsible for a separate modular function. ROS2 has a total of four communication methods to achieve interaction between nodes, namely topics, services, actions, and parameters. Topics are a communication mechanism based on the publish-subscribe mode in ROS2. In this mode, nodes can publish messages to a topic, and other nodes can subscribe to this topic to receive messages.
[0025] The collected actual medication dispensing action videos need to be preprocessed, including normalization and denoising, frame selection, image enhancement and other operations, to improve the accuracy and efficiency of subsequent action detection.
[0026] S3, using the improved YOLOv10-pose deep learning model to perform hand action recognition on the actual medication dispensing action video, locate the key points of the hand, and crop and save the model output frame, filter out the background, and retain the hand area image; The improved YOLOv10-pose deep learning model is the RD-YOLOv10 model, such as Figure 3 As shown, Figure 3 Schematic diagram of the RD-YOLOv10 network structure. The RD-YOLOv10 model uses an improved RD-Pose detection head to replace the original Pose detection head ( Figure 4 Schematic diagram of RD-Pose structure), and use Focal Loss to replace the original classification loss function to improve the accuracy and performance of the model in detection tasks.
[0027] In addition, the improved deep learning algorithm in the embodiment of the present invention can significantly reduce the number of parameters by using shared convolution, which makes the model more lightweight, especially on resource-constrained devices, but the shared parameters may limit the expressiveness of the model because different features may require different convolution kernels to capture complex patterns. Since shared parameters may not be able to fully capture these differences, in order to make up for the negative impact of shared convolutions used to achieve lightweight, reparameterized convolutions are used. By introducing more learnable parameters, the network can more effectively extract features from the data, thereby compensating for the loss of precision that may be caused by the lightweight model, and reparameterized convolutions can greatly improve parameter utilization, which is no different from ordinary convolutions in the inference stage, bringing a lossless optimization solution to the model.
[0028] While using shared convolution, in order to address the problem of inconsistent scales of targets detected by each detection head, the improved deep learning algorithm in the embodiment of the present invention uses a Scale layer to scale features and improve the loss function, wherein the improved loss function includes: The improved IoU loss and GIoU loss are used to measure the overlap and encirclement between the predicted box and the true box to better optimize the shape and position of the bounding box.
[0029] Focal Loss is introduced to deal with the problem of class imbalance. By adjusting the weights, the model pays more attention to samples that are difficult to classify. The Focal Loss formula is expressed as follows: ; In the formula, is the model’s predicted probability of the target class, and γ is the focus factor, which is used to adjust the weight of difficult and easy samples; Represents the Focal Loss of the model. Figure 5 This is the Focal Loss effect diagram corresponding to different focus factors.
[0030] Comprehensive loss function: The improved IoU loss, GIoU loss and classification loss are combined to form a comprehensive target detection loss function, and the key point loss is added to complete the key point detection in the detection task.
[0031] Through experimental comparison, the comparison results before and after the model improvement are referenced Figure 6 , Figure 6In , YOLOv10n-pose is the original model, RD-YOLOv10 is the improved deep learning model in the embodiment of the present invention, Class represents the category, A and B categories represent the left and right hands respectively, all is the case of not distinguishing between left and right hands, P is the precision of the model detection, R is the recall rate of the model detection, mAP@.5 is the average precision of the detection model when the IoU threshold is 0.5, Pose / P is the precision of the model pose estimation, Pose / R is the recall rate of the model pose estimation, and Pose / mAP@.5 is the mAP (average precision) of the key point detection of the model under the IoU=0.5 threshold, as shown in Figure 6 As shown in the figure, the improved deep learning model RD-YOLOv10 has numerical improvements in most of the data. Although the improvement in detection tasks is not comprehensive, considering that the focus of the experiment is on key point tasks, it can be considered that the model improvement is effective.
[0032] The embodiment of the present invention performs action recognition and classification on the collected images or videos based on the RD-YOLOv10 model. Based on pre-training, the RD-YOLOv10 model is fine-tuned through transfer learning to adapt to the action scenes of medical staff in dispensing medicine. RD-YOLOv10 is used to detect the key points (such as joint positions) when dispensing medicine on the hand, analyze the spatial distribution and motion trajectory of the corresponding key points, and the classification results and the relative positions of the key points of the hand are sent to the comparison module through the ROS2 node.
[0033] The output of the RD-YOLOv10 model is a rectangular box with a hand category label on it, and the box contains the predicted key point information. Among them, the category label is "hand". Because it is a detection of whether the hand movement is compliant, it is necessary to ensure that it is a hand before subsequent operations can be performed. The key points are the joints of the hand bones, including the joints of the fingers, wrists, etc. According to the pixel coordinates of the upper left and lower right points of the rectangular box predicted by the model, a rectangular area can be determined. According to this rectangular area, the prediction box output by the model is cropped to extract the image of the hand area. The cropped image only contains the hand area, filtering out the interference of the background and other objects, minimizing the amount of data, and the cropped hand area image can be saved for subsequent processing and analysis. By separating and extracting specific hand area images from the prediction box output by the model, interference information can be effectively filtered out, which can improve the efficiency of data storage, transmission or subsequent computing tasks.
[0034] Furthermore, in step S3, before using the improved YOLOv10-pose deep learning model to perform hand motion recognition on the actual medication dispensing action video, a series of preliminary preparations are required, including model construction, image or video acquisition, annotation and preprocessing, and model training. Specifically, they include: S301, collecting images or videos of hand movements during medicine dispensing, including standard and non-standard movements, and annotating the collected images or videos, wherein the annotated contents include category labels and key points; Standard actions are correct dispensing actions, and non-standard actions are incorrect or need to be improved. The collected images or videos should cover all dispensing actions and scenes. The annotation content should include category labels - hands, and key points of the hand - various skeletal joints. The annotation work can be done by professionals or with the help of semi-automatic or automatic annotation tools.
[0035] S302, preprocessing the hand motion image or video using data enhancement technology; Preprocessing includes data normalization, rotation, scaling, flipping, cropping, denoising, etc.
[0036] The collected hand motion images or videos are resized to a uniform preset size to ensure that the input data meets the input requirements of the YOLOv10-pose model. The preset size is usually determined based on the input layer design of the model and actual application requirements.
[0037] The resized pictures of hand movements of dispensing medicine are then standardized and data enhanced to simulate images taken in complex environments, expand the data scale, and complete the production of the data set. Among them, standardization usually includes pixel value normalization, that is, scaling the pixel value to a specific range (such as 0-1) to eliminate the differences between different images due to factors such as lighting and contrast, and improve the efficiency and accuracy of the model. Data enhancement processing includes denoising and clarity improvement. Denoising refers to removing noise from an image, improving image clarity, and reducing interference, which can be achieved using a Gaussian filter. Image processing techniques (such as sharpening, contrast enhancement, etc.) can improve image clarity, and histogram equalization or gamma correction methods can be used to improve image quality. Data enhancement can enhance data diversity and prevent model overfitting. By introducing different enhancement methods, various situations that may be encountered in actual applications can be simulated, thereby improving the generalization ability of the model.
[0038] S303, dividing the preprocessed image or video into a training set, a validation set, and a test set; The data can be divided into training set, validation set and test set in 6:2:2, 8:1:1 or other ratios to train and evaluate the model.
[0039] S304, building an improved YOLOv10-pose deep learning model, loading the training set, validation set, and test set to train and evaluate the model, and obtaining the optimal model; During the training process, the training set is used to continuously adjust the model parameters and optimize the algorithm so that the model can accurately identify hand movements and locate the key points of the hand; the model is verified using the validation set to evaluate the performance and accuracy of the model, and the model is tuned according to the verification results until the optimal model is obtained; the optimal model is evaluated using the test set, and the model that passes the evaluation is the final optimal model.
[0040] S305: Using the optimal model to perform hand motion recognition during actual medication dispensing.
[0041] The hand movement image or video to be recognized is input into the trained RD-YOLOv10 optimal model. The model will output the position and category information of the hand key points to realize the recognition and evaluation of hand movements.
[0042] S4, analyzing the spatial distribution and motion trajectory of each key point based on the located hand key points; performing template matching between the cropped actual medication dispensing action and the standard action in the expert database, comparing the similarity between the actual medication dispensing action and the standard action in spatial distribution and motion trajectory, and judging whether the action of the medical staff meets the standard; The improved YOLOv10-pose deep learning model is used to perform hand motion recognition on actual medicine dispensing action videos. According to the hand motion classification results, the expert library actions of each category are matched, the relative positions of the hand key points are compared, and judgment is made based on the spatial distribution relationship of adjacent key points. An error threshold is set, and actions within the threshold range are considered standard, while actions outside the threshold range are considered non-standard.
[0043] Specifically, S4 includes: S41, calculating a first spatial relative position relationship between two adjacent key points according to the located hand key points; The spatial relative position relationship includes distance, angle, etc. Based on the detected position of the hand key points, the distance and angle between adjacent key points in the actual action are calculated, which can be achieved through geometric formulas (such as Euclidean distance formula, cosine theorem, etc.). Figure 7 , Figure 7 Schematic diagram of the relative position relationship of the key points of the hand.
[0044] S42, going through each standard action in the expert database one by one, comparing the actual medication dispensing action with each standard action in the expert database in turn, and finding the standard action that best matches the actual medication dispensing action according to the hand action characteristics of the actual medication dispensing action; This step involves feature matching. When processing hand motion recognition, although everyone's hands are of different sizes, the method of the present invention is more concerned with the spatial distribution and motion trajectory of the action itself, rather than the specific size of the hand. Therefore, through feature matching, we can scale or normalize hand images of different sizes to the same scale, and then compare their motion features, such as Euclidean distance, cosine similarity, etc., set matching thresholds to perform similarity calculations and judgments on the medical staff's medication dispensing actions, traverse each standard action in the expert library, and compare and calculate with the standard action template in turn through the index to find the standard action that is most similar to the hand motion features of the actual medication dispensing action, that is, the most matching standard action.
[0045] S43, obtaining a second spatial relative position relationship between corresponding key points in the most matching standard action; The second spatial relative position relationship is the distance and angle between corresponding key points in the most matching standard action, which is used to compare with the actual action to evaluate whether the actual action is accurate and standardized.
[0046] S44, calculating a relative error between an actual action and a standard action according to the first spatial relative position relationship and the second spatial relative position relationship; Calculate the distance error between the actual distance between key points (i.e., the distance between adjacent key points in the first spatial relative position relationship) and the standard distance (i.e., the distance between key points in the corresponding second spatial relative position relationship), or calculate the angle error between the actual angle between key points and the standard angle. These errors reflect the similarity between the actual action and the standard action in terms of spatial distribution and motion trajectory.
[0047] S45, determining whether the relative error is less than or equal to a preset error threshold; The preset error threshold can be set according to experience or experimental data, and the calculated relative error between the actual action and the standard action is compared with the preset error threshold.
[0048] S46, if the relative error is less than or equal to the preset error threshold, it is considered that the action of the medical staff meets the standard; otherwise, it is considered that the action of the medical staff does not meet the standard.
[0049] If the distance error between the actual action and the standard action is less than or equal to the preset distance error threshold, and the angle error between the actual action and the standard action is less than or equal to the preset angle error threshold, the action of the medical staff is considered to meet the standard. Otherwise, it does not meet the standard.
[0050] S5, if the actions of the medical staff do not meet the standards, improvement suggestions are generated based on the standard actions in the expert database to guide the medical staff to standardize the non-standard actions based on the standard actions in the expert database.
[0051] The improvement suggestions include graphical interfaces, voice prompts or reports. If the medical staff's actions do not meet the standards, they will be prompted to correct through graphical interfaces or voice, such as "adjust the elbow height" or "pay attention to the way the medicine bottle is held", or provide standard action reference videos or action decomposition diagrams for medical staff to learn and improve bad actions. Among them, voice prompts can be generated using rclpy in ROS2 and combined with TTS (text-to-speech) tools.
[0052] In addition, after detecting that the medical staff has completed the medication dispensing, an operation evaluation report can also be generated, which includes the proportion of standard actions of medical staff in medication dispensing, the types of non-standard actions and improvement suggestions, and a historical trend chart of action optimization. The operation evaluation report can be stored in PDF format and can be viewed or printed through the web interface.
[0053] The motion detection algorithm based on the RD-YOLOv10 model in the embodiment of the present invention has efficient and accurate detection capabilities, and can analyze the motion behaviors of medical staff in real time in complex scenarios. By real-time monitoring and evaluating the dispensing actions of medical staff, irregular operations can be discovered in a timely manner, helping medical staff to correct incorrect actions and develop standardized operating habits; through automated detection and feedback of actions, the need for manual supervision is reduced, the dispensing process is optimized, and the work efficiency of medical staff is improved; by prompting medical staff to optimize the dispensing posture, fatigue and occupational disease risks caused by improper actions are reduced, and long-term health protection is improved.
[0054] The method of the present invention can be applied to hospitals, pharmacies and medical training institutions, has good promotion value, and provides technical support for the standardization and intelligent development of the medical industry.
[0055] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0056] Embodiment 2 of the present invention:
[0057] In this embodiment, a medical action training device based on a machine vision model is also provided to implement the above-mentioned embodiments and preferred implementation modes, which have been described and will not be repeated. As used below, the terms "module" and "unit" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0058] Figure 8 is a structural block diagram of a medical action training device based on a machine vision model according to an embodiment of the present invention, such as Figure 8 As shown, the device includes: an expert library construction module 100, a data acquisition module 200, an action recognition module 300, an image cropping module 400, an action comparison module 500, and a feedback module 600, wherein: An expert database building module 100 is used to collect standard action videos of dispensing medicines and build an expert database as a detection standard; The data collection module 200 is used to film the actual medication dispensing process of medical staff and collect the actual medication dispensing action video; The action recognition module 300 is used to use the improved YOLOv10-pose deep learning model to perform hand action recognition on the actual medication dispensing action video and locate the key points of the hand; An image cropping module 400 is used to crop and save the model output frame, filter out the background, and retain the hand area image; The action comparison module 500 is used to analyze the spatial distribution and motion trajectory of each key point according to the located hand key points; perform template matching between the cropped actual medication dispensing action and the standard action in the expert database, compare the similarity between the actual medication dispensing action and the standard action in spatial distribution and motion trajectory, and determine whether the action of the medical staff meets the standard; The feedback module 600 is used to generate improvement suggestions based on the standard actions in the expert database if the actions of the medical staff do not meet the standards, and guide the medical staff to standardize the non-standard actions according to the standard actions.
[0059] Optionally, the data acquisition module is also used to collect images or videos of hand movements during medication dispensing, including standard and non-standard movements for medication dispensing; The device also includes: The data annotation module is used to annotate the collected images or videos, including category labels and key points; A data preprocessing module, used to preprocess the hand action image or video using data enhancement technology; and divide the preprocessed image or video into a training set, a validation set, and a test set; The model building module is used to build an improved YOLOv10-pose deep learning model, load the training set, validation set and test set to train and evaluate the model, and obtain the optimal model; use the optimal model to perform hand motion recognition during actual medication dispensing.
[0060] Optionally, the data preprocessing module is also used to: adjust the collected hand motion images or videos to a uniform preset size to ensure that the input data meets the input requirements of the YOLOv10-pose model; and perform standardization and data enhancement processing on the size-adjusted medicine dispensing hand motion images, wherein the data enhancement processing includes denoising and clarity improvement.
[0061] Optionally, the improved YOLOv10-pose deep learning model is an RD-YOLOv10 model, and the RD-YOLOv10 model uses an improved RD-Pose detection head to replace the original Pose detection head, and uses Focal Loss to replace the original classification loss function.
[0062] Optionally, the action comparison module is also used to: calculate a first spatial relative position relationship between two adjacent key points in the actual dispensing action based on the located hand key points; traverse each standard action in the expert library one by one, compare the actual dispensing action with each standard action in the expert library in turn, and find the standard action that best matches the actual dispensing action based on the hand action characteristics of the actual dispensing action; obtain a second spatial relative position relationship between corresponding key points in the best matching standard action; calculate a relative error between the actual action and the standard action based on the first spatial relative position relationship and the second spatial relative position relationship; determine whether the relative error is less than or equal to a preset error threshold; if the relative error is less than or equal to the preset error threshold, it is considered that the medical staff's action meets the standard; otherwise, it is considered that the medical staff's action does not meet the standard.
[0063] Optionally, the improvement suggestions include graphical interfaces, voice prompts or report forms, and the feedback module is also used to: if the medical staff's movements do not meet the standards, prompt corrections through a graphical interface or voice, or provide standard movement reference videos or movement decomposition diagrams for medical staff to learn.
[0064] Optionally, the device also includes: a report generation module, which is used to detect whether the medical staff has completed the medication dispensing, and generate an operation evaluation report after the medication dispensing is completed. The operation evaluation report includes the proportion of standard actions of the medical staff in medication dispensing, the types of non-standard actions and improvement suggestions, and a historical trend chart of action optimization.
[0065] Optionally, the data acquisition module is further used to capture the dispensing action using a structured light camera, the resolution of the structured light camera is at least 1080p, the frame rate is at least 30fps, and the collected image or video data is published to the ROS2 topic in the sensor_msgs / Image format.
[0066] An embodiment of the present invention further provides a storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.
[0067] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store computer programs.
[0068] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0069] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0070] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0071] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0072] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0073] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0074] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0075] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk, etc., which can store program code.
[0076] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A medical action training method based on a machine vision model, characterized in that: include: Collect videos of standard drug dispensing actions and build an expert database as a testing standard; Film the actual medication dispensing process of medical staff and collect actual medication dispensing action videos; The improved YOLOv10-pose deep learning model is used to identify hand movements in actual medication dispensing action videos, locate the key points of the hands, crop and save the model output frame, filter out the background, and retain the hand area image; Analyze the spatial distribution and motion trajectory of each key point based on the located hand key points; perform template matching between the cropped actual medication dispensing action and the standard action in the expert database, compare the similarity between the actual medication dispensing action and the standard action in spatial distribution and motion trajectory, and judge whether the medical staff's action meets the standard; If the medical staff's actions do not meet the standards, improvement suggestions are generated based on the standard actions in the expert database to guide the medical staff to standardize the non-standard actions based on the standard actions in the expert database.
2. The medical action training method based on machine vision model according to claim 1 is characterized in that: The improved YOLOv10-pose deep learning model is used to perform hand motion recognition on the actual medication dispensing action video. Before locating the key points of the hand, the method includes: Collect images or videos of hand movements during medication dispensing, including standard and non-standard movements, and annotate the collected images or videos, including category labels and key points; Preprocessing the hand action image or video using data enhancement technology; Split the preprocessed images or videos into training sets, validation sets, and test sets; Build an improved YOLOv10-pose deep learning model, load the training set, validation set, and test set to train and evaluate the model, and obtain the optimal model; The optimal model is used to recognize hand movements during actual medication dispensing.
3. The medical action training method based on machine vision model according to claim 2 is characterized in that: The preprocessing of the hand motion image or video using data enhancement technology includes: Resize the collected hand motion images or videos to a uniform preset size to ensure that the input data meets the input requirements of the YOLOv10-pose model; The resized medication dispensing hand action images are standardized and data enhanced, wherein the data enhancement processing includes denoising and clarity improvement.
4. The medical action training method based on machine vision model according to claim 1, characterized in that: The improved YOLOv10-pose deep learning model is the RD-YOLOv10 model, which uses an improved RD-Pose detection head to replace the original Pose detection head and uses Focal Loss to replace the original classification loss function.
5. The medical action training method based on machine vision model according to claim 1, characterized in that: The spatial distribution and motion trajectory of each key point are analyzed based on the located key points of the hand; the actual medication dispensing action after clipping is matched with the standard action in the expert database for template matching, and the similarity between the actual medication dispensing action and the standard action in spatial distribution and motion trajectory is compared to determine whether the action of the medical staff meets the standard, including: Calculating the first spatial relative position relationship between two adjacent key points in the actual medication dispensing action according to the located hand key points; Go through each standard action in the expert database one by one, compare the actual medication dispensing action with each standard action in the expert database in turn, and find the standard action that best matches the actual medication dispensing action according to the hand movement characteristics of the actual medication dispensing action; Acquire a second spatial relative position relationship between corresponding key points in the best matching standard action; Calculating a relative error between an actual action and a standard action according to the first spatial relative position relationship and the second spatial relative position relationship; Determine whether the relative error is less than or equal to a preset error threshold; If the relative error is less than or equal to the preset error threshold, it is considered that the action of the medical staff meets the standard; otherwise, it is considered that the action of the medical staff does not meet the standard.
6. The medical action training method based on machine vision model according to claim 1 is characterized in that: The improvement suggestions include graphical interfaces, voice prompts or report forms. If the medical staff's actions do not meet the standards, corrections will be given through a graphical interface or voice, or standard action reference videos or action decomposition diagrams will be provided for medical staff to learn.
7. The medical action training method based on machine vision model according to claim 1, characterized in that: After generating improvement suggestions according to the standard actions in the expert database, the method further includes: Detect whether the medical staff has completed the medication dispensing, and generate an operation evaluation report after the medication dispensing is completed. The operation evaluation report includes the proportion of standard actions of medical staff in medication dispensing, the types of non-standard actions and improvement suggestions, and a historical trend chart of action optimization.
8. The medical action training method based on machine vision model according to claim 1 is characterized in that: The method uses a structured light camera to capture the medication dispensing action. The resolution of the structured light camera is at least 1080p and the frame rate is at least 30fps. The collected image or video data is published to the ROS2 topic in the sensor_msgs / Image format.
9. A medical action training device based on a machine vision model, characterized in that: include: The expert database building module is used to collect standard action videos of dispensing medicines and build an expert database as a detection standard; The data collection module is used to film the actual medication dispensing process of medical staff and collect the actual medication dispensing action video; The action recognition module is used to use the improved YOLOv10-pose deep learning model to perform hand action recognition on the actual medication dispensing action video and locate the key points of the hand; The image cropping module is used to crop and save the model output frame, filter out the background, and retain the hand area image; The action comparison module is used to analyze the spatial distribution and motion trajectory of each key point based on the located hand key points; the actual medication dispensing action after clipping is matched with the standard action in the expert database for template matching, and the similarity between the actual medication dispensing action and the standard action in spatial distribution and motion trajectory is compared to determine whether the medical staff's action meets the standard; The feedback module is used to generate improvement suggestions based on the standard actions in the expert database if the medical staff's actions do not meet the standards, and guide the medical staff to standardize the non-standard actions based on the standard actions in the expert database.
10. A storage medium, characterized in that: The storage medium stores a computer program, which implements the steps of the method according to any one of claims 1 to 8 when executed by a processor.
Citation Information
Patent Citations
Motion evaluation guidance method and system based on deep learning
CN112101315A
Power distribution operation behavior normalization evaluation method and system, and computing device
CN116976721A
Stacked mature strawberry fruit stem posture detection method based on YOLOv10-pose
CN119027936A
Human body action recognition method based on machine learning
CN119152572A
Assistive guidance method and system for exercise, and computer terminal
WO2024138780A1
Cited By
Nursing quality whole-process closed-loop management method and system based on lens enabling
CN121169191A