Multi-camera cooperation method and device for medicine dispensing training and electronic equipment

By acquiring multiple video data of pharmacists' medication dispensing operations through a multi-camera collaborative method, selecting target frame images, and evaluating pharmacists' medication dispensing behavior, the waste and blind spots caused by manual judgment in drug preparation are solved, and automated and high-precision medication dispensing training is achieved.

CN121366447APending Publication Date: 2026-01-20THE FIRST AFFILIATED HOSPITAL OF SOOCHOW UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511638612.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

During the drug preparation process, relying on manual judgment to determine whether pharmacists have made mistakes leads to a waste of human and material resources, and single-angle video data has blind spots.

Method used

By using a multi-camera collaborative method, multiple video data points are acquired during the pharmacist's medication dispensing process. Target frame images are selected from different angles, drug label data is extracted, the accuracy of the pharmacist's medication dispensing behavior is evaluated, and prompt information is generated.

Benefits of technology

It has enabled automated identification of the pharmacist's medication dispensing process, avoiding blind spots, improving identification accuracy, and reducing waste of manpower and resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121366447A_ABST
    Figure CN121366447A_ABST
Patent Text Reader

Abstract

The invention provides a multi-camera cooperation method and device for medicine dispensing training and electronic equipment, and relates to the technical field of medical instruments and image recognizing.The method comprises the steps that after a step starting instruction is obtained, multi-angle video data of medicine dispensing operation of a pharmacist is collected; and after an acquisition step completion instruction, judging the video data to generate prompt information for evaluating whether a medicine dispensing behavior is correct or not, and displaying the prompt information to a human-computer interaction interface, namely extracting a frame image to be evaluated, identifying a hand and a medicine bottle, determining the used medicine bottle through distance calculation, and comparing a target medicine name to judge whether the medicine is correct or not. If the medicine is correct, generating a human body skeleton space-time diagram based on the video, identifying the human body skeleton space-time diagram to obtain an action identification result, comparing a target action label to judge whether the action is correct or not, and finally outputting corresponding prompt information which can be displayed through images, videos or voice. The problem that manpower and material resources are wasted due to the fact that medicine dispensing actions of pharmacists are judged mistakenly manually is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of medical devices and image recognition technology, and in particular to a multi-camera collaborative method and device for dispensing training and electronic equipment. BACKGROUND

[0002] In the process of drug configuration, each drug must be configured according to the specified process. Drug configuration requires professional pharmacists to undergo a lot of training and examination before they can be certified to work as pharmacists. In the conventional training of pharmacists, experienced teachers need to guide and inspect them, which consumes a lot of manpower and material resources. Therefore, a dispensing training device is needed. SUMMARY

[0003] Therefore, the present application provides a multi-camera collaborative method and device for dispensing training and electronic equipment. By obtaining multiple video data of different angles during the dispensing operation of the pharmacist, the accuracy of the pharmacist's dispensing behavior is evaluated based on the label data, and the corresponding prompt information is generated. The present application can solve the problem of relying on manual judgment of the pharmacist's action to determine whether an error has occurred, which causes waste of manpower and material resources.

[0004] Some embodiments of the present application provide a multi-camera collaborative method, device and electronic equipment for dispensing training. The following aspects of the present application are introduced as follows. The embodiments and advantages of the following aspects can be mutually referenced.

[0005] In a first aspect, the present application provides a multi-camera collaborative method for dispensing training, comprising:

[0006] an acquisition step starts an instruction;

[0007] in response to the start instruction, multiple video data of the pharmacist's dispensing operation process are obtained from different angles;

[0008] an acquisition step completes an instruction;

[0009] in response to the completion instruction, target frame images are selected from the video data, wherein the target frame images include the pharmacist's hand and the medicine container, and the distance between the hand and the medicine container meets the preset distance;

[0010] label data of the medicine corresponding to the medicine container is obtained from the target frame image;

[0011] based on the label data, the accuracy of the pharmacist's dispensing behavior is evaluated and corresponding prompt information is generated, and media information is output based on the prompt information.

[0012] According to the embodiments of the present application, by acquiring multiple video data of different angles in the pharmacist dispensing operation process, target frame images are filtered from the video data, and based on the label data of the medicine corresponding to the medicine container in the target frame images, the accuracy of the pharmacist dispensing behavior is evaluated and corresponding prompt information is generated. The automatic identification of the pharmacist dispensing process is realized, and the problem of wasting manpower and material resources caused by relying on the traditional way to judge whether the action performed by the pharmacist is wrong is solved; single-angle video data acquisition of the pharmacist dispensing process may have visual dead angle problems, which may cause some actions in the pharmacist dispensing process to be unable to be identified, and by acquiring multiple video data of different angles in the pharmacist dispensing operation process, the visual dead angle problem can be effectively avoided.

[0013] In a possible implementation of the first aspect, the media information is output based on the prompt information, and the media information includes at least one of the following:

[0014] displaying a first image, wherein the first image includes the prompt information;

[0015] playing a first video, wherein the first video includes the prompt information;

[0016] or, playing a first voice, wherein the first voice includes the prompt information.

[0017] In a possible implementation of the first aspect, the target frame images are filtered from the video data, and the method includes the following steps.

[0018] acquiring a template image and a similarity threshold, wherein the template image includes a hand of the pharmacist and a medicine container, and the distance between the hand and the medicine container meets a preset distance;

[0019] for each frame image in the video data, the similarity between the image and the template image is calculated respectively, and all object frame images with a similarity less than the similarity threshold are filtered from the video data, and the object frame images are integrated to obtain the object frame image set;

[0020] the object frame image set is accurately filtered to obtain the target frame images.

[0021] In a possible implementation of the first aspect, the object frame image set is accurately filtered to obtain the target frame images, and the method includes the following steps.

[0022] the object frame image set is accurately filtered to obtain the target frame images, and the method includes the following steps.

[0023] Feature data of each region of interest is extracted, and each feature data is identified respectively to obtain an identification result corresponding to each feature data, and based on each identification result, the object frame image set is filtered to obtain a target frame image.

[0024] In a possible implementation of the first aspect, the target frame image is filtered from the video data, including:

[0025] Based on the video data, a frame image in which a distance between a hand and a medicament container in the video stream reaches a preset distance is identified, and the frame image is determined as the target frame image.

[0026] In a possible implementation of the first aspect, when there are multiple frame images in which the distance between the hand and the medicament container in the video stream reaches the preset distance, equal-length sub-video data is cut from the video data respectively with each frame image as a starting point.

[0027] From all the sub-video data, target sub-video data containing a behavior of a pharmacist grabbing the medicament container is filtered, and a first frame image in the target sub-video data is taken as the target frame image.

[0028] In a possible implementation of the first aspect, based on the label data, accuracy of the pharmacist dispensing behavior is evaluated and corresponding prompt information is generated, including:

[0029] Target label data is obtained, and the target label data is used to represent a name of a drug that should be used in a current step performed by the pharmacist;

[0030] It is determined whether the label data is same as the target label data, and when the target label data is same as the label data, an operation action of the pharmacist in each video data is identified respectively to obtain an action identification result;

[0031] A target action label is obtained, and it is determined whether the action identification result is same as the target action label, and when the action identification result is correct, prompt information of successful operation is output; and when the action identification result is incorrect, prompt information of incorrect operation is output.

[0032] In a possible implementation of the first aspect, when the target label data is different from the label data, prompt information of incorrect drug use is output.

[0033] In a possible implementation of the first aspect, the operation action of the pharmacist in each video data is identified respectively to obtain the action identification result, including:

[0034] A human skeleton space-time graph of the operation action in a same time period in each video data is extracted;

[0035] Identify each human skeleton space-time graph respectively to obtain an action prediction result corresponding to each human skeleton space-time graph.

[0036] Count the occurrence times of each action prediction result, and determine the action prediction result with the most occurrence times as the action recognition result.

[0037] In a second aspect, the present application provides a multi-camera collaborative system for pharmacy training, comprising:

[0038] A start instruction acquisition module is configured to acquire a start instruction.

[0039] An acquisition module is configured to acquire a plurality of video data of a pharmacist's dispensing operation process captured from different angles in response to the start instruction.

[0040] An end instruction acquisition module is configured to acquire a completion instruction.

[0041] A response module is configured to filter out a target frame image from the video data in response to the completion instruction, the target frame image including the pharmacist's hand and a medicine container, and the distance between the hand and the medicine container meeting a preset distance.

[0042] The response module is further configured to acquire label data of a medicine corresponding to the medicine container from the target frame image.

[0043] The response module is further configured to evaluate the accuracy of the pharmacist's dispensing behavior based on the label data and generate corresponding prompt information, and output media information based on the prompt information.

[0044] In a third aspect, the present application provides an electronic device, comprising: a processor; and a memory having computer program instructions stored therein, wherein when the computer program instructions are run by the processor, the processor executes the method disclosed in the first aspect and any one of the possible implementation manners of the first aspect.

[0045] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is run by a processor, the processor executes the method in the first aspect and any one of the possible implementation manners of the first aspect.

[0046] In a fifth aspect, the present application discloses a device, comprising:

[0047] a memory configured to store instructions for execution by one or more processors of the device, and

[0048] a processor, one of the processors of the device, configured to execute the method disclosed in the first aspect and any one of the possible implementation manners of the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0049] Figure 1 A flow chart of the multi-camera cooperative method for pharmacy training of an embodiment of the present application is shown;

[0050] Figure 2 A man-machine interface diagram of a pharmacy training system of an embodiment of the present application is shown;

[0051] Figure 3 A flow chart of screening target frame images from video data in step S140 of an embodiment of the present application is shown;

[0052] Figure 4 A flow chart of step S143 of an embodiment of the present application is shown;

[0053] Figure 5 A flow chart of screening target frame images from multiple frame images of an embodiment of the present application is shown;

[0054] Figure 6 A flow chart of evaluating the accuracy of pharmacist dispensing behavior of an embodiment of the present application is shown;

[0055] Figure 7 A flow chart of identifying the operation actions of the pharmacist in each video data in step S162 of an embodiment of the present application is shown;

[0056] Figure 8 A block diagram of the device of an embodiment of the present application is shown;

[0057] Figure 9 A block diagram of the SoC (System on Chip) of an embodiment of the present application is shown.

[0058] Explanation of reference numerals:

[0059] 1: Man-machine interface of a pharmacy training system; 2: End training button; 3: Step completion button; 4: Training name display area; 5: Language broadcast information display area; 6: Matters needing attention display area; 7: Guiding video display area; 8: Prompt information display area; 9: Real-time video display area; 10: Execution step display area; 11: Execution sub-step display area. DETAILED DESCRIPTION

[0060] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the protection scope of the present application.

[0061] In order to facilitate the understanding of the technical solutions of the present application, first, the technical problems to be solved by the present application are described.

[0062] Since professional pharmacists need a lot of training and examination to obtain a certificate to perform drug configuration, however, in the training of regular pharmacists, experienced teachers need to guide and inspect them, which consumes a lot of manpower and material resources, and therefore the drug dispensing training device is born.

[0063] To solve the above problems, the present application provides a multi-camera cooperative method for drug dispensing training, which can obtain multiple video data of the pharmacist dispensing operation process from different angles, and screen target frame images from the above video data, then extract drug label data corresponding to the drug container from the target frame images, and finally evaluate the accuracy of the pharmacist dispensing behavior based on the label data and generate corresponding prompt information. The traditional way of relying on manual judgment of whether the action performed by the pharmacist is wrong, thereby causing waste of manpower and material resources. By collecting multiple video data of different angles during the pharmacist dispensing operation process, the visual dead angle problem can be effectively avoided, and the action recognition performed by the pharmacist during the dispensing process is more accurate.

[0064] The multi-camera cooperative method for drug dispensing training of the embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0065] Reference Figure 1 and Figure 2 , Figure 1 The flow chart of the multi-camera cooperative method for drug dispensing training of the embodiment of the present application is shown. Figure 2 The man-machine interface diagram of the drug dispensing training system of the embodiment of the present application is shown.

[0066] As Figure 1 shown, the method steps include S110-S140.

[0067] S110, start the acquisition step instruction.

[0068] In the embodiment of the present application, taking the parenteral nutrition liquid mixing and dispensing training as an example, as Figure 2 shown, the man-machine interface 1 of the drug dispensing training system includes multiple modules, such as an end training button 2, a completion button 3, a training name display area 4, a language broadcast information display area 5, a notice display area 6, a guide video display area 7, a prompt information display area 8, a real-time video display area 9 and an execution step display area 10. Among them, as Figure 2The training name display area 4 shows the name of this medication preparation training – Parenteral Nutrition Solution Mixing and Preparation Training. When the pharmacist is about to perform the step of drawing sodium glycerophosphate solution in the parenteral nutrition solution mixing and preparation training, the medication preparation training system can announce a prompt message asking the pharmacist to open a 20ml syringe to draw sodium glycerophosphate, thus notifying the pharmacist that the step is about to begin. At this time, the user can enter voice input or click a specific button on the display screen to confirm the start of the training. In some embodiments, the user may not directly enter information as a step start command; instead, the step start command can be considered based on the user not making an exit operation within 5 seconds.

[0069] S120 responds to the step-by-step start command and acquires multiple video data of the pharmacist's medication dispensing operation process captured from different angles.

[0070] like Figure 2 As shown, after the medication training system announces the prompt to open a 20ml syringe to draw sodium glycerophosphate, it displays this prompt in the voice announcement display area 5. At this time, the controller generates a first control signal to control cameras at different angles to begin capturing the entire process of the pharmacist's medication preparation.

[0071] In an embodiment of the invention, the pharmacist performs medication dispensing operations at a workbench equipped with a top camera, a side camera, and a head-mounted camera. A controller is communicatively connected to each of these cameras. After the medication dispensing training system completes its voice broadcast, the controller generates control signals to control each camera to begin capturing the entire medication dispensing process.

[0072] It should be noted that the first control signal mentioned above is a high level. When the controller receives the step-start command, it generates a high level to control each camera to turn on.

[0073] S130, obtain the step completion instruction.

[0074] like Figure 2 As shown, after the pharmacist completes the medication dispensing operation, they can click the "Step Complete" button 3 on the human-computer interaction interface 1 of the medication dispensing training system to indicate that the steps performed by the pharmacist have been completed.

[0075] S140, in response to the step completion instruction, a target frame image is selected from the video data, wherein the target frame image includes the pharmacist's hand and the medicine container, and the distance between the hand and the medicine container conforms to a preset distance; the label data of the medicine corresponding to the medicine container is obtained from the target frame image; based on the label data, the accuracy of the pharmacist's dispensing behavior is evaluated and corresponding prompt information is generated, and media information is output based on the prompt information.

[0076] like Figure 2When the pharmacist triggers the step completion button 3, the controller receives the step completion instruction and generates a second control signal to control each camera to stop collecting video data, then displays the collected video data at the real-time video display area 9, and finally the processor analyzes each video data to obtain the prompt information.

[0077] In the embodiment of the present application, the label data of the medicine corresponding to the medicine container is obtained from the target frame image by the OCR technology.

[0078] It should be noted that the above-mentioned second control signal is a low level, and when the controller receives the step completion instruction, a low level is generated to control each camera to turn off. Subsequently, the processor processes and analyzes the video data collected by the overhead camera, the side-view camera and the head-mounted camera to obtain the prompt information for evaluating whether the pharmacist's dispensing behavior is correct, and stores the prompt information in the database. The human-computer interaction interface 1 obtains the prompt information in the above-mentioned database through the interface, and displays the prompt information in the prompt information display area 8, and modifies the state of each dispensing operation sub-step in the execution step display area 10.

[0079] The video data collected from a single angle of the pharmacist's dispensing process has a visual dead angle problem, which makes the collected pharmacist's dispensing operation process incomplete. If the processor only analyzes the video data from a single angle, the above-mentioned prompt information is inaccurate. If the processor analyzes multiple video data from different angles to obtain the prompt information, the problem of inaccurate prompt information can be effectively avoided.

[0080] In the embodiment of the present application, the media information is output based on the prompt information, including at least one of the following items:

[0081] displaying a first image, wherein the first image includes the prompt information;

[0082] playing a first video, wherein the first video includes the prompt information;

[0083] or, playing a first voice, wherein the first voice includes the prompt information.

[0084] As Figure 2As shown, when the system has not yet identified a certain medication preparation operation sub-step, the status of that sub-step is displayed as "Not Started". When the system is detecting the medication preparation operation sub-steps of opening the 20ml syringe and dropping the empty ampoule body into the blue box, the prompt information display area 8 will display an error message indicating that a 20ml syringe should be used and the empty ampoule body should not be dropped into the blue box, and then change the status of that sub-step to "Detecting". When the system does not detect the operation of drawing sodium glycerophosphate with a 20ml syringe, the prompt information display area 8 will display an error message indicating that sodium glycerophosphate was not drawn with a 20ml syringe, and then change the status of that sub-step to "Not Operated". When the system detects the operation of breaking the ampoule, the system only changes its corresponding status to "Completed" to indicate that the operation is correct. When the user chooses to skip the detection and drops the broken ampoule head into the sharps box, the system only changes its corresponding status to "Skipped".

[0085] like Figure 1 As shown, the following describes the "End Training" button 2, the precautions display area 6, the instruction video display area 7, and the execution steps display area 10 in the human-computer interaction interface 1 of the medication dispensing training system. Specifically, if a pharmacist needs to exit the training due to an unforeseen situation during medication dispensing operation training, they can exit the training by clicking the "End Training" button 2. The precautions display area 6 displays precautions during the medication dispensing operation, such as avoiding touching the needle plunger when pushing and withdrawing the syringe with the right hand. The instruction video display area 7 displays a standardized video of the medication dispensing operation. Pharmacists can watch the standardized video of the medication dispensing operation in the instruction video display area 7 before or after completing the medication dispensing operation training to facilitate learning.

[0086] The execution step display area 10 is used to display which drug preparation operation sub-steps are included in the current drug preparation training, the status of each drug preparation operation sub-step, and the operations that can be performed on each drug preparation operation sub-step. As shown in the execution sub-step display area 11, when identifying the step of drawing sodium glycerophosphate with a syringe, it is necessary to simultaneously identify the following drug preparation operation sub-steps: opening the 20ml syringe, breaking open the ampoule, discarding the broken ampoule head into the sharps box, drawing sodium glycerophosphate with the 20ml syringe, and discarding the empty ampoule body into the blue box.

[0087] Below, on Figure 3 The S140 step is described in further detail.

[0088] refer to Figure 3 , Figure 4 This illustrates a flowchart of an embodiment of the present invention for filtering target frame images from video data, wherein filtering target frame images from video data includes:

[0089] S141, acquire a template image and a similarity threshold, wherein the template image comprises a hand of a pharmacist and a medicine container, and a distance between the hand and the medicine container meets a preset distance.

[0090] It should be noted that the medicine container in the template image is a medicine container corresponding to a medicine used in the current dispensing operation.

[0091] S142, for each frame image in the video data, similarity between the frame image and the template image is calculated respectively, and all object frame images with similarity less than the similarity threshold are filtered out from the video data, and the object frame image set is obtained by integrating the object frame images.

[0092] S143, the target frame image is obtained by accurately screening the object frame image set.

[0093] It should be noted that the similarity is essentially a quantification of the surface characteristics such as image pixel distribution and contour matching, and the object frame is screened by the similarity between the image and the template image, which cannot accurately verify the core technical condition of the target frame, that is, whether the actual distance between the hand and the medicine container strictly meets the preset distance, so the object frame image set needs to be accurately screened.

[0094] First, the object frame image set is screened by calculating the similarity between each frame image in the video data and the template image, and then the object frame image set is accurately screened, which can eliminate a large number of noise images before accurate screening, effectively avoiding too many noise images from entering accurate identification, and wasting computing resources.

[0095] Reference Figure 4 , Figure 5 The flowchart of step S143 of the embodiment of the application is shown, wherein the target frame image is obtained by accurately screening the object frame image set, comprising:

[0096] S143-1, each object frame image in the object frame image set is regionally divided to obtain a region of interest in each object frame image, wherein the region of interest comprises a hand of a pharmacist and a medicine container, and a distance between the hand and the medicine container meets a preset distance part of the image.

[0097] It should be noted that by dividing the region of interest of each object frame image, the background noise in the object frame image can be effectively eliminated, so as to reduce the waste of computing resources while improving the recognition accuracy.

[0098] S143-2, feature data of each region of interest is extracted, and each feature data is identified respectively to obtain an identification result corresponding to each feature data, and the object frame image set is screened based on each identification result to obtain the target frame image.

[0099] It should be noted that the above identification result is a numerical result for judging the probability that the distance between the hand and the medicine container in the region of interest meets the preset distance. Since the region of interest itself has limited the region where the distance meets the preset distance, the feature recognition can further verify this condition through detailed feature verification, such as the relative position of the hand edge and the container edge, the pixel distance distribution, to avoid false positive frames in the preliminary screening that are visually similar but actually do not meet the distance.

[0100] By extracting the exclusive features of the hand and the medicine container, such as the finger contour, the holding posture, the bottle shape, and the label position, it is ensured that the screened frames indeed contain the pharmacist's hand and the medicine container, and the risk of misjudging other containers as target objects is excluded.

[0101] In an embodiment of the present application, an AlexNet network model is used to extract feature data of each region of interest, and each feature data is identified to obtain an identification result corresponding to each feature data. It can be understood that the AlexNet network model solves the gradient vanishing problem of deep network by introducing the ReLU activation function; it breaks through the algorithm power bottleneck by first large-scale application of GPU parallel computing; and it can effectively retain more image details in the region of interest by designing local response normalization and overlapping pooling. At the same time, this mode surpasses the traditional mode of combining traditional manual feature extraction and traditional classifiers, and realizes end-to-end automatic learning from raw pixels to high-order features.

[0102] In another embodiment of the present application, a FEMNet network model is used to extract feature data of each region of interest, and each feature data is identified to obtain an identification result corresponding to each feature data. It can be understood that the FEMNet network model enhances the extraction capability of global and local features of the region of interest by proposing a feature context embedding module; and improves the recognition accuracy of the target frame image by realizing flexible matching of features between the target frame image and the template image based on a matching network.

[0103] In an embodiment of the present application, the target frame image is screened from the video data, including: based on the video data, identifying a frame image in which the distance between the hand and the medicine container in the video stream reaches a preset distance, and determining the frame image as the target frame image.

[0104] It should be noted that when the pharmacist takes the medicine container, the hand of the pharmacist will inevitably approach the medicine container. When the distance between the hand and the medicine container is greater than the preset distance, it is impossible to determine which medicine container the pharmacist wants to take. When the distance between the hand and the medicine container is less than the preset distance, the label data of the medicine on the medicine container may be blocked by the hand, so that the label data cannot be extracted in the subsequent step.

[0105] The preset distance needs to be less than half of the distance between the medicine containers. The target frame image is screened by judging whether the distance between the hand and the medicine container in the video stream frame image reaches the preset distance. Through the distance setting judgment, the medicine container to be taken by the operator can be predicted, so that the label of the medicine container is obtained in advance. Not only the acquisition of the medicine label can be ensured, but also the technical realization threshold is low, easy to quickly land, and the redundant frames can be accurately filtered.

[0106] In the embodiment of the application, the target frame image is determined by judging whether the Euclidean distance between the center point of the hand region and the center point of the medicine container region in the frame image reaches the preset distance.

[0107] In the embodiment of the application, the target frame image is determined by judging whether the Euclidean distance between the particle of the hand region and the particle of the medicine container region in the frame image reaches the preset distance.

[0108] Reference Figure 5 , Figure 6 The flowchart of the embodiment of the application is shown, which screens the target frame image from multiple frame images. When there are multiple frame images in which the distance between the hand and the medicine container in the video stream reaches the preset distance, the target frame image is screened from the multiple frame images, including:

[0109] S151, respectively taking the same length of sub-video data from the video data as the starting point of each frame image.

[0110] S152, from all the sub-video data, the target sub-video data containing the behavior of the pharmacist grabbing the medicine container is screened, and the first frame image in the target sub-video data is taken as the target frame image.

[0111] It should be noted that when there are multiple frame images in which the distance between the hand and the medicine container in the video stream reaches the preset distance, the target sub-video data containing the behavior of the pharmacist grabbing the medicine container is screened by identifying the sub-video data, thereby avoiding the problem that when the pharmacist takes the medicine, the pharmacist's hand passes by other medicine containers without taking the medicine, but this frame image is identified as the target frame image.

[0112] In the embodiment of the application, each sub-video data is identified by a spatial-temporal graph convolutional network (STGCN), and then the target sub-video data containing the behavior of the pharmacist grabbing the medicine container is screened from all the sub-video data.

[0113] In another embodiment of the present application, each sub-video data is identified by a two-stream network (Two-Stream CNN), and then target sub-video data containing the behavior of the pharmacist grabbing the medicine container is screened out from all sub-video data.

[0114] Reference Figure 6 , Figure 7 A flowchart for evaluating the accuracy of the pharmacist dispensing behavior is shown, wherein the accuracy of the pharmacist dispensing behavior is evaluated based on the label data and corresponding prompt information is generated, including:

[0115] S161, target label data is obtained, wherein the target label data is used to represent the name of the medicine that should be used in the current step performed by the pharmacist.

[0116] S162, it is judged whether the label data is the same as the target label data, and when the target label data is the same as the label data, the operation action of the pharmacist in each video data is identified respectively to obtain an action recognition result.

[0117] In an embodiment of the present application, it is judged whether the label data is the same as the target label data by the method of character-by-character sequential comparison.

[0118] In an embodiment of the present application, when the target label data is different from the label data, medicine use error prompt information is output.

[0119] S163, target action label is obtained, and it is judged whether the action recognition result is the same as the target action label, when the action recognition result is correct, operation success prompt information is output; when the action recognition result is incorrect, operation error prompt information is output.

[0120] It should be noted that when the target action label is to open a 20ml syringe, if the action recognition result at this time is also to open a 20ml syringe, operation success prompt information is output; if the action recognition result at this time is not to open a 20ml syringe, operation error prompt information is output.

[0121] By comparing the target label data with the label data, the risk of misusing medicine is directly intercepted, for example, the basic error of mistakenly taking sodium glycerophosphate as water-soluble vitamins is intercepted from the first step of dispensing, and the core risk point is blocked. Under the premise of correct medicine, it is further verified whether the operation action conforms to the standard to avoid the implicit risk of correct medicine but incorrect operation, and a double safety barrier of medicine compliance and action compliance is formed.

[0122] Reference Figure 7 , Figure 3A flowchart of identifying the operation actions of the pharmacists in each video data in step S162 is shown, wherein the operation actions of the pharmacists in each video data are identified respectively to obtain action recognition results, including:

[0123] S171, extracting the human skeleton space-time graph of the operation actions in the same time period in each video data.

[0124] It should be noted that before step S171 is performed, each video data also needs to be segmented by using a time window with the same size and at the same time step.

[0125] In the embodiment of the present application, the human skeleton space-time graph of each operation action is extracted by using the OpenPose model.

[0126] S172, identifying each human skeleton space-time graph respectively to obtain an action prediction result corresponding to each human skeleton space-time graph.

[0127] In the embodiment of the present application, each human skeleton space-time graph is identified by using a spatial-temporal graph convolutional network model (STGCN) to obtain an action prediction result corresponding to each human skeleton space-time graph.

[0128] S173, counting the occurrence times of each action prediction result, and determining the action prediction result with the most occurrence times as the action recognition result.

[0129] It should be noted that the human skeleton data only retains the joint positions and connection relationships of the pharmacists, directly filters irrelevant visual information such as the color of the medicine package and the change of the light, and avoids the confusion of action features caused by these factors. The human skeleton space-time graph simultaneously fuses the spatial correlation and the time dynamics, and compared with a single frame skeleton or pure spatial data, can more completely restore the action logic of the dispensing operation and avoid the action misjudgment caused by the loss of time sequence information. By using the multi-prediction result statistical voting mechanism, the action prediction result with the most occurrence times is determined as the action recognition result, which avoids the error of the final action recognition result caused by the visual dead angle of the single-angle video data.

[0130] The present application provides a multi-camera cooperative system for dispensing training, comprising:

[0131] A start instruction acquisition module is configured to acquire a start instruction.

[0132] An acquisition module is configured to acquire a plurality of video data of the dispensing operation process of the pharmacists from different angles in response to the start instruction.

[0133] An end instruction obtaining module, configured to obtain a step completion instruction;

[0134] A response module, configured to respond to the step completion instruction, and filter a target frame image from the video data, the target frame image including the hand of the pharmacist and the medicine container, and the distance between the hand and the medicine container meeting a preset distance;

[0135] The response module is further configured to obtain label data of the medicine corresponding to the medicine container from the target frame image;

[0136] The response module is further configured to evaluate the accuracy of the pharmacist's dispensing behavior based on the label data and generate corresponding prompt information, and output media information based on the prompt information.

[0137] It should be noted that the functions and effects of the modules of the multi-camera collaborative system for dispensing training are the same as those of the above-mentioned embodiments, and specific reference can be made to the above-mentioned embodiments. Figure 8 The functions and effects of the modules of the multi-camera collaborative system for dispensing training are the same as those of the above-mentioned embodiments, and specific reference can be made to the above-mentioned embodiments.

[0138] Reference is now made to the drawing of FIG. 12, Figure 8 which shows a block diagram of a device 1200 according to one embodiment of the present application. The device 1200 can include one or more processors 1201 coupled to a controller hub 1203. For at least one embodiment, the controller hub 1203 communicates with the processor(s) 1201 via a bus, such as a Peripheral Component Interconnect (PCI) bus, a HyperTransport® bus, or industry standard architecture (ISA). Additionally, the controller hub 1203 can communicate with a graphics processor 1205 via a bus, such as an Accelerated Graphics Port (AGP) bus or a HyperTransport® bus. Also, the controller hub 1203 can be coupled to a memory 1209, a graphics processor 1205, a display 1211, a wireless transceiver 1213, a speaker / microphone 1215, a keyboard and mouse 1217, and a disk storage 1219 via an I / O controller 1221. In one embodiment, the I / O controller 1221 can include a storage interface (e.g., a SATA or SCSI interface) for the disk storage 1219. In another embodiment, the I / O controller 1221 can include a USB interface, a Bluetooth® interface, or an IEEE 1394 interface for the wireless transceiver 1213, the keyboard and mouse 1217, and the speaker / microphone 1215. In some embodiments, the I / O controller 1221 can include an audio interface for the speaker / microphone 1215. In some embodiments, the I / O controller 1221 can include a wireless transceiver interface for the wireless transceiver 1213. In some embodiments, the I / O controller 1221 can include a storage interface for the disk storage 1219.

[0139] The device 1200 can also include a coprocessor 1202 coupled to the controller hub 1203. Alternatively, one or both of the memory and the GMCH can be integrated into the processor (as described in the present application), the memory 1204 and the coprocessor 1202 are directly coupled to the processor 1201 and the controller hub 1203 is integrated into the IOH. The memory 1204 can be, for example, a dynamic random access memory (DRAM) such as a synchronous dynamic random access memory (SDRAM), a PCMA or a combination of both. In one embodiment, the coprocessor 1202 is a special-purpose processor, such as for example a high-throughput MIC processor (Many Integrated Core), a network or communication processor, compression engine, graphics processor, a general purpose computing on GPU (GPGPU) processor, embedded processor, or the like. The optional nature of the coprocessor 1202 is denoted in Figure 8

[0140] The memory 1204, as a computer-readable storage medium, can include one or more tangible, non-transitory, computer-readable media used to store data and / or instructions for use by or in connection with the computer system 1200. In the illustrated example, memory 1204 includes a volatile memory 1208 and a non-volatile memory 1210. Common forms of non-volatile memory include, for example, a floppy disk, a flexible disk, hard disk, solid-state drive, magnetic tape, or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any other physical and / or tangible computer-readable medium and / or storage device(s) that can be used to store desired program code means in the form of computer-executable instructions or data structures and that can be accessed by a general purpose or special purpose computer system, or a combination thereof. In the illustrated example, memory 1204 includes a volatile memory 1208 and a non-volatile memory 1210. Common forms of volatile memory include, for example, a random access memory (RAM), dynamic random access memory (DRAM), or static random access memory (SRAM). Common forms of non-volatile memory include, for example, a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or flash memory. The memories 1208, 1210 can be used to store data and / or instructions for use when executing programs by the processor 1201. Memories 1208, 1210 can also be used to store temporary variables or other intermediate information during program execution.

[0141] In one embodiment, the device 1200 can further include a network interface controller (NIC) 1206. The network interface 1206 can include a transceiver to provide a radio interface to the device 1200 to enable communications with any other suitable device (e.g., a front end module, an antenna, etc.). In various embodiments, the network interface 1206 can be integrated with other components of the device 1200. The network interface 1206 can implement the functionality of the communication units in the above-described embodiments.

[0142] ​The device 1200 can further include an Input / Output (I / O) device 1205. The I / O 1205 can include a user interface designed to enable a user to interact with the device 1200, a peripheral component interface designed to enable peripheral components to interact with the device 1200, and / or a sensor designed to determine environmental conditions and / or location information related to the device 1200.

[0143] Notably, Figure 8 are merely exemplary. That is, although Figure 8 In some embodiments, the device 1200 includes a processor 1201, a controller hub 1203, a memory 1204, and the like, as shown in FIG. 12. However, in actual applications, a device using the methods of the present application can include only a portion of the components of the device 1200, for example, the processor 1201 and the NIC 1206. Figure 9 The optional nature of various components is represented by dashed lines in FIG. 12. According to some embodiments of the present application, the memory 1204, which is a computer readable storage medium, stores instructions thereon that, when executed by a computer, cause the device 1200 to perform a multi-camera coordination method for pharmacy training according to one of the above embodiments, which can be specifically referred to the method of the above embodiments, and thus will not be described here again.

[0144] Reference is now made to Figure 9 FIG. 13, which shows a block diagram of a System on Chip (SoC) 1300 according to an embodiment of the present application. In Figure 9 Similar elements in FIG. 13 bear like reference numerals. Additionally, dashed lined boxes represent optional features. In ​ In one embodiment, the SoC 1300 includes an interconnect unit 1350 coupled to an application processor 1310; a system agent unit 1380; a bus controller unit 1390; an integrated memory controller unit 1340; a set or one or more coprocessors 1320, which can include integrated graphics logic, an image processor, an audio processor, and a video processor; a static random access memory (SRAM) unit 1330; and a direct memory access (DMA) unit 1360. In one embodiment, the coprocessors 1320 include a special-purpose processor, such as for example a network or communication processor, compression engine, GPGPU, a high- throughput MIC processor, embedded processor, or the like.

[0145] The static random access memory (SRAM) unit 1330 can include one or more computer-readable media for storing data and / or instructions. The computer-readable storage media can store instructions, in particular, a transient or a permanent copy of the instructions. The instructions can include causing the Soc 1300 to perform a multi-camera coordination method for pharmacy training according to one of the above-described embodiments, in particular, the method of the above-described embodiments, which will not be repeated here.

[0146] Embodiments of the mechanisms disclosed herein can be implemented in hardware, software, firmware, or any combination thereof. Embodiments of the application can be implemented as computer programs or program code executing on programmable systems comprising at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0147] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices, in known fashion. For purposes of this application, a processing system includes any system that has a processor, such as a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0148] The program code can be implemented in a high level procedural or object oriented programming language to communicate with a processing system. The program code can be implemented in assembly or machine language, if desired. In fact, the mechanisms described herein are not limited in scope to any particular programming language. In any case, the language can be a compiled or interpreted language.

[0149] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments can also be implemented as instructions carried by or stored on one or more transitory or non-transitory machine- readable (e.g., computer-readable) media, which can be read and executed by one or more processors. For example, the instructions can be distributed over the network or by other computer readable media. Thus, a machine-readable medium can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including without limitation, a floppy disk, an optical disc, an optical disc read only memory (CD-ROM), a magnetic disk read only memory (DVD-ROM), a read only memory (ROM), a random access memory (RAM), an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM), a magnetic or optical card, a flash memory, or a tangible, machine-readable storage used in the transmission of information over the Internet with a propagation signal in a form of an electromagnetic wave, a sound wave, or other form of propagated signal. Accordingly, a machine-readable medium includes any type of mechanical or optical

[0150] In the drawings, some of the structural or methodological features can be shown in particular arrangements and / or orders. However, it should be understood that such particular arrangements and / or orders can not be required. Instead, in some embodiments, the features can be arranged in a different manner and / or order than that shown in the figures of the specification. Additionally, inclusion of a structural or methodological feature in a particular figure is not meant to imply that such feature is required in all embodiments, and in some embodiments, the features can not be included or can be combined with other features.

[0151] It should be noted that each unit / module mentioned in each device embodiment of the present application is a logical unit / module, and in physical, one logical unit / module can be a physical unit / module, or a part of a physical unit / module, or be realized in a combination of multiple physical unit / modules, and the physical realization of these logical units / modules is not the most important, and the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce the units / modules which are not closely related to solving the technical problems proposed in the present application, which does not mean that the above-mentioned device embodiments do not have other units / modules.

[0152] It should be noted that in the examples and descriptions of the present patent, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including one" does not exclude the presence of other identical elements in the process, method, article or device including the element.

[0153] Although the present application has been illustrated and described with reference to certain preferred embodiments thereof, it should be understood by those skilled in the art that various changes in form and details can be made therein without departing from the spirit and scope of the present application.

Claims

1. A multi-camera coordinated method for pharmacy training, characterized in that, The method comprises the following steps: an acquisition step start instruction is obtained; in response to the step start instruction, a plurality of video data of pharmacist dispensing operation processes taken at different angles are acquired; an acquisition step completion instruction is obtained; in response to the step completion instruction, target frame images are filtered out from the video data, wherein the target frame images include a hand of a pharmacist and a medicine container, and the distance between the hand and the medicine container meets a preset distance; label data of a medicine corresponding to the medicine container is acquired from the target frame images; based on the label data, the accuracy of the pharmacist dispensing behavior is evaluated, and corresponding prompt information is generated, and media information is output based on the prompt information.

2. The method of claim 1, wherein, The media information output based on the prompt information comprises at least one of the following: a first image is displayed, wherein the first image comprises the prompt information; a first video is played, wherein the first video comprises the prompt information; or a first voice is played, wherein the first voice comprises the prompt information.

3. The method of claim 1, wherein, The filtering of the target frame images from the video data comprises the following steps: a template image and a similarity threshold are acquired, wherein the template image includes a hand of a pharmacist and a medicine container, and the distance between the hand and the medicine container meets a preset distance; for each frame image in the video data, the similarity between the image and the template image is calculated respectively, and all object frame images with a similarity less than the similarity threshold are filtered out from the video data, and the object frame images are integrated to obtain an object frame image set; the object frame image set is accurately filtered to obtain the target frame images.

4. The method of claim 3, wherein, The accurate filtering of the object frame image set to obtain the target frame images comprises the following steps: each object frame image in the object frame image set is regionally divided to obtain a region of interest in each object frame image, wherein the region of interest includes a hand of a pharmacist and a medicine container, and the distance between the hand and the medicine container meets a preset distance part of the image; feature data of each region of interest is extracted, and each feature data is identified respectively to obtain an identification result corresponding to each feature data, and the object frame image set is filtered based on each identification result to obtain the target frame images.

5. The method of claim 1, wherein, The filtering of the target frame images from the video data comprises the following steps: based on the video data, frame images in which the distance between the hand and the medicine container in the video stream reaches a preset distance are identified, and the frame images are determined as the target frame images.

6. The method of claim 5, wherein, When there are multiple frame images in which the distance between the hand and the medicine container in the video stream reaches a preset distance, equal-length sub-video data are intercepted from the video data respectively starting from each frame image; target sub-video data containing the behavior of the pharmacist grabbing the medicine container are filtered out from all the sub-video data, and a first frame image in the target sub-video data is taken as the target frame image.

7. The method of claim 1, wherein, The evaluation of the accuracy of the pharmacist dispensing behavior based on the label data and the generation of corresponding prompt information comprises the following steps: Obtaining target label data, the target label data is used to represent a drug name that should be used in a current step performed by the pharmacist; Determining whether the label data is the same as the target label data, when the target label data is the same as the label data, identifying the operation action of the pharmacist in each video data respectively to obtain an action recognition result; Obtaining a target action label, determining whether the action recognition result is the same as the target action label, when the action recognition result is correct, outputting a prompt information of successful operation, and when the action recognition result is incorrect, outputting a prompt information of incorrect operation.

8. The method of claim 7, wherein, When the target label data is different from the label data, outputting a prompt information of incorrect drug use.

9. The method of claim 7, wherein, The identifying the operation action of the pharmacist in each video data respectively to obtain an action recognition result comprises: Extracting a human skeleton space-time graph of the operation action in the same time period in each video data; Identifying each human skeleton space-time graph respectively to obtain an action prediction result corresponding to each human skeleton space-time graph; Counting the occurrence times of each action prediction result, and determining the action prediction result with the most occurrence times as the action recognition result.

10. A multi-camera coordinated system for dispensing training, characterized in that, Comprise: A start instruction obtaining module for obtaining a step start instruction; A collection module for obtaining a plurality of video data of pharmacist dispensing operation process shot from different angles in response to the step start instruction; An end instruction obtaining module for obtaining a step completion instruction; A response module for filtering out a target frame image from the video data in response to the step completion instruction, the target frame image comprising a hand of the pharmacist and a drug container, and the distance between the hand and the drug container meeting a preset distance; The response module is further used to obtain label data of a drug corresponding to the drug container from the target frame image; The response module is further used to evaluate the accuracy of the pharmacist dispensing behavior based on the label data and generate corresponding prompt information, and output media information based on the prompt information.

11. An electronic device, comprising: Comprise: A processor; And a memory, the memory has computer program instructions stored therein, When the computer program instructions are run by the processor, the processor executes the method of any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is run by the processor to make the processor execute the method of any one of claims 1-9. The computer readable storage medium stores a computer program, and the computer program is run by the processor to make the processor execute the method of any one of claims 1-9.