A real-time drug identification method and system based on deep learning

Through the real-time drug recognition method based on deep learning, the images collected by the camera are used to detect drugs and text, and the appearance and text information are integrated, which solves the problems of slow recognition speed and low accuracy of existing drug recognition methods, real-time accurate identification and efficient verification of drug types.

CN114937176BActive Publication Date: 2025-05-06HUNAN SHENFAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210560976.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-05-06
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

Existing drug recognition methods rely on manual identification and appearance information, resulting in slow recognition speed, low accuracy, and prone to misjudgment of drugs with similar appearance, increasing the recognition error rate.

Method used

Real-time drug recognition method based on deep learning is adopted, and images collected by loading the camera are used to detect drugs and text using pre-trained models, integrating appearance and text information to achieve real-time and accurate recognition of drug types.

Benefits of technology

It improves the work efficiency of drug identification, reduces the complexity and intensity of work, increases the reliability of the identification and verification links, reduces the psychological pressure of medical staff, and supports the identification of multiple drugs at a time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114937176B_ABST
    Figure CN114937176B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time drug recognition method and system based on deep learning, which processes the loaded first image and the second image by loading a first image captured by a first camera and a second image captured by a second camera, respectively, to obtain drug and text detection results; fuses the acquired drug classification and drug area, outputs the classification confidence based on appearance and the fused drug area; performs text area cutting according to the acquired text area and the output fused drug area; performs text recognition on the first image and the second image after the text area is cut, respectively, to identify the text information of the text area cutting; fuses the output classification confidence based on appearance and the identified text information, and finally outputs the recognition result based on appearance and text information. The present invention can realize the real-time verification of prescription drugs in the dispensing link, increase the verification accuracy, and reduce the burden on staff.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical technology, and in particular discloses a real-time drug recognition method and system based on deep learning. Background Art

[0002] With the development of medical technology and the increase in urban population, the types and quantities of injectable drugs required by hospitals every day are very huge, and the identification and verification of drug preparation are facing great pressure. The existing drug identification method mainly relies on medical staff to manually identify and verify the types of drugs. At present, the workload of medical staff in hospitals is heavy, and the unreasonable site configuration leads to chaotic management of the pharmacy, which greatly increases the recognition error rate. There are many drugs from various manufacturers that look similar but have completely different functions, which increases the difficulty of drug identification. As the last checkpoint for medication, drug identification and verification is extremely important.

[0003] As the core force of a new round of scientific and technological revolution and industrial transformation, deep learning is promoting the upgrading of traditional industries, driving the rapid development of the "unmanned economy", and having a positive impact on people's livelihood areas such as smart transportation, smart home, and smart medical care. At present, the target detection network based on deep neural networks can accurately locate and identify targets in thousands of different natural scenes. At the same time, the scene text detection and recognition network based on deep learning can locate and recognize text at a near real-time speed. The use of real-time drug recognition methods based on deep learning can greatly improve the work efficiency of drug recognition, reduce work complexity, reduce work intensity, increase the reliability of drug recognition and verification links, and reduce the psychological pressure of medical staff.

[0004] At present, most intelligent drug recognition methods mainly rely on taking pictures to obtain drug images, and then rely on appearance information to identify the type of drug. For example, patent CN113837070 mainly relies on appearance information to process drug images to achieve drug inspection and classification. However, this method of relying on taking pictures to obtain drug images and then using appearance information to identify drugs has the following two problems. First, the operator needs to manually take pictures after placing the drugs, which increases the burden on the staff, and the recognition speed cannot meet the real-time requirements. Second, relying on appearance information to identify drugs is prone to misjudgment of drugs with similar appearances, and the recognition accuracy cannot meet actual needs.

[0005] Therefore, the above-mentioned defects of the existing drug identification method are a technical problem that needs to be solved urgently. Summary of the invention

[0006] The present invention provides a real-time drug identification method and system based on deep learning, aiming to solve the technical problems of the above-mentioned defects in existing drug identification methods.

[0007] One aspect of the present invention relates to a real-time drug recognition method based on deep learning, comprising the following steps:

[0008] Load the first image captured by the first camera and the second image captured by the second camera respectively;

[0009] Using the pre-trained model to process the loaded first image and the second image respectively, and obtain drug and text detection results in the first image and the second image respectively, wherein the drug and text detection results include drug classification, drug area and text area;

[0010] The acquired drug classification and drug area are fused, and the appearance-based classification confidence and fused drug area are output;

[0011] Performing text region cutting on the first image and the second image according to the acquired text region and the output fused medicine region;

[0012] Using a pre-trained model, respectively perform text recognition on the first image and the second image after the text region is cut, and recognize text information in the first image and the second image after the text region is cut;

[0013] The output appearance-based classification confidence and the recognized text information are fused, and finally the recognition result based on appearance and text information is output.

[0014] Furthermore, before the step of loading the first image captured by the first camera and the second image captured by the second camera respectively, the step includes:

[0015] In the training phase, the model is trained, and the model includes a drug and text detection network model and a text recognition network model;

[0016] The steps to train the drug and text detection network model include:

[0017] Cut the image to be trained into sample images of the same size and load relevant annotations;

[0018] Perform horizontal or flipping actions on sample images loaded with relevant annotations to perform data expansion;

[0019] Pre-trained model loaded on COCO dataset;

[0020] During the training process, different data enhancement methods are used to increase the amount of data, and the drug and text detection network model is trained on the training set until convergence;

[0021] The steps to train a text recognition network model include:

[0022] Use text synthesis programs to synthesize text images of commonly used drugs to increase the amount of training data;

[0023] Load a pre-trained model trained on a dataset;

[0024] A text recognition training set is cut out from the real training set according to the text annotation, and the text recognition training set is merged with the synthetic drug name dataset;

[0025] Train the text recognition network model in the merged text recognition training set until convergence.

[0026] Furthermore, the steps of respectively loading the first image captured by the first camera and the second image captured by the second camera include:

[0027] Real-time drug identification video collected by a video acquisition box is obtained. The inner wall of the video acquisition box is made of non-reflective white material. A strip light source for providing stable lighting conditions is provided on the upper part of the video acquisition box, and a light shielding cover is covered on the top of the video acquisition box to prevent the light of the strip light source from directly shining into the eyes of the operator and causing discomfort. A drug placement area is provided at the bottom of the video acquisition box, and the first camera and the second camera are symmetrically installed on both sides of the drug placement area and the installation height is half of the average height of commonly used drugs.

[0028] Furthermore, the steps of fusing the acquired drug classification and drug region and outputting the appearance-based classification confidence and the fused drug region include:

[0029] Calculate the intersection-and-union ratio of all detection frames to obtain the association relationship between each drug in the first image and the second image;

[0030] If objects with an IoU ratio greater than a preset IoU threshold and belonging to the same category are identified, they are considered to be the same object, and the final appearance-based classification confidence is the product of the classification confidence of the same drug in the first image and the second image.

[0031] Furthermore, the steps of fusing the output appearance-based classification confidence and the recognized text information and finally outputting the recognition result based on the appearance and text information include:

[0032] Detection appearance classification confidence P i and text recognition confidence;

[0033] Assume that p i is the confidence of the i-th character, if p i >α, the character prediction result is considered correct; where J is the correctly predicted character set, the total number of characters is n, the number of correct characters is m, and the final text recognition confidence is the product of the proportion of predicted correct characters and the confidence of all predicted correct characters, which is The final result is P = P t β ×Pi 1-β , α is the preset character confidence threshold, and β is the confidence coefficient.

[0034] Another aspect of the present invention relates to a real-time drug recognition system based on deep learning, comprising:

[0035] A video acquisition module, used to load a first image acquired by the first camera and a second image acquired by the second camera respectively;

[0036] A drug and text detection module, which processes the loaded first image and the second image respectively using a pre-trained model, and obtains drug and text detection results in the first image and the second image respectively, wherein the drug and text detection results include drug classification, drug area and text area;

[0037] The detection result fusion module is used to fuse the acquired drug classification and drug area, and output the appearance-based classification confidence and the fused drug area;

[0038] A text region cutting module, used for cutting the text region of the first image and the second image according to the acquired text region and the output fused medicine region;

[0039] A text recognition module, used to perform text recognition on the first image and the second image after the text region is cut out by using a pre-trained model, and recognize text information in the first image and the second image after the text region is cut out;

[0040] The recognition result fusion module is used to fuse the output appearance-based classification confidence and the recognized text information, and finally output the recognition result based on appearance and text information.

[0041] Furthermore, the real-time drug recognition system based on deep learning also includes:

[0042] A training module is used to train the model in the training phase, and the model includes a drug and text detection network model and a text recognition network model;

[0043] The training modules include drug and text detection network model training module and text recognition network model training module.

[0044] The drug and text detection network model training module includes:

[0045] A cutting unit is used to cut the image to be trained into sample images of the same size and load relevant annotations;

[0046] A data expansion unit, used to perform horizontal or flipping actions on sample images loaded with relevant annotations to perform data expansion work;

[0047] The first loading unit is used to load the pre-trained model on the COCO dataset;

[0048] A first training unit is used to increase the amount of data by using different data enhancement methods during the training process, and train the drug and text detection network model on the training set until convergence;

[0049] The text recognition network model training module includes:

[0050] A synthesis unit, used to use a text synthesis program to synthesize text images of commonly used drugs to increase the amount of training data;

[0051] The second loading unit loads the pre-trained model trained on other large text recognition datasets;

[0052] A merging unit, used for cutting out a text recognition training set from the real training set according to the text annotation, and merging the text recognition training set with the synthesized drug name dataset;

[0053] The second training unit is used to train the text recognition network model in the merged text recognition training set until convergence.

[0054] Furthermore, the video acquisition module includes:

[0055] The video acquisition unit is used to acquire real-time drug identification video collected by a video acquisition box. The inner wall of the video acquisition box is made of non-reflective white material. A strip light source for providing stable lighting conditions is provided on the upper part of the video acquisition box, and a light shielding cover is covered on the top of the video acquisition box to prevent the light of the strip light source from directly shining into the eyes of the operator and causing discomfort. A drug placement area is provided at the bottom of the video acquisition box, and a first camera and a second camera are symmetrically installed on both sides of the drug placement area and the installation height is half of the average height of commonly used drugs.

[0056] Furthermore, the detection result fusion module includes:

[0057] A calculation unit, used for calculating the intersection-and-union ratio of all detection frames to obtain the association relationship between each drug in the first image and the second image;

[0058] The object recognition unit is used to identify objects of the same category if their intersection-over-union ratio is greater than a preset intersection-over-union threshold. The final appearance-based classification confidence is the product of the classification confidence of the same drug in the first image and the second image.

[0059] Furthermore, the recognition result fusion module includes:

[0060] Detection unit, used to detect the appearance classification confidence P i and text recognition confidence;

[0061] Character recognition unit, used to assume p i is the confidence of the i-th character, if p i >α, the character prediction result is considered correct; where J is the correctly predicted character set, the total number of characters is n, the number of correct characters is m, and the final text recognition confidence is the product of the proportion of predicted correct characters and the confidence of all predicted correct characters, which is The final result is P = P t β ×P i 1-β , α is the preset character confidence threshold, and β is the confidence coefficient.

[0062] The beneficial effects achieved by the present invention are:

[0063] The present invention provides a real-time drug recognition method and system based on deep learning, which respectively loads a first image captured by a first camera and a second image captured by a second camera; pre-processes the loaded first image and the second image using a pre-trained model, and obtains drug and text detection results in the first image and the second image, respectively, wherein the drug and text detection results include drug classification, drug area and text area; fuses the obtained drug classification and drug area, and outputs the classification confidence based on appearance and the fused drug area; performs text area cutting on the first image and the second image according to the obtained text area and the output fused drug area; performs text recognition on the first image and the second image after the text area is cut using the pre-trained model, and recognizes the text information in the first image and the second image after the text area is cut; fuses the output classification confidence based on appearance and the recognized text information, and finally outputs the recognition result based on appearance and text information. The deep learning-based real-time drug identification method and system provided by the present invention only requires the operator to roughly place the drugs in a designated area, and can automatically and accurately identify the types of drugs in real time through appearance information and text information, and supports identifying multiple drugs at a time; combined with the prescription QR code recognition system, real-time verification of prescription drugs in the dispensing process can be achieved, the verification accuracy can be increased, and the burden on staff can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 A flowchart of a first example of a real-time drug identification method based on deep learning provided by the present invention;

[0065] Figure 2 A schematic flow chart of a second example of a real-time drug identification method based on deep learning provided by the present invention;

[0066] Figure 3 for Figure 2A detailed flowchart of a first example of the steps of training a model in a training phase is shown in FIG.

[0067] Figure 4 for Figure 2 A detailed flowchart of a second example of the steps of training a model in a training phase is shown in FIG.

[0068] Figure 5 A schematic diagram of the structure of an example of a target detection network in a real-time drug recognition method based on deep learning provided by the present invention;

[0069] Figure 6 A schematic diagram of the three-dimensional structure of an example of a video acquisition box in the real-time drug recognition method based on deep learning provided by the present invention;

[0070] Figure 7 for Figure 1 A detailed flow chart of an example of step 1 of fusing the acquired drug classification and drug region and outputting the appearance-based classification confidence and the fused drug region;

[0071] Figure 8 This is a functional block diagram of the first embodiment of the real-time drug identification system based on deep learning provided by the present invention;

[0072] Fig. 9 This is a functional block diagram of a second embodiment of a real-time drug identification system based on deep learning provided by the present invention;

[0073] Fig.10 for Fig. 9 The functional module diagram of the first embodiment of the training module shown in FIG.

[0074] Fig.11 for Fig. 9 The functional module diagram of the second embodiment of the training module shown in FIG.

[0075] Fig.12 for Figure 8 A functional module diagram of an example of a detection result fusion module shown in FIG.

[0076] Fig.13 for Figure 8 Schematic diagram of functional modules of an example of a recognition result fusion module shown in FIG.

[0077] Description of Figure Numbers:

[0078] 10. Video acquisition module; 20. Drug and text detection module; 30. Detection result fusion module; 40. Text area cutting module; 50. Text recognition module; 60. Recognition result fusion module; 70. Training module; 71. Cutting unit; 72. Data expansion unit; 73. First loading unit; 74. First training unit; 75. Synthesis unit; 76. Second loading unit; 77. Merging unit; 78. Second training unit; 31. Calculation unit; 32. Object recognition unit; 61. Detection unit; 62. Character recognition unit; 100. Strip light source; 200. Shading cover; 300. Drug placement area; 400. First camera; 500. Second camera. DETAILED DESCRIPTION

[0079] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0080] like Figure 1 As shown, the first example of the present invention proposes a real-time drug recognition method based on deep learning, comprising the following steps:

[0081] Step S100: loading a first image captured by a first camera and a second image captured by a second camera respectively.

[0082] Get the real-time drug recognition video captured by the video capture box. Figure 6 The video acquisition box is high at the back and low at the front for easy operation. The inner wall of the video acquisition box is made of non-reflective white material. A strip light source 100 for providing stable lighting conditions is provided on the upper part of the video acquisition box, and a light shielding cover 200 is provided on the upper part of the video acquisition box to prevent the light of the strip light source 100 from directly shining into the eyes of the operator and causing discomfort. A medicine placement area 300 is provided at the bottom of the video acquisition box. The first camera 400 and the second camera 500 are symmetrically installed on both sides of the medicine placement area 300 and the installation height is half of the average height of commonly used medicines. The video acquisition box requires that the video acquisition environment be controlled under the premise of convenient operation, so that the video acquisition environment is relatively stable, thereby improving the robustness of drug recognition.

[0083] Step S200: Use the pre-trained model to process the loaded first image and the second image respectively, and obtain drug and text detection results in the first image and the second image respectively, wherein the drug and text detection results include drug classification, drug area and text area.

[0084] The preparation of training data sets mainly includes three aspects: first, collecting training images, second, labeling training images, and third, data set segmentation.

[0085] (1) Collect training images. First, two high-definition cameras, namely the first camera 400 and the second camera 50, are used to continuously collect videos of operators placing drugs in different combinations, directions, positions, and randomly placed horizontally in the drug placement area. Secondly, when all drugs are collected, the detection model of medicine bottles and hands trained in the public data set is used to detect and identify medicine bottles and hands in the entire video. Images with medicine bottles but no hands are selected as training images at certain intervals. This method has two advantages. First, the method of collecting training images is exactly the same as the method when operators use this system to identify drugs, eliminating the situation where the images used for training and reasoning are inconsistent. Second, this method of video collection does not require manual intervention in taking pictures, and the collection speed is fast and the workload is small.

[0086] (2) Labeling training images. Labeling images requires not only labeling the drug’s bounding box and type, but also labeling the drug’s name text box and the visible text. The labeling box should tightly surround the entire drug area or the drug’s name text area, and the type name in the text box is the text in the visible text area.

[0087] (3) Dataset segmentation. The labeled training set is divided into a training set and a validation set according to a certain ratio. The training set is mainly used to train the model, while the validation set is mainly responsible for verifying the effect of the model and guiding the adjustment of related parameters.

[0088] Processing the loaded first image and the second image mainly includes three tasks: First, detecting the location of the drug. Second, identifying the type of drug based on the appearance information and outputting the appearance classification confidence. Third, outputting the location of the drug name text box to provide text slices for text recognition. In this example, a general real-time target detection network is used, such as YOLOX (YoloX: Exceeding yolo series in 2021: arXiv preprint arXiv: 2107.08430. 2021.). The structure of the network is as follows Figure 5 As shown in the figure, data augmentation methods such as Mosaic and MixUp are used during the training process to increase the expressiveness of the model. All text boxes are treated as one type of target so that the network can locate the position of the text box.

[0089] Step S300: Fuse the acquired drug classification and drug region, and output the appearance-based classification confidence and the fused drug region.

[0090] The detection and recognition results from two cameras, namely the first camera 400 and the second camera 500, are integrated to output the classification confidence based on appearance. During the inference process, the first image and the second image respectively captured by the first camera 400 and the second camera 500 are processed simultaneously. Since the first camera 400 and the second camera 500 are installed symmetrically, the second image captured by the second camera can be horizontally flipped to roughly align the position of the drug in the image, and then the IOU (Intersection over Union) of all detection frames is calculated to obtain the association relationship of each drug in the first image and the second image. Objects with an IOU greater than 0.5 and belonging to the same category are considered to be the same object, and the final classification confidence based on appearance is the product of the classification confidence of the same drug in the two images.

[0091] Step S400: cutting the text area of ​​the first image and the second image according to the acquired text area and the output fused medicine area.

[0092] According to the detection results, the text area is cut in the first image and the second image as the input of the text recognition module. According to the fused detection results, the text box inside the same drug area in the first image and the second image is found and associated with the drug category, and then the text slice is cut out in the original image according to the position of the text box and the slice height is normalized to 32 pixels.

[0093] Step S500: Use a pre-trained model to perform text recognition on the first image and the second image after the text region is cut, respectively, to recognize text information in the first image and the second image after the text region is cut.

[0094] Recognize the words in the drug name to further improve the accuracy of the entire drug recognition. The text area slices output by the detection result fusion module, the height of all slices is normalized to 32 pixels, and the output is the recognized Chinese and English characters and confidence. The text recognition module can use the CRNN algorithm (An End-to-End Trainable Neural Network for Image-based Sequence Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 2017).

[0095] Step S600: Fusing the output appearance-based classification confidence and the recognized text information, and finally outputting a recognition result based on appearance and text information.

[0096] The final recognition result of the drug is output by fusing the appearance and text recognition information. It is the appearance classification confidence P output by the detection result fusion module. i And the text recognition confidence output by the text recognition module. For any placed medicine, the drug name may appear in the images obtained by the two cameras, so it is necessary to fuse the text recognition results from the two images. Assume that p i is the confidence of the i-th character, if p i >α, the character prediction result is considered correct, J is the set of correctly predicted characters, where the total number of characters is n and the number of correct characters is m. The final text recognition confidence is the product of the proportion of correctly predicted characters and the confidence of all correctly predicted characters, which is The final result is P = P t β ×P i 1-β , where β is the confidence coefficient, in order to balance the impact of appearance and text recognition on the final confidence.

[0097] The real-time drug recognition method based on deep learning provided in this example is performed by respectively loading a first image captured by a first camera and a second image captured by a second camera; using a pre-trained model to process the loaded first image and the second image respectively, and obtaining drug and text detection results in the first image and the second image respectively, wherein the drug and text detection results include drug classification, drug area and text area; fusing the obtained drug classification and drug area, and outputting the classification confidence based on appearance and the fused drug area; performing text area cutting on the first image and the second image according to the obtained text area and the output fused drug area; using a pre-trained model to perform text recognition on the first image and the second image after the text area cutting, and identifying the text information in the first image and the second image after the text area cutting; fusing the output classification confidence based on appearance and the identified text information, and finally outputting the recognition result based on appearance and text information. The deep learning-based real-time drug identification method provided in this example only requires the operator to place the drugs roughly in the designated area, and it can automatically and accurately identify the type of drugs in real time through appearance information and text information, and supports identifying multiple drugs at a time; combined with the prescription QR code recognition system, it can realize real-time verification of prescription drugs in the dispensing process, increase verification accuracy, and reduce the burden on staff.

[0098] Further, see Figure 2 , Figure 2 This is a flow chart of a second example of the real-time drug recognition method based on deep learning provided by the present invention. Based on the first example, the real-time drug recognition method based on deep learning provided by this example includes before step S100:

[0099] Step S100A: training the model in the training phase, wherein the model includes a drug and text detection network model and a text recognition network model.

[0100] The whole system is divided into two stages: training and reasoning. In the training stage, two networks are mainly trained, one is the drug and text detection network, and the other is the text recognition network.

[0101] Please see Figure 3 , Figure 3 The flowchart of training a drug and text detection network is shown in step S100A.

[0102] Step S110a: cutting the image to be trained into sample images of the same size and loading relevant annotations.

[0103] Chop the image into sample images of consistent size to facilitate network processing, and load the image and related annotations.

[0104] Step S120a: horizontally or flipping the sample image loaded with relevant annotations to perform data expansion.

[0105] Perform data augmentation tasks such as Mosaic and MixUp on images horizontally or flipping them.

[0106] Step S130a: load the pre-trained model on the COCO dataset.

[0107] Step S140a: Use different data enhancement methods to increase the amount of data during the training process, and train the drug and text detection network model on the training set until convergence.

[0108] Train the text detection network model on the training set until convergence.

[0109] See also Figure 4 , Figure 4 The flowchart of training a text recognition network is shown in step S100A.

[0110] Step S110b: Use a text synthesis program to synthesize text images of commonly used medicines to increase the amount of training data.

[0111] Use a text synthesis program to synthesize text images of commonly used drugs to increase the amount of training data.

[0112] Step S120b: Load a pre-trained model trained on other large text recognition datasets.

[0113] Load a pre-trained model trained on the Syn90k dataset.

[0114] Step S130b: cut out a text recognition training set from the real training set according to the text annotation, and merge the text recognition training set with the synthesized drug name dataset.

[0115] A text recognition training set is cut out from the real training set based on text annotation, and this dataset is merged with the synthetic drug name dataset.

[0116] Step S140b: training the text recognition network model in the merged text recognition training set until convergence.

[0117] Train the text recognition network model until convergence.

[0118] The real-time drug recognition method based on deep learning provided in this example, by training the drug and text detection network and the text recognition network, uses the pre-trained model to pre-process the loaded first image and the second image respectively, and obtains the drug and text detection results in the first image and the second image respectively. The real-time drug recognition method based on deep learning provided in this example only requires the operator to roughly place the drugs in the designated area, and can automatically realize the real-time and accurate recognition of the drug type through the appearance information and text information, and supports the recognition of multiple drugs at a time; combined with the prescription QR code recognition system, it can realize the real-time verification of prescription drugs in the dispensing process, increase the verification accuracy, and reduce the burden on staff.

[0119] Preferably, see Figure 7 , Figure 7 for Figure 1 FIG. 1 is a detailed flow chart of an example of step S300, in which step S300 includes:

[0120] Step S310: Calculate the intersection-over-union ratio of all detection frames to obtain the association relationship between each drug in the first image and the second image.

[0121] The IOU (Intersection over Union) of all detection boxes is calculated to obtain the association relationship between each drug in the first image and the second image.

[0122] Step S320: If the intersection-and-union ratio is greater than P i Objects that meet the preset intersection threshold and are of the same category are considered to be the same object, and the final appearance-based classification confidence is the product of the classification confidence of the same drug in the first image and the second image.

[0123] Objects with an IOU greater than 0.5 and belonging to the same category are considered to be the same object. The final appearance-based classification confidence is the product of the classification confidence of the same drug in the two images.

[0124] This example provides a real-time drug recognition method based on deep learning. It obtains the association relationship between each drug in the first image and the second image by calculating the intersection and union ratio of all detection frames. If the intersection and union ratio is greater than P i Objects that meet the preset intersection threshold and are of the same category are considered to be the same object, and the final classification confidence based on appearance is the product of the classification confidence of the same drug in the first image and the second image. The real-time drug recognition method based on deep learning provided in this example only requires the operator to roughly place the drugs in the designated area, and can automatically and accurately identify the type of drugs in real time through appearance information and text information, and supports the recognition of multiple drugs at a time; combined with the prescription QR code recognition system, it can realize real-time verification of prescription drugs in the dispensing process, increase verification accuracy, and reduce the burden on staff.

[0125] Furthermore, in the real-time drug recognition method based on deep learning provided in this example, step S600 includes:

[0126] Step S610: Detect appearance classification confidence and text recognition confidence.

[0127] The final recognition result of the drug is output by fusing the appearance and text recognition information. It is the appearance classification confidence P output by the detection result fusion module. i And the text recognition confidence output by the text recognition module. For any placed medicine, the name of the medicine may appear in the images obtained by the two cameras, so it is necessary to fuse the text recognition results from the two images.

[0128] Step S620: Assume that p i is the confidence of the i-th character, if p i >α, the character prediction result is considered correct; where J is the correctly predicted character set, the total number of characters is n, the number of correct characters is m, and the final text recognition confidence is the product of the proportion of predicted correct characters and the confidence of all predicted correct characters, which is The final result is P = P t β ×P i 1-β , α is the preset character confidence threshold, and β is the confidence coefficient.

[0129] Assume that p i is the confidence of the i-th character, if p i >α, the character prediction result is considered correct, J is the set of correctly predicted characters, where the total number of characters is n and the number of correct characters is m. The final text recognition confidence is the product of the proportion of predicted correct characters and the confidence of all predicted correct characters, which is The final result is P = P t β ×Pi 1-β , where β is the confidence coefficient, in order to balance the impact of appearance and text recognition on the final confidence.

[0130] This example provides a real-time drug recognition method based on deep learning, which detects the appearance classification confidence and text recognition confidence; assuming that p i is the confidence of the i-th character, if p i >α, the character prediction result is considered correct; where J is the correctly predicted character set, the total number of characters is n, the number of correct characters is m, and the final text recognition confidence is the product of the proportion of predicted correct characters and the confidence of all predicted correct characters, which is The final result is P = P t β ×P i 1-β , α is the preset character confidence threshold, and β is the confidence coefficient. The real-time drug recognition method based on deep learning provided in this example only requires the operator to roughly place the drugs in the designated area, and can automatically and accurately recognize the types of drugs in real time through appearance information and text information, and supports the recognition of multiple drugs at a time; combined with the prescription QR code recognition system, it can realize the real-time verification of prescription drugs in the dispensing process, increase the verification accuracy, and reduce the burden on staff.

[0131] like Figure 8 As shown, Figure 8The functional block diagram of the first example of the real-time drug recognition system based on deep learning provided by the present invention, in this example, the real-time drug recognition system based on deep learning includes a video acquisition module 10, a drug and text detection module 20, a detection result fusion module 30, a text area cutting module 40, a text recognition module 50 and a recognition result fusion module 60, wherein the video acquisition module 10 is used to load the first image captured by the first camera and the second image captured by the second camera respectively; the drug and text detection module 20 is used to use a pre-trained model to process the loaded first image and the second image respectively, and obtain the drug and text detection results in the first image and the second image respectively, wherein the drug and text detection results include drug classification, drug product area and text area; a detection result fusion module 30, used to fuse the acquired drug classification and drug area, and output the classification confidence based on appearance and the fused drug area; a text area cutting module 40, used to perform text area cutting on the first image and the second image according to the acquired text area and the output fused drug area; a text recognition module 50, used to use a pre-trained model to perform text recognition on the first image and the second image after the text area is cut, respectively, and recognize the text information in the first image and the second image after the text area is cut; a recognition result fusion module 60, used to fuse the output classification confidence based on appearance and the recognized text information, and finally output the recognition result based on appearance and text information.

[0132] The video acquisition module 10 obtains the training samples collected by the video acquisition box and the input real-time drug recognition video. Figure 6 , the video acquisition box is high at the back and low at the front for easy operation. The inner wall of the video acquisition box is made of non-reflective white material. A strip light source 100 for providing stable lighting conditions is provided on the upper part of the video acquisition box, and a light shielding cover 200 is provided on the top of the video acquisition box to prevent the light of the strip light source 100 from directly shining into the eyes of the operator and causing discomfort. A medicine placement area 300 is provided at the bottom of the video acquisition box, and the first camera 400 and the second camera 500 are symmetrically installed on both sides of the medicine placement area 300 and the installation height is half of the average height of commonly used medicines. The video acquisition box requires that the video acquisition environment be controlled under the premise of convenient operation, so that the video acquisition environment is relatively stable, thereby improving the robustness of drug recognition.

[0133] Training data preparation mainly includes three aspects: first, collecting training images, second, labeling training images, and third, data set segmentation.

[0134] (1) Collect training images. First, two high-definition cameras, namely the first camera 400 and the second camera 50, are used to continuously collect videos of operators placing drugs in different combinations, directions, positions, and randomly placed horizontally in the drug placement area. Secondly, when all drugs are collected, the detection model of medicine bottles and hands trained in the public data set is used to detect and identify medicine bottles and hands in the entire video. Images with medicine bottles but no hands are selected as training images at certain intervals. This method has two advantages. First, the method of collecting training images is exactly the same as the method when operators use this system to identify drugs, eliminating the situation where the images used for training and reasoning are inconsistent. Second, this method of video collection does not require manual intervention in taking pictures, and the collection speed is fast and the workload is small.

[0135] (2) Labeling training images. Labeling images requires not only labeling the drug’s bounding box and type, but also labeling the drug’s name text box and the visible text. The labeling box should tightly surround the entire drug area or the drug’s name text area, and the type name in the text box is the text in the visible text area.

[0136] (3) Dataset segmentation. The labeled training set is divided into a training set and a validation set according to a certain ratio. The training set is mainly used to train the model, while the validation set is mainly responsible for verifying the effect of the model and guiding the adjustment of related parameters.

[0137] The drug and text detection module 20 preprocesses the loaded first image and the second image, which mainly includes three tasks: First, detect the location of the drug. Second, identify the type of drug based on the appearance information and output the appearance classification confidence. Third, output the location of the drug name text box to provide text slices for text recognition. In this example, a general real-time target detection network is used, such as YOLOX (YoloX: Exceeding yolo series in 2021: arXiv preprintarXiv: 2107.08430, 2021.). The structure of the network is as follows Figure 5 As shown in the figure, data augmentation methods such as Mosaic and MixUp are used during the training process to increase the expressiveness of the model. All text boxes are treated as one type of target so that the network can locate the position of the text box.

[0138] The detection result fusion module 30 fuses the detection and recognition results from two cameras, namely the first camera 400 and the second camera 500, and outputs the classification confidence based on appearance. During the reasoning process, the first image and the second image respectively captured by the first camera 400 and the second camera 500 are processed simultaneously. Since the first camera 400 and the second camera 500 are installed symmetrically, the second image captured by the second camera can be horizontally flipped to roughly align the position of the drug in the image, and then the IOU (Intersection over Union) of all detection frames is calculated to obtain the association relationship of each drug in the first image and the second image. Objects with an IOU greater than 0.5 and belonging to the same category are considered to be the same object, and the final classification confidence based on appearance is the product of the classification confidence of the same drug in the two images.

[0139] The text region cutting module 40 cuts the text region in the first image and the second image according to the detection results as the input of the text recognition module. According to the fused detection results, the text box inside the same drug region in the first image and the second image is found and associated with the drug category, and then the text slice is cut out in the original image according to the position of the text box and the slice height is normalized to 32 pixels.

[0140] The text recognition module 50 recognizes the text of the drug name, further improving the accuracy of the entire drug recognition. The input of this module is the text area slice output by the detection result fusion module. The height of all slices is normalized to 32 pixels, and the output is the recognized Chinese and English characters and confidence. The text recognition module can use the CRNN algorithm (An End-to-End Trainable Neural Network for Image-based Sequence Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 2017).

[0141] The recognition result fusion module 60 fuses the appearance and text recognition information to output the final recognition result of the drug. i And the text recognition confidence output by the text recognition module. For any placed medicine, the drug name may appear in the images obtained by the two cameras, so it is necessary to fuse the text recognition results from the two images. Assume that p i is the confidence of the i-th character, if p i>α, the character prediction result is considered correct, J is the set of correctly predicted characters, where the total number of characters is n and the number of correct characters is m. The final text recognition confidence is the product of the proportion of predicted correct characters and the confidence of all predicted correct characters, which is The final result is P = P t β ×P i 1-β , where β is the confidence coefficient, in order to balance the impact of appearance and text recognition on the final confidence.

[0142] The real-time drug recognition system based on deep learning provided in this example loads a first image captured by a first camera and a second image captured by a second camera respectively; uses a pre-trained model to process the loaded first image and the second image respectively, and obtains drug and text detection results in the first image and the second image respectively, wherein the drug and text detection results include drug classification, drug area and text area; fuses the obtained drug classification and drug area, and outputs the classification confidence based on appearance and the fused drug area; performs text area cutting on the first image and the second image according to the obtained text area and the output fused drug area; uses the pre-trained model to perform text recognition on the first image and the second image after the text area is cut, and recognizes the text information in the first image and the second image after the text area is cut; fuses the output classification confidence based on appearance and the recognized text information, and finally outputs the recognition result based on appearance and text information. The deep learning-based real-time drug recognition system provided in this example only requires the operator to place the drugs roughly in the designated area, and it can automatically and accurately identify the type of drugs in real time through appearance information and text information, and supports identifying multiple drugs at a time; combined with the prescription QR code recognition system, it can realize real-time verification of prescription drugs in the dispensing process, increase verification accuracy, and reduce the burden on staff.

[0143] Further, see Fig. 9 , Fig. 9 This is a functional block diagram of a second example of a real-time drug identification system based on deep learning provided by the present invention. On the basis of the first implementation, the real-time drug identification system based on deep learning also includes a training module 70. The training module 70 is used to train the model in the training stage. The model includes a drug and text detection network model and a text recognition network model.

[0144] The training module 70 cuts the image into sample images of uniform size to facilitate network processing, and loads the image and related annotations.

[0145] Please see Fig.10 and Fig.11In the real-time drug recognition system based on deep learning provided in this example, the training module 70 includes a drug and text detection network model training module and a text recognition network model training module. The drug and text detection network model training module includes a cutting unit 71, a data expansion unit 72, a first loading unit 73 and a first training unit 74, wherein the cutting unit 71 is used to cut the image to be trained into sample images of the same size and load relevant annotations; the data expansion unit 72 is used to perform horizontal or flipping actions on the sample images loaded with relevant annotations to perform data expansion; the first loading unit 73 is used to load the pre-trained model on the COCO data set; the first training unit 74 is used to use different data enhancement methods to increase the data volume during the training process, and train the drug and text detection network model on the training set until convergence. The text recognition network model training module includes a synthesis unit 75, a second loading unit 76, a merging unit 77 and a second training unit 78, wherein the synthesis unit 75 is used to use a text synthesis program to synthesize text images of commonly used drugs to increase the amount of training data; the second loading unit 76 is used to load a pre-trained model trained on other large text recognition data sets; the merging unit 77 is used to cut out a text recognition training set from a real training set according to text annotations, and merge the text recognition training set with the synthesized drug name data set; the second training unit 78 is used to train the text recognition network model in the merged text recognition training set until convergence.

[0146] The real-time drug recognition system based on deep learning provided in this example, by training the drug and text detection network and the text recognition network, uses the pre-trained model to pre-process the loaded first image and the second image respectively, and obtains the drug and text detection results in the first image and the second image respectively. The real-time drug recognition system based on deep learning provided in this example only requires the operator to roughly place the drugs in the designated area, and can automatically realize the real-time and accurate recognition of the drug type through the appearance information and text information, and supports the recognition of multiple drugs at a time; combined with the prescription QR code recognition system, it can realize the real-time verification of prescription drugs in the dispensing process, increase the verification accuracy, and reduce the burden on staff.

[0147] Preferably, see Fig.12 , Fig.12 for Figure 8The functional module diagram of an example of the detection result fusion module shown in the figure, the detection result fusion module 30 includes a calculation unit 31 and an object recognition unit 32, wherein the calculation unit 31 is used to calculate the intersection-and-union ratio of all detection frames to obtain the association relationship between each drug in the first image and the second image; the object recognition unit 32 is used to identify objects with an intersection-and-union ratio greater than a preset intersection-and-union threshold and belonging to the same category, and then consider them to be the same object, and finally the classification confidence based on appearance is the product of the classification confidence of the same drug in the first image and the second image.

[0148] The calculation unit 31 calculates the IOU (Intersection over Union) of all detection frames to obtain the association relationship between each drug in the first image and the second image.

[0149] The object recognition unit 32 considers objects with an IOU greater than 0.5 and belonging to the same category to be the same object, and the final appearance-based classification confidence is the product of the classification confidences of the same drug in the two images.

[0150] The real-time drug recognition system based on deep learning provided in this example obtains the association relationship between each drug in the first image and the second image by calculating the intersection and union ratio of all detection frames; if the intersection and union ratio is greater than P i Objects that meet the preset intersection threshold and are of the same category are considered to be the same object, and the final classification confidence based on appearance is the product of the classification confidence of the same drug in the first image and the second image. The real-time drug recognition system based on deep learning provided in this example only requires the operator to roughly place the drugs in the designated area, and can automatically and accurately identify the type of drugs in real time through appearance information and text information, and supports the recognition of multiple drugs at a time; combined with the prescription QR code recognition system, it can realize real-time verification of prescription drugs in the dispensing process, increase verification accuracy, and reduce the burden on staff.

[0151] Further, see Fig.13 , Fig.13 for Figure 8 The functional module diagram of an example of the recognition result fusion module shown in FIG. 6 is a schematic diagram of a recognition result fusion module 60. In this example, the recognition result fusion module 60 includes a detection unit 61 and a character recognition unit 62, wherein the detection unit 61 is used to detect the appearance classification confidence P i and character recognition confidence; character recognition unit 62, for assuming p i is the confidence of the i-th character, if p i >α, the character prediction result is considered correct; where J is the correctly predicted character set, the total number of characters is n, the number of correct characters is m, and the final text recognition confidence is the product of the proportion of predicted correct characters and the confidence of all predicted correct characters, which is The final result is P = P t β ×P i 1-β , α is the preset character confidence threshold, and β is the confidence coefficient.

[0152] The detection unit 61 integrates the appearance and text recognition information to output the final recognition result of the drug. The appearance classification confidence P output by the detection result fusion module is i And the text recognition confidence output by the text recognition module. For any placed medicine, the name of the medicine may appear in the images obtained by the two cameras, so it is necessary to fuse the text recognition results from the two images.

[0153] The character recognition unit 62 assumes that p i is the confidence of the i-th character, if p i >α, the character prediction result is considered correct, J is the set of correctly predicted characters, where the total number of characters is n and the number of correct characters is m. The final text recognition confidence is the product of the proportion of predicted correct characters and the confidence of all predicted correct characters, which is The final result is P = P t β ×P t 1-β , where β is the confidence coefficient, in order to balance the impact of appearance and text recognition on the final confidence.

[0154] The real-time drug recognition system based on deep learning provided in this example detects the appearance classification confidence and text recognition confidence; assuming that p i is the confidence of the i-th character, if p i >α, the character prediction result is considered correct; where J is the correctly predicted character set, the total number of characters is n, the number of correct characters is m, and the final text recognition confidence is the product of the proportion of predicted correct characters and the confidence of all predicted correct characters, which is The final result is P = P t β ×P i 1-β , α is the preset character confidence threshold, and β is the confidence coefficient. The real-time drug recognition system based on deep learning provided in this example only requires the operator to roughly place the drugs in the designated area, and can automatically and accurately recognize the types of drugs in real time through appearance information and text information, and supports the recognition of multiple drugs at a time; combined with the prescription QR code recognition system, it can realize the real-time verification of prescription drugs in the dispensing process, increase the verification accuracy, and reduce the burden on staff.

[0155] Although preferred embodiments of the present invention have been described, additional changes and modifications may be made to these embodiments by those skilled in the art once the basic inventive concept is known. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A real-time drug identification method based on deep learning, characterized in that: The following steps are involved: Load the first image captured by the first camera and the second image captured by the second camera respectively; Using the pre-trained model to process the loaded first image and the second image respectively, and obtain drug and text detection results in the first image and the second image respectively, wherein the drug and text detection results include drug classification, drug area and text area; Fuse the acquired drug classification and drug area, and output the appearance-based classification confidence and the fused drug area; Performing text region cutting on the first image and the second image according to the acquired text region and the output fused medicine region; Using a pre-trained model, respectively perform text recognition on the first image and the second image after the text region is cut, and recognize text information in the first image and the second image after the text region is cut; The output appearance-based classification confidence and the recognized text information are integrated, and finally the recognition result based on appearance and text information is output; The step of fusing the outputted classification confidence based on appearance and the recognized text information and finally outputting the recognition result based on appearance and text information comprises: Detection appearance classification confidence and text recognition confidence; Assumptions for The confidence of characters, if , then the character prediction result is considered correct; among them, is the set of correctly predicted characters, the total number of characters is n, the number of correct characters is m, and the final text recognition confidence is the product of the proportion of predicted correct characters and the confidence of all predicted correct characters, which is , the final result , is the preset character confidence threshold, is the confidence coefficient.

2. The method for real-time drug identification based on deep learning according to claim 1, characterized in that: The step of loading the first image captured by the first camera and the second image captured by the second camera respectively includes: Training the model in a training phase, wherein the model includes a drug and text region detection network model and a text recognition network model; The steps of training the drug and text detection network model include: Cut the image to be trained into sample images of the same size and load relevant annotations; Perform horizontal or flipping actions on sample images loaded with relevant annotations to perform data expansion; Load the pre-trained model on the COCO dataset; During the training process, different data enhancement methods are used to increase the amount of data, and the drug and text detection network model is trained on the training set until convergence; The steps of training the text recognition network model include: Use text synthesis programs to synthesize text images of commonly used drugs to increase the amount of training data; Load pre-trained models trained on other large text recognition datasets; A text recognition training set is cut out from the real training set according to the text annotation, and the text recognition training set is merged with the synthetic drug name dataset; The text recognition network model is trained in the merged text recognition training set until convergence.

3. The real-time drug identification method based on deep learning according to claim 1, characterized in that: The steps of respectively loading a first image captured by a first camera and a second image captured by a second camera include: Real-time medicine video captured by a video acquisition box is obtained, the inner wall of the video acquisition box is made of non-reflective white material, a strip light source for providing stable lighting conditions is provided on the upper part of the video acquisition box, and a light shielding cover is covered on the top of the video acquisition box to prevent the light of the strip light source from directly shining into the eyes of the operator and causing discomfort, a medicine placement area is provided at the bottom of the video acquisition box, the first camera and the second camera are symmetrically installed on both sides of the medicine placement area and the installation height is half of the average height of commonly used medicines.

4. The method for real-time drug identification based on deep learning according to claim 1, characterized in that: The step of fusing the acquired drug classification and drug region and outputting the appearance-based classification confidence and the fused drug region comprises: Calculate the intersection-and-union ratio of all detection frames to obtain the association relationship between each drug in the first image and the second image; If objects with an IoU ratio greater than a preset IoU threshold and belonging to the same category are identified, they are considered to be the same object, and the final appearance-based classification confidence is the product of the classification confidence of the same drug in the first image and the second image.

5. A real-time drug recognition system based on deep learning, characterized in that: include: A video acquisition module (10), used to load a first image acquired by a first camera and a second image acquired by a second camera respectively; A drug and text detection module (20), used to process the loaded first image and the second image respectively using a pre-trained model, and obtain drug and text detection results in the first image and the second image respectively, wherein the drug and text detection results include drug classification, drug area and text area; A detection result fusion module (30) is used to fuse the acquired drug classification and drug area, and output the classification confidence based on appearance and the fused drug area; A text region cutting module (40) is used to cut the text region of the first image and the second image according to the acquired text region and the output fused medicine region; A text recognition module (50) is used to use a pre-trained model to perform text recognition on the first image and the second image after the text region is cut, respectively, to recognize text information in the first image and the second image after the text region is cut; A recognition result fusion module (60) is used to fuse the outputted classification confidence based on appearance and the recognized text information, and finally output a recognition result based on appearance and text information; The recognition result fusion module (60) comprises: A detection unit (61) for detecting the appearance classification confidence and text recognition confidence; Character recognition unit (62), used to assume for The confidence of characters, if , then the character prediction result is considered correct; among them, is the set of correctly predicted characters, the total number of characters is n, the number of correct characters is m, and the final text recognition confidence is the product of the proportion of predicted correct characters and the confidence of all predicted correct characters, which is , the final result , is the preset character confidence threshold, is the confidence coefficient.

6. The real-time drug identification system based on deep learning according to claim 5, characterized in that: The real-time drug recognition system based on deep learning also includes: A training module (70), used to train the model in a training phase, wherein the model includes a drug and text detection network model and a text recognition network model; The training module (70) includes a drug and text detection network model training module and a text recognition network model training module. The drug and text detection network model training module includes: A cutting unit (71), used for cutting the image to be trained into sample images of uniform size and loading relevant annotations; A data expansion unit (72), used for performing a horizontal or flipping action on the sample image loaded with relevant annotations to perform data expansion; A first loading unit (73) loads a pre-trained model on the COCO dataset; A first training unit (74) is used to increase the amount of data using different data enhancement methods during the training process, and train the drug and text detection network model on the training set until convergence; The text recognition network model training module includes: A synthesis unit (75), used for synthesizing text images of commonly used drugs using a text synthesis program to increase the amount of training data; A second loading unit (76) is used to load a pre-trained model trained on other large text recognition datasets; A merging unit (77), used for cutting out a text recognition training set from the real training set according to the text annotation, and merging the text recognition training set with the synthesized drug name dataset; The second training unit (78) is used to train the text recognition network model in the merged text recognition training set until convergence.

7. The real-time drug identification system based on deep learning as claimed in claim 5, characterized in that: The video acquisition module (10) comprises: The video acquisition unit is used to acquire real-time drug identification video captured by a video acquisition box, the inner wall of the video acquisition box is made of non-reflective white material, a strip light source for providing stable lighting conditions is provided on the upper part of the video acquisition box, and a light shielding cover is covered on the top of the video acquisition box to prevent the light of the strip light source from directly shining into the eyes of the operator and causing discomfort, a drug placement area is provided at the bottom of the video acquisition box, and the first camera and the second camera are symmetrically installed on both sides of the drug placement area and the installation height is half of the average height of commonly used drugs.

8. The real-time drug identification system based on deep learning as claimed in claim 5, characterized in that: The detection result fusion module (30) comprises: A calculation unit (31), used for calculating the intersection-over-union ratio of all detection frames to obtain the association relationship between each drug in the first image and the second image; The object recognition unit (32) is used to identify objects of the same category if their intersection-over-union ratio is greater than a preset intersection-over-union threshold and they are considered to be the same object, and the final classification confidence based on appearance is the product of the classification confidence of the same drug in the first image and the second image.

Citation Information

Patent Citations

  • A method for detecting and recognizing sensitive characters in natural scene images

    CN109447078A

  • Patient assistance intelligent auditing system based on deep learning

    CN111353445A