Unmanned pharmacy prescription drug matching confirmation warning system based on deep learning and matching confirmation method thereof
Through a deep learning-based method, combined with the SAM segmentation model, YOLOv8 model and double-head network structure, the problem of insufficient accuracy and large prescription image processing of prescription matching confirmation in unmanned pharmacies is solved, and the matching effect is achieved with high accuracy and robustness, which is suitable for multiple data formats and heterogeneous data.
Patent Information
- Application Number
- CN202510002611.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art has problems such as insufficient accuracy, difficulty in processing large prescription images, multi-category imbalance, lack of fine-grained feature matching capabilities and insufficient ability to process heterogeneous data in unmanned pharmacies.
Using a deep learning-based method, drug and prescription images are obtained through practical cameras, data preprocessing and complex data set construction are carried out, target segmentation and detection are used using SAM segmentation model and YOLOv8 model, combined with sliding window detection and global NMS technology, a double-head network structure is used for position and category matching, and the model is optimized through multiple loss functions and knowledge distillation techniques.
It achieves high-precision matching of drugs and prescriptions, solves the problem of large prescription image processing, improves the matching accuracy and robustness of the system, is suitable for a variety of data formats and heterogeneous data, and has good versatility and scalability.
Smart Images

Figure CN119942154A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unmanned automated intelligence, and in particular to an unmanned pharmacy prescription drug matching confirmation warning system based on deep learning and a matching confirmation method thereof. Background Art
[0002] In medical scenarios such as public hospitals or comprehensive pharmacies, automatic drug sorting and dispensing technology is a key means to improve the efficiency of manual dispensing and reduce the error rate. In the medical industry, accurate distribution of drugs is crucial to patient safety.
[0003] Traditional manual drug prescription comparison and review methods mainly rely on manual operations or simple process control. Manual review usually requires experienced pharmacists or physicians to conduct, which is not only inefficient but also prone to missed detection or misjudgment. Although electronic prescription systems, barcodes, and RFID technology methods have introduced computer software and hardware assistance, they are mostly limited to finished boxed medicinal materials and are difficult to cope with complex prescription changes. These methods are unable to cope with non-standard, multi-brand, Chinese medicinal materials, and different dosages of drugs, and are difficult to meet the needs of unmanned pharmacies and retail compliance.
[0004] In recent years, with the development of computer vision technology, prescription and drug matching methods based on image processing have gradually attracted attention. However, these methods still have significant shortcomings in practical applications:
[0005] 1. Limitations of single-stage image matching: Most existing image matching methods use single-stage models, which usually only consider the global or local features of the image. This method is easily affected by deformation, scale change, rotation and other factors when facing complex prescriptions, resulting in insufficient matching accuracy. In addition, single-stage methods have difficulty in handling the fine-grained features of different targets in drugs or prescriptions, especially in multi-target and multi-category scenarios.
[0006] 2. Challenges in processing large prescription images: Traditional Chinese medicine prescriptions are often multiple and large, and existing image processing algorithms and models are difficult to apply directly. In object detection in large-size images, traditional global detection methods cannot effectively balance accuracy and efficiency, and simple local detection methods (such as sliding window detection) are prone to redundant detection and missed detection problems during splicing and global processing, further affecting the accuracy of detection and matching.
[0007] 3. Multi-class imbalance problem: In special prescriptions, targets of different categories usually have highly unbalanced distributions. For example, in a complex Chinese medicine prescription, the frequency of some medicinal materials (such as licorice and cinnamon twigs) is much higher than that of other medicinal materials. Existing image processing and machine learning methods are difficult to maintain high matching and classification accuracy when the distribution of sample categories is severely unbalanced.
[0008] 4. Lack of fine-grained feature matching capabilities: In the TCM scenario, Chinese medicinal materials contain a lot of detailed information, such as processing technology, medicinal material quality, medicinal material color properties, etc. Existing methods often only focus on geometric features or shape features and lack the ability to effectively process these details. Therefore, it is impossible to accurately judge the consistency between prescriptions and medicinal materials in scenarios with large differences in details.
[0009] 5. Insufficient ability to process heterogeneous data: In actual engineering applications, prescriptions have various formats and layouts, and the sources, origins, methods and brands of medicinal materials are diverse. There are both handwritten prescriptions and receipts, both electronic and paper. Existing technologies can usually only process one or several formats, and lack the ability to uniformly process multiple formats, which affects the applicability of the automated drug prescription matching and confirmation system. Summary of the invention
[0010] Purpose of the invention: To provide a method for confirming prescription and drug matching in an unmanned pharmacy based on deep learning, and further to provide a confirmation warning system applied to the above-mentioned method for confirming prescription and drug matching in an unmanned pharmacy based on deep learning, so as to solve the above-mentioned problems existing in the prior art.
[0011] Technical solution: A method for confirming prescription and drug matching in an unmanned pharmacy based on deep learning, comprising the following steps:
[0012] S1. Use a camera to obtain the images of the medicines and prescriptions, and perform data preprocessing on the collected images and construct a complex data set;
[0013] S2. Perform single-target segmentation on the multi-target drug images in the constructed complex data set, and classify and sort the segmented images; simultaneously identify and process prescription images, and use optical character recognition text information;
[0014] S3, identifying the segmented drug information, searching the drug database, and obtaining the corrected drug information;
[0015] S4, performing block detection on the aligned image by sliding window method to ensure that the target in each local area can be effectively detected;
[0016] S5, restore the global coordinates of the local targets detected by the sliding window, and remove redundant detection frames by global non-maximum suppression;
[0017] S6. Matching the positions of the drugs in the two images by using the weighted Hungarian algorithm to obtain possible matching pairs;
[0018] S7, using a double-head matching network to further verify the above matching pair and calculate the matching score by cosine distance;
[0019] S8. Perform similarity analysis and difference analysis on all matching pairs, and finally output the results and generate a JSON file.
[0020] In a further embodiment, the specific operations of data preprocessing and constructing a complex data set in step S1 are as follows:
[0021] S101, highlighting the detail features in the image by enhancing the local contrast of the image;
[0022] S102, performing data enhancement processing by random rotation, translation, scaling, cropping, and color jittering to generate diversified training data;
[0023] S103. Balance the data by synthesizing minority class oversampling technology and random undersampling technology to avoid the impact of class imbalance during model training.
[0024] In a further embodiment, in step S2, the SAM segmentation model is used to accurately segment the target in the image, and the image in the complex scene is segmented by a multi-stage segmentation method;
[0025] The multi-stage segmentation is a sliding window segmentation, specifically, an image of a predetermined size is divided into multiple small windows, each small window is input into the SAM model for segmentation with a corresponding size, and the segmentation result includes the precise outline of the target and the segmentation mask. The specific steps are as follows:
[0026] S201, sliding window definition: Assume that the image size is W×H, the sliding window size is w×h, and the sliding step size is s x and y ; The sliding window moves on the image according to the following formula:
[0027] x i =i·s x ,y j =j·s y
[0028] Where i, j are integers such that 0≤x i ≤Ww and 0≤y j ≤Hh;
[0029] S202, local segmentation: for each sliding window segmentation result D k , after local optimization, it is transformed in the global coordinate system;
[0030] S203, global fusion: Finally, fusion is performed through global optimization to obtain the global segmentation result D final .
[0031] In a further embodiment, the local optimization is to refine the segmentation result within each sliding window; and the global optimization is to merge the segmentation results of all sliding windows in a global scope to ensure the consistency and accuracy of the overall segmentation.
[0032] In a further embodiment, the specific method of block detection in step S4 is to use the YOLOv8 model to perform multi-target detection on drug packaging and medicinal material images; when processing images of a predetermined size, processing is performed by a multi-stage detection method, including sliding window detection, local non-maximum suppression, and global non-maximum suppression;
[0033] The sliding window detection specifically involves dividing a drug or medicinal material image of a predetermined size into multiple small windows, each small window of a suitable size is input into a YOLOv8 model for detection, and the detection result includes a target position and a confidence score;
[0034] The local non-maximum suppression and global non-maximum suppression are to perform local non-maximum suppression in each sliding window to suppress redundant detection in the window; merge the detection results of all sliding windows globally, and perform global non-maximum suppression to remove overlapping targets. The specific formula is as follows:
[0035] Set the image size to W×H, the sliding window size to w×h, and the sliding step size to s x and y ; The sliding window moves on the image according to the following formula:
[0036] x i =i·s x ,y j =j·s y
[0037] Where i, j are integers such that 0≤x i ≤Ww and 0≤y j ≤Hh. For each sliding window detection result D k After local NMS processing, it is transformed in the global coordinate system and finally fused through global NMS to obtain the global detection result D final .
[0038] In a further embodiment, the dual-head matching network structure in step S7 improves the recognition and matching capabilities of the model by combining position matching and category matching; multiple loss functions are used in the matching process, and loss calculations are performed in a targeted manner.
[0039] In a further embodiment, the dual-head network structure includes a position matching branch and a category matching branch, and after the input image is subjected to feature extraction by the convolution layer, a matching feature vector and a category prediction are output respectively;
[0040] The target allocation of the various loss functions is as follows:
[0041] Type A loss function: used for position matching, measures the similarity between samples by cosine distance, and combines mining strategies to optimize feature vectors so that the distance between positive samples is minimized and the distance between negative samples is maximized:
[0042] L triplet =max(0,cos(f a ,f n )-cos(f a ,f p )+margin);
[0043] Type B loss function: used for category matching. By introducing angle intervals, it enhances the aggregation of similar samples in high-dimensional space and optimizes the accuracy of category prediction:
[0044]
[0045] Type C loss function: used in the contrastive learning stage to optimize feature distribution and enhance the distinction between different categories. This loss adjusts the feature distribution through a temperature scaling mechanism to make the distance between similar samples as close as possible:
[0046]
[0047] D-type loss function: It is used in the category matching stage to solve the problem of category imbalance and ensure that the model still has good robustness when the category distribution is uneven:
[0048] L CE =-∑ i w yi logp yi ;
[0049] Type E loss function: During the student network training phase, the knowledge of the teacher network is transferred to the student network through knowledge distillation of the final output logits and feature vector features, achieving lightweight model while maintaining high performance:
[0050]
[0051] F-type loss function: used for feature extraction of the teacher network, enhancing the expressiveness of unlabeled data through self-supervised learning and improving the generalization performance of the model:
[0052]
[0053] G loss function: used to suppress the overfitting problem of complex networks during complex task training:
[0054] L L2 =λ∑k‖θ k || 2 ;
[0055] Through the above-mentioned multiple loss functions, the practical comprehensive loss of the teacher model is finally obtained as follows:
[0056] L teacher =α1L triplet +α2L Arcface +α3L NTXent +α4L CE +α5L BYOL +α6L L2 ;
[0057] The comprehensive loss used by the student model is:
[0058] L student =α1L triplet +α2L Arcface +α3L NTXent +α4L CE +α5L distill +α6L L2 .
[0059] In a further embodiment, the unmanned pharmacy prescription drug matching confirmation method based on deep learning is
[0060] Features:
[0061] In step S7, the cosine distance calculation matching is specifically that the system inputs the extracted feature vector into the matching judgment module, and uses the cosine distance and F1 norm to make the final matching judgment. The specific calculation formula is as follows:
[0062] S701, the system uses cosine distance as a matching metric to determine the similarity of two targets:
[0063]
[0064] S702, the system calculates the F1 norm of positive and negative samples and selects the optimal threshold to ensure the accuracy and stability of the matching results:
[0065]
[0066] A confirmation warning system for an unmanned pharmacy prescription drug matching confirmation method based on deep learning, including a data preprocessing and complex data set construction model, a target segmentation model, a target detection model, a target matching model, and a model reasoning and matching judgment model.
[0067] Data preprocessing and complex data set model building: Preprocess images of medicinal materials, drugs, and prescriptions to build a balanced and diverse training data set.
[0068] Target segmentation model: Use the SAM model to segment the total drug image and separate individual drugs from the image for subsequent recognition and matching processing.
[0069] Target detection model: Use the YOLOv8 model for multi-target detection, combined with sliding window detection and global non-maximum suppression technology, to process images of a predetermined size.
[0070] Target matching model: A dual-head network structure is used to achieve accurate position matching and category matching of targets.
[0071] Model reasoning and matching judgment model: Optimize feature extraction and matching decisions through multiple loss functions, and use cosine distance and F1 norm for final matching judgment.
[0072] Beneficial effects: The present invention relates to an unmanned pharmacy prescription drug matching confirmation warning system based on deep learning and a matching confirmation method thereof, which relates to the field of unmanned automated intelligence. Through the combination of sliding window detection, local NMS and global NMS, the processing problem of oversized prescriptions and drug images is solved, and high-precision target detection and matching are achieved, avoiding the redundant detection and missed detection problems in traditional methods. The present invention adopts a double-headed network structure to independently process position matching and category matching, so that category information does not interfere with position matching, and at the same time, category information is used to assist matching decisions, thereby improving the matching accuracy and robustness of the system. Through the combined use of multiple loss functions, the feature expression ability, matching accuracy and classification performance of the model are comprehensively improved, especially when processing complex samples and category imbalances. Through the knowledge distillation technology, the lightweight design of the model is realized, so that the system can still maintain efficient reasoning performance in a resource-constrained environment. The F1 norm is used to calculate the optimal matching threshold to ensure the accuracy and stability of the matching results, especially to achieve a good balance between recall rate and accuracy. The system supports the processing of multiple data formats, adapts to the needs of multi-source heterogeneous data, has good versatility and scalability, and is suitable for a variety of industrial and engineering application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 This is the structure and training flowchart of the dual-head image matching network.
[0074] Figure 2 This is the flowchart for prescription and drug image matching confirmation and alarm. DETAILED DESCRIPTION
[0075] In the following description, a large number of specific details are provided to provide a more thorough understanding of the present invention. However, it is apparent to those skilled in the art that the present invention can be implemented without one or more of these details. In other examples, in order to avoid confusion with the present invention, some technical features well known in the art are not described.
[0076] The unmanned pharmacy prescription drug matching confirmation warning system based on deep learning involved in the present invention includes the following main modules:
[0077] 1. Data preprocessing and complex data set construction: Preprocess images of medicinal materials, drugs, and prescriptions to build a balanced and diverse training data set.
[0078] 2. Target segmentation model: Use the SAM model to segment the total drug image and separate individual drugs from the image for subsequent recognition and matching processing.
[0079] 3. Target detection model: Use the YOLOv8 model for multi-target detection, combined with sliding window detection and global non-maximum suppression (NMS) technology to process oversized images.
[0080] 4. Target matching model: A dual-head network structure is used to achieve accurate position matching and category matching of the target.
[0081] 5. Model reasoning and matching judgment: Optimize feature extraction and matching decisions through multiple loss functions, and use cosine distance and F1 norm for final matching judgment.
[0082] Data preprocessing is an important foundation of the system and directly affects the robustness and generalization ability of the model. In view of the characteristics of prescriptions, medicinal materials and drug images, the present invention proposes the following processing methods:
[0083] 1. CLAHE (Adaptive Histogram Equalization): It highlights the detail features in the image by enhancing the local contrast of the image.
[0084] 2. Data enhancement: Generate diverse training data through various means including random rotation, translation, scaling, cropping, color jittering, etc.
[0085] 3. Sample balance and complex data set construction: Use SMOTE (synthetic minority oversampling technology) and RandomUnderSampler (random undersampling technology) to balance the data to avoid the impact of class imbalance during model training.
[0086] Through the above steps, the system generates a high-quality and balanced dataset, providing a solid foundation for subsequent model training.
[0087] Object segmentation model training:
[0088] The system uses the SAM segmentation model to accurately segment the objects in the image. When dealing with image segmentation tasks in complex scenes, a multi-stage segmentation method is adopted:
[0089] Sliding window segmentation:
[0090] The oversized image is divided into multiple small windows, and each small window is input into the SAM model for segmentation with a suitable size. The segmentation result includes the precise outline of the target and the segmentation mask. The specific steps are as follows:
[0091] (1) Sliding window definition: Assume that the image size is W×H, the sliding window size is w×h, and the sliding step size is s x and y The sliding window moves on the image according to the following formula:
[0092] x i =i·s x ,y j =j·s y
[0093] Where i, j are integers such that 0≤x i ≤Ww and 0≤y j ≤Hh.
[0094] (2) Local segmentation: For each sliding window, the segmentation result D k , after local optimization, it is transformed in the global coordinate system.
[0095] (3) Global fusion: Finally, fusion is performed through global optimization to obtain the global segmentation result D final .
[0096] Local optimization and global optimization:
[0097] (1) Local optimization: First, local optimization is performed within each sliding window to refine the segmentation results within the window.
[0098] (2) Global optimization: Then, the segmentation results of all sliding windows are merged globally and globally optimized to ensure the consistency and accuracy of the overall segmentation.
[0099] The innovative combination of sliding window segmentation and global optimization technology in target segmentation ensures accurate recognition of targets in very large images. Sliding window segmentation enables the model to process very large images, while global optimization avoids errors caused by inconsistent local segmentation.
[0100] Through the above steps, the SAM model can be effectively trained for the object segmentation task. After the training is completed, the model will be able to perform accurate object segmentation on new images.
[0101] Object detection model training:
[0102] The system uses the YOLOv8 model to perform multi-target detection on drug packaging and medicinal material images. When processing oversized images, an innovative multi-stage detection method is adopted:
[0103] Sliding window detection: The oversized drug and medicinal material images are divided into multiple small windows, and each small window is input into the YOLOv8 model for detection with a suitable size. The detection results include the target location and confidence score.
[0104] Local NMS and global NMS: First, local NMS is performed in each sliding window to suppress redundant detections in the window. Then, the detection results of all sliding windows are merged globally, and global NMS is performed to remove overlapping targets. The specific explanation is as follows:
[0105] Assume that the image size is W×H, the sliding window size is w×h, and the sliding step size is s x and y The sliding window moves on the image according to the following formula:
[0106] x i =i·s x ,y j =j·s y
[0107] Where i, j are integers such that 0≤x i ≤Ww and 0≤y j ≤Hh. For each sliding window detection result D k After local NMS processing, it is transformed in the global coordinate system and finally fused through global NMS to obtain the global detection result D final .
[0108] The innovative combination of sliding window detection and global NMS technology in target detection ensures accurate recognition of targets in very large images. Sliding window detection enables the model to process very large images, while global NMS avoids misjudgment caused by overlapping detection.
[0109] Target matching model training and reasoning:
[0110] The target matching model adopts a dual-head network structure, which improves the recognition and matching capabilities of the model by combining position matching and category matching. A variety of loss functions are used in the matching process. The specific steps are as follows:
[0111] Two-head network architecture: The model structure includes a position matching branch and a category matching branch. After the input image is extracted through the convolution layer, the matching feature vector and category prediction are output respectively.
[0112] Triplet Loss with Hard Negative Mining: used for position matching, measures the similarity between samples by cosine distance, and combines the hard example mining strategy to optimize the feature vector so that the distance between positive samples is minimized and the distance between negative samples is maximized:
[0113] L triplet =max(0,cos(f a ,f n )-cos(f a ,f p )+margin)
[0114] ArcFace loss: used for category matching. By introducing angle intervals, it enhances the aggregation of similar samples in high-dimensional space and optimizes the accuracy of category prediction:
[0115]
[0116] NTXent loss: used in the contrastive learning stage to optimize feature distribution and enhance the distinction between different categories. This loss adjusts the feature distribution through a temperature scaling mechanism to make the distance between similar samples as close as possible:
[0117]
[0118] Category-sensitive cross entropy loss: This is applied in the category matching stage to address the category imbalance problem, ensuring that the model remains robust even when category distribution is imbalanced:
[0119]
[0120] Knowledge distillation loss: During the student network training phase, the knowledge of the teacher network is transferred to the student network through knowledge distillation of the final output logits and feature vectors, achieving lightweight model while maintaining high performance:
[0121]
[0122] BYOL self-supervised learning loss: used for feature extraction of the teacher network, enhancing the expressiveness of unlabeled data through self-supervised learning and improving the generalization performance of the model:
[0123]
[0124] Network parameter L2 regularization loss: used to suppress the overfitting problem of complex networks during complex task training:
[0125] L L2 =λ∑k‖θ k ||2
[0126] Finally, we get the comprehensive loss used by the teacher model:
[0127] L teacher =α1L triplet +α2L Arcface +α3L NTXent +α4L CE +α5L BYOL +α6L L2
[0128] The comprehensive loss used by the student model is:
[0129] L student =α1L triplet +α2L Arcface +α3L NTXent +α4L CE +α5L distill +α6L L2
[0130] Through the combined effect of multiple losses, the accuracy and stability of the matching model are significantly improved. Triplet loss and ArcFace loss ensure the accuracy of position matching and category matching, while BYOL and NTXent losses enhance the robustness and versatility of features. Knowledge distillation loss effectively realizes the knowledge transfer between the teacher-student network, making the lightweight model have high performance.
[0131] Model reasoning and matching judgment:
[0132] In the inference stage, the system inputs the extracted feature vector into the matching judgment module and uses the cosine distance and F1 norm to make the final matching judgment.
[0133] Cosine distance calculation: The system uses cosine distance as a matching metric to determine the similarity between two targets:
[0134]
[0135] Calculation of the optimal matching threshold: In order to determine the best matching threshold, the system calculates the F1 norm of positive and negative samples and selects the optimal threshold to ensure the accuracy and stability of the matching results:
[0136]
[0137] In view of the defects and shortcomings of the above-mentioned prior art, the present invention has the following advantages:
[0138] 1. Multi-stage detection and matching mechanism: Through the combination of sliding window detection, local NMS and global NMS, the processing problem of oversized prescriptions and drug images is solved, high-precision target detection and matching is achieved, and redundant detection and missed detection problems in traditional methods are avoided.
[0139] 2. Efficient matching of dual-head network structure: The present invention adopts a dual-head network structure to process position matching and category matching independently, so that category information does not interfere with position matching. At the same time, category information is used to assist matching decisions, thereby improving the matching accuracy and robustness of the system.
[0140] 3. Combination optimization of multiple loss functions: Through the combined use of multiple loss functions (such as triplet loss, self-supervised learning loss, distillation loss, ArcFace loss, etc.), the model's feature expression ability, matching accuracy and classification performance are comprehensively improved, especially when dealing with complex samples and category imbalance.
[0141] 4. Efficient knowledge distillation and model lightweighting: Through knowledge distillation technology, the high-performance features of the teacher network are transferred to the student network, realizing the lightweight design of the model, so that the system can still maintain efficient reasoning performance in a resource-constrained environment.
[0142] 5. Optimal matching judgment guided by F1 norm: The present invention adopts F1 norm to calculate the optimal matching threshold to ensure the accuracy and stability of the matching results, especially to achieve a good balance between recall rate and accuracy rate.
[0143] 6. Versatility and scalability: The system supports the processing of multiple data formats (such as PDF, images, vector graphics), adapts to the needs of multi-source heterogeneous data, has good versatility and scalability, and is suitable for a variety of industrial and engineering application scenarios.
[0144] In a further preferred embodiment, it can be specifically divided into two embodiments:
[0145] Implementation 1: Training process of dual-head matching network
[0146] The dual-head matching network in the present invention adopts a teacher-student architecture, combines multiple loss functions, and improves the performance of the student model through a knowledge distillation strategy. Figure 1 shown.
[0147] During the training process, the teacher network and the student network are trained using the following loss functions:
[0148] Teacher network loss function:
[0149] Triplet loss with hard example mining: By optimizing the similarity of feature vectors, it ensures that samples of the same class are closer and samples of different classes are farther away.
[0150] BYOL self-supervised loss minimizes the mean square error loss of two different transform graphs derived from the same source image by the network model, and strengthens the network's ability to learn the intrinsic features of the image.
[0151] ArcFace loss: By introducing angle intervals, the aggregation of similar samples in high-dimensional space is enhanced.
[0152] NTXent loss: Optimizing feature distribution via a temperature scaling mechanism in contrastive learning.
[0153] Category-sensitive cross entropy loss: introduces category weights to enhance the robustness of the model in the case of unbalanced samples.
[0154] L2 regularization loss: suppress overfitting by performing L2 regularization on model parameters.
[0155] Student network loss function:
[0156] The loss is the same as the teacher network, but the BYOL self-supervised loss is not used and the distillation loss is added.
[0157] Distillation loss: includes logits distillation and feature vector distillation, ensuring that the student network acquires the knowledge of the teacher network while maintaining high efficiency and light weight.
[0158] Implementation Method 2: Overall Matching Process of Drugs, Medicinal Materials and Prescription Images
[0159] For the task of matching medicines, medicinal materials and prescription images, the present invention has designed a complete set of processes that can efficiently and accurately detect and match elements in prescriptions and medicine images. Figure 2 shown.
[0160] The process is divided into the following key steps:
[0161] 1. Image acquisition: First, use the camera to obtain the images of the drugs and prescriptions.
[0162] 2. Perform single target segmentation on multi-target drug images.
[0163] 3. The segmented images are classified and sorted according to a specific algorithm.
[0164] 4. Identify and process prescription images, and use OCR to recognize text information.
[0165] 5. Drug library retrieval alignment: By identifying the segmented drug information, search the drug library to obtain the corrected drug information.
[0166] 6. Sliding window detection: The aligned image is divided into blocks for detection using a sliding window method to ensure that the target in each local area can be effectively detected.
[0167] 7. Global coordinate recovery and NMS: The global coordinates of the local targets detected by sliding window are restored, and redundant detection boxes are removed by global non-maximum suppression (NMS).
[0168] 8. Dual-image drug prescription position matching based on the Hungarian algorithm: The weighted Hungarian algorithm is used to match the drug positions in the two images to obtain possible matching pairs.
[0169] 9. Drug matching based on cosine distance matching network: A double-head matching network is used to further verify the above matching pairs, and the matching score is calculated by cosine distance.
[0170] 10. Similarity analysis and difference analysis: Perform similarity analysis and difference analysis on all matching pairs, and finally output the results and generate a JSON file.
[0171] As described above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as limiting the present invention itself. Various changes may be made to it in form and detail without departing from the spirit and scope of the present invention as defined in the appended claims.
Claims
1. A method for confirming prescription and drug matching in an unmanned pharmacy based on deep learning, characterized by: The following steps are involved: S1. Use a camera to obtain the images of the medicines and prescriptions, and perform data preprocessing on the collected images and construct a complex data set; S2. Perform single-target segmentation on the multi-target drug images in the constructed complex data set, and classify and sort the segmented images; simultaneously identify and process prescription images, and use optical character recognition text information; S3, identifying the segmented drug information, searching the drug database, and obtaining the corrected drug information; S4, performing block detection on the aligned image by sliding window method to ensure that the target in each local area can be effectively detected; S5, restore the global coordinates of the local targets detected by the sliding window, and remove redundant detection frames by global non-maximum suppression; S6. Matching the positions of the drugs in the two images by using the weighted Hungarian algorithm to obtain possible matching pairs; S7, using a double-head matching network to further verify the above matching pair and calculate the matching score by cosine distance; S8. Perform similarity analysis and difference analysis on all matching pairs, and finally output the results and generate a JSON file.
2. According to claim 1, a method for confirming prescription and drug matching in an unmanned pharmacy based on deep learning, characterized in that: The specific operations of data preprocessing and building complex data sets in step S1 are as follows: S101, highlighting the detail features in the image by enhancing the local contrast of the image; S102, performing data enhancement processing by random rotation, translation, scaling, cropping, and color jittering to generate diversified training data; S103. Balance the data by synthesizing minority class oversampling technology and random undersampling technology to avoid the impact of class imbalance during model training.
3. The method for confirming prescription and drug matching in an unmanned pharmacy based on deep learning according to claim 2, characterized in that: In step S2, the SAM segmentation model is used to accurately segment the target in the image, and the image in the complex scene is segmented by a multi-stage segmentation method; The multi-stage segmentation is a sliding window segmentation, specifically, an image of a predetermined size is divided into multiple small windows, each small window is input into the SAM model for segmentation with a corresponding size, and the segmentation result includes the precise outline of the target and the segmentation mask. The specific steps are as follows: S201, sliding window definition: Assume that the image size is W×H, the sliding window size is w×h, and the sliding step size is s x and y ; The sliding window moves on the image according to the following formula: x i =i·s x ,y j =j·s y Where i, j are integers such that 0≤x i ≤Ww and 0≤y j ≤Hh; S202, local segmentation: for each sliding window segmentation result D k , after local optimization, it is transformed in the global coordinate system; S203, global fusion: Finally, fusion is performed through global optimization to obtain the global segmentation result D final .
4. The method for confirming prescription and drug matching in an unmanned pharmacy based on deep learning according to claim 3, characterized in that: The local optimization is to refine the segmentation results within each sliding window; the global optimization is to merge the segmentation results of all sliding windows in a global scope to ensure the consistency and accuracy of the overall segmentation.
5. The method for confirming prescription and drug matching in an unmanned pharmacy based on deep learning according to claim 3, characterized in that: The specific method of block detection in step S4 is to use the YOLOv8 model to perform multi-target detection on drug packaging and medicinal material images; when processing images of a predetermined size, a multi-stage detection method is used for processing, including sliding window detection, local non-maximum suppression, and global non-maximum suppression; The sliding window detection specifically involves dividing a drug or medicinal material image of a predetermined size into multiple small windows, each small window of a suitable size is input into a YOLOv8 model for detection, and the detection result includes a target position and a confidence score; The local non-maximum suppression and global non-maximum suppression are to perform local non-maximum suppression in each sliding window to suppress redundant detection in the window; merge the detection results of all sliding windows globally, and perform global non-maximum suppression to remove overlapping targets. The specific formula is as follows: Set the image size to W×H, the sliding window size to w×h, and the sliding step size to s x and y ; The sliding window moves on the image according to the following formula: x i =i·s x ,y j =j·s y Where i, j are integers such that 0≤x i ≤Ww and 0≤y j ≤Hh. For each sliding window detection result D k After local NMS processing, it is transformed in the global coordinate system and finally fused through global NMS to obtain the global detection result D final .
6. The method for confirming prescription and drug matching in an unmanned pharmacy based on deep learning according to claim 1, characterized in that: The dual-head matching network structure in step S7 improves the recognition and matching capabilities of the model by combining position matching and category matching; a variety of loss functions are used in the matching process, and loss calculations are performed in a targeted manner.
7. The method for confirming prescription and drug matching in an unmanned pharmacy based on deep learning according to claim 6, characterized in that: The dual-head network structure includes a position matching branch and a category matching branch. After the input image is passed through the convolution layer to extract features, the matching feature vector and category prediction are output respectively. The target allocation of the various loss functions is as follows: Type A loss function: used for position matching, measures the similarity between samples by cosine distance, and combines mining strategies to optimize feature vectors so that the distance between positive samples is minimized and the distance between negative samples is maximized: L triplet =max(0,cos(f a ,f n )-cos(f a ,f p )+margin); Type B loss function: used for category matching. By introducing angle intervals, it enhances the aggregation of similar samples in high-dimensional space and optimizes the accuracy of category prediction: Type C loss function: used in the contrastive learning stage to optimize feature distribution and enhance the distinction between different categories. This loss adjusts the feature distribution through a temperature scaling mechanism to make the distance between similar samples as close as possible: D-type loss function: It is used in the category matching stage to solve the problem of category imbalance and ensure that the model still has good robustness when the category distribution is uneven: Type E loss function: During the student network training phase, the knowledge of the teacher network is transferred to the student network through knowledge distillation of the final output logits and feature vector features, achieving lightweight model while maintaining high performance: F-type loss function: used for feature extraction of the teacher network, enhancing the expressiveness of unlabeled data through self-supervised learning and improving the generalization performance of the model: G loss function: used to suppress the overfitting problem of complex networks during complex task training: L L2 =λ∑k||θ k || 2 ; Through the above-mentioned multiple loss functions, the practical comprehensive loss of the teacher model is finally obtained as follows: L teacher =α1L triplet +α2L Arcface +α3L NTXent +α4L CE +α5L BYOL +α6L L2 ; The comprehensive loss used by the student model is: L student =α1L triplet +α2L Arcface +α3L NTXent +α4L CE +α5L distill +α6L L2 。 8. The method for confirming prescription and drug matching in an unmanned pharmacy based on deep learning according to claim 7, characterized in that: In step S7, the cosine distance calculation matching is specifically that the system inputs the extracted feature vector into the matching judgment module, and uses the cosine distance and F1 norm to make the final matching judgment. The specific calculation formula is as follows: S701, the system uses cosine distance as a matching metric to determine the similarity of two targets: S702, the system calculates the F1 norm of positive and negative samples and selects the optimal threshold to ensure the accuracy and stability of the matching results:
9. A confirmation warning system for a method for confirming prescription and drug matching in an unmanned pharmacy based on deep learning as claimed in any one of claims 1 to 8, characterized in that include: Data preprocessing and complex data set model building: Preprocess images of medicinal materials, drugs, and prescriptions to build a balanced and diverse training data set. Target segmentation model: Use the SAM model to segment the total drug image and separate individual drugs from the image for subsequent recognition and matching processing. Target detection model: Use the YOLOv8 model for multi-target detection, combined with sliding window detection and global non-maximum suppression technology, to process images of a predetermined size. Target matching model: A dual-head network structure is used to achieve accurate position matching and category matching of targets. Model reasoning and matching judgment model: Optimize feature extraction and matching decisions through multiple loss functions, and use cosine distance and F1 norm for final matching judgment.