Rotary medicine box identification and correction method and device based on AI
By constructing a medicine box data set with rotation angle labeling, training the neural network model and performing affine transformation correction, combined with the multimodal drug matching model, the problem of degradation of recognition accuracy caused by the rotation of the medicine box image is solved, and efficient drug information recognition and matching in complex backgrounds is achieved.
Patent Information
- Application Number
- CN202510368087.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the rotation angle problem of the medicine box image leads to a decrease in the accuracy of drug recognition, especially in complex backgrounds, which is difficult to adapt to the angle estimation requirements, and the traditional detection model lacks the feature extraction ability of the rotation target.
By constructing a medicine box data set with rotation angle labeling, the neural network model is trained, the medicine box detection model is obtained, and the affine transformation algorithm is used to correct the rotation angle of the medicine box to the standard horizontal direction, and cross-modal search is carried out in combination with the multimodal drug matching model to improve the recognition accuracy of drug information.
It effectively solves the problem of inaccurate drug identification caused by the angle of the medicine box image, improves the recognition accuracy of drug information and the adaptability of system, especially under complex shooting conditions, ensuring efficient matching and recognition of drug information.
Smart Images

Figure CN120340037A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the medical field, and in particular, to an AI-based rotating medicine box recognition and correction method and device. Background Art
[0002] With the in-depth application of artificial intelligence technology in the field of medical health, the medicine recognition system based on computer vision plays an increasingly important role in scenarios such as pharmacy automation management, intelligent prescription review systems, and patient medication guidance. In the prior art, the automatic recognition of medicine box images mainly relies on target detection algorithms (such as YOLO, Faster R-CNN, etc.) combined with OCR text recognition technology to extract medicine information.
[0003] However, the medicine box images collected in actual application scenarios often have the problem of multi-angle rotation. The main reasons include: (1) non-professional shooting angles when patients take selfies; (2) manipulator operation deviations during the medicine-taking process of automatic equipment; (3) natural tipping of medicine boxes during transportation and storage. Through experimental verification, when the rotation angle of the medicine box exceeds ±15°, the recognition accuracy of traditional detection models will decrease by 32%-45%. The fundamental reasons are as follows: First, existing models are mostly based on training data labeled in the horizontal positive direction, and have insufficient feature extraction ability for rotating targets; second, conventional affine transformation preprocessing schemes are difficult to meet the angle estimation requirements in complex backgrounds, especially when dealing with special scenarios such as reflective material packaging, densely arranged patterns, or partial occlusion, it is easy to generate angle misjudgments.
[0004] For the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of the present invention provide an AI-based rotating medicine box recognition and correction method and device to at least solve the technical problem that the medicine cannot be accurately recognized due to the angle inclination of the medicine box picture.
[0006] According to one aspect of the embodiments of the present invention, an AI-based rotating medicine box recognition and correction method is provided, including: obtaining a target medicine box picture to be recognized; using a medicine box detection model to recognize the target medicine box picture to be recognized, and correcting the rotation angle of the recognized target medicine box to the standard horizontal direction; wherein, the medicine box detection model is obtained by: training a neural network model based on a medicine box data set with rotation angle annotations to obtain the medicine box detection model.
[0007] According to another aspect of the embodiments of the present invention, there is also provided an AI-based rotating medicine box recognition and correction device, including: an acquisition module configured to acquire a target medicine box picture to be recognized; a recognition module configured to recognize the target medicine box picture to be recognized by using a medicine box detection model and correct the rotation angle of the recognized target medicine box to the standard horizontal direction; wherein, the medicine box detection model is obtained by: training a neural network model based on a medicine box data set with rotation angle annotations to obtain the medicine box detection model.
[0008] In the embodiments of the present invention, a target medicine box picture to be recognized is acquired; the target medicine box picture to be recognized is recognized by using a medicine box detection model, and the rotation angle of the recognized target medicine box is corrected to the standard horizontal direction; wherein, the medicine box detection model is obtained by: training a neural network model based on a medicine box data set with rotation angle annotations to obtain the medicine box detection model. Through the above solution, the technical problem that medicines cannot be accurately recognized due to the angle inclination of the medicine box picture is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and the illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0010] Figure 1 is a flowchart of an intelligent medication consultation method based on a large model according to an embodiment of the present invention;
[0011] Figure 2 is a flowchart of a method for identifying and correcting a medicine box by using AI according to an embodiment of the present invention;
[0012] Figure 3 is a flowchart of a cross-modal retrieval and matching method for medicine information based on AI according to an embodiment of the present invention;
[0013] Figure 4 is a flowchart of an intelligent medicine box recognition and medication consultation aging-friendly method for a multi-modal large model according to an embodiment of the present invention;
[0014] Figure 5 is a training process diagram of a medicine box detection model and a multi-modal medicine matching model according to an embodiment of the present invention;
[0015] Figure 6 is an interface diagram of medication consultation and a medicine electronic instruction manual according to an embodiment of the present invention;
[0016] Figure 7 is an interface diagram of multi-modal retrieval of a medicine vector library according to an embodiment of the present invention;
[0017] Figure 8 is the structure diagram of the ChineseClip model according to an embodiment of the present invention;
[0018] Figure 9 is the schematic diagram of the corrected picture according to an embodiment of the present invention;
[0019] Figure 10 is the flowchart of user interaction according to an embodiment of the present invention;
[0020] Figure 11 is the interface diagram of medicine box recognition and aging adaptation according to an embodiment of the present invention;
[0021] Figure 12 is the interface diagram of the electronic medicine instruction manual and intelligent medication consultation according to an embodiment of the present invention;
[0022] Figure 13 is the interface diagram of medication reminder and expiration reminder setting according to an embodiment of the present invention;
[0023] Figure 14 shows the schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. Detailed implementation manners
[0024] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0026] According to an embodiment of the present invention, a method embodiment of an intelligent medication consultation method based on a large model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0027] Figure 1 is a flowchart of an intelligent medication consultation method based on a large model according to an embodiment of the present invention, as Figure 1 shown, the method includes the following steps:
[0028] Step S102, obtain user input information.
[0029] Step S104, in the case where the user input information is text, use the pre-trained language model BGE to extract features from the text to obtain a first feature representation, perform chunked retrieval on the electronic drug instructions based on the first feature representation, and perform context integration on the chunked retrieval results to generate medication consultation information for responding to the user input information.
[0030] In the case where the text is recognized as a question, perform QA feature extraction on the text to obtain the first feature representation, perform chunked retrieval on the electronic drug instructions based on the first feature representation, and perform context integration on the retrieval results to generate medication consultation information for responding to the user input information. In this embodiment, the first feature representation of the question text is obtained through QA feature extraction, and chunked retrieval is performed in the electronic drug instructions based on this feature, effectively improving the pertinence and accuracy of the retrieval. At the same time, the retrieval results are optimized through context integration, making the generated medication consultation information more complete and coherent, and avoiding understanding deviations caused by information fragmentation.
[0031] In the case where the text is intended to be recognized as a non-question, cross-modal feature extraction is performed on the text to obtain the first feature representation. Based on the first feature representation, cross-modal retrieval is performed to match the associated medicine box image and text description from the medicine picture and text database, obtaining a cross-modal retrieval result. Based on the cross-modal retrieval result, block retrieval is performed on the electronic medicine instructions, and the retrieval results are contextually integrated to generate medicine use consultation information for responding to the user input information. For example, based on the first feature representation, the cosine similarity algorithm is used to retrieve relevant content segments in the instruction vector library generated based on the electronic medicine instructions, and through semantic similarity ranking, the most relevant content paragraphs are screened out from the relevant content segments; the most relevant content paragraphs are integrated as context information, and based on the context information, a large language model is used to generate medicine use consultation information for responding to the user input information. This embodiment combines cross-modal feature extraction and retrieval technologies. When the user inputs non-question text, it can intelligently match the medicine box image and text description, and further combine with the block retrieval of the electronic medicine instructions to ensure the comprehensiveness and accuracy of the information. Through the cosine similarity algorithm and semantic similarity ranking, the screening of relevant content segments is optimized, improving the precision of the retrieval results.
[0032] In some embodiments, the instruction vector library is obtained as follows: all relevant fields of the electronic medicine instructions are concatenated into a complete document in logical order, and the concatenated document is preprocessed for standardization; for the preprocessed document, it is blocked with a preset number of characters as the basic unit, and an overlapping area with a preset overlap degree is set between adjacent blocks, where the preset number of characters is 1024 characters, and the preset overlap degree is 10% of the number of characters in adjacent blocks. Based on the blocked document, the instruction vector library is constructed. This embodiment improves the integrity and retrieval accuracy of vector library construction by preprocessing the electronic medicine instructions for standardization and using fixed-length blocking and overlapping area setting. The blocking method ensures information coherence, avoids content fragmentation, and improves the accuracy and semantic matching effect of subsequent retrieval.
[0033] Step S106, in the case where the user input information is a picture, perform medicine box detection on the picture, identify the target medicine box, and perform cross-modal feature extraction based on the target medicine box to obtain a second feature representation. Based on the second feature representation, perform cross-modal retrieval to generate medicine use consultation information for responding to the user input information.
[0034] 1) Perform medicine box detection on the picture, identify the target medicine box, and correct the target medicine box.
[0035] For example, identify the main body position of the target medicine box in the picture; when the rotation angle of the main body position deviates from the standard horizontal direction, use the affine transformation algorithm to correct the rotation angle of the main body position to the standard horizontal direction. In the embodiment of the present application, the rotation angle of the target medicine box is corrected by the affine transformation algorithm to align it to the standard horizontal direction, improving the accuracy of image normalization processing, ensuring the stability and accuracy of subsequent image recognition and information extraction, and being applicable to the image processing of medicine boxes photographed at different angles.
[0036] In some embodiments, AI can be used to identify and correct the medicine box. For example, as Figure 2 shown, the method for identifying and correcting a rotated medicine box based on AI includes the following steps:
[0037] Step S202, obtain a picture of the target medicine box to be identified.
[0038] Step S204, use the medicine box detection model to identify the picture of the target medicine box to be identified, and correct the rotation angle of the identified target medicine box to the standard horizontal direction.
[0039] In some embodiments, the medicine box detection model is obtained as follows: based on a medicine box data set with rotation angle annotations, train a neural network model to obtain the medicine box detection model. The specific training process can be as follows:
[0040] First, obtain medicine box sample data, where the medicine box sample data includes medicine box sample pictures.
[0041] Next, perform rotation angle annotation on each medicine box sample picture in the medicine box sample data to construct the medicine box data set with rotation angle annotations, and perform data augmentation on the medicine box data set with rotation angle annotations, where the data augmentation includes at least one of the following: random angle rotation, scale transformation, and brightness contrast adjustment. For example, annotate the four vertex coordinates of the minimum bounding rectangle and the rotation angle relative to the horizontal axis of each medicine box sample picture on each medicine box sample picture to obtain the medicine box data set with rotation angle annotations; perform data augmentation on the medicine box data set by at least one of the following: perform the random angle rotation within ±90°, perform the scale transformation between 0.8 and 1.2 times, and perform the brightness contrast adjustment within ±30°. In this embodiment, by performing rotation angle annotation on the medicine box sample pictures to construct a medicine box data set with rotation angle information, the recognition ability of the model for medicine box images at different angles is improved. The data augmentation strategy further enhances the generalization performance and robustness of the model, enabling it to stably recognize medicine box information under different shooting conditions, improving the adaptability and accuracy of image processing, and being applicable to automatic recognition and information extraction of drug packaging in complex scenarios.
[0042] Finally, based on the enhanced medicine box dataset, the neural network model is trained to obtain the medicine box detection model. Specifically, using the oriented bounding box (OBB), redundant rotated detection boxes in each medicine box sample image in the enhanced medicine box dataset are removed to obtain target detection boxes. For example, for each medicine box sample image, when calculating the intersection over union (IoU) between the rotated detection boxes using the OBB, the calculation method of the rotated IoU is adopted, and the final overlap degree is obtained by dividing the intersection area of the polygons by the union area; the final overlap degree is used to filter out redundant rotated detection boxes. Then, an affine transformation is used to correct the target detection boxes and adjust the target detection boxes to the 0° standard horizontal direction; the adjusted angle of the target detection boxes is compared with the labeled angle of each medicine box sample image, and based on the comparison result, the loss function of the neural network model is adjusted to train the neural network model. In this application, the neural network model is trained with the enhanced medicine box dataset, and the oriented bounding box (OBB) is used to optimize the rotated detection boxes, remove redundant boxes, and improve the accuracy of target detection. The rotated IoU is used to calculate the overlap degree to ensure that the selected detection boxes are optimal. The affine transformation corrects the detection boxes to the standard horizontal direction, and the loss function is adjusted in combination with the labeled angle to optimize the model training effect. This method improves the stability and robustness of the medicine box detection model, making it more suitable for the medicine packaging recognition task under different shooting angles and complex environments.
[0043] 2) Perform cross-modal feature retrieval.
[0044] Based on the second feature representation, the multi-modal medicine matching model is used to perform cross-modal matching on the medicine text-image vector database to obtain medicine information; wherein, the medicine text-image vector database is constructed by at least one of the following: combining the brand and the medicine name as the main text description, and combining the key features of the medicine to form a complete text description information; performing at least one of the following data augmentation processes on the medicine sample images: random angle rotation within the range of ±30°, brightness contrast adjustment, and horizontal and vertical flipping; performing at least one of the following data augmentation processes on the training sample words: synonym replacement, and swapping the brand and the generic name. In this application, the multi-modal medicine matching model is used to perform cross-modal matching on the medicine text-image vector database, improving the accuracy and comprehensiveness of medicine information retrieval. Combining the brand and the medicine name and the key features to construct a complete text description makes the matching result more targeted. At the same time, through various data augmentation processes on the medicine sample images and the training text data, the robustness and generalization ability of the model are improved, ensuring accurate matching of medicine information under different lighting, angles, and text variant conditions, and enhancing the adaptability and reliability of cross-modal retrieval.
[0045] In some embodiments, AI can be utilized for cross-modal feature retrieval. For example, as Figure 3 shown, the cross-modal retrieval matching method for drug information based on AI may include the following steps:
[0046] Step S302, obtain a target medicine box picture to be recognized.
[0047] Step S304, use a multi-modal drug matching model to extract features from the target medicine box picture to obtain a feature representation.
[0048] Step S306, based on the feature representation, use the multi-modal drug matching model for cross-modal matching to match drug information related to the target medicine box.
[0049] Among them, the multi-modal drug matching model is obtained as follows: Based on a training set of text-image pairs in the drug field, perform domain adaptation training on a pre-trained multi-modal network model to obtain the multi-modal drug matching model. The specific training process is as follows:
[0050] First, obtain the text-image pair training set. Obtain drug sample data, where the drug sample data includes drug sample pictures and training text data; perform at least one of the following enhancement processes on the drug sample pictures: perspective transformation, lighting adjustment, and local magnification, and expand the training text data by means of synonym replacement and interchange of brand names and generic names; use the enhanced drug sample pictures and the expanded training text data as the text-image pair training set. This embodiment effectively improves the training effect of the cross-modal model, enabling it to more accurately and stably match and identify drug information in practical applications.
[0051] Next, based on the text-image pair training set, pre-train the multi-modal network model in the field of pharmaceuticals to obtain the multi-modal pharmaceutical matching model. Pre-train the multi-modal network model on the general Chinese text-image dataset; based on the text-image pair training set, pre-train the pre-trained multi-modal network model in the field of pharmaceuticals to obtain the multi-modal pharmaceutical matching model. For example, based on the attention mechanism enhanced by pharmaceutical knowledge, use the model to extract the key visual features in the pharmaceutical sample images in the text-image pair training set through an additional pharmaceutical attribute supervision signal; through the attention mechanism enhanced by the embedding of the pharmaceutical knowledge graph, combined with the pharmaceutical attribute classification loss function, use the multi-modal network model to extract key visual features; when encoding the training text data in the text-image pair training set, fuse the prior information of entity relationships in the pharmaceutical knowledge graph; use the key visual features and the text encoding that has fused the prior information of entity relationships to perform the pre-training in the field of pharmaceuticals. In the embodiments of the present application, by pre-training the multi-modal network model on the general Chinese text-image dataset and further pre-training it in combination with the text-image pair training set in the field of pharmaceuticals, the accuracy of the model in the pharmaceutical matching task is improved. In addition, the attention mechanism enhanced by pharmaceutical knowledge and the pharmaceutical attribute supervision signal are used to optimize the visual feature extraction ability; finally, the prior information of entity relationships in the pharmaceutical knowledge graph is fused to improve the semantic understanding ability of text encoding, thereby enhancing the accuracy and generalization ability of cross-modal matching and realizing efficient pharmaceutical information retrieval and matching.
[0052] Finally, based on the feature representation, use the multi-modal pharmaceutical matching model to perform cross-modal matching to match the pharmaceutical information related to the target medicine box. Specifically, based on the feature representation, use the multi-modal pharmaceutical matching model to perform cross-modal matching on the text-image vector database to obtain the pharmaceutical information. For example, use the multi-modal pharmaceutical matching model to perform similarity retrieval on the text-image vector database and calculate the cosine similarity; sort based on the cosine similarity and select the top K candidate results with the highest similarity as the pharmaceutical information; wherein, the text-image vector database stores the image feature vectors of medicine box pictures and the feature vectors of the corresponding drug names, brands, indications, dosages, and instruction manual texts; the pharmaceutical information includes at least one of the following: the type, brand, indication, and instruction manual of the drug. The present application realizes accurate pharmaceutical information retrieval through cross-modal matching of the text-image vector database by the multi-modal pharmaceutical matching model.
[0053] The embodiment of the present invention provides an intelligent medicine box recognition and medication consultation system based on a multi-modal large model, aiming to solve the problems of insufficient accuracy and inconvenient information acquisition in the processes of medicine recognition, information query, and medication consultation. The system utilizes the pre-training in the medicine field of the multi-modal model, cross-modal retrieval, and the intelligent question-answering technology of the large model to provide efficient and accurate medicine information processing capabilities and offer a comprehensive intelligent solution for users' medication management.
[0054] This system includes a pre-training and cross-modal retrieval module, and a database construction and intelligent medication consultation module.
[0055] The pre-training and cross-modal retrieval module is used for the pre-training in the medicine field and cross-modal retrieval of the multi-modal large model. First, pre-training in the medicine field is carried out. For example, based on data of Western medicines, traditional Chinese medicines, medical devices, and health products, the ChineseCLIP model is pre-trained for the medicine field. The data covers pictures, names, and brands of medicines. Through constructing picture-text pairs for training, the cross-modal image and text understanding ability of the model in the medicine field is improved, realizing the efficient matching of medicine box pictures and medicine information. Then, the construction of a picture-text vector database and cross-modal retrieval are carried out. For example, the system vectorizes the medicine box picture and its corresponding medicine name, brand, and other descriptive texts and stores them in the picture-text vector database. When the user inputs by taking a picture of the medicine box or entering the medicine name, the system quickly matches the corresponding medicine information based on cross-modal retrieval technology, including the type, brand, indications, and other details of the medicine, thus realizing accurate and convenient medicine recognition and information query.
[0056] The database construction and intelligent medication consultation module is used for the construction of the electronic medicine instruction vector database and intelligent medication consultation. First, the construction of the electronic medicine instruction vector database is carried out. The electronic medicine instructions are stored in blocks by paragraph, covering complete information such as indications, usage and dosage, contraindications, adverse reactions, drug interactions, etc. of the medicine. The semantic vectors of each paragraph are generated through the BGE vector model and stored in the instruction vector database. Then, intelligent question-answering based on the knowledge base and vector matching is carried out. A medicine knowledge base is constructed, combining the structured information in the instructions with other authoritative data sources (such as pharmacopoeias, clinical guidelines) to form a comprehensive medicine knowledge base. When the user asks questions (such as "What are the indications of this medicine?", "Is it applicable to children?", "What are the adverse reactions?", etc.), the system first matches the relevant information in the instruction vector database through the user's query, and integrates it with the context semantic analysis technology and the knowledge base. The LLM large model generates accurate natural language answers. The intelligent question-answering function supports dynamic multi-round interactions and can generate coherent answers based on historical questions, providing a complete medicine consultation service for users. Cross-modal retrieval and intelligent question-answering.
[0057] The user initiates a retrieval request by taking a picture of the medicine box, entering the medicine name, or querying by voice. The system retrieves all the information of the corresponding medicine based on the matching results in the image-text vector database.
[0058] When the user queries specific questions such as the usage method, indications, and contraindications of a medicine, the system quickly matches relevant content through the electronic instruction manual vector database and combines the context semantic analysis ability of the multi-modal large model to generate accurate and natural intelligent answers to ensure comprehensive and accurate information.
[0059] The present invention applies the image and text understanding capabilities of the multi-modal large model to the scenarios of medicine identification and medication consultation, and combines the age-friendly design to significantly improve the efficiency of medicine information query and the convenience of medication management, especially having wide application value among the elderly user group.
[0060] The embodiment of the present application also provides an age-friendly method for intelligent medicine box identification and medication consultation of a multi-modal large model. This method is applied to the above-mentioned intelligent medicine box identification and medication consultation system based on the multi-modal large model. This method is as Figure 4 shown and includes the following steps:
[0061] S402, data preprocessing and model pre-training.
[0062] The overall process of data preprocessing and model pre-training is as Figure 5 shown.
[0063] The construction of the medicine vector library will be described in detail below.
[0064] The medicine vector library constructed by the present invention is a large-scale multi-source heterogeneous data set, which contains structured data records of multiple health products, medical devices, western medicines, and traditional Chinese medicines, etc. Each record contains multiple data fields, covering basic information such as product name, brand, specification, scope of application, storage conditions, etc., as well as regulatory information such as approval number, validity period, and manufacturer. At the same time, the record also contains clinical medication information such as main ingredients, precautions, usage and dosage, and contraindication symptoms, and provides medication guidance for special populations, such as children's medication, medication for pregnant and lactating women, etc. In addition, for different types of medicines, their specific professional information is also stored, such as the pharmacology and toxicology, pharmacokinetics of western medicines, and the structural composition, registration certificate number of medical devices, etc.
[0065] In the construction of multi-modal training samples, the present invention adopts an innovative strategy for aligning images and texts. First, the brand and drug name are combined as the main text description, and key features such as specifications and properties are incorporated to form a complete text description information. To improve the generalization ability and robustness of the model, the present invention performs all-round data augmentation on drug images, covering enhancement techniques such as random angle rotation within the range of 0° to 360°, brightness and contrast adjustment, horizontal and vertical flipping, etc. At the same time, multi-dimensional enhancement processing such as synonym replacement and interchange of brand and generic names is also carried out at the text level, effectively expanding the diversity of training samples.
[0066] For the processing of drug instructions, the present invention adopts a refined text chunking strategy. As Figure 6 shown, first, all relevant fields are concatenated into a complete electronic instruction document in logical order. After standardized preprocessing, chunks are made with 512 characters as the basic unit. To ensure semantic coherence and integrity, a 10%-15% overlapping area is set between adjacent chunks. This strategy not only ensures the independence of the content after chunking but also maintains the semantic association of the context. During the chunking process, special attention is paid to avoiding truncation of important information to ensure that each chunk has a complete semantic expression.
[0067] As Figure 7 shown, in the vector feature extraction and storage section, the present invention generates high-dimensional feature vectors for drug images, text descriptions, and instruction chunks respectively. The dimension of the image feature vector extracted by the pre-trained model is 768, the dimension of the text description vector is 768, and the dimension of the instruction chunk vector is 1024. These feature vectors are stored in the Milvus vector database, establishing an efficient multi-modal index structure to support fast similarity retrieval and approximate nearest neighbor search. At the same time, a perfect data quality control mechanism is implemented, including functions such as data integrity verification, outlier detection, and incremental update, ensuring the reliability and maintainability of the vector library.
[0068] The multi-modal drug matching model will be described in detail below.
[0069] The present invention adopts an improved ChineseClip multi-modal model as the core architecture, which is specifically optimized for the Chinese drug field based on the original CLIP model. As Figure 8 shown, the model as a whole consists of two major parts: an image encoder and a text encoder, and realizes the unified representation space mapping of image and text features through the contrastive learning method.
[0070] Among them, the image encoder adopts the Vision Transformer (ViT) architecture, which contains 12 layers of Transformer blocks, each layer has 12 attention heads, and the input resolution is uniformly adjusted to 224×224 pixels; the text encoder is based on the RoBERTa architecture, consisting of 12 layers of bidirectional Transformer, the vocabulary expands the professional vocabulary in the pharmaceutical field, and the maximum sequence length is set to 128 tokens.
[0071] In the model pre-training stage, we first constructed a training set of image and text pairs using 142,307 pieces of drug data. Considering the particularity of drug images, we improved the standard data enhancement strategy: on the image side, we introduced drug-specific enhancement methods such as perspective transformation, lighting adjustment, and local magnification; on the text side, we expanded the training samples by synonym replacement and brand generic name conversion. The training process adopted a contrastive learning strategy, and used the InfoNCE loss function to optimize the model parameters. The batch size was set to 1024, and the learning rate adopted a cosine annealing strategy, which gradually decayed from 1e-4.
[0072] In order to improve the recognition accuracy of the model in the pharmaceutical field, the present invention designs a special training framework. First, pre-training is performed on a general Chinese image and text dataset, and then domain adaptability fine-tuning is performed using the collected pharmaceutical dataset. In the fine-tuning stage, an attention mechanism for pharmaceutical knowledge enhancement is introduced, and the model is guided to focus on key visual features in pharmaceutical images, such as packaging features, dosage form features, and identification features, through additional pharmaceutical attribute supervision signals. At the same time, the prior information of the pharmaceutical knowledge graph is incorporated into the text encoding, which enhances the model's ability to understand pharmaceutical terminology.
[0073] The training objective design of the model adopts a multi-task learning framework. In addition to the basic image-text comparison loss, it also includes drug classification loss and feature consistency loss. Among them, the classification loss helps the model learn the category information of the drug, and the feature consistency loss ensures that the feature representation of the same drug from different perspectives is consistent. Through this multi-objective optimization strategy, the model can take into account the two key tasks of image-text alignment and drug feature extraction at the same time.
[0074] The optimized multimodal model demonstrates excellent feature extraction capabilities. In practical applications, the model can accurately understand the visual information of drug images and establish precise semantic associations with text descriptions, providing a reliable feature representation basis for subsequent smart medicine box recognition and medication consultation. At the same time, through the knowledge-enhanced feature extraction mechanism, the model has a strong ability to distinguish subtle differences in drugs and can effectively handle the identification of similar drugs.
[0075] The pill box detection model will be described below.
[0076] In view of the actual problem that there are diverse shooting angles for the medicine box pictures uploaded by users, the present invention innovatively proposes a rotating medicine box detection scheme based on OBB (Oriented Bounding Box). Although the samples have been augmented by angle rotation in the aforementioned multi-modal training stage, due to the randomness and complexity of the medicine box placement angles in the actual application scenarios, the recognition accuracy may still be insufficient. Therefore, the present invention deeply improves the traditional YOLO model, introduces a rotating object detection mechanism, and realizes the accurate positioning and angle correction of medicine boxes at any angle.
[0077] To improve the training effect of the model, the present invention specifically constructs a medicine box data set with annotated angles. Through a semi-automatic annotation tool, the four vertex coordinates of the minimum bounding rectangle and the main direction angle are annotated for each medicine box. At the same time, the training samples are augmented through data augmentation techniques, including random angle rotation, scale transformation, brightness and contrast adjustment, etc., to improve the adaptability of the model to various shooting conditions.
[0078] In the post-processing stage, in order to effectively filter redundant rotating detection boxes, the present invention designs a rotating non-maximum suppression algorithm (R-NMS). Different from the traditional non-maximum suppression (NMS) which is only applicable to horizontal rectangular boxes, R-NMS is optimized according to the characteristics of rotating objects. When calculating the IoU between detection boxes, the calculation method of rotating IoU is adopted, and the final overlap degree is obtained by dividing the intersecting area of polygons by the union area. The traditional IoU calculation only considers the intersection and union areas of horizontal rectangles, while the present invention calculates the overlap degree between rotating detection boxes, that is, the rotated intersection over union (Rotated IoU, R-IoU). Specifically, the present invention adopts the following steps:
[0079] 1) Parametric representation of the rotating detection box.
[0080] Each detection box is represented by a five-element (x c , y c , w, h, θ) group, where {(x} c , y c ) is the center coordinate of the box, w and h are the width and height of the box respectively, and θ is the main direction angle of the box (in radians, with a range of [-π / 2, π / 2]).
[0081] 2) Calculate the rotating IoU.
[0082] For two rotating detection boxes B i and B j , the formula for calculating R-IoU is:
[0083]
[0084] 3) Perform suppression.
[0085] Sort all the detected bounding boxes in descending order of confidence scores. Traverse the list of detected bounding boxes. For the current bounding box B with the highest score max , calculate its R-IoU with the remaining bounding boxes. If R-IoU > τ (τ is a preset threshold, such as 0.3), then suppress the bounding box with a lower score. Repeat the above steps until all the detected bounding boxes are processed.
[0086] Through the above method, R-NMS can effectively eliminate redundant rotated bounding boxes, while retaining key targets and improving the accuracy of the detection results. Compared with traditional NMS, R-NMS shows higher adaptability when dealing with rotated targets such as medicine boxes, especially in dense scenes.
[0087] In some other embodiments, redundant rotated bounding boxes can also be filtered based on the following improved Rotated Non-Maximum Suppression (R-NMS):
[0088] 1) Preprocessing.
[0089] Input the original set of rotated detected bounding boxes {B i = (x c , y c , w, h, θ, s)}. Generate a direction vector based on the angle θ: v x = cosθ, v y = sinθ, and output the standardized parameter set {B i ′ = (x c , y c , w, h, v x , v y , s)}.
[0090] Compensate the confidence based on the visible area:
[0091]
[0092] The compensated confidence queue Q = {B i ′ | sorted in descending order of s′}.
[0093] 2) Perform hierarchical screening.
[0094] Based on the current bounding box B with the highest confidence max and the candidate bounding box B j perform fast collision detection. The inputs are B max , B j . Apply the Separating Axis Theorem (SAT) to detect whether the projections are separated, and output a boolean value {True / False} indicating whether to continue the calculation.
[0095] 3) Calculate the multi-precision R-IoU.
[0096] Based on the center distance Select a calculation mode. When d > 2 × max(w max , h max ), estimate the overlapping area at the sampling points based on the Monte Carlo method to obtain an approximate R-IoU value. In some embodiments, the Greiner-Hormann algorithm can also be used to calculate the vertices of the exact intersection polygon to obtain the exact R-IoU value.
[0097] 4) Dynamic suppression.
[0098] Adjust the threshold based on the angle difference Δθ = ∣θ max - θ j ∣. Calculate the angle-sensitive threshold based on the base threshold τ0 and Δθ:
[0099] τ = τ0 + 0.1τ0(1 - cosΔθ)
[0100] Perform a secondary adjustment based on the local density . Calculate the final threshold based on the current threshold τ, density ρ, the number of points within a circle with a radius of r and the number of neighboring boxes, and ρ0 as the reference density:
[0101] τ final = τ × (1 - 0.2tanh(ρ / ρ0))
[0102] Judge whether to suppress based on R-IoU > τ final and output the updated detection box queue Q'. Based on the matrix-based parallel computing of GPU. Input the set of rotated box parameters, and convert each box into a homogeneous coordinate matrix, where M i represents the homogeneous coordinate matrix of the i-th rotated box, θ i represents the rotation angle of the i-th rotated box, and (x i , y i ) represents the center coordinates of the i-th rotated box. The calculation formula of Mi is as follows:
[0103]
[0104] Use CUDA thread blocks to perform parallel computing on the projection of 16 pairs of boxes and output the batch-computed R-IoU matrix.
[0105] Input the candidate box pair (B max , B j ), where F max represents the Fourier descriptor of the candidate box B max , and F j represents the Fourier descriptor of the candidate box B j . Calculate the Fourier descriptor distance, i.e., the shape similarity score S:
[0106]
[0107] Based on the overlapping region image patches, extract the shallow features of the pre-trained CNN and calculate the cosine similarity to obtain the texture similarity score T. Final suppression condition: (R - IoU > τ final ) ∧ (S < 0.2) ∧ (T > 0.7).
[0108] After detecting the rotation angle of the medicine box, the present invention further corrects the image by using affine transformation to adjust the medicine box to the standard horizontal direction. During the correction process, considering the preservation of image quality, based on the geometric parameters of the detection frame, combined with the bicubic interpolation algorithm for image resampling, the detail information of the image is retained to the greatest extent. The image after angle correction will be input into the aforementioned multi-modal model for feature extraction and recognition, thereby significantly improving the overall recognition accuracy of the system. The specific steps are as follows:
[0109] 1) Construct the transformation matrix.
[0110] For the detected center (x c , y c ) of the medicine box and the rotation angle θ, construct the affine transformation matrix M for rotating the image to the horizontal direction. The transformation matrix is defined as:
[0111]
[0112] This matrix rotates the image counterclockwise by θ degrees around the center (xc, yc) to align the main direction of the medicine box with the horizontal axis.
[0113] 2) Map the coordinates.
[0114] For each pixel point (x, y) in the original image, calculate the corrected coordinates (x′, y′) through M
[0115]
[0116] 3) Optimize the bicubic interpolation.
[0117] During the correction process, to maintain the image quality, the bicubic interpolation algorithm is used for image resampling. Where W(t) represents the weight function.
[0118]
[0119] When a = -0.5, it provides the best edge preservation and detail reconstruction ability. t represents the pixel position offset during resampling, usually in the range of [-1, 1], and is used to calculate the interpolation weight.
[0120] The optimized rotating medicine box detection model can effectively handle the angular changes of the medicine box under complex shooting conditions, ensuring that the system can maintain high-precision recognition ability in different scenarios. Through this solution, the accuracy of medicine box detection and angle correction has been significantly improved, further enhancing the practicability and reliability of the intelligent medicine box recognition system.
[0121] The specific optimization effects are as Figure 9 shown. The optimized rotating medicine box detection model can effectively handle the angular changes of the medicine box under complex shooting conditions, ensuring that the system can maintain high-precision recognition ability in different scenarios. Through this solution, the accuracy of medicine box detection and angle correction has been significantly improved, further enhancing the practicability and reliability of the intelligent medicine box recognition system.
[0122] Step S404, user interaction.
[0123] As Figure 10 shown, the process of user interaction includes the following steps:
[0124] Step S1002, obtain user input.
[0125] Obtain user input information. For example, pictures and text. Users can upload pictures of the medicine box, and the system preprocesses the pictures, including size adjustment, illumination normalization, and noise elimination. Then, through the rotation angle detection algorithm, the main body position of the medicine box is identified and its tilt angle is corrected to ensure that the image is upright in subsequent processing. Users can also input text. For example, users input the name of the medicine, and the system cleans, segments, and standardizes the text.
[0126] Step S1004, determine whether the user input is text or a picture.
[0127] If the user input is a picture, execute step S1006; otherwise, execute step S1012.
[0128] Step S1006, use the medicine box detection model for medicine box recognition and correction.
[0129] Step S1008, cross-modal feature extraction.
[0130] In the case of a picture input, use the pre-trained ChineseClip model (medicine box detection model) to extract the picture feature vector. In the case of text input, convert the medicine brand and name into a feature vector through the ChineseClip text encoder. Then, generate a unified feature representation for the feature vectors for subsequent retrieval.
[0131] Step S1010, cross-modal retrieval and matching.
[0132] Use a vector database (such as Milvus or Faiss) for efficient similarity retrieval. Perform Top-K recall based on cosine similarity calculation to obtain the drug information most similar to the query and its corresponding instruction manual ID.
[0133] After cross-modal retrieval matching, the process can be ended to directly output the retrieval result, or it can jump to step S1016.
[0134] In step S1012, determine whether the intent recognition is a question.
[0135] If it is a question, execute step S1014; otherwise, execute step S1008.
[0136] In step S1014, perform QA feature extraction.
[0137] In step S1016, perform block retrieval of the electronic instruction manual.
[0138] Convert the user's specific consultation question into a query vector through the BGE model. Retrieve relevant content fragments in the instruction manual vector database. Screen out the most relevant content paragraphs through semantic similarity ranking.
[0139] In step S1018, integrate context information.
[0140] Integrate basic drug information (such as name, specification, usage, etc.). Integrate relevant instruction manual content (such as dosage and administration, contraindications, etc.). Construct a structured prompt template to provide complete context information for the large language model.
[0141] Input the integrated context information into the large language model (LLM). Guide the model to generate accurate and professional answers through optimized prompt engineering. Verify the reliability and integrity of the generated answers.
[0142] Finally, format the generated answer to make it more understandable. Necessary warning information and disclaimers can also be added. Finally, the final consultation result can be presented in an age-friendly manner (such as enlarged font, voice broadcast, etc.).
[0143] Specifically, the prompt words for the large model to generate medication results are as follows:
[0144]
[0145] In the medication consultation process, the present invention adopts a large language model architecture based on retrieval enhancement. Through multi-level context information fusion and intelligent dialogue generation, it provides professional and accurate medication guidance for users. The system uses the drug instruction manual vector library as the knowledge basis and combines optimized prompt engineering to achieve intelligent medication consultation services.
[0146] In terms of prompt design, the present invention constructs a multi-level prompt template system. First, a clear role positioning is set, defining the model as a "professional drug consultant". This role positioning ensures the professionalism and rigor of the model when answering questions. Secondly, at the instruction level, the answering strategy is clarified: when the answer to the question can be found in the context, the model will generate an accurate answer based on the context; if the question goes beyond the context scope, the model needs to answer politely and guide the user to ask answerable questions or suggest seeking medical advice outside the scope of assisted medical care. This strategy design not only ensures the reliability of the answer but also reflects the service awareness of the system.
[0147] In the context construction process, the system adopts a multi-level retrieval strategy. First, it obtains the most relevant drug information fragments to the user's question through vector similarity retrieval, and then sorts these fragments by importance. The sorting considers multiple dimensions: the semantic relevance to the question, the importance of the content (such as the importance of contraindications is higher than general instructions), the timeliness of the information, etc. The selected context content is organized in the logical order of "basic information - usage and dosage - precautions - professional information" to ensure that the large model can obtain complete and structured reference information.
[0148] To improve the answer quality, the present invention designs a multiple constraint mechanism at the model output end. First is the reliability constraint, which requires the model to generate answers based on the given context to avoid generating false information; secondly is the integrity constraint, ensuring that the answer covers all aspects that the user cares about; finally is the security constraint, in issues related to medication safety, the model will emphasize the importance of referring to medical advice. In addition, the system also implements an intelligent answer expansion mechanism. Besides directly answering the user's question, it will also supplement relevant medication precautions or possible risk warnings. The specific interface is as Figures 11 to 13 shown.
[0149] Through the above methods, the present invention can provide more accurate, professional and safe medication consultation services, ensuring that users receive correct guidance and suggestions during the medication process, thereby effectively improving the safety and accuracy of medication.
[0150] The present application also provides an intelligent medication consultation device based on a large model, including an acquisition module configured to acquire user input information; a text processing module configured to, when the user input information is text, perform feature extraction on the text to obtain a first feature representation, perform block retrieval on the electronic instruction manual based on the first feature representation, and perform context integration on the block retrieval results to generate medication consultation information for responding to the user input information; and an image processing module configured to, when the user input information is an image, perform medicine box detection on the image to identify a target medicine box, perform cross-modal feature extraction based on the target medicine box to obtain a second feature representation, and perform cross-modal retrieval based on the second feature representation to generate medication consultation information for responding to the user input information.
[0151] The present application also provides an intelligent medication consultation device based on a large model, including an acquisition module configured to acquire user input information; a text processing module configured to, when the user input information is text, perform feature extraction on the text to obtain a first feature representation, perform block retrieval on the electronic instruction manual based on the first feature representation, and perform context integration on the block retrieval results to generate medication consultation information for responding to the user input information; and an image processing module configured to, when the user input information is an image, perform medicine box detection on the image to identify a target medicine box, perform cross-modal feature extraction based on the target medicine box to obtain a second feature representation, and perform cross-modal retrieval based on the second feature representation to generate medication consultation information for responding to the user input information.
[0152] The present application also provides a cross-modal retrieval and matching device for drug information based on AI, including: an acquisition module configured to acquire a target medicine box image to be recognized; an extraction module configured to perform feature extraction on the target medicine box image by using a multi-modal drug matching model to obtain a feature representation; and a matching module configured to perform cross-modal matching based on the feature representation by using the multi-modal drug matching model to match drug information related to the target medicine box; wherein, the multi-modal drug matching model is obtained by: performing domain adaptation training on a pre-trained multi-modal network model based on a graphic-text pair training set in the drug field to obtain the multi-modal drug matching model.
[0153] The present application also provides an AI-based rotating medicine box recognition and correction device, including: an acquisition module configured to acquire a target medicine box picture to be recognized; a recognition module configured to recognize the target medicine box picture to be recognized by using a medicine box detection model and correct the rotation angle of the recognized target medicine box to the standard horizontal direction; wherein, the medicine box detection model is obtained by: training a neural network model based on a medicine box data set with rotation angle annotations to obtain the medicine box detection model.
[0154] It should be noted that: for the device provided in the above embodiment, only the division of the above functional modules is used for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the corresponding method embodiment belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be elaborated here.
[0155] Figure 14 The structural schematic diagram of an electronic device suitable for implementing the embodiments of the present disclosure is shown. It should be noted that Figure 14 The shown electronic device is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.
[0156] As Figure 14 shown, the electronic device includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage part 1008 into the random access memory (RAM) 1003. In the RAM 1003, various programs and data required for system operation are also stored. The CPU 1001, ROM 1002, and RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.
[0157] The above are only the preferred embodiments of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. An AI-based method for identifying and calibrating a rotating medicine box, characterized in that, Including: Obtain a target medicine box picture to be recognized; Use a medicine box detection model to recognize the target medicine box picture to be recognized, and correct the rotation angle of the recognized target medicine box to the standard horizontal direction; Among them, the medicine box detection model is obtained through the following: Based on a medicine box data set with rotation angle annotations, train a neural network model to obtain the medicine box detection model.
2. The method according to claim 1, wherein Based on a medicine box data set with rotation angle annotations, training a neural network model to obtain the medicine box detection model includes: Obtain medicine box sample data, where the medicine box sample data includes medicine box sample pictures; Perform rotation angle annotation on each medicine box sample picture in the medicine box sample data, construct the medicine box data set with rotation angle annotations, and perform data augmentation on the medicine box data set with rotation angle annotations, where the data augmentation includes at least one of the following: random angle rotation, scale transformation, and brightness contrast adjustment; Based on the medicine box data set after data augmentation, train the neural network model to obtain the medicine box detection model.
3. The method according to claim 2, wherein Performing angle annotation on each medicine box sample picture in the medicine box sample data, constructing a medicine box data set with rotation angle annotations, and performing data augmentation on the medicine box data set with rotation angle annotations includes: Annotate the four vertex coordinates of the minimum bounding rectangle and the rotation angle relative to the horizontal axis of each medicine box sample picture on each medicine box sample picture to obtain the medicine box data set with rotation angle annotations; Perform data augmentation on the medicine box data set through at least one of the following: perform the random angle rotation within the range of ±90°, perform the scale transformation between 0.8 and 1.2 times, and perform the brightness contrast adjustment within the range of ±30°.
4. The method according to claim 2, wherein Based on the medicine box data set after data augmentation, training the neural network model includes: Use an oriented bounding box OBB to remove redundant rotation detection boxes in each medicine box sample picture in the medicine box data set after data augmentation to obtain a target detection box; Perform correction on the target detection box using affine transformation, and adjust the target detection box to the 0° standard horizontal direction; Compare the adjustment angle of the target detection box with the annotation angle of each medicine box sample picture, and adjust the loss function of the neural network model based on the comparison result to train the neural network model.
5. The method according to claim 4, wherein Using an oriented bounding box OBB to remove redundant rotation detection boxes in each medicine box sample picture in the medicine box data set after data augmentation includes: For each medicine box sample picture, when calculating the intersection over union between the rotation detection boxes using the OBB, adopt the calculation method of the rotated intersection over union, and obtain the final overlap degree by dividing the intersection area of the polygons by the union area; Use the final overlap degree to filter out redundant rotation detection boxes.
6. The method according to claim 1, characterized in that, After performing angle correction on the recognized target medicine box, the method further includes: Use the multi-modal medicine matching model of the pre-trained vision-text alignment model to extract features from the target medicine box after angle correction to obtain a feature representation; Based on the feature representation, use the multi-modal drug matching model for cross-modal matching to match information related to the target medicine box; Among them, the multi-modal drug matching model is obtained as follows: based on the text-image pair training set, perform vertical domain pre-training on the multi-modal network model to obtain the multi-modal drug matching model.
7. An AI-based rotating medicine box recognition and calibration device, characterized in that, Including: An acquisition module, configured to acquire a target medicine box image to be recognized; An identification module, configured to use the medicine box detection model to identify the target medicine box image to be recognized, and correct the rotation angle of the recognized target medicine box to the standard horizontal direction; Among them, the medicine box detection model is obtained as follows: based on the medicine box data set with rotation angle annotation, train the neural network model to obtain the medicine box detection model.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of claims 1 to 6.
9. A computer device, characterized in that, Including: A memory and a processor, The memory stores a computer program; The processor is used to execute the computer program stored in the memory, and when the computer program runs, it causes the processor to execute the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.