Cross-modal retrieval matching method and device for drug information

Through the combination of multimodal drug matching model and large language model, the problems of insufficient text and image matching accuracy and low voice interaction recognition in the drug information retrieval system are solved, and efficient and intuitive drug information query and drug use consultation services are achieved.

CN120336606AInactive Publication Date: 2025-07-18SHENZHEN BODE RUIJIE HEALTH TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510368086.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing drug information retrieval system has insufficient text and image matching accuracy, low voice interaction recognition accuracy, and lacks intuitive feedback, making it difficult to meet the drug needs of the elderly.

Method used

The multimodal drug matching model is adopted, and the pre-trained multimodal network model is trained through graphics and text on the training set, and the feature extraction and cross-modal matching of the training set is performed in combination with graphics and text on the drug field. The large language model is used for intelligent Q&A and drug consultation.

Benefits of technology

It improves the accuracy and comprehensiveness of drug information retrieval, provides efficient and intuitive drug use guidance, especially suitable for the elderly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336606A_ABST
    Figure CN120336606A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal retrieval matching method and device for medicine information, and the method comprises the steps: obtaining a to-be-recognized target medicine box picture; performing feature extraction on the target medicine box picture by using a multi-modal medicine matching model to obtain feature representation; based on the feature representation, performing cross-modal matching by using the multi-modal drug matching model to match drug information related to the target drug box; wherein the multi-modal drug matching model is obtained by the following steps: based on an image-text pair training set of a drug field, performing field adaptation training on a pre-trained multi-modal network model to obtain the multi-modal drug matching model. The technical problem of inaccurate drug retrieval in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the medical field, and in particular, to a cross-modal retrieval and matching method and device for drug information Background Art

[0002] Currently, the drug information retrieval and matching methods on the market still have significant deficiencies. First, the matching accuracy between text and image is insufficient. Many drug information retrieval systems rely on the text and image information on drug packages. However, due to the complex design of drug packages and small fonts, it is difficult for the system to accurately identify and extract key information. Second, the voice interaction function is not intelligent enough. Although some systems support voice retrieval functions, their recognition accuracy and interaction experience are poor. Especially when the voices of the elderly are blurred or have accents, the system often cannot accurately understand the user's needs. In addition, voice retrieval results are usually presented in text form, lacking intuitive image or voice feedback and unable to meet the needs of the elderly for simple and direct information

[0003] For the above problems, no effective solutions have been proposed yet Summary of the Invention

[0004] Embodiments of the present invention provide a cross-modal retrieval and matching method and device for drug information to at least solve the technical problem of inaccurate drug retrieval in the prior art

[0005] According to one aspect of the embodiments of the present invention, a cross-modal retrieval and matching method for drug information based on AI is provided, including: obtaining a target medicine box picture to be recognized; using a multi-modal drug matching model to extract features from the target medicine box picture to obtain a feature representation; based on the feature representation, using the multi-modal drug matching model to perform cross-modal matching to match drug information related to the target medicine box; wherein, the multi-modal drug matching model is obtained by: performing vertical domain pre-training on a multi-modal network model based on a text-image pair training set to obtain the multi-modal drug matching model; performing domain adaptation training on the pre-trained multi-modal network model based on a text-image pair training set in the drug field to obtain the multi-modal drug matching model

[0006] According to another aspect of the embodiments of the present invention, there is also provided a cross-modal retrieval and matching device for drug information based on AI, including an acquisition module configured to acquire a target medicine box picture to be recognized; an extraction module configured to extract features from the target medicine box picture by using a multi-modal drug matching model to obtain a feature representation; a matching module configured to perform cross-modal matching based on the feature representation by using the multi-modal drug matching model to match drug information related to the target medicine box; wherein, the multi-modal drug matching model is obtained by: performing domain adaptation training on a pre-trained multi-modal network model based on a text-image pair training set in the drug field to obtain the multi-modal drug matching model. Wherein, the multi-modal drug matching model is obtained by: performing vertical domain pre-training on a multi-modal network model based on a text-image pair training set to obtain the multi-modal drug matching model.

[0007] In the embodiments of the present invention, a target medicine box picture to be recognized is acquired; features are extracted from the target medicine box picture by using a multi-modal drug matching model to obtain a feature representation; cross-modal matching is performed based on the feature representation by using the multi-modal drug matching model to match drug information related to the target medicine box; wherein, the multi-modal drug matching model is obtained by: performing vertical domain pre-training on a multi-modal network model based on a text-image pair training set to obtain the multi-modal drug matching model. Performing domain adaptation training on a pre-trained multi-modal network model based on a text-image pair training set in the drug field to obtain the multi-modal drug matching model. This application solves the technical problem of inaccurate drug retrieval in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0009] Figure 1 is a flowchart of an intelligent medication consultation method based on a large model according to an embodiment of the present invention;

[0010] Figure 2 is a flowchart of a method for identifying and correcting a medicine box by using AI according to an embodiment of the present invention;

[0011] Figure 3 is a flowchart of a cross-modal retrieval and matching method for drug information based on AI according to an embodiment of the present invention;

[0012] Figure 4 It is a flowchart of an age-friendly method for intelligent medicine box recognition and medication consultation of a multimodal large model according to an embodiment of the present invention;

[0013] Figure 5 It is a process diagram of the training of a medicine box detection model and a multimodal drug matching model according to an embodiment of the present invention;

[0014] Figure 6 It is an interface diagram of medication consultation and an electronic medicine instruction according to an embodiment of the present invention;

[0015] Figure 7 It is an interface diagram of multimodal retrieval of a drug vector library according to an embodiment of the present invention;

[0016] Figure 8 It is a structure diagram of the ChineseClip model according to an embodiment of the present invention;

[0017] Figure 9 It is a schematic diagram of a corrected picture according to an embodiment of the present invention;

[0018] Figure 10 It is a flowchart of user interaction according to an embodiment of the present invention;

[0019] Figure 11 It is an interface diagram of medicine box recognition and age-friendly adaptation according to an embodiment of the present invention;

[0020] Figure 12 It is an interface diagram of an electronic medicine instruction and intelligent medication consultation according to an embodiment of the present invention;

[0021] Figure 13 It is an interface diagram of medication reminder and expiration reminder setting according to an embodiment of the present invention;

[0022] Figure 14 It shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. Detailed implementation manners

[0023] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0024] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0025] According to an embodiment of the present invention, a method embodiment of an intelligent medication consultation method based on a large model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from that here.

[0026] Figure 1 is a flowchart of an intelligent medication consultation method based on a large model according to an embodiment of the present invention, as Figure 1 shown, the method includes the following steps:

[0027] Step S102, obtain user input information.

[0028] Step S104, in the case where the user input information is text, use the pre-trained language model BGE to extract features from the text to obtain a first feature representation, perform block retrieval on the electronic drug instructions based on the first feature representation, and perform context integration on the block retrieval results to generate medication consultation information for responding to the user input information.

[0029] In the case where the text is recognized as a question, perform QA feature extraction on the text to obtain the first feature representation, perform block retrieval on the electronic drug instructions based on the first feature representation, and perform context integration on the retrieval results to generate medication consultation information for responding to the user input information. In this embodiment, the first feature representation of the question text is obtained through QA feature extraction, and block retrieval is performed in the electronic drug instructions based on this feature, effectively improving the pertinence and accuracy of the retrieval. At the same time, the retrieval results are optimized through context integration, making the generated medication consultation information more complete and coherent, and avoiding understanding deviations caused by information fragmentation.

[0030] In the case where the text is intended to be recognized as non-question, cross-modal feature extraction is performed on the text to obtain the first feature representation. Based on the first feature representation, cross-modal retrieval is carried out to match the associated medicine box image and text description from the medicine graphic database, obtaining the cross-modal retrieval result. Based on the cross-modal retrieval result, block retrieval is performed on the electronic medicine instructions, and the retrieval results are contextually integrated to generate medicine usage consultation information for responding to the user input information. For example, based on the first feature representation, the cosine similarity algorithm is used to retrieve relevant content segments in the instruction vector library generated based on the electronic medicine instructions, and through semantic similarity ranking, the most relevant content paragraphs are screened out from the relevant content segments; the most relevant content paragraphs are integrated as context information, and based on the context information, a large language model is used to generate medicine usage consultation information for responding to the user input information. This embodiment combines cross-modal feature extraction and retrieval technologies. When the user inputs non-question text, it can intelligently match the medicine box image and text description, and further combine the block retrieval of the electronic medicine instructions to ensure the comprehensiveness and accuracy of the information. Through the cosine similarity algorithm and semantic similarity ranking, the screening of relevant content segments is optimized, improving the precision of the retrieval results.

[0031] In some embodiments, the instruction vector library is obtained as follows: all relevant fields of the electronic medicine instructions are concatenated into a complete document in logical order, and the concatenated document is preprocessed for standardization; for the preprocessed document, it is segmented with a preset number of characters as the basic unit, and an overlapping area with a preset overlap degree is set between adjacent segments, where the preset number of characters is 1024 characters, and the preset overlap degree is 10% of the number of characters in adjacent segments. Based on the segmented document, the instruction vector library is constructed. This embodiment improves the integrity and retrieval accuracy of vector library construction through standardizing and preprocessing the electronic medicine instructions and adopting fixed-length segmentation and overlapping area setting. The segmentation method ensures information coherence, avoids content breakage, and improves the accuracy and semantic matching effect of subsequent retrieval.

[0032] Step S106, in the case where the user input information is a picture, perform medicine box detection on the picture, identify the target medicine box, and based on the target medicine box, perform cross-modal feature extraction to obtain the second feature representation. Based on the second feature representation, perform cross-modal retrieval to generate medicine usage consultation information for responding to the user input information.

[0033] 1) Perform medicine box detection on the picture, identify the target medicine box, and correct the target medicine box.

[0034] For example, identify the main body position of the target medicine box in the picture; when the rotation angle of the main body position deviates from the standard horizontal direction, use the affine transformation algorithm to correct the rotation angle of the main body position to the standard horizontal direction. In the embodiment of the present application, the rotation angle of the target medicine box is corrected by the affine transformation algorithm to align it to the standard horizontal direction, improving the accuracy of image normalization processing, ensuring the stability and accuracy of subsequent image recognition and information extraction, and being applicable to the image processing of medicine boxes photographed at different angles.

[0035] In some embodiments, AI can be used to identify and correct the medicine box. For example, as Figure 2 shown, the method for identifying and correcting a rotated medicine box based on AI includes the following steps:

[0036] Step S202, obtain a picture of the target medicine box to be identified.

[0037] Step S204, use the medicine box detection model to identify the picture of the target medicine box to be identified, and correct the rotation angle of the identified target medicine box to the standard horizontal direction.

[0038] In some embodiments, the medicine box detection model is obtained as follows: based on a medicine box data set with rotation angle annotations, train a neural network model to obtain the medicine box detection model. The specific training process can be as follows:

[0039] First, obtain medicine box sample data, where the medicine box sample data includes medicine box sample pictures.

[0040] Next, perform rotation angle annotation on each medicine box sample picture in the medicine box sample data to construct the medicine box data set with rotation angle annotations, and perform data augmentation on the medicine box data set with rotation angle annotations, where the data augmentation includes at least one of the following: random angle rotation, scale transformation, and brightness contrast adjustment. For example, annotate the four vertex coordinates of the minimum circumscribed rectangle and the rotation angle relative to the horizontal axis of each medicine box sample picture on each medicine box sample picture to obtain the medicine box data set with rotation angle annotations; perform data augmentation on the medicine box data set through at least one of the following: perform the random angle rotation within ±90°, perform the scale transformation between 0.8 and 1.2 times, and perform the brightness contrast adjustment within ±30°. In this embodiment, by performing rotation angle annotation on the medicine box sample pictures to construct a medicine box data set with rotation angle information, the recognition ability of the model for medicine box images at different angles is improved. The data augmentation strategy further enhances the generalization performance and robustness of the model, enabling it to stably recognize medicine box information under different shooting conditions, improving the adaptability and accuracy of image processing, and being applicable to automatic recognition and information extraction of medicine packaging in complex scenarios.

[0041] Finally, based on the enhanced medicine box dataset, the neural network model is trained to obtain the medicine box detection model. Specifically, using the oriented bounding box (OBB), redundant rotated detection boxes in each medicine box sample image in the enhanced medicine box dataset are removed to obtain target detection boxes. For example, for each medicine box sample image, when calculating the intersection over union (IoU) between the rotated detection boxes using the OBB, the calculation method of the rotated IoU is adopted, and the final overlap degree is obtained by dividing the intersection area of the polygons by the union area; the final overlap degree is used to filter out redundant rotated detection boxes. Then, affine transformation is used to correct the target detection boxes, and the target detection boxes are adjusted to the 0° standard horizontal direction; the adjusted angle of the target detection boxes is compared with the labeled angle of each medicine box sample image, and based on the comparison result, the loss function of the neural network model is adjusted to train the neural network model. In this application, the neural network model is trained using the enhanced medicine box dataset, and the oriented bounding box (OBB) is used to optimize the rotated detection boxes, remove redundant boxes, and improve the accuracy of target detection. The rotated IoU is used to calculate the overlap degree to ensure that the selected detection boxes are optimal. Affine transformation is used to correct the detection boxes to the standard horizontal direction, and the loss function is adjusted in combination with the labeled angle to optimize the model training effect. This method improves the stability and robustness of the medicine box detection model, making it more suitable for the medicine packaging recognition task under different shooting angles and complex environments.

[0042] 2) Perform cross-modal feature retrieval.

[0043] Based on the second feature representation, the multi-modal medicine matching model is used to perform cross-modal matching on the medicine text-image vector database to obtain medicine information; wherein, the medicine text-image vector database is constructed by at least one of the following: combining the brand and the medicine name as the main text description, and combining the key features of the medicine to form complete text description information; performing at least one of the following data augmentation processes on the medicine sample images: random angle rotation within the range of ±30°, brightness and contrast adjustment, and horizontal and vertical flipping; performing at least one of the following data augmentation processes on the training sample words: synonym replacement, and interchange of brand and generic name. In this application, the multi-modal medicine matching model is used to perform cross-modal matching on the medicine text-image vector database, improving the accuracy and comprehensiveness of medicine information retrieval. Combining the brand and the medicine name and the key features to construct a complete text description makes the matching result more targeted. At the same time, through various data augmentation processes on the medicine sample images and the training text data, the robustness and generalization ability of the model are improved, ensuring accurate matching of medicine information under different lighting, angles, and text variant conditions, and enhancing the adaptability and reliability of cross-modal retrieval.

[0044] In some embodiments, AI can be utilized for cross-modal feature retrieval. For example, as Figure 3 shown, the cross-modal retrieval and matching method for drug information based on AI may include the following steps:

[0045] Step S302: Obtain a target medicine box picture to be recognized.

[0046] Step S304: Use a multi-modal drug matching model to extract features from the target medicine box picture to obtain a feature representation.

[0047] Step S306: Based on the feature representation, use the multi-modal drug matching model for cross-modal matching to match drug information related to the target medicine box.

[0048] Among them, the multi-modal drug matching model is obtained as follows: Based on a training set of text-image pairs in the drug field, perform domain adaptation training on a pre-trained multi-modal network model to obtain the multi-modal drug matching model. The specific training process is as follows:

[0049] First, obtain the text-image pair training set. Obtain drug sample data, where the drug sample data includes drug sample pictures and training text data; perform at least one of the following enhancement processes on the drug sample pictures: perspective transformation, illumination adjustment, and local magnification, and expand the training text data by means of synonym replacement and interchange between brand names and generic names; use the enhanced drug sample pictures and the expanded training text data as the text-image pair training set. This embodiment effectively improves the training effect of the cross-modal model, enabling it to more accurately and stably match and identify drug information in practical applications.

[0050] Next, based on the text-image pair training set, pre-train the multi-modal network model in the field of pharmaceuticals to obtain the multi-modal pharmaceutical matching model. Pre-train the multi-modal network model on the general Chinese text-image dataset; based on the text-image pair training set, pre-train the pre-trained multi-modal network model in the field of pharmaceuticals to obtain the multi-modal pharmaceutical matching model. For example, based on the attention mechanism enhanced by pharmaceutical knowledge, use the model to extract the key visual features in the pharmaceutical sample images in the text-image pair training set through an additional pharmaceutical attribute supervision signal; through the attention mechanism enhanced by pharmaceutical knowledge graph embedding, combined with the pharmaceutical attribute classification loss function, use the multi-modal network model to extract key visual features; when encoding the training text data in the text-image pair training set, fuse the prior information of entity relationships in the pharmaceutical knowledge graph; use the key visual features and the text encoding that has fused the prior information of entity relationships to perform the pre-training in the field of pharmaceuticals. In the embodiments of the present application, by pre-training the multi-modal network model on the general Chinese text-image dataset and further pre-training it in combination with the text-image pair training set in the field of pharmaceuticals, the accuracy of the model in the pharmaceutical matching task is improved. In addition, the attention mechanism enhanced by pharmaceutical knowledge and the pharmaceutical attribute supervision signal are used to optimize the visual feature extraction ability; finally, the prior information of entity relationships in the pharmaceutical knowledge graph is fused to improve the semantic understanding ability of text encoding, thereby enhancing the accuracy and generalization ability of cross-modal matching and realizing efficient pharmaceutical information retrieval and matching.

[0051] Finally, based on the feature representation, use the multi-modal pharmaceutical matching model to perform cross-modal matching to match the pharmaceutical information related to the target medicine box. Specifically, based on the feature representation, use the multi-modal pharmaceutical matching model to perform cross-modal matching on the text-image vector database to obtain the pharmaceutical information. For example, use the multi-modal pharmaceutical matching model to perform similarity retrieval on the text-image vector database and calculate the cosine similarity; sort based on the cosine similarity and select the top K candidate results with the highest similarity as the pharmaceutical information; where the text-image vector database stores the image feature vectors of medicine box pictures and the feature vectors of the corresponding drug names, brands, indications, dosages, and instruction manual texts; the pharmaceutical information includes at least one of the following: the type, brand, indication, and instruction manual of the drug. In this application, the multi-modal pharmaceutical matching model is used to perform cross-modal matching on the text-image vector database to achieve accurate pharmaceutical information retrieval.

[0052] An embodiment of the present invention provides an intelligent medicine box recognition and medication consultation system based on a multimodal large model, aiming to solve problems such as insufficient accuracy and inconvenient information acquisition in the processes of medicine identification, information query, and medication consultation. The system utilizes the pre-training in the medicine field of the multimodal model, cross-modal retrieval, and the intelligent question-answering technology of the large model to provide efficient and accurate medicine information processing capabilities, and provide a comprehensive intelligent solution for users' medication management.

[0053] This system includes a pre-training and cross-modal retrieval module, and a database construction and intelligent medication consultation module.

[0054] The pre-training and cross-modal retrieval module is used for pre-training in the medicine field of the multimodal large model and cross-modal retrieval. First, pre-training in the medicine field is carried out. For example, based on data of Western medicines, traditional Chinese medicines, medical devices, and health products, the ChineseCLIP model is pre-trained for the medicine field. The data covers pictures, names, and brands of medicines. Through constructing picture-text pairs for training, the cross-modal image and text understanding ability of the model in the medicine field is improved, realizing efficient matching between medicine box pictures and medicine information. Then, the construction of a picture-text vector database and cross-modal retrieval are carried out. For example, the system vectorizes the medicine box picture and its corresponding medicine name, brand, and other descriptive texts and stores them in the picture-text vector database. When the user takes a picture of the medicine box or enters the medicine name, the system quickly matches the corresponding medicine information based on cross-modal retrieval technology, including the type, brand, indications, and other details of the medicine, thus realizing accurate and convenient medicine identification and information query.

[0055] The database construction and intelligent medication consultation module is used for the construction of an electronic medicine instruction vector database and intelligent medication consultation. First, the construction of the electronic medicine instruction vector database is carried out. The electronic medicine instructions are stored in blocks by paragraph, covering complete information such as indications, usage and dosage, contraindications, adverse reactions, and drug interactions of the medicine. The semantic vectors of each paragraph are generated through the BGE vector model and stored in the instruction vector database. Then, intelligent question-answering based on the knowledge base and vector matching. A medicine knowledge base is constructed, combining the structured information in the instructions with other authoritative data sources (such as pharmacopoeias, clinical guidelines) to form a comprehensive medicine knowledge base. When the user asks questions (such as "What are the indications of this medicine?", "Is it applicable to children?", "What are the adverse reactions?", etc.), the system first matches relevant information in the instruction vector database through the user query, and integrates it with the context semantic analysis technology and the knowledge base. The LLM large model generates accurate answers in natural language. The intelligent question-answering function supports dynamic multi-round interactions and can generate coherent answers based on historical questions, providing a complete medicine consultation service for users. Cross-modal retrieval and intelligent question-answering.

[0056] The user initiates a retrieval request by taking a picture of the medicine box, entering the medicine name, or making a voice query. Based on the matching results in the graphic and text vector database, the system retrieves all the information of the corresponding medicine.

[0057] When the user queries specific issues such as the usage method, indications, and contraindications of a medicine, the system quickly matches relevant content through the electronic instruction manual vector database and combines the context semantic analysis ability of the multi-modal large model to generate accurate and natural intelligent answers to ensure comprehensive and accurate information.

[0058] The present invention applies the image and text understanding capabilities of the multi-modal large model to the scenarios of medicine identification and medication consultation, and combines the design for the elderly, significantly improving the efficiency of medicine information query and the convenience of medication management, especially having wide application value among the elderly user group.

[0059] The embodiment of this application also provides an elderly-friendly method for intelligent medicine box identification and medication consultation using a multi-modal large model. This method is applied to the above-mentioned intelligent medicine box identification and medication consultation system based on a multi-modal large model. This method is as Figure 4 shown and includes the following steps:

[0060] S402, data preprocessing and model pre-training.

[0061] The overall process of data preprocessing and model pre-training is as Figure 5 shown.

[0062] The construction of the medicine vector library will be described in detail below.

[0063] The medicine vector library constructed by the present invention is a large-scale multi-source heterogeneous data set, containing structured data records of multiple health products, medical devices, Western medicines, and Chinese patent medicines, etc. Each record contains multiple data fields, covering basic information such as product name, brand, specification, scope of application, storage conditions, etc., as well as regulatory information such as approval number, validity period, and manufacturer. At the same time, the record also contains clinical medication information such as main ingredients, precautions, usage and dosage, and contraindication symptoms, and provides medication guidance for special populations, such as children's medication, medication for pregnant and lactating women, etc. In addition, for different types of medicines, their specific professional information is also stored, such as the pharmacology and toxicology, pharmacokinetics of Western medicines, the structural composition and registration certificate number of medical devices, etc.

[0064] In the construction of multi-modal training samples, the present invention adopts an innovative strategy for aligning images and texts. First, the brand and drug name are combined as the main text description, and key features such as specifications and characteristics are incorporated to form complete text description information. To improve the generalization ability and robustness of the model, the present invention performs all-round data augmentation on drug images, covering enhancement techniques such as random angle rotation within the range of 0° to 360°, brightness and contrast adjustment, horizontal and vertical flipping, etc. At the same time, multi-dimensional enhancement processing such as synonym replacement and interchange of brand and generic names is also carried out at the text level, effectively expanding the diversity of training samples.

[0065] For the processing of drug instructions, the present invention adopts a fine-grained text chunking strategy. As Figure 6 shown, first, all relevant fields are concatenated into a complete electronic instruction document in logical order. After standardized preprocessing, chunks are made with 512 characters as the basic unit. To ensure the coherence and integrity of semantics, a 10%-15% overlapping area is set between adjacent chunks. This strategy not only ensures the independence of the content after chunking but also maintains the semantic association of the context. During the chunking process, special attention is paid to avoiding truncation of important information to ensure that each chunk has a complete semantic expression.

[0066] As Figure 7 shown, in the vector feature extraction and storage section, the present invention generates high-dimensional feature vectors for drug images, text descriptions, and instruction chunks respectively. The dimension of the image feature vector extracted by the pre-trained model is 768, the dimension of the text description vector is 768, and the dimension of the instruction chunk vector is 1024. These feature vectors are stored in the Milvus vector database, establishing an efficient multi-modal index structure to support fast similarity retrieval and approximate nearest neighbor search. At the same time, a perfect data quality control mechanism is implemented, including functions such as data integrity verification, outlier detection, and incremental update, ensuring the reliability and maintainability of the vector library.

[0067] The multi-modal drug matching model will be described in detail below.

[0068] The present invention adopts an improved ChineseClip multi-modal model as the core architecture, which is specifically optimized for the Chinese drug field based on the original CLIP model. As Figure 8 shown, the overall model consists of two main parts: an image encoder and a text encoder, and realizes the unified representation space mapping of image and text features through the contrastive learning method.

[0069] Among them, the image encoder adopts the Vision Transformer (ViT) architecture, which contains 12 layers of Transformer blocks, each layer has 12 attention heads, and the input resolution is uniformly adjusted to 224×224 pixels; the text encoder is based on the RoBERTa architecture, consisting of 12 layers of bidirectional Transformer, the vocabulary expands the professional vocabulary in the pharmaceutical field, and the maximum sequence length is set to 128 tokens.

[0070] In the model pre-training stage, we first constructed a training set of image and text pairs using 142,307 pieces of drug data. Considering the particularity of drug images, we improved the standard data enhancement strategy: on the image side, we introduced drug-specific enhancement methods such as perspective transformation, lighting adjustment, and local magnification; on the text side, we expanded the training samples by synonym replacement and brand generic name conversion. The training process adopted a contrastive learning strategy, and used the InfoNCE loss function to optimize the model parameters. The batch size was set to 1024, and the learning rate adopted a cosine annealing strategy, which gradually decayed from 1e-4.

[0071] In order to improve the recognition accuracy of the model in the pharmaceutical field, the present invention designs a special training framework. First, pre-training is performed on a general Chinese image and text dataset, and then domain adaptability fine-tuning is performed using the collected pharmaceutical dataset. In the fine-tuning stage, an attention mechanism for pharmaceutical knowledge enhancement is introduced, and the model is guided to focus on key visual features in pharmaceutical images, such as packaging features, dosage form features, and identification features, through additional pharmaceutical attribute supervision signals. At the same time, the prior information of the pharmaceutical knowledge graph is incorporated into the text encoding, which enhances the model's ability to understand pharmaceutical terminology.

[0072] The training objective design of the model adopts a multi-task learning framework. In addition to the basic image-text comparison loss, it also includes drug classification loss and feature consistency loss. Among them, the classification loss helps the model learn the category information of the drug, and the feature consistency loss ensures that the feature representation of the same drug from different perspectives is consistent. Through this multi-objective optimization strategy, the model can take into account the two key tasks of image-text alignment and drug feature extraction at the same time.

[0073] The optimized multimodal model demonstrates excellent feature extraction capabilities. In practical applications, the model can accurately understand the visual information of drug images and establish precise semantic associations with text descriptions, providing a reliable feature representation basis for subsequent smart medicine box recognition and medication consultation. At the same time, through the knowledge-enhanced feature extraction mechanism, the model has a strong ability to distinguish subtle differences in drugs and can effectively handle the identification of similar drugs.

[0074] The pill box detection model will be described below.

[0075] In view of the actual problem that there are diverse shooting angles for the medicine box pictures uploaded by users, the present invention innovatively proposes a rotating medicine box detection scheme based on OBB (Oriented Bounding Box). Although the samples have been augmented by angle rotation in the aforementioned multi-modal training stage, due to the randomness and complexity of the medicine box placement angles in the actual application scenarios, the recognition accuracy may still be insufficient. Therefore, the present invention deeply improves the traditional YOLO model, introduces a rotating object detection mechanism, and realizes the precise positioning and angle correction of medicine boxes at any angle.

[0076] To improve the training effect of the model, the present invention specifically constructs a medicine box data set with annotated angles. Through a semi-automatic annotation tool, the four vertex coordinates of the minimum bounding rectangle and the main direction angle are annotated for each medicine box. At the same time, the training samples are augmented through data augmentation techniques, including random angle rotation, scale transformation, brightness and contrast adjustment, etc., to improve the adaptability of the model to various shooting conditions.

[0077] In the post-processing stage, in order to effectively filter redundant rotating detection boxes, the present invention designs a rotating non-maximum suppression algorithm (R-NMS). Different from the traditional non-maximum suppression (NMS) which is only applicable to horizontal rectangular boxes, R-NMS is optimized according to the characteristics of rotating objects. When calculating the IoU between detection boxes, the calculation method of rotating IoU is adopted, and the final overlap degree is obtained by dividing the intersection area of polygons by the union area. The traditional IoU calculation only considers the intersection and union areas of horizontal rectangles, while the present invention calculates the overlap degree between rotating detection boxes, that is, the rotated intersection over union (Rotated IoU, R-IoU). Specifically, the present invention adopts the following steps:

[0078] 1) Parametric representation of the rotating detection box.

[0079] Each detection box is represented by a five-element (x c ,y c ,w,h,θ) group, where {(x} c ,y c ) is the center coordinate of the box, w and h are the width and height of the box respectively, and θ is the main direction angle of the box (in radians, with a range of [-π / 2,π / 2]).

[0080] 2) Calculate the rotating IoU.

[0081] For two rotating detection boxes B i and B j , the formula for calculating R-IoU is:

[0082]

[0083] 3) Perform suppression.

[0084] Sort all the detected bounding boxes in descending order of confidence scores. Traverse the list of detected bounding boxes. For the currently highest-score box B max , calculate its R-IoU with the remaining boxes. If R-IoU > τ (τ is a preset threshold, e.g., 0.3), then suppress the box with a lower score. Repeat the above steps until all detected bounding boxes are processed.

[0085] Through the above method, R-NMS can effectively eliminate redundant rotated detection bounding boxes, while retaining key targets and improving the accuracy of the detection results. Compared with traditional NMS, R-NMS shows higher adaptability when dealing with rotated targets such as medicine boxes, especially in dense scenes.

[0086] In some other embodiments, redundant rotated detection bounding boxes can also be filtered based on the following improved Rotated Non-Maximum Suppression (R-NMS):

[0087] 1) Preprocessing.

[0088] Input the original set of rotated detection bounding boxes {B i = (x c , y c , w, h, θ, s)}. Generate a direction vector based on the angle θ: v x = cosθ, v y = sinθ, and output the standardized parameter set {B i ′ = (x c , y c , w, h, v x , v y , s)}.

[0089] Compensate the confidence based on the visible area:

[0090]

[0091] The compensated confidence queue Q = {B i ′ | sorted in descending order of s′}.

[0092] 2) Perform hierarchical screening.

[0093] Based on the currently highest-confidence box B max and the candidate box B j , perform a fast collision detection. The inputs are B max , B j . Apply the Separating Axis Theorem (SAT) to detect whether the projections are separated, and output a boolean value {True / False} indicating whether to continue the calculation.

[0094] 3) Calculate the multi-precision R-IoU.

[0095] Based on center distance Select the calculation mode when d>2×max(w max ,h max ), based on the Monte Carlo method The sampling points estimate the overlapping area to obtain an approximate R-IoU value. In some embodiments, the Greiner-Hormann algorithm can also be used to calculate the exact intersecting polygon vertices to obtain an accurate R-IoU value.

[0096] 4) Dynamic inhibition.

[0097] Based on the angle difference Δθ=|θ max -θ j | Adjust the threshold and calculate the angle-sensitive threshold based on the basic threshold τ0 and Δθ:

[0098] τ=τ0+0.1τ0(1-cos△θ)

[0099] Based on local density Perform secondary adjustments, and calculate the final threshold based on the current threshold τ, density ρ, the number of neighboring boxes as the number of points in a circle with a radius of r, and ρ0 as the reference density:

[0100] τ final =τ×(1-0.2tanh(ρ / ρ0))

[0101] Based on R-IoU>τ final Determine whether to suppress and output the updated detection box queue Q′. Matrix parallel computing based on GPU. Input the rotation box parameter set and convert each box into a homogeneous coordinate matrix, where M i represents the homogeneous coordinate matrix of the i-th rotation frame, θ i represents the rotation angle of the i-th rotation frame, (x i ,y i ) represents the center coordinates of the i-th rotation frame. The calculation formula of Mi is as follows:

[0102]

[0103] Use CUDA thread blocks to parallelize the projections of 16 pairs of boxes and output the R-IoU matrix calculated in batches.

[0104] Input candidate box pair (B max ,B j ), where F max Represents candidate box B max The Fourier descriptor, F j Represents candidate box B j The Fourier descriptor represented by . Calculate the Fourier descriptor distance, that is, the shape similarity score S:

[0105]

[0106] Based on the overlapping region image patches, extract the shallow features of the pre-trained CNN and calculate the cosine similarity to obtain the texture similarity score T. Final suppression condition: (R - IoU > τ final ) ∧ (S < 0.2) ∧ (T > 0.7).

[0107] After detecting the rotation angle of the medicine box, the present invention further corrects the image by using an affine transformation to adjust the medicine box to the standard horizontal direction. During the correction process, considering the preservation of image quality, based on the geometric parameters of the detection frame and combined with the bicubic interpolation algorithm for image resampling, the detailed information of the image is retained to the greatest extent. The image after angle correction will be input into the aforementioned multi-modal model for feature extraction and recognition, thereby significantly improving the overall recognition accuracy of the system. The specific steps are as follows:

[0108] 1) Construct a transformation matrix.

[0109] For the detected center (x c , y c ) of the medicine box and the rotation angle θ, construct an affine transformation matrix M for rotating the image to the horizontal direction. The transformation matrix is defined as:

[0110]

[0111] This matrix rotates the image counterclockwise by θ degrees around the center (xc, yc) to align the main direction of the medicine box with the horizontal axis.

[0112] 2) Map the coordinates.

[0113] For each pixel point (x, y) in the original image, calculate the corrected coordinates (x′, y′) through M

[0114]

[0115] 3) Optimize the bicubic interpolation.

[0116] During the correction process, to maintain the image quality, the bicubic interpolation algorithm is used for image resampling. Where W(t) represents the weight function.

[0117]

[0118] When a = -0.5, it provides the best edge preservation and detail reconstruction ability. t represents the pixel position offset during resampling, usually in the range of [-1, 1], and is used to calculate the interpolation weight.

[0119] The optimized rotating medicine box detection model can effectively handle the angular changes of the medicine box under complex shooting conditions, ensuring that the system can maintain high-precision recognition ability in different scenarios. Through this solution, the accuracy of medicine box detection and angle correction has been significantly improved, further enhancing the practicability and reliability of the intelligent medicine box recognition system.

[0120] The specific optimization effects are as Figure 9 shown. The optimized rotating medicine box detection model can effectively handle the angular changes of the medicine box under complex shooting conditions, ensuring that the system can maintain high-precision recognition ability in different scenarios. Through this solution, the accuracy of medicine box detection and angle correction has been significantly improved, further enhancing the practicability and reliability of the intelligent medicine box recognition system.

[0121] Step S404, user interaction.

[0122] As Figure 10 shown, the process of user interaction includes the following steps:

[0123] Step S1002, obtain user input.

[0124] Obtain user input information. For example, pictures and text. The user can upload a picture of the medicine box, and the system preprocesses the picture, including size adjustment, illumination normalization, and noise elimination. Then, through the rotation angle detection algorithm, the main body position of the medicine box is identified and its tilt angle is corrected to ensure that the image is upright in subsequent processing. The user can also input text. For example, the user inputs the name of the medicine, and the system cleans, segments, and standardizes the text.

[0125] Step S1004, determine whether the user input is text or a picture.

[0126] If the user input is a picture, execute step S1006; otherwise, execute step S1012.

[0127] Step S1006, use the medicine box detection model for medicine box recognition and correction.

[0128] Step S1008, cross-modal feature extraction.

[0129] In the case of a picture input, use the pre-trained ChineseClip model (medicine box detection model) to extract the picture feature vector. In the case of text input, convert the medicine brand and name into a feature vector through the ChineseClip text encoder. Then, generate a unified feature representation for the feature vectors for subsequent retrieval.

[0130] Step S1010, cross-modal retrieval and matching.

[0131] Use a vector database (such as Milvus or Faiss) for efficient similarity retrieval. Perform Top-K recall based on cosine similarity calculation to obtain the drug information most similar to the query and its corresponding instruction manual ID.

[0132] After cross-modal retrieval matching, the process can end directly with the retrieval result output, or it can jump to step S1016.

[0133] In step S1012, determine whether the intent recognition is a question.

[0134] If it is a question, execute step S1014; otherwise, execute step S1008.

[0135] In step S1014, perform QA feature extraction.

[0136] In step S1016, perform block retrieval of the electronic instruction manual.

[0137] Convert the user's specific consultation question into a query vector through the BGE model. Retrieve relevant content fragments in the instruction manual vector database. Screen out the most relevant content paragraphs through semantic similarity ranking.

[0138] In step S1018, integrate context information.

[0139] Integrate basic drug information (such as name, specification, usage, etc.). Integrate relevant instruction manual content (such as dosage and usage, contraindications, etc.). Construct a structured prompt template to provide complete context information for the large language model.

[0140] Input the integrated context information into the large language model (LLM). Guide the model to generate accurate and professional answers through optimized prompt engineering. Verify the reliability and integrity of the generated answers.

[0141] Finally, format the generated answer to make it more understandable. Necessary warning information and disclaimers can also be added. Finally, the final consultation result can be presented in an age-friendly manner (such as enlarged font, voice broadcast, etc.).

[0142] Specifically, the prompt words for the large model to generate medication results are as follows:

[0143]

[0144] In the medication consultation link, the present invention adopts a large language model architecture based on retrieval enhancement. Through multi-level context information fusion and intelligent dialogue generation, it provides professional and accurate medication guidance for users. The system uses the drug instruction manual vector library as the knowledge basis and combines optimized prompt engineering to achieve intelligent medication consultation services.

[0145] In terms of prompt design, the present invention constructs a multi-level prompt template system. First, a clear role positioning is set, defining the model as a "professional drug consultant". This role positioning ensures the professionalism and rigor of the model when answering questions. Secondly, at the instruction level, the answering strategy is clarified: when the answer to the question can be found in the context, the model will generate an accurate answer based on the context; if the question goes beyond the context scope, the model needs to answer politely and guide the user to ask answerable questions or suggest seeking medical advice outside the scope of assisted medical care. This strategy design not only ensures the reliability of the answer but also reflects the service awareness of the system.

[0146] In the context construction process, the system adopts a multi-level retrieval strategy. First, it retrieves the drug information fragments most relevant to the user's question through vector similarity retrieval, and then sorts these fragments by importance. The sorting considers multiple dimensions: the semantic relevance to the question, the importance of the content (such as the importance of contraindications is higher than general instructions), the timeliness of the information, etc. The selected context content is organized in the logical order of "basic information - dosage and administration - precautions - professional information" to ensure that the large model can obtain complete and structured reference information.

[0147] To improve the answer quality, the present invention designs a multiple constraint mechanism at the model output end. First is the reliability constraint, which requires the model to generate answers based on the given context to avoid generating false information; second is the integrity constraint, ensuring that the answer covers all aspects that the user cares about; finally is the security constraint, in questions related to medication safety, the model will emphasize the importance of referring to medical advice. In addition, the system also implements an intelligent answer expansion mechanism. In addition to directly answering the user's questions, it will also supplement relevant medication precautions or possible risk warnings. The specific interface is as Figures 11 to 13 shown.

[0148] Through the above methods, the present invention can provide more accurate, professional and safe medication consultation services, ensuring that users receive correct guidance and suggestions during the medication process, thereby effectively improving the safety and accuracy of medication.

[0149] The present application also provides an intelligent medication consultation device based on a large model, including an acquisition module configured to acquire user input information; a text processing module configured to, when the user input information is text, perform feature extraction on the text to obtain a first feature representation, perform block retrieval on the electronic instruction manual based on the first feature representation, and perform context integration on the block retrieval results to generate medication consultation information for responding to the user input information; and an image processing module configured to, when the user input information is an image, perform medicine box detection on the image, identify a target medicine box, perform cross-modal feature extraction based on the target medicine box to obtain a second feature representation, and perform cross-modal retrieval based on the second feature representation to generate medication consultation information for responding to the user input information.

[0150] The present application also provides an intelligent medication consultation device based on a large model, including an acquisition module configured to acquire user input information; a text processing module configured to, when the user input information is text, perform feature extraction on the text to obtain a first feature representation, perform block retrieval on the electronic instruction manual based on the first feature representation, and perform context integration on the block retrieval results to generate medication consultation information for responding to the user input information; and an image processing module configured to, when the user input information is an image, perform medicine box detection on the image, identify a target medicine box, perform cross-modal feature extraction based on the target medicine box to obtain a second feature representation, and perform cross-modal retrieval based on the second feature representation to generate medication consultation information for responding to the user input information.

[0151] The present application also provides a cross-modal retrieval and matching device for drug information based on AI, including: an acquisition module configured to acquire a target medicine box image to be recognized; an extraction module configured to perform feature extraction on the target medicine box image by using a multi-modal drug matching model to obtain a feature representation; and a matching module configured to perform cross-modal matching based on the feature representation by using the multi-modal drug matching model to match drug information related to the target medicine box; wherein, the multi-modal drug matching model is obtained by: performing domain adaptation training on a pre-trained multi-modal network model based on a graphic-text pair training set in the drug field to obtain the multi-modal drug matching model.

[0152] The present application also provides an AI-based rotating medicine box recognition and calibration device, including: an acquisition module configured to acquire a target medicine box image to be recognized; a recognition module configured to recognize the target medicine box image to be recognized by using a medicine box detection model and correct the rotation angle of the recognized target medicine box to the standard horizontal direction; wherein, the medicine box detection model is obtained through the following: based on a medicine box data set with rotation angle annotations, training a neural network model to obtain the medicine box detection model.

[0153] It should be noted that: for the device provided in the above embodiment, only the above division of each functional module is used for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the corresponding method embodiment belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be repeated here.

[0154] Figure 14 The structural schematic diagram of the electronic device suitable for implementing the embodiments of the present disclosure is shown. It should be noted that Figure 14 The shown electronic device is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.

[0155] As Figure 14 shown, the electronic device includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage part 1008 into the random access memory (RAM) 1003. In the RAM 1003, various programs and data required for system operation are also stored. The CPU 1001, ROM 1002, and RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.

[0156] The above are only the preferred embodiments of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and retouches can be made, and these improvements and retouches should also be regarded as the protection scope of the present application.

Claims

1. A cross-modal retrieval and matching method for drug information based on AI, characterized in that, Including: Obtain a target medicine box picture to be recognized; Use a multimodal medicine matching model to extract features from the target medicine box picture to obtain a feature representation; Based on the feature representation, use the multimodal medicine matching model for cross-modal matching to match medicine information related to the target medicine box; Among them, the multimodal medicine matching model is obtained as follows: Based on a training set of text-image pairs in the medicine field, perform domain adaptation training on a pre-trained multimodal network model to obtain the multimodal medicine matching model.

2. The method according to claim 1, wherein The training set of text-image pairs is obtained as follows: Obtain medicine sample data, where the medicine sample data includes medicine sample pictures and training text data; Perform at least one of the following enhancement processes on the medicine sample pictures: perspective transformation, lighting adjustment, and local magnification, and expand the training text data by means of synonym replacement and interchange of brand names and generic names; Use the enhanced medicine sample pictures and the expanded training text data as the training set of text-image pairs.

3. The method according to claim 2, characterized in that, Based on the training set of text-image pairs, perform medicine field pre-training on the multimodal network model to obtain the multimodal medicine matching model, including: Perform pre-training on the multimodal network model on a general Chinese text-image dataset; Based on the training set of text-image pairs, perform medicine field pre-training on the pre-trained multimodal network model to obtain the multimodal medicine matching model.

4. The method according to claim 3, wherein Based on the training set of text-image pairs, perform medicine field pre-training on the pre-trained multimodal network model, including: Based on an attention mechanism enhanced by medicine knowledge, use an additional medicine attribute supervision signal to guide the model to extract key visual features in the medicine sample pictures in the training set of text-image pairs; Through an attention mechanism enhanced by medicine knowledge graph embedding, combined with a medicine attribute classification loss function, guide the multimodal network model to extract key visual features; When encoding the training text data in the training set of text-image pairs, fuse the prior information of entity relationships in the medicine knowledge graph; Use the key visual features and the text encoding fused with the prior information of entity relationships to perform the medicine field pre-training.

5. The method according to claim 1, wherein Based on the feature representation, use the multimodal medicine matching model for cross-modal matching to match medicine information related to the target medicine box, including: Based on the feature representation, use the multimodal medicine matching model to perform cross-modal matching on a text-image vector database to obtain the medicine information; among them, the text-image vector database stores image feature vectors of medicine box pictures, as well as feature vectors corresponding to medicine names, brands, indications, dosages, and instruction manual texts; The medicine information includes at least one of the following: the type, brand, indication, and instruction manual of the medicine.

6. The method according to claim 5, wherein Use the multimodal medicine matching model to perform cross-modal matching on a text-image vector database to obtain the medicine information, including: Use the multimodal medicine matching model to perform similarity retrieval on the text-image vector database and calculate the cosine similarity; Sort based on the cosine similarity, and select the top K candidate results with the highest similarity as the drug information.

7. A cross-modal retrieval and matching device for drug information based on AI, characterized in that, It includes: An acquisition module, configured to acquire a target medicine box picture to be recognized; An extraction module, configured to extract features from the target medicine box picture by using a multi-modal drug matching model to obtain a feature representation; A matching module, configured to perform cross-modal matching by using the multi-modal drug matching model based on the feature representation to match drug information related to the target medicine box; Among them, the multi-modal drug matching model is obtained as follows: Based on a text-image pair training set in the drug field, perform domain adaptation training on a pre-trained multi-modal network model to obtain the multi-modal drug matching model.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of claims 1 to 6.

9. A computer device, characterized in that, It includes: A memory and a processor, The memory stores a computer program; The processor is used to execute the computer program stored in the memory, and when the computer program runs, it causes the processor to execute the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Intelligent box type recommendation method based on multi-modal retrieval

    CN121071003A

  • Drug information automatic identification and matching warehousing method, device and equipment, medium and product

    CN121116998A

  • Drug package multi-scale OCR (optical character recognition) and matching method oriented to complex background

    CN121121718A