Drug query method, point reading equipment and storage medium

Through the sensors and audio output components of the point-reading device, combined with voice conversion and semantic analysis, the readability and interactivity problems of traditional drug instructions are solved, accurate acquisition of drug information and real-time guidance are achieved, and the medication error rate for the elderly and visually impaired population is reduced.

CN120705268APending Publication Date: 2025-09-26BEIJING JISHUITAN HOSPITAL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510856783.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Traditional drug instructions have too small fonts, complex medical terminology, and lack of interactive functions, making them difficult for the elderly and visually impaired to read and unable to answer medication questions in real time.

Method used

A point-reading device is provided that collects interactive information from drug instructions through sensors, combines vibration and audio output components to achieve voice conversion and semantic analysis, supports AI multi-round question and answer, obtains drug information and provides personalized feedback.

Benefits of technology

It improves the readability and interactivity of drug information, reduces the reading difficulty for the elderly and visually impaired, provides real-time medication guidance, and reduces the rate of medication errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705268A_ABST
    Figure CN120705268A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of point reading equipment, and discloses a medicine query method, point reading equipment and a storage medium, the medicine query method comprises the following steps: in response to a query instruction for a medicine, under the condition that the query instruction is a voice query instruction input by a user, performing voice conversion on the voice query instruction to obtain a voice query text; performing semantic analysis on the voice query text to determine a target entity and a semantic relationship queried by the user; performing lightweight query in a drug database according to the obtained semantic relationship and the target entity to obtain a query result associated with the drug; and synthesizing a character answer according to the query result, and outputting the character answer in a voice form. According to the drug query method, the user can obtain accurate drug information and clear answering guidance, drug use errors caused by misunderstanding of a specification are avoided, and the wrong drug use rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of point-reading devices, and in particular to a drug query method, a point-reading device, and a storage medium. Background Art

[0002] Traditional drug instructions have many limitations, especially in terms of readability and interactivity. Specifically, they include: Small font size: For the elderly and visually impaired, small fonts make reading difficult, increasing barriers to obtaining drug information. Complex medical terminology: Drug instructions contain a large amount of professional terminology that is difficult for ordinary users (especially the elderly) to understand, potentially leading to medication errors or misunderstandings. Lack of interactivity: Traditional instructions are static text and cannot provide personalized answers or real-time guidance based on user needs, which is particularly inconvenient when medication questions arise frequently. Summary of the Invention

[0003] In view of this, the embodiments of the present application provide a drug query method, a point-reading device, and a storage medium, which can effectively solve the problems of lack of readability and interactivity in traditional drug instructions.

[0004] In a first aspect, an embodiment of the present application provides a drug query method, applied to a point-reading device, the method comprising: In response to a query instruction for a drug, if the query instruction is a voice query instruction input by a user, performing voice conversion on the voice query instruction to obtain a voice query text; Performing semantic analysis on the voice query text to determine the target entity and semantic relationship of the user query; Performing a lightweight query in a drug database based on the acquired semantic relationship and the target entity to obtain a query result associated with the drug; A text answer is synthesized according to the query result, and the text answer is output in the form of voice.

[0005] In a first possible embodiment of the first aspect, the further comprising: In the case where the query instruction is an interactive trigger instruction for the user to click and read the drug instructions, the query feedback information that can be perceived by the user is controlled and output according to the interactive trigger instruction.

[0006] In a second possible embodiment of the first aspect, the drug instruction sheet is provided with an optical recognition area, the optical recognition area is provided with multiple tactile codes, and each tactile code corresponds to unique medication information; the point reading device includes a sensor, a vibration output component, and an audio output component; and the interaction trigger instruction includes: interaction information collected when the point reading device clicks the optical recognition area of ​​the drug instruction sheet; Outputting the query feedback information perceivable by the user according to the interaction trigger instruction includes: According to the interactive trigger instruction, the vibration information corresponding to the optical recognition area is provided to the user through the vibration output component, and the audio information corresponding to the medication information of the optical recognition area is provided to the user through the audio output component.

[0007] In a third possible embodiment of the first aspect, the present invention further includes: If the target entity cannot be obtained, triggering a follow-up question mechanism to output related questions related to the voice query text through a preset follow-up question template; In response to the user's associated answer to the associated question, the voice query text is modified to re-determine a target entity based on the modified voice query text.

[0008] In a fourth possible embodiment of the first aspect, converting the voice query instruction to obtain voice query text includes: performing language classification on the voice query instruction through a language classification branch of an acoustic hybrid recognition model to determine a language type; The voice query instruction is converted into a voice query text according to a language model corresponding to the language type.

[0009] In a fifth possible embodiment of the first aspect, performing language classification on the voice query instruction and determining the language type by the language classification branch of the acoustic hybrid recognition model includes: Extracting Mel-spectrogram features of the voice query instruction through a shared encoder of the acoustic hybrid recognition model; The voice query instruction is language-classified according to the Mel-spectrogram feature to determine the language type.

[0010] In a sixth possible embodiment of the first aspect, semantically parsing the voice query text to determine the target entity and semantic relationship of the user query includes: Convert the voice query text into a query vector using the Wenxin model, and perform similarity matching between the query vector and the entity vector in the knowledge graph to determine the target entity; The semantic relationship of the target entity in the voice query text is extracted through the Wenxin model.

[0011] In a seventh possible embodiment of the first aspect, the drug database includes all information of each entity, and performing a lightweight query on the drug database based on the acquired semantic relationship and the target entity to obtain a query result includes: generating a query statement according to the semantic relationship and the target entity; Performing convolution optimization based on the lightweight backbone network in the lightweight model to determine the mapping relationship between relevant entities in the drug database; All redundant attention heads that are not related to the query statement in the multi-head attention mechanism of the lightweight model are eliminated to query the information of the target entity in the drug database based on the multi-head attention mechanism, and to capture the mapping relationship between entities related to the target entity to determine the query result.

[0012] In a second aspect, an embodiment of the present application provides a point reading device, comprising: A sensor for collecting interactive information in an optical recognition area of ​​a drug package insert; A sound collection device is used to collect voice query commands input by users; a vibration output component, configured to provide the user with vibration information corresponding to the optical recognition area; an audio output device for providing audio information corresponding to medication information of the optical recognition area and outputting a voice answer; A processor is used to implement the above-mentioned drug query method.

[0013] In a third aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which implements the above-mentioned drug query method when executed on a processor.

[0014] The embodiments of the present application have the following beneficial effects: A drug query method of this embodiment is applied to a point-reading device. The drug query method includes: responding to a query instruction for drugs, when the query instruction is a voice query instruction input by the user, converting the voice query instruction into voice to obtain a voice query text; semantically parsing the voice query text to determine the target entity and semantic relationship of the user's query; performing a lightweight query in the drug database based on the acquired semantic relationship and target entity to obtain a query result associated with the drug; synthesizing a text answer based on the query result, and outputting the text answer in the form of voice. Based on the above scheme, the drug query method of this application can obtain accurate drug information and avoid medication errors caused by misunderstanding of the instructions. This application provides a question-and-answer interaction of a point-reading device, providing users with clear guidance and answers, and reducing the rate of incorrect medication. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0016] Figure 1 A schematic diagram of the structure of a point reading device according to an embodiment of the present application is shown; Figure 2 A first flow chart of the drug query method according to an embodiment of the present application is shown; Figure 3 A second flow chart of the drug query method according to an embodiment of the present application is shown; Figure 4 A third flow chart of the drug query method according to an embodiment of the present application is shown.

[0017] Description of main component symbols: 100 - point reading device; 110 - sensor; 120 - vibration output component; 130 - audio output device; 140 - sound collection device; 150 - processor. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments.

[0019] The components of the embodiments of the present application generally described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but rather merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present application.

[0020] Hereinafter, the terms "including", "having" and their cognates used in various embodiments of the present application are intended only to indicate specific features, numbers, steps, operations, elements, components or combinations of the aforementioned items, and should not be understood as excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the aforementioned items or adding the possibility of one or more features, numbers, steps, operations, elements, components or combinations of the aforementioned items. In addition, the terms "first", "second", "third" and the like are only used to distinguish descriptions and should not be understood as indicating or implying relative importance.

[0021] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art to which the various embodiments of the present application belong. The terms (such as those defined in generally used dictionaries) will be interpreted as having the same meaning as in the context of the relevant technical field and will not be interpreted as having an idealized meaning or an overly formal meaning unless clearly defined in the various embodiments of the present application.

[0022] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.

[0023] When patients review medication inserts, they often encounter small fonts and complex medical terminology, making them difficult for the elderly and visually impaired to read. Furthermore, they lack interactive features, preventing real-time answers to medication questions. Touch-reading devices are electronic tools that play corresponding audio or explanations by touching text or images on paper. However, traditional touch-reading devices are primarily used in education and have not been integrated into medication safety scenarios. Furthermore, their AI voice interaction systems lack multimodal feedback, resulting in a poor user experience.

[0024] In response to the above problems, the present application provides a drug query method, a point reading device and a storage medium, which can obtain corresponding feedback information based on the interactive information of the point reading device in the drug instructions, so that patients can clearly understand the corresponding drug information in the drug instructions. The point reading device 100 also supports AI multi-round question and answer to answer questions about actual medication for patients.

[0025] First, an embodiment of the present application provides a point-to-point reading device 100. The point-to-point reading device 100 includes a sensor 110, a vibration output component 120, an audio output device 130, a sound collection device 140, and a processor 150. The processor 150 can be directly or indirectly electrically connected to the sensor 110, the vibration output component 120, the audio output device 130, and the sound collection device 140 to enable data transmission and interaction. For example, these components can be electrically connected to each other via a bus and / or signal lines.

[0026] In one embodiment, sensor 110 is used to collect user interaction information in the optical recognition area of ​​a medication insert. Sensor 110 includes an optical sensor and a piezoelectric film sensor. In this embodiment, the optical sensor is used to illuminate the tactilely encoded bump surface in the optical recognition area using a near-infrared light source. A CMOS camera is used to capture an image of the reflected light from the optical recognition area, which is then recognized to determine the medication information corresponding to the optical recognition area.

[0027] The piezoelectric film sensor is used to provide a pressure signal indicating the pressure of the point reading device 100 pressing the optical recognition area when the point reading device 100 contacts the optical recognition area, wherein different pressure signals correspond to different interactive trigger instructions. The present application provides the user with audio information and tactile information corresponding to the optical recognition area according to the interactive trigger instruction. For example, a larger pressure signal can output only the audio information corresponding to the encoding area, and a smaller pressure signal can output all the audio information corresponding to the entire optical recognition area. The present application can provide a vibration prompt for medication information, and distinguish the level of medication information by the difference in vibration frequency. For example, high-frequency vibration is a contraindication warning, and low-frequency vibration is for routine usage.

[0028] In one embodiment, the vibration output component 120 is used to provide vibration information corresponding to the optical recognition area to the user. The audio output device 130 is used to provide audio information corresponding to the optical recognition area according to the interaction trigger instruction, and output a voice answer. The sound collection device 140 is used to collect the voice query instruction input by the user.

[0029] In one embodiment, the vibration output component 120 may be a tactile vibration motor that provides vibration prompts and feedback on medication level information. The sound collection device 140 may be a microphone for receiving user voice questions and capturing voice query instructions. The audio output device 130 may be a speaker for playing audio information corresponding to the optical recognition area and playing voice answers.

[0030] In one embodiment, the drug instructions are provided with an optical recognition area, and the optical recognition area is provided with multiple tactile codes, each tactile code corresponding to unique medication information; the drug instructions are used in conjunction with the above-mentioned point reading device 100.

[0031] In the embodiment of the present application, the point reading device 100 can obtain medication information corresponding to the optical recognition area, such as drug dosage, indications, etc., by scanning and identifying the optical recognition area. Tactile coding can be perceived directly through finger touch without relying on vision, enabling blind users to quickly locate key information.

[0032] In one embodiment, the optical recognition area can be configured as an OID (Optical Identification) micro-dot matrix. The OID micro-dot matrix is ​​an optical identification-based encoding technology that stores drug information by printing tiny dots on the surface of the drug insert. The position and color of each dot can represent binary data, forming a high-density information encoding area. The optical recognition area can also be configured as a QR code. The point reading device 100 can scan and identify the QR code to obtain the corresponding medication information.

[0033] The processor 150 may process information and / or data related to the drug query method to perform one or more functions described herein. For example, in response to a query instruction for drugs, if the query instruction is a voice query instruction input by a user, the processor 150 may convert the voice query instruction into speech to obtain a voice query text; perform semantic parsing on the voice query text to determine the target entity and semantic relationship of the user's query; perform a lightweight query in the drug database based on the obtained semantic relationship and target entity to obtain query results associated with the drug; synthesize a text answer based on the query results, and output the text answer in the form of speech.

[0034] The processor 150 may be an integrated circuit chip with signal processing capabilities. The processor 150 may be a general-purpose processor, including at least one of a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor may be a microprocessor or any conventional processor, and may implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0035] For ease of understanding, the following examples of this application will be described in Figure 1 Taking the illustrated point reading device 100 as an example, in conjunction with FIG1 , the drug query method provided in the embodiment of the present application is described.

[0036] Figure 2 A flow chart of a drug query method according to an embodiment of the present application is shown. Exemplarily, the drug query method includes the following steps: S110 , in response to a query instruction for drugs, if the query instruction is a voice query instruction input by a user, performing voice conversion on the voice query instruction to obtain a voice query text.

[0037] Exemplarily, the present application can obtain the user's voice query instructions, perform voice conversion and semantic analysis on the voice query instructions, match the drug database according to the semantic analysis results, and provide the user with corresponding answers.

[0038] In this embodiment, Figure 3 As shown, this application converts the voice query instruction into voice through voice recognition, including the following steps: S111 , classifying the voice query instruction by language through the language classification branch of the acoustic hybrid recognition model to determine the language type.

[0039] In one embodiment, the acoustic hybrid recognition model is an improved DeepSpeech2 architecture that can support multiple dialects. A dialect classification branch is added after the shared encoder to achieve end-to-end multi-task learning. In this application, the shared encoder is used to extract features from the audio signal of the input voice query command. The shared encoder is a neural network structure that can learn feature representations from audio signals and output text transcription results. The neural network structure includes but is not limited to a convolutional neural network (CNN), a recurrent neural network (RNN), etc.

[0040] In one embodiment, the present application extracts mel-spectrogram features of a voice query instruction through a shared encoder of an acoustic hybrid recognition model; and classifies the voice query instruction into a language based on the mel-spectrogram features to determine the language type.

[0041] In this embodiment, the application uses the first 0.5 seconds of speech signal data (corresponding to the first 10 frames of the mel-spectrogram) as input. The mel-spectrogram captures the time-frequency characteristics of the speech signal by performing a short-time Fourier transform on the audio signal and calculating its energy distribution on the mel-scale. Based on the statistical features of the mel-spectrogram (such as energy distribution and frequency change rate), the application extracts dialect-related features (such as tone patterns and stress placement). A deep learning model (such as a convolutional neural network or a recurrent neural network) is used to classify the mel-spectrogram and output the predicted dialect type.

[0042] In one embodiment, when the query instruction is an interactive trigger instruction for the user to click and read the drug instructions, the query feedback information perceivable by the user is output according to the interactive trigger instruction. The feedback information may be audio information and tactile information corresponding to the optical recognition area.

[0043] In one embodiment, a drug insert includes an optical recognition area with multiple tactile codes, each of which corresponds to unique medication information. The interaction trigger instruction includes interaction information collected when a point-and-click device clicks on the optical recognition area of ​​the drug insert. For example, the interaction information may be a user lightly scanning or pressing a tactile code with the point-and-click device.

[0044] In one embodiment, the present application provides vibration information corresponding to the optical recognition area to the user through the vibration output component according to the interaction trigger instruction, and provides audio information of the corresponding medication information of the optical recognition area to the user through the audio output component. For example, when the user presses a certain tactile code hard, the audio of the medication information corresponding to the tactile code can be provided. When the user gently touches and scans the tactile code, the audio of the medication information corresponding to the entire optical recognition area can be provided.

[0045] S112. Convert the voice query instruction into a voice query text according to the language model corresponding to the language type.

[0046] In one embodiment, the language model of the present application adopts the n-gram language model. The n-gram language model is a statistics-based language modeling method that predicts the probability of the next word by analyzing the historical word sequence. In the voice input stage of the present application, dialect prediction is first completed through the voice recognition module. According to the prediction result, the n-gram language model containing the corresponding dialect is selected. The n-gram language model for each dialect is optimized for a specific dialect and contains high-frequency vocabulary and grammatical structures of that dialect. For example, the n-gram language model for Cantonese dialect contains high-frequency words such as "唔该" (excuse me) and "点解" (why). When converting the voice query instruction in Cantonese input, high-frequency words such as "唔该" and "点解" are preferentially matched.

[0047] In one embodiment, while converting the voice query instruction through the shared encoder, the high-frequency words in the loaded n-gram language model are preferentially matched. When the Cantonese model is loaded, the possibilities of words such as "唔该" and "点解" are preferentially considered, and the voice query text is output. For example, the Cantonese sentence "呢只药同降压药可唔可以一齐食" (Can this medicine and antihypertensive medicine be taken together?) is converted into Mandarin "这种药和降压药能不能一起吃".

[0048] S120. Semantically parse the voice query text to determine the target entity and semantic relationship queried by the user.

[0049] In one embodiment, the target entity includes the actual drug name, disease name, etc., and the semantic relationship is the semantic relationship between the target entities, such as drug interaction, contraindication, indication, etc. For example, for the voice query text: "阿司匹林能不能和氯吡格雷一起吃" (Can aspirin and clopidogrel be taken together?), the named entity recognition module recognizes "阿司匹林" (aspirin) and "氯吡格雷" (clopidogrel) as the target entities, and "一起吃" (be taken together) as the target entity interaction relationship.

[0050] Demonstratively, since the patient cannot accurately state the actual drug name or disease name and can only state a part of the drug name or disease name, the present application can perform similarity matching according to the fuzzy semantics of the patient to determine the specific drug name, disease name, etc. that the patient wants to query.

[0051] In one embodiment, the present application converts the voice query text into a query vector through the Wenxin Big Model, performs similarity matching between the query vector and the entity vector in the knowledge graph to determine the target entity; and extracts the semantic relationship of the target entity in the voice query text through the Wenxin Big Model.

[0052] In one embodiment, the present application identifies target entities and entity relationships through a medical domain fine-tuning sub-model of the Wenxin model. The Wenxin model is a large-scale pre-trained language model with strong natural language understanding and generation capabilities. In order to adapt to application scenarios in the medical field, the basic model will be fine-tuned to improve its performance in this field. In this application, high-quality data is extracted from a large amount of medical literature, drug instructions, clinical guidelines, patient question and answer records and other medical-related corpora, and the Wenxin model is fine-tuned to the field so that the model can better understand medical terms, drug names, disease names and the relationships between them.

[0053] In this embodiment, to quickly identify the target entity in a user's question, this application uses a semantic vectorization and similarity matching method based on the medical field fine-tuning sub-model of the Wenxin Grand Model. The user's voice query text is converted into a high-dimensional semantic vector (i.e., a query vector), and then similarity is calculated with the entity vectors in the knowledge graph to find the closest entity name as a matching result. This application pre-constructs a knowledge graph containing entities such as drugs and diseases, and each entity and its description is converted into a corresponding entity vector. By learning from large-scale corpus, the medical field fine-tuning sub-model of the Wenxin Grand Model can understand the deep semantics of natural language and map the voice query text to a vector representation in a high-dimensional space. In this application, the medical field fine-tuning sub-model of the Wenxin Grand Model converts the above-mentioned voice query text into a fixed-dimensional query vector. This application can measure the similarity between the query vector and the entity vector using cosine similarity, sort based on the similarity, and select the closest entity as the target entity.

[0054] For example, the voice query text is "Interaction between enteric-coated aspirin tablets and clopidogrel". When converted into a query vector, the similarity between the query vector and the entity vector of enteric-coated aspirin tablets is 0.85, and the similarity between the query vector and the entity vector of clopidogrel is 0.92. Both the entity aspirin and the entity clopidogrel are successfully matched as target entities.

[0055] In another embodiment, when the target entity cannot be determined, a follow-up question mechanism is triggered, and related questions related to the voice query text are output through a preset follow-up question template; in response to the user's related answers to the related questions, the voice query text is corrected to re-determine a target entity based on the corrected voice query text.

[0056] In this embodiment, the medical domain fine-tuning sub-model of the Wenxin Grand Model analyzes the user's question. If it cannot extract sufficient key target entities (such as drug names, disease names, etc.), the question is determined to be ambiguous, triggering a follow-up question mechanism. Based on the semantic features of the user's question and predefined rules, this application categorizes ambiguous questions into different types (such as ambiguous drug names, ambiguous disease names, etc.). For example, if a user's question lacks a drug name, it is classified as "Ambiguous drug name"; if a user's question lacks a disease name, it is classified as "Ambiguous disease name." Each question type corresponds to a set of preset follow-up question templates, and this application selects the most appropriate template based on the current question type. For example, if a user asks, "Can this medicine be taken with that medicine?" and the current question type is "Ambiguous drug name," the follow-up question mechanism can trigger the corresponding preset follow-up question template, asking, "What specific drug are you referring to?'" The user responds, "I'm talking about aspirin." Then, if the user responds, "What specific drug are you referring to?'" The user responds, "I'm talking about clopidogrel."

[0057] In one implementation, this application uses Dialogue State Tracking (DST) to respond to user responses, recording and updating the conversation state, including key entities mentioned by the user (e.g., drug names, disease names, etc.) and the current conversation goal. This application uses the DST module to dynamically modify query conditions, ensuring that each response is based on complete context.

[0058] It is understandable that since the patient did not provide the actual drug name description text when asking the question, for example, the patient's query text is "Can it be taken with antihypertensive drugs?" At this time, the named entity recognition module cannot determine the actual drug name. This application can trigger the follow-up question mechanism to ask the patient what antihypertensive drugs are and what drugs can be taken with antihypertensive drugs. The handling of fuzzy questions guides users to supplement missing information through the follow-up question mechanism, which can accurately understand the user's intentions and provide accurate answers. This application uses friendly and clear follow-up question language to help users gradually supplement missing information, which is particularly suitable for the elderly, visually impaired people and people with low cultural levels.

[0059] S130, performing a lightweight query in the drug database based on the acquired semantic relationship and target entity to obtain a query result.

[0060] In this embodiment, the drug database is the core data storage component, containing all the information about each entity. This application constructs multiple table structures based on this entity information to record information such as drug name, drug interactions, contraindications, usage and dosage. The following is the design of the tables and fields within the drug database: The drug basic information table (drug_info) stores basic information about each drug. The fields include: drug_id, which is the unique drug identifier (e.g., ASPIRIN for aspirin and CLOPIDOGREL for clopidogrel); drug_name, which is the generic name (e.g., aspirin, clopidogrel); brand_name, which is the trade name (e.g., Aspirin, Plavix); and category, which is the drug category (e.g., antihypertensive, antithrombotic).

[0061] The drug interaction table (drug_interaction) records the interaction between two drugs. Fields include: drugA_id, which represents the ID of the first drug; drugB_id, which represents the ID of the second drug; interaction_level, which represents the risk level of the interaction (e.g., low risk, moderate risk, high risk); and description, which represents the specific description of the interaction (e.g., increased risk of bleeding).

[0062] The contraindications table (drug_contraindications) records the contraindications between a drug and a specific disease or population. Fields include: drug_id (drug ID); disease_id (disease ID (e.g., HYPERTENSION for hypertension, DIABETES for diabetes); population_id (population ID (e.g., ELDERLY for the elderly); and contraindication_description (contraindication description (e.g., contraindication for pregnant women).

[0063] The drug_usage table records the usage and dosage information for a particular drug. Fields include: drug_id (drug ID); dosage (single dose, e.g., 100mg); frequency (number of times per day, e.g., twice per day); and timing (time of use, e.g., before meals, before bed). This table is linked to the medication information in the drug insert.

[0064] The drug-disease association table stores the relationships between drugs and other entities (such as diseases and symptoms). Fields include: entity1_id, the ID of the first entity (such as a drug ID); entity2_id, the ID of the second entity (such as a disease ID); and relation_type, the relationship type between the entities (such as indications and contraindications).

[0065] In one embodiment, if Figure 4 As shown, this application performs lightweight query in the drug database based on semantic relationships and target entities, including the following steps: S131: Generate a query statement based on the semantic relationship and the target entity.

[0066] In one embodiment, the present application automatically generates SQL query statements that conform to the database structure based on target entities and semantic relationships. SQL (Structured Query Language) is a standard language for managing and operating relational databases. In this application, SQL query statements are the core tool for retrieving data from a database by writing specific grammatical structures. SQL query statements consist of the following parts: SELECT: specifies the fields to be searched; FROM: specifies the table from which the data comes; WHERE: sets the search conditions (optional); ORDER BY: sorts the results (optional); LIMIT: limits the number of records returned (optional).

[0067] For example, if the user queries the text instruction "Interaction between enteric-coated aspirin tablets and clopidogrel", the corresponding SQL query statement is: SELECT interaction_level, description FROM drug_interaction WHERE (drugA_id = 'ASPIRIN' AND drugB_id = 'CLOPIDOGREL') OR (drugA_id = 'CLOPIDOGREL' AND drugB_id = 'ASPIRIN').

[0068] The interaction_level field represents the risk level of the drug interaction (e.g., low risk, medium risk, high risk). The description field describes the specific impact of the drug-drug interaction (e.g., increased bleeding risk). The search criteria ensure that the interaction relationship is correctly matched regardless of the order of the drugs.

[0069] It can be understood that the role of SQL query statements is to extract specific information related to the user's question (such as drug interactions, contraindications, usage and dosage, etc.) from the drug database, so as to provide users with accurate answers.

[0070] S132, performing convolution optimization based on the lightweight backbone network in the lightweight model to determine the mapping relationship between various related entities in the drug database.

[0071] In one embodiment, in order to efficiently process the complex mapping relationships between various entities in the drug database (such as drug names, disease names, etc.), this application uses a lightweight backbone network combined with depthwise separable convolution (DSC) to perform efficient convolution optimization to determine the most relevant mapping relationships between various entities.

[0072] In this application, in order to meet the real-time and resource constraints of the point-reading device 100, this application adopts a lightweight model and achieves efficient computing through a lightweight backbone network and depth-wise separable convolution. A lightweight model refers to a deep learning model that can run efficiently in a resource-constrained environment by reducing the number of model parameters and computational complexity through technical means such as optimized architecture design or pruning. The lightweight backbone network is the core component of the lightweight model. It is based on architecture designs such as MobileNet and EfficientNet-Lite. These networks significantly reduce computational complexity while ensuring performance by reducing the number of channels and introducing efficient convolution operations. Depthwise separable convolution is an efficient convolution optimization method that decomposes the standard convolution into two steps: depthwise convolution and pointwise convolution. This decomposition method can determine the most relevant mapping relationships between entities. For example, aspirin is most relevant to the treatment of coronary heart disease. Depthwise separable convolution can simplify aspirin and coronary heart disease to one or two mapping relationships, improving the correlation between the two and significantly reducing computational complexity and memory usage.

[0073] S133, remove all redundant attention heads that are irrelevant to the query statement in the multi-head attention mechanism of the lightweight model, query the information of the target entity in the drug database based on the multi-head attention mechanism, and capture the mapping relationship between entities related to the target entity to determine the query result.

[0074] For example, the multi-head attention mechanism is one of the core components of the Transformer model. It decomposes the input data into multiple "heads", each of which independently calculates attention weights to capture feature relationships in different subspaces. Traditional multi-head attention mechanisms calculate all possible attention heads, even if some heads do not contribute to the final task (such as target entity mapping).

[0075] In one embodiment, some attention heads may be irrelevant to the query statement, resulting in redundant calculations and resource waste. This application removes redundant attention heads that are irrelevant to all query statements and focuses on capturing the mapping relationship between target entities through an optimized multi-head attention mechanism, thereby improving the accuracy of query results. For example, patients are more concerned about the usage, dosage, and medication time of drugs, rather than the manufacturer of the drugs. This application can remove the attention heads related to the drug manufacturers, thereby improving query efficiency and reducing latency.

[0076] In one embodiment, the present application stores frequently occurring voice queries and their corresponding voice answers based on their frequency of occurrence. In response to these frequently occurring voice queries, the pre-stored voice answers are output. In this embodiment, when a user issues the same frequently occurring query again, the answer can be directly retrieved from the cache without recalculation or retrieval, improving response speed and avoiding repetitive processing of the same task.

[0077] In one embodiment, the weights of frequently accessed voice query instructions are cached in the tightly coupled memory (TCM) of the MCU (Microcontroller Unit). The weights are used here to represent the frequency of use of the voice query instructions. The present application stores the frequently accessed weights in the tightly coupled memory (TCM) of the MCU to quickly determine which voice query instructions require priority caching. The present application arranges and stores tensor data according to the bit width (e.g., 64 / 128 bits) supported by the hardware SIMD (Single Instruction, Multiple Data) instructions. The tensor data can be the weight of the voice query instruction, thereby improving SIMD instruction utilization.

[0078] In one embodiment, the present application obtains user medication information and encrypts and stores it; upon user authorization, the medication information is synchronized to the cloud. In this embodiment, medication records are divided into two parts: structured data: non-sensitive information such as time and drug ID; and sensitive data: information that requires protection, such as dosage and patient name. The present application encrypts the storage and transmission of sensitive data to ensure that even if the structured data is leaked, the sensitive information cannot be directly accessed.

[0079] In one embodiment, the key management system (three-level key structure) of this application includes a master key, a data key, and a session key. The master key is scoped to the device level, stored in the SE security chip, and has a rotation policy of being fixed at the time the device leaves the factory. The data key is scoped to the user level, stored in the TEE trusted execution environment, and has a monthly rotation policy. The session key is scoped to be transmitted word-by-word, dynamically generated in memory, and updated with each connection. In this application, the scope describes the key's application scope or applicable scenario, the storage location specifies the secure environment or hardware facility where the key is stored, and the rotation policy describes the frequency and rules for key updates or replacements.

[0080] In one embodiment, this application ensures system security through hardware-level protection, preventing software-level security measures from being bypassed. The security chip is one of the core components for implementing hardware-level protection. The security chip's functional components include a key generation module, a cryptographic operation module, and a physical attack protection module. The key generation module uses a true random number generator (TRNG) to generate encryption keys; the cryptographic operation module performs encryption and decryption operations using an independent cryptographic coprocessor 150; and the physical attack protection module utilizes an active shielding layer and voltage anomaly detection mechanisms to prevent external physical attacks.

[0081] In one embodiment, the application obtains user authorization through an authentication mechanism. After confirming the user's identity, a short-lived access token is generated to allow medication information to be synchronized to the cloud. The authentication mechanism involves verifying a biometric (such as a fingerprint) using the FIDO2 protocol and generating a short-lived access token (JSON Web Token, JWT).

[0082] The FIDO2 protocol is an authentication standard based on public key cryptography, designed to provide enhanced security and a better user experience. FIDO2 combines the WebAuthn API and the CTAP protocol to support multiple authentication methods (such as fingerprint and facial recognition). To ensure security, short-term access tokens are typically set with a short validity period (such as 15 minutes). After expiration, the user needs to re-authenticate to obtain a new token. Short-term access tokens are used to authorize access to protected resources, avoiding authentication for each request. Medication information can be retained locally for 30 days, encrypted and uploaded to the cloud (retention for 7 years), and automatically erased.

[0083] S140, synthesize the text answer according to the query result, and output the text answer in the form of voice. In one embodiment, a speech speed adjustment button is set on the side of the reading device 100. A single press increases or decreases the speech speed by 10%, and a long press for 3 seconds restores the default. The present application adjusts the playback speed of the audio through the time domain linear interpolation method, while keeping the pitch unchanged as much as possible. The time domain linear interpolation method is used for audio signal processing, and the effect of acceleration or deceleration is achieved by sampling or interpolating the audio frames. For example, deleting the silent frame and the sound segment frame sampling can accelerate the audio; inserting the transition frame generated by linear interpolation, and extending the audio time axis by inserting transition frames between adjacent frames, thereby achieving a deceleration effect.

[0084] In this embodiment, the application provides users with accurate medication guidance through the dual modes of "point reading + AI question and answer", solving the problem of one-way information transmission in traditional instructions and greatly reducing the rate of incorrect medication. Blind users can quickly locate key information through tactile coding. Combining the synchronous trigger mechanism of tactile vibration rules and voice broadcast, the efficiency of information acquisition is improved and barrier-free adaptation is achieved. The dialect recognition function lowers the threshold for use. The user's medication records are stored locally in encrypted form and synchronized to the cloud only after authorization to ensure the security of the user's personal medication information.

[0085] This application also provides a computer-readable storage medium for storing the computer program used in the aforementioned point-to-point reading device. For example, the computer-readable storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0086] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in an alternative implementation, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the structure diagram and / or flowchart, and the combination of boxes in the structure diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.

[0087] In addition, the functional modules or units in the various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0088] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a smart phone, personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.

[0089] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.

Claims

1. A drug query method, characterized in that: Applied to a point reading device, the method comprises: In response to a query instruction for a drug, if the query instruction is a voice query instruction input by a user, performing voice conversion on the voice query instruction to obtain a voice query text; Performing semantic analysis on the voice query text to determine the target entity and semantic relationship of the user query; Performing a lightweight query in a drug database based on the acquired semantic relationship and the target entity to obtain a query result associated with the drug; A text answer is synthesized according to the query result, and the text answer is output in the form of voice.

2. The drug query method according to claim 1, characterized in that: Also includes: In the case where the query instruction is an interactive trigger instruction for the user to click and read the drug instructions, the query feedback information that can be perceived by the user is controlled and output according to the interactive trigger instruction.

3. The drug query method according to claim 2, characterized in that: The drug instructions are provided with an optical recognition area, and the optical recognition area is provided with a plurality of tactile codes, each of the tactile codes corresponding to unique medication information; The point reading device includes a sensor, a vibration output component, and an audio output component; the interaction trigger instruction includes: interaction information collected when the point reading device clicks the optical recognition area of ​​the drug instruction sheet; Outputting the query feedback information perceivable by the user according to the interaction trigger instruction includes: According to the interactive trigger instruction, the vibration information corresponding to the optical recognition area is provided to the user through the vibration output component, and the audio information corresponding to the medication information of the optical recognition area is provided to the user through the audio output component.

4. The drug query method according to claim 1, characterized in that: Also includes: If the target entity cannot be obtained, triggering a follow-up question mechanism to output related questions related to the voice query text through a preset follow-up question template; In response to the user's associated answer to the associated question, the voice query text is modified to re-determine a target entity based on the modified voice query text.

5. The drug query method according to claim 1, characterized in that: The voice query instruction is voice-converted to obtain a voice query text, including: performing language classification on the voice query instruction through a language classification branch of an acoustic hybrid recognition model to determine a language type; The voice query instruction is converted into a voice query text according to a language model corresponding to the language type.

6. The drug query method according to claim 5, characterized in that: The language classification branch of the acoustic hybrid recognition model classifies the voice query instruction to determine the language type, including: Extracting Mel-spectrogram features of the voice query instruction through a shared encoder of the acoustic hybrid recognition model; The voice query instruction is language-classified according to the Mel-spectrogram feature to determine the language type.

7. The drug query method according to claim 1, characterized in that: The semantic parsing of the voice query text to determine the target entity and semantic relationship of the user query includes: Convert the voice query text into a query vector using the Wenxin model, and perform similarity matching between the query vector and the entity vector in the knowledge graph to determine the target entity; The semantic relationship of the target entity in the voice query text is extracted through the Wenxin model.

8. The drug query method according to claim 1, characterized in that: The drug database includes all information of each entity. The lightweight query is performed in the drug database based on the acquired semantic relationship and the target entity to obtain the query results, including: generating a query statement according to the semantic relationship and the target entity; Performing convolution optimization based on the lightweight backbone network in the lightweight model to determine the mapping relationship between relevant entities in the drug database; All redundant attention heads that are not related to the query statement in the multi-head attention mechanism of the lightweight model are eliminated to query the information of the target entity in the drug database based on the multi-head attention mechanism, and to capture the mapping relationship between entities related to the target entity to determine the query result.

9. A point reading device, characterized in that: include: A sensor for collecting interactive information in an optical recognition area of ​​a drug package insert; A sound collection device is used to collect voice query commands input by users; a vibration output component, configured to provide the user with vibration information corresponding to the optical recognition area; an audio output device for providing audio information corresponding to medication information of the optical recognition area and outputting a voice answer; A processor for implementing the drug query method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that It stores a computer program, which, when executed on a processor, implements the drug query method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Method and system for adapting credential host to deep learning framework based on edge calculation

    CN121031726A