Drug information retrieval method, device and equipment based on multi-scene information fusion
Patent Information
- Application Number
- CN202610726054.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本申请中,通过获取用户的访问权限,不具备访问权限的用户则被拒之门外,极大的提高了访问的安全性;通过场景问题分类器对用户问题进行分类,向用户问题划分对应的问题场景,在用户问题对应的问题场景内的知识库中寻找药物信息,该方法可快速得到更加准确的药物信息,解决了现有技术中并未对应用场景进行划分,直接基于用户问题向全部的知识库中寻找答案,容易导致药物信息冗余,不准确以及效率低下的技术问题;由于各问题场景相互独立,单一问题场景的变更不会对其他问题场景产生任何干扰,即使某一特定问题场景在运行过程中出现异常或错误,也不会将错误传导至其他独立场景,从而保证了其他功能的正常运行,有效防止了单点故障的扩散,显著提升了系统整体的容错能力和持续运行的稳定性;此外,当药物业务需求变化需要新增问题场景时,系统无需进行整体重构或停机更新,仅需增量添加新的问题场景模块即可,极大的提升了系统的可扩展性与灵活部署能力
[0018]1)本申请实施例通过条件分支控制器确定用户的访问权限,不具备访问权限的用户则被条件分支器拒之门外,极大的提高了访问的安全性;通过场景问题分类器对用户问题进行分类,向用户问题划分对应的问题场景,在用户问题对应的问题场景内的知识库中寻找药物信息,该方法可快速得到更加准确的药物信息,解决了现有技术中并未对应用场景进行划分,直接基于用户问题向全部的知识库中寻找答案,容易导致药物信息冗余,不准确以及效率低下的技术问题;由于各问题场景相互独立,单一问题场景的变更不会对其他问题场景产生任何干扰,即使某一特定问题场景在运行过程中出现异常或错误,也不会将错误传导至其他独立场景,从而保证了其他功能的正常运行,有效防止了单点故障的扩散,显著提升了系统整体的容错能力和持续运行的稳定性;此外,当药物业务需求变化需要新增问题场景时,系统无需进行整体重构或停机更新,仅需增量添加新的问题场景模块即可,极大的提升了系统的可扩展性与灵活部署能力。
Smart Images

Figure CN122838461A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of intelligent drug information processing technology, and relates to a drug information retrieval method, device and equipment based on multi-scenario information fusion. Background Technology
[0002] Against the backdrop of the deepening digital transformation of enterprises, data has established its strategic position as a core production factor. To drive precise decision-making and refined operations, the need for efficient extraction of business insights from massive amounts of data is increasingly urgent at all levels of enterprises. However, traditional data analysis paradigms, which directly search the entire knowledge base for answers to user questions, suffer from significant bottlenecks in processing timeliness and analytical accuracy, making it difficult to meet the agile response requirements of modern enterprises. Therefore, improving the accuracy and efficiency of drug information retrieval and analysis has become a pressing technical problem that needs to be solved. Summary of the Invention
[0003] This application provides a drug information retrieval method, apparatus, and device based on multi-scenario information fusion, which can improve the accuracy and efficiency of drug information query and processing.
[0004] Firstly, this application provides a drug information retrieval method based on multi-scenario information fusion. The method includes: obtaining a URL parameter representing user identity information; inputting a user question into an input window of a question-and-answer service platform based on the access permissions; classifying the user question using a scenario-based question classifier within the question-and-answer service platform to determine the corresponding question scenario; retrieving a knowledge base within the question scenario corresponding to the user question, and recalling matching professional terms from the knowledge base that match the user question; wherein the number of question scenarios corresponding to the user question is one, and the number of knowledge bases within the question scenario is multiple; and then, in the knowledge bases matched within the question scenario... The process involves: recalling a first professional term that matches the user's question; sorting all the first professional terms corresponding to the question scenario to obtain a sorted second professional term; selecting a preset number of third professional terms from the second professional terms based on the actual matching score, and using a variable aggregator to fuse the third professional terms to obtain a fused matching professional term; performing semantic parsing on the matching professional term to generate a structured query language executable by the knowledge base; inputting the structured query language into the knowledge base for querying to obtain the result data output by the knowledge base corresponding to the user's question; and determining the drug information corresponding to the user's question based on the result data.
[0005] In this application, by granting users access permissions and excluding users without such permissions, access security is greatly improved. A scenario-based problem classifier categorizes user problems, assigning them to corresponding problem scenarios. Drug information is then searched for within the knowledge base of each problem scenario. This method quickly obtains more accurate drug information, solving the technical problem in existing technologies where the application scenarios are not divided, leading to redundant, inaccurate, and inefficient drug information searches across the entire knowledge base. Since each problem scenario is independent, changes to a single scenario do not interfere with other scenarios. Even if an anomaly or error occurs in a specific scenario, it will not propagate to other independent scenarios, ensuring the normal operation of other functions and effectively preventing the spread of single-point failures. This significantly improves the overall fault tolerance and continuous operational stability of the system. Furthermore, when drug business needs change and new problem scenarios are required, the system does not require overall reconstruction or downtime updates; only incremental addition of new problem scenario modules is needed, greatly enhancing the system's scalability and flexible deployment capabilities.
[0006] In one implementation of the first aspect, word segmentation is performed on the user question to generate multiple word segmentation sequences; the word segmentation sequences are mapped into high-dimensional vectors through an embedding layer to form a vector sequence that can be processed by the correction model; utilizing the bidirectional contextual understanding capability of the correction model, global context modeling is performed on each vector sequence through a self-attention mechanism to analyze the semantic association between the word segmentation sequences corresponding to each vector sequence, calculate the contextual confidence of the word segmentation sequences, and identify potential erroneous candidate word segments of the word segmentation sequences; the correct word segmentation of the erroneous candidate word segments is predicted based on a masked language model to complete the automatic correction of the erroneous candidate word segments and obtain the corrected user question.
[0007] In one implementation of the first aspect, recalling matching professional terms that match the user's question in the knowledge base includes: determining a preset matching score; vectorizing the professional terms in the knowledge base based on a vectorization model to obtain name vectors; determining the actual matching score of the name vectors that match the user's question in the knowledge base; and recalling the professional terms corresponding to the name vectors whose actual matching scores are greater than the preset matching scores as matching professional terms.
[0008] In one implementation of the first aspect, the expression corresponding to the actual matching score of the name vector matching the user question in the knowledge base is:
[0009]
[0010] Where a represents the user question vector, and b represents the name vector. Represents the first element in the user question vector. vector elements, Represents the first in the name vector There are 1 vector elements, where n represents the number of vector elements in the user question vector and the name vector.
[0011] In one implementation of the first aspect, determining the structured query language corresponding to the user question based on the matching technical terminology includes: constructing a prompt information template for the corresponding question scenario; inputting the matching technical terminology into the prompt information template to obtain a matching information template; and determining the structured query language based on the matching information template, the context information of the user question and / or the structured query language for a preset time period, and the table structure of the database within the corresponding question scenario.
[0012] In one implementation of the first aspect, obtaining user access permissions includes: obtaining a URL parameter representing user identity information, wherein the URL parameter includes at least one of a username and an organization address; and determining the access permissions corresponding to the URL parameter based on a conditional branch controller.
[0013] In one implementation of the first aspect, determining the access permissions corresponding to the URL parameter based on the conditional branch controller includes: determining whether the user has permission to access the question service platform based on the URL parameter; if yes, then inputting a user question into the input window of the question service platform; if no, then not inputting a user question into the input window of the question service platform.
[0014] In one implementation of the first aspect, determining the drug information corresponding to the user's problem based on the result data includes: if the result data output is abnormal, outputting the abnormal reason corresponding to the result data; if the result data output is normal, performing template conversion on the result data to obtain the drug information corresponding to the result data, and outputting the drug information to the user terminal.
[0015] Secondly, this application provides a drug information retrieval device based on multi-scenario information fusion, characterized in that the device includes: an access permission acquisition module for acquiring user access permissions; a user question input module for inputting user questions into an input window of a question service platform based on the access permissions; a question scenario determination module for classifying the user questions based on a scenario question classifier within the question service platform, determining the question scenario corresponding to the user questions, wherein the number of question scenarios corresponding to the user questions is at least one; and a drug information determination module for retrieving a knowledge base within the question scenario corresponding to the user questions, and retrieving matching professional terms that match the user questions from the knowledge base; wherein the number of question scenarios corresponding to the user questions is one, and the question... If there are multiple knowledge bases within a given problem scenario, then the first professional term matching the user's question is retrieved from the knowledge bases matched within the problem scenario. All the first professional terms corresponding to the problem scenario are sorted to obtain the sorted second professional terms. Based on the actual matching score, a preset number of third professional terms are selected from the second professional terms, and a variable aggregator is used to fuse the third professional terms to obtain the fused matching professional terms. Semantic parsing is performed on the matching professional terms to generate a structured query language executable by the knowledge base. The structured query language is input into the knowledge base for querying, and the result data output by the knowledge base corresponding to the user's question is obtained. Based on the result data, the drug information corresponding to the user's question is determined.
[0016] Thirdly, embodiments of this application provide an electronic device, the electronic device comprising: a memory storing a computer program; and a processor communicatively connected to the memory, which, when the computer program is invoked, executes the drug information retrieval method based on multi-scenario information fusion as described in any one of the first aspects of this application.
[0017] As described above, the drug information retrieval method, apparatus, and device based on multi-scenario information fusion described in this application have the following beneficial effects:
[0018] 1) This application embodiment determines user access permissions through a conditional branch controller. Users without access permissions are denied access by the conditional branch controller, greatly improving access security. A scenario problem classifier categorizes user problems, assigning corresponding problem scenarios. Drug information is then searched in the knowledge base within the corresponding problem scenario. This method quickly obtains more accurate drug information, solving the technical problem in existing technologies where application scenarios are not divided, and answers are directly searched in the entire knowledge base based on the user problem, leading to redundant, inaccurate, and inefficient drug information. Since each problem scenario is independent, changes in a single problem scenario will not interfere with other problem scenarios. Even if an anomaly or error occurs in a specific problem scenario during operation, the error will not propagate to other independent scenarios, ensuring the normal operation of other functions and effectively preventing the spread of single-point failures. This significantly improves the overall fault tolerance and continuous operational stability of the system. Furthermore, when drug business requirements change and new problem scenarios need to be added, the system does not require overall reconstruction or downtime updates; only incremental addition of new problem scenario modules is needed, greatly improving the system's scalability and flexible deployment capabilities.
[0019] 2) This application developed an intelligent data query solution based on the Dify platform to meet actual business needs and has been implemented. It has realized the core AI function of text understanding and data acquisition, which has lowered the threshold for ordinary users to obtain data and has important practical significance.
[0020] 3) This application leverages the powerful orchestration capabilities of the Dify platform to construct and deliver an intelligent data analysis solution deeply aligned with practical business applications. The system has been successfully implemented and integrated with core natural language processing technologies, achieving precise mapping from natural language intent to structured data queries. This application effectively breaks down technical barriers to data acquisition, significantly lowering the threshold for end-users to access data, and has significant practical implications. When determining the drug information corresponding to the user's question, this embodiment sorts all the first professional terms corresponding to the question scenario to obtain sorted second professional terms. Based on the actual matching score, a preset number of third professional terms are selected from the second professional terms, and a variable aggregator is used to fuse these third professional terms to obtain fused matching professional terms, providing an accurate foundation for determining the drug information corresponding to the user's question.
[0021] 4) This application embodiment determines a preset matching score and the actual matching score of the vector in the knowledge base that matches the user question. It recalls name vectors with a match score greater than the preset matching score, effectively filtering all vectors in the knowledge base. This avoids recalling vectors in the knowledge base that do not match the user question, greatly improving the accuracy of filtering matching professional terms related to the user question in the knowledge base, and providing accurate matching professional terms for obtaining drug information corresponding to the user question in the future.
[0022] 5) The matching information template in this application embodiment can prompt or guide the specific execution process of the model. For different matching professional terms, the corresponding matching information template may be different, which reflects the uniqueness and exclusivity of the matching information template, so that the corresponding model can quickly and accurately obtain the structured query language under the guidance of the matching information template. Attached Figure Description
[0023] Figure 1 The flowchart shown is a drug information retrieval method based on multi-scenario information fusion provided in an embodiment of this application.
[0024] Figure 2 The flowchart shown is for determining matching technical terms as provided in the embodiments of this application.
[0025] Figure 3 The flowchart shown is a process for determining a structured query language as provided in an embodiment of this application.
[0026] Figure 4 The flowchart shown is a process for determining access permissions as provided in an embodiment of this application.
[0027] Figure 5 The flowchart shown is provided in an embodiment of this application for determining the drug information corresponding to the user problem.
[0028] Figure 6 The diagram shows the overall flowchart of the drug information retrieval method based on multi-scenario information fusion provided in the embodiments of this application.
[0029] Figure 7 The diagram shown is a structural diagram of a drug information retrieval device based on multi-scenario information fusion provided in an embodiment of this application.
[0030] Figure 8 The diagram shown is a structural diagram of an electronic device provided in an embodiment of this application.
[0031] Component designation explanation
[0032] S11~S17 step 74 Drug Information Determination Module S21~S24 step 75 Structured Query Language Generation Module S31~S33 step 76 Result Data Acquisition Module S41~S43 step 77 Drug Information Determination Module S51~S52 step 80 electronic devices 70 Drug information retrieval device based on multi-scenario information fusion 81 processor 71 Access Permission Acquisition Module 82 Non-volatile storage media 72 User Question Input Module 83 System bus 73 Problem scenario determination module 84 Internal memory 85 Network interface Detailed Implementation
[0033] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.
[0034] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. Therefore, the drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0035] The following embodiments of this application provide a drug information retrieval method, apparatus, and device based on multi-scenario information fusion, including but not limited to the hardware application scenarios listed in these embodiments.
[0036] The technical solutions in the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0037] like Figure 1 As shown in the figure, this application provides a flowchart of a drug information retrieval method based on multi-scenario information fusion, as follows: Figure 1 As shown, the drug information retrieval method based on multi-scenario information fusion provided in this application includes the following steps S11 to S17.
[0038] S11, Obtain user access permissions.
[0039] In some embodiments, obtaining user access permissions includes: obtaining a URL parameter representing user identity information, wherein the URL parameter includes at least one of a username and an organization address; and determining the access permissions corresponding to the URL parameter based on a conditional branch controller.
[0040] The URL parameter includes at least one of the following: username and organization address.
[0041] For example, the URL parameter corresponding to the username is represented as zhangsan, and the URL parameter corresponding to the organization address is represented as dept_tech_001.
[0042] It should be noted that the corresponding URL parameters are different for different users, and this application does not impose any restrictions on this.
[0043] Specifically, the conditional branch controller is configured to determine whether a user has the permission to access the data query service. If the user has the permission, the conditional branch controller allows the user to enter the subsequent process; if not, a response is returned, for example, an output: Hello user, you have not activated the data query permission, please contact the administrator.
[0044] In some embodiments, the method further comprises: performing a word segmentation operation on the user question to generate a plurality of word segmentation sequences;
[0045] mapping the word segmentation sequences into high-dimensional vectors through an embedding layer to form vector sequences processable by a correction model; utilizing the bidirectional context understanding capability of the correction model, performing global context modeling on each of the vector sequences through a self-attention mechanism, analyzing the semantic association of the word segmentation sequences corresponding to each of the vector sequences, calculating the context confidence of the word segmentation sequences, and identifying potential error candidate word segmentations in the word segmentation sequences; predicting correct word segmentations for the error candidate word segmentations based on a masked language model, so as to complete automatic correction of the error candidate word segmentations and obtain a corrected user question.
[0046] For example, the correction model may be Bidirectional Encoder Representations from Transformers (BERT) based on Transformer, or a generative pre-trained large model (such as GPT-4, ChatGLM, Qwen, etc.).
[0047] It should be noted that when the input user question contains typos, the typos in the user question can be corrected based on the technical solution in the foregoing embodiments.
[0048] For example, if the initially input user question is "Is there any stock of yellow flag particles", the typos in the user question can be corrected based on the technical solution in the foregoing embodiments, and the specific steps are as follows:
[0049] 1) Input preprocessing stage (word segmentation and tokenization)
[0050] the correction model first performs a word segmentation operation on the original input character string "Is there any stock of yellow flag particles" to generate an initial word segmentation sequence ["Huang", "Qi", "Ke", "Li", "You", "Mei", "You", "Ku", "Cun"]. Subsequently, a classification token <[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]> and a separator token [SEP] are spliced at the head and tail of the sequence respectively. The output is: <[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]> Huang Qi Ke Li You Mei You Ku Cun [SEP].[end]]
[0051] 2) Vector representation conversion stage (embedding layer mapping)
[0052] Input the above word segmentation sequence containing 11 elements into the embedding layer. Through a table lookup operation, the correction model maps each word segmentation sequence into a word embedding vector of a fixed dimension (e.g., 768 dimensions). Meanwhile, the position information of each word segmentation (1st position, 2nd position, ...) is converted into a position embedding vector, and the sentence attribution information is converted into a word segmentation type embedding vector. The three are added to obtain a high-dimensional vector integrating semantics, position and structure. This step outputs a floating-point matrix with a shape of (11, 768), which is sent to the Transformer encoder.
[0053] 3) Contextual semantic analysis stage (self-attention and error detection)
[0054] The vector sequence passes through a multi-layer Transformer network, and the correction model calculates the association strength between word segments through a self-attention mechanism, and outputs the probability distribution (confidence) of each word segment in the current context. Through global context capture, the self-attention mechanism enables the vector of the Chinese character "qi" to not only contain its own information, but also deeply integrate the information of the preceding and following context. The correction model finds that "ke", "li" and "stock" constitute a strong context of "commodity / drug name + price inquiry".
[0055] Confidence evaluation: the correction model calculates the rationality score of the Chinese character "qi" in the current context at the output layer. Since in medical domain corpora, the most common "Huang X Ke Li" is "Huang Qi Ke Li", and "Huang Qi Ke Li" (where Qi means flag) almost does not exist in the corpus (or has an extremely low probability), the correction model determines that the contextual confidence of the Chinese character "qi" is lower than a preset threshold. Output of this step: the error candidate word segmentation is accurately positioned as the Chinese character "qi" at the 3rd position in the sequence.
[0056] 4) Error correction and output stage (mask prediction mechanism)
[0057] Execution action: the correction model logically "masks" the error candidate position (regarded as [MASK]), and performs predictive correction in combination with the Masked Language Model (MLM) mechanism.
[0058] Mask context construction: the correction model uses the hidden layer state that has integrated correct context information in step 3 to calculate the conditional probability for the 3rd position. The implicit context at this time is equivalent to: <[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]> Huang [MASK] ke li you mei you ku cun <[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]>.
[0059] Probabilistic decoding: The MLM head of the correction model maps the hidden state vector at this position back to the softmax probability distribution of the entire vocabulary (e.g., 21128 Chinese characters and symbols). The calculation result shows that: given the condition of "Huang…qi ke li you mei you ku cun" (literally: "Yellow…qi granules are in stock or not"), the probability that the next character is "Qi (Astragalus)" (e.g., 98.5%) is much higher than that of the character "Qi (flag)" (e.g., 0.01%).
[0060] Word segmentation replacement: The model replaces the original character "Qi (flag)" with the character "Qi (Astragalus)", which has the highest probability, and strips special tokens.
[0061] Finally, the corrected user question "Is there any stock of Astragalus granules?" is generated, thereby eliminating the interference of typos on the subsequent inventory retrieval system and accurately identifying the user's real query intention.
[0062] It should be noted that the embodiments of the present application further include: establishing a closed-loop mechanism from "user interaction -> data query -> result feedback -> correction model fine-tuning", recording the user's correction behaviors and satisfaction, and continuously optimizing the effect of the correction model based on this.
[0063] S12, inputting the user question into the input window of the query service platform based on the access permission.
[0064] Wherein, the query service platform is the Dify platform.
[0065] For example, the user questions include: "Is there any stock of Astragalus granules?", "What is the warehouse location and inventory quantity of Astragalus granules?", "Which manufacturers produce Astragalus granules?" and the like.
[0066] It should be noted that in practical applications, users can input any appropriate question into the input window of the query service platform based on specific application requirements, which is not limited in the present application.
[0067] S13, classifying the user question based on the scene question classifier in the query service platform, and determining the question scene corresponding to the user question.
[0068] In some embodiments, question scenes include but are not limited to inventory scene, accounts receivable collection query scene and sales order query scene.
[0069] For example, the question scenes include: 1. Inventory scene: query dimensional data including product batch / lot number, specification, cargo owner, inventory amount / quantity, warehouse, manufacturer, expiration date, dosage form, supplier, days in stock, as well as unit inventory cost, (tax-inclusive / tax-exclusive) purchase price and other dimensional data. 2. Accounts receivable collection query scene: query data including customer receivables, overdue receivables, receivables overdue for more than one year, current month collection amount, current year collection amount, account age and other data. 3. Sales order query scene: query customer, product, order status, quantity, amount, product name, specification, manufacturer.
[0070] For example, if a user's question is "the manufacturer and inventory quantity of Astragalus membranaceus granules", the scenario question classifier will automatically classify the user's question and determine that the corresponding scenario is: inventory scenario.
[0071] For example, if the user's question is "order status of Astragalus Granules", the scenario question classifier will automatically classify the user's question and determine that the application scenario corresponding to the user's question is: sales order scenario.
[0072] It should be noted that the problem scenarios listed in the above examples are only for illustrative purposes. In actual applications, the number of problem scenarios that the scenario problem classifier can classify for user problems can be increased or decreased, and this application does not impose any restrictions on this.
[0073] S14, retrieve the knowledge base within the problem scenario corresponding to the user's question, and recall the matching professional terms that match the user's question in the knowledge base.
[0074] In some embodiments, there are multiple knowledge bases within the problem scenario; then, a first professional term matching the user's question is recalled from the knowledge bases matched within the problem scenario; all the first professional terms corresponding to the problem scenario are sorted to obtain a sorted second professional term; based on the actual matching score, a preset number of third professional terms are selected from the second professional terms, and a variable aggregator is used to fuse the third professional terms to obtain a fused matching professional term.
[0075] It should be noted that the knowledge base in the problem scenario includes a knowledge base for company names, customer names, and a name knowledge base for abbreviations of full drug names. Considering that the database stores complete names, it is necessary to use the knowledge base as a tool to match and obtain the complete names.
[0076] For example, the method provided in this application can be applied to accounts receivable management scenarios in medical institutions. Taking the accounts receivable processing of a target medical institution (e.g., "The First Affiliated Hospital of Zhejiang University") as an example: If the user's question is "The accounts receivable status of the First Affiliated Hospital of Zhejiang University", and the corresponding question scenario is an accounts receivable inquiry scenario, then the specific steps for recalling matching professional terms that match the user's question "The accounts receivable status of the First Affiliated Hospital of Zhejiang University" in the knowledge base within the accounts receivable inquiry scenario are as follows:
[0077] Step 1: Conduct multi-path recall within multiple knowledge bases related to the accounts receivable issue to obtain the primary technical terminology.
[0078] After identifying the "accounts receivable inquiry scenario," text retrieval is performed on the user's question "Accounts receivable status of the First Affiliated Hospital of Zhejiang University" based on multiple knowledge bases configured for this scenario (including: company name knowledge base, customer name knowledge base, and name knowledge base for abbreviations of drug full names, etc.). Retrieval in the "Customer Name Knowledge Base": Based on string similarity or semantic vectors, candidate words matching "First Affiliated Hospital of Zhejiang University" are retrieved, such as "The First Affiliated Hospital of Zhejiang University School of Medicine (Qingchun Campus)" and "The First Affiliated Hospital of Zhejiang University School of Medicine (Headquarters)". Retrieval in the "Company Name Knowledge Base": Abbreviations of related companies with accounts receivable settlement relationships with this customer are retrieved, such as "Zhejiang University Hospital Management Co., Ltd." Retrieval in the "Name Knowledge Base for Abbreviations of Drug Full Names, etc.": Although word segmentation detection does not hit the core drug term, low-scoring noise words may be generated due to contextual general retrieval, such as the abbreviation of a drug called "The Same Anti-inflammatory Drug A as at the First Affiliated Hospital of Zhejiang University" (assuming it exists). The following are the first professional terms: ["The First Affiliated Hospital of Zhejiang University School of Medicine (Qingchun Campus)", "The First Affiliated Hospital of Zhejiang University School of Medicine (Headquarters)", "Zhejiang University Hospital Management Co., Ltd.", "Anti-inflammatory drug A of the same type as Zhejiang University First Affiliated Hospital"].
[0079] Step 2: Sort all the first-level technical terms to obtain the second-level technical terms.
[0080] The system sorts the above first professional terms in descending order based on the actual matching score (taking into account edit distance, word frequency weight, knowledge base category weight, etc., with "customer name knowledge base" having the highest weight in the accounts receivable scenario):
[0081] The results are: “The First Affiliated Hospital of Zhejiang University School of Medicine (Qingchun Campus)” (matching score: 0.95), “The First Affiliated Hospital of Zhejiang University School of Medicine (Headquarters)” (matching score: 0.92), “Zhejiang University Hospital Management Co., Ltd.” (matching score: 0.75), “The same anti-inflammatory drug A as used by the First Affiliated Hospital of Zhejiang University” (matching score: 0.30), which are the sorted “second professional terms” (arranged from highest to lowest score).
[0082] Step 3: Perform Top-K filtering based on matching scores to obtain the third professional term.
[0083] The system sets a preset number (assuming K=3) and performs hard truncation based on the actual matching score in step two, removing noise words with too low a score and irrelevant to the intention of "receivables": removing "the same anti-inflammatory drug A used by the First Affiliated Hospital of Zhejiang University" with a score of 0.30.
[0084] The following "third professional terms" were obtained through filtering: ["The First Affiliated Hospital of Zhejiang University School of Medicine (Qingchun Campus)", "The First Affiliated Hospital of Zhejiang University School of Medicine (Headquarters)", "Zhejiang University Hospital Management Co., Ltd."].
[0085] Step 4: Use a variable aggregator to perform fusion processing to obtain the fused matching professional terms.
[0086] Input the aforementioned third-party terminology into the "Variable Aggregator" for logical fusion. The aggregator executes the following rules:
[0087] Synonymous / Hierarchical Entity Merging: It was identified that "Qingchun Campus" and "Headquarters" are different campuses of "The First Affiliated Hospital of Zhejiang University School of Medicine". In the case of accounts receivable, it is usually necessary to merge the queries to obtain the total accounts receivable of the institution. Therefore, they are merged into a standard institutional entity.
[0088] Remove non-target entities: "Zhejiang University Hospital Management Co., Ltd." is identified as a management company rather than an entity that directly generates medical receivables, and is downgraded or removed based on the intent of the scenario.
[0089] The variable aggregator ultimately outputs the merged, fully consistent name that conforms to the underlying database storage format. The resulting "Merged Matching Terminology" is: ["The First Affiliated Hospital of Zhejiang University School of Medicine"].
[0090] By combining the above-mentioned knowledge bases, the user's colloquial abbreviation "The First Affiliated Hospital of Zhejiang University" can be accurately converted into the full standard name "The First Affiliated Hospital of Zhejiang University School of Medicine" stored in the database. Subsequently, this fused matching professional term can be directly used as an SQL query condition (such as WHERE customer_name = 'The First Affiliated Hospital of Zhejiang University School of Medicine'), thereby accurately recalling the hospital's accounts receivable data.
[0091] For example, suppose a user inputs the question "What medications are available for treating hypertension?" into the question data service platform. After analysis by the scenario question classifier, this question corresponds to question scenario B: "symptom-medication" association (the association between accompanying symptoms of hypertension, such as "dizziness," and medications). It is necessary to independently retrieve the first professional terms that match the user's question from the knowledge base corresponding to question scenario B. These first professional terms are the "local matching results" within question scenario B.
[0092] Specifically, in the multiple knowledge bases configured in problem scenario B, the following recall is performed: In the "Disease and Symptom Standard Name Knowledge Base," words related to diseases and accompanying symptoms directly related to the question are recalled, resulting in: hypertension (match score 0.9), dizziness (match score 0.7); In the "Drug Full Name and Abbreviation Name Knowledge Base," considering that the underlying database stores the standard full name of the drug, while users often use common names or abbreviations in their questions, the system performs semantic mapping recall through this knowledge base, resulting in: antihypertensive drugs (match score 0.85), antihypertensive drugs (match score 0.75). Therefore, the sum of the first professional terms matching the user's question recalled by multiple knowledge bases in problem scenario B is: hypertension (0.9), antihypertensive drugs (0.85), antihypertensive drugs (0.75), and dizziness (0.7).
[0093] All first-order terms are globally sorted according to their actual matching scores (reflecting their relevance to the user's question), resulting in the ordered second-order terms. Assume the actual matching scores are sorted from highest to lowest as follows: hypertension (0.9) > antihypertensive drugs (0.85) > antihypertensive medications (0.75) > dizziness (0.7). The resulting list of ordered second-order terms is: [hypertension, antihypertensive drugs, antihypertensive medications, dizziness].
[0094] Based on the actual matching score, a preset number of third professional terms are selected from the second professional terms. Assuming the preset number is 3, the top 3 are selected: hypertension, antihypertensive drugs, and antihypertensive drugs. That is, the third professional terms are: hypertension, antihypertensive drugs, and antihypertensive drugs.
[0095] Finally, the variable aggregator is used to merge the third professional terms. For example, the synonyms "antihypertensive drugs" and "antihypertensive drugs" are deduplicated and merged to obtain the merged matching professional terms: "hypertension, antihypertensive drugs", which can be directly input into the underlying database for accurate querying.
[0096] It should be noted that the values listed in the above examples are merely illustrative and this application does not impose any limitations on them.
[0097] It should be noted that the system automatically identifies the user's query scenario based on a large classifier model. Currently, three scenarios are available. Compared to the previous approach of putting all scenarios into a single prompt project, the classifier-based architecture not only reduces processing time (by about 3 seconds), but more importantly, it allows for the addition of multiple scenarios.
[0098] S15, perform semantic parsing on the matched professional terms to generate a structured query language executable by the knowledge base.
[0099] Wherein, Structured Query Language can be represented by (Structured Query Language, sql).
[0100] For example, semantic parsing can be performed on the matched professional nouns based on the Text2SQL model to obtain an initial structured query language. Since the Text2SQL model outputs the initial structured query language based on probability, it often contains grammatical flaws or structures that do not conform to specific database dialects, while the knowledge base engine requires strict grammatical rules. After obtaining the initial structured query language, a Python script can be used to detect and repair grammatical errors in the initial structured query language, and output a structured query language executable by the knowledge base.
[0101] For example, if the user's question is: "Manufacturer and stock quantity of Astragalus Granules", the matched professional nouns are: "Astragalus Granules", "Manufacturer", "Stock quantity".
[0102] The Text2SQL model (e.g., Qwen3) performs semantic parsing according to the user's intention and possible context, and generates the following SQL statement.
[0103] Initial Structured Query Language:
[0104] SELECT manufacturer,stock_qty FROM t_drug_inventory WHERE drug_name = Astragalus Granules, which has the following grammatical flaws: the model correctly recognized that the query intention is to obtain "manufacturer" and "stock quantity", and correctly mapped the table name and field name. However, since the model generates results based on probability, it ignores the string reference rules in SQL syntax specifications. Error point: WHERE drug_name = Astragalus Granules. Consequence: In the SQL standard, if not wrapped by single quotes, the database engine will mistakenly parse "Astragalus Granules" as a column name or keyword, resulting in a query error (Unknown column 'Astragalus Granules').
[0105] The Python script intervenes as a post-processor to perform syntax tree analysis or regular matching on the initial SQL and execute the repair logic.
[0106] Detection and Repair Process: Type Detection: The script queries the database metadata and finds that the `drug_name` field is of type VARCHAR (string). Syntax Validation: The script finds that the name "Huangqi Granules" to the right of the equals sign in the WHERE clause is not enclosed in quotation marks and does not conform to column name naming conventions (containing Chinese characters that are not reserved words). Automatic Repair: The script automatically adds single quotes to both sides of the string value and performs escaping to prevent injection. Format Standardization: A semicolon ";" is added at the end of the statement to ensure its completeness. Output: Executable structured query language for the knowledge base: SELECT manufacturer, stock_qty FROM t_drug_inventory WHERE drug_name = 'Huangqi Granules'.
[0107] It should be noted that the Text2SQL model for outputting the initial structured query language and the Python script for detecting and correcting syntax errors in the initial structured query language listed in the above examples are merely illustrative. In practical applications, other suitable models can be selected to output the initial structured query language and other suitable scripts can be used to detect and correct syntax errors in the initial structured query language based on specific application requirements. This application does not impose any restrictions on this.
[0108] S16, input the structured query language into the knowledge base to perform a query operation, and obtain the result data output by the knowledge base corresponding to the user's question.
[0109] For example, the knowledge base parses the structured query language above and determines that a search needs to be performed in the t_drug_inventory table. The engine matches the string 'Astragalus Granules' in the drug_name column, locates the row record that meets the condition, and extracts the data from the manufacturer and stock_qty columns of that row. The knowledge base outputs the query results.
[0110] S17, Determine the drug information corresponding to the user's problem based on the result data.
[0111] For example, the steps for determining the drug information corresponding to the user's problem based on the result data are as follows:
[0112] 1. Results Data Acquisition and Analysis
[0113] The system receives the raw result set returned by the knowledge base, which is usually in the form of key-value pairs or data structures. For example, the raw data format is: { "drug_name": "Astragalus Granules", "manufacturer": "Yunnan Baiyao Group", "stock_qty": 1500}.
[0114] 2. Field semantic mapping
[0115] Since database field names (such as manufacturer, stock_qty) are usually technical definitions, the system needs to map them to business terms that users can understand (i.e., attribute labels for "drug information"). For example, map the field manufacturer to the business attribute "manufacturer"; map the field stock_qty to the business attribute "inventory quantity"; and retain the query keyword drug_name as the core entity "drug name".
[0116] 3. Data validation and null value handling
[0117] Before determining drug information, the system verifies the integrity of the data:
[0118] If the result data is empty (i.e. no record was found in the knowledge base), a prompt message will be generated: "No relevant inventory information for 'Astragalus Granules' was found at present."
[0119] If the result data exists, its validity is further checked. For example, if stock_qty is NULL, it is converted to "No data available" or the default value of 0 to avoid abnormal information display.
[0120] 4. Structured encapsulation of drug information
[0121] The system encapsulates the parsed data according to a pre-defined drug information display model, generating a final drug information object. This object contains not only data values but also a semantic description of the data.
[0122] 5. Results Output and Presentation
[0123] Finally, based on the above structured encapsulation, the system generates drug information in natural language and feeds it back to the user:
[0124] The search results are as follows: The manufacturer of the medicine 'Astragalus Granules' is 'Yunnan Baiyao Group', and the current inventory is 1,500 boxes.
[0125] This application provides a drug information retrieval method based on multi-scenario information fusion. In this method, by determining user access permissions, users without access rights are denied access, greatly improving access security. A scenario problem classifier categorizes user questions, assigning them to corresponding problem scenarios. Drug information is then searched in the knowledge base within the corresponding problem scenario. This method can quickly obtain more accurate drug information, solving the technical problems of existing technologies that directly search the entire knowledge base based on user questions, easily leading to redundant, inaccurate, and inefficient drug information. Since each problem scenario is independent, changes in a single problem scenario will not interfere with other problem scenarios. Even if an anomaly or error occurs in a specific problem scenario during operation, the error will not be reported. This mechanism extends to other independent scenarios, ensuring the normal operation of other functions, effectively preventing the spread of single-point failures, and significantly improving the overall fault tolerance and continuous operational stability of the system. By sorting all the first professional terms corresponding to the problem scenario, a sorted second professional term is obtained. Based on the actual matching score, a preset number of third professional terms are selected from the second professional terms, and a variable aggregator is used to fuse the third professional terms to obtain the fused matching professional terms, providing an accurate matching professional terminology basis for subsequently determining the drug information corresponding to the user's problem. In addition, when changes in drug business requirements necessitate the addition of new problem scenarios, the system does not require overall reconstruction or downtime updates; only incremental addition of new problem scenario modules is needed, greatly improving the system's scalability and flexible deployment capabilities.
[0126] like Figure 2 As shown in the figure, this application provides a flowchart for determining matching technical terms, such as... Figure 2 As shown, the method for determining matching technical terms provided in this application includes the following steps S21 to S24.
[0127] S21, Determine the preset matching score.
[0128] For example, the preset matching score can be set to any suitable value such as 0.2 or 0.3, and this application does not limit it.
[0129] S22, Based on the vectorization model, the professional terms in the knowledge base are vectorized to obtain name vectors.
[0130] For example, technical terms can be stored in vector form based on the text-embedding-v3 vectorization model.
[0131] It should be noted that the vectorization models listed in the above examples are only for illustrative purposes. In practical applications, other suitable vectorization models can be selected based on specific application requirements, and this application does not impose any restrictions on this.
[0132] S23, in the knowledge base, determine the actual matching score of the name vector that matches the user question.
[0133] Specifically, within the knowledge base of the problem scenario corresponding to the user's question, the actual matching score of the vector matching the user's question can be automatically calculated.
[0134] Specifically, the knowledge base stores specific business knowledge, such as the full names of all companies. However, in practice, the full name of a company is retrieved by indexing its abbreviation. In other words, the user's question is matched with the knowledge base for index retrieval. For example, to find the inventory status of AA Co., Ltd., the question needs to retrieve the full company name as AA Co., Ltd. so that the generated SQL statement is accurate.
[0135] In some embodiments, the expression for determining the actual matching score of the name vector matching the user question in the knowledge base is:
[0136]
[0137] Where a represents the user question vector, and b represents the name vector. Represents the first element in the user question vector. vector elements, Represents the first in the name vector There are 1 vector elements, where n represents the number of vector elements in the user question vector and the name vector.
[0138] For example, if the user's question vector is Hebei First Hospital, a = [0.7, 0.5, 0.2], the corresponding name vectors in the knowledge base are: Guangzong County Tangtuan Central Health Center A = [0.9, 0.1, 0.2]; Cangzhou People's Hospital B = [0.1, 0.9, 0.1]; Tangshan Fengrun District Traditional Chinese Medicine Hospital C = [0.3, 0.4, 0.8]; Hebei Medical University First Hospital D = [0.7, 0.5, 0.2]; Quzhou County Quzhou Town Health Center E = [0.1, 0.7, 0.3].
[0139] Based on the above expression, the actual matching scores of the user question vector with the name vectors A, B, C, D and E are 0.008, 0.672, 0.684, 0.882 and 0.708, respectively.
[0140] Through matching and calculation, the complete name of Hebei First Hospital in the user's question was finally retrieved: Hebei Medical University First Hospital.
[0141] For example, when actually searching for the vector corresponding to a user's question, the user's question can be embedded (vectorized) into a vector. For instance, when searching for the inventory status of Company AA, the user's question can be vectorized into a question vector a=(0.5, 0.8, 0.9). All the full names of companies in the knowledge base are stored as vectors. For example, AA Co., Ltd. is stored as vector b=(1, 2, 3) in the knowledge base, and BB Co., Ltd. is stored as vector c=(4, 5, 6). Based on the above expressions, the actual matching score corresponding to AA Co., Ltd. is 0.9839, and the actual matching score corresponding to BB Co., Ltd. is 0.9960.
[0142] It should be noted that the actual matching score ranges from 0 to 1.
[0143] S24, the name vector whose actual matching score is greater than the preset matching score is used as the matching professional term for recall.
[0144] For example, if the actual matching score of a name vector in the knowledge base is 0.1, then the actual matching score of the name vector is less than the preset matching score, and the system considers the name vector to be a mismatch with the user's question, and the name vector will not be recalled. Conversely, if the actual matching score of a name vector in the knowledge base is 0.3, then the actual matching score of the name vector is greater than the preset matching score, and the system considers the name vector to be a match with the user's question, and the professional term corresponding to the name vector will be recalled as a matching professional term.
[0145] This application provides a method for determining matching professional terms. In this method, by determining a preset matching score and the actual matching score of the name vectors in the knowledge base that match the user's question, name vectors with a matching score greater than the preset matching score are recalled. This effectively filters all name vectors in the knowledge base, avoiding the recall of name vectors in the knowledge base that do not match the user's question. This significantly improves the accuracy of filtering matching professional terms related to the user's question in the knowledge base, and provides accurate matching professional terms for obtaining drug information corresponding to the user's question.
[0146] like Figure 3 As shown, this application provides a flowchart for determining a structured query language, such as... Figure 3 As shown, the method for determining a structured query language provided in this application includes the following steps S31 to S33.
[0147] S31, Construct a prompt message template for the corresponding problem scenario.
[0148] For example, in a pharmaceutical scenario, the prompt template would be: "When writing SQL, please pay attention to the relationship between the order table and the user table."
[0149] S32, input the matching technical terms into the prompt information template to obtain the matching information template.
[0150] For example, if the matching technical terms are "XX warehouse", "inventory quantity", and "batch number", the matching technical terms are entered into the prompt information template. The resulting matching information template will import the inventory table and warehouse table, clarify the connection logic, and add a default filter condition of "inventory greater than 0".
[0151] For example, based on matching the technical terms "XX warehouse", "inventory quantity", and "batch number", the corresponding matching information template is as follows:
[0152] ### Please generate the SQL statement to query the inventory status based on the user's question.
[0153] ### Database table structure
[0154] Table name: drug_inventory (Drug inventory table)
[0155] Field description:
[0156] - inventory_id (Inventory ID, INT)
[0157] - drug_id (drug ID, INT)
[0158] - warehouse_id (warehouse ID, INT)
[0159] - batch_no (batch number, VARCHAR)
[0160] - quantity (inventory quantity, INT)
[0161] - expiry_date (expiration date, DATE)
[0162] Table name: warehouse_info (warehouse information table)
[0163] Field description:
[0164] - warehouse_id (warehouse ID, INT)
[0165] - warehouse_name (warehouse name, VARCHAR)
[0166] ### Business Rule Constraints
[0167] 1. When querying inventory, by default only records with quantity > 0 are filtered, excluding out-of-stock records.
[0168] 2. If the user does not specify a warehouse, query the inventory status of all warehouses.
[0169] 3. The results must include the drug name and warehouse name. Please join the corresponding dimension tables.
[0170] ### User Issues
[0171] {user_question}
[0172] ### Output Requirements
[0173] Output the standard SQL query statement.
[0174] It should be noted that the matching technical terms and matching information templates listed in the above examples are merely illustrative and this application does not impose any limitations on them.
[0175] S33. Determine the structured query language based on the matching information template, the user question and / or the context information of the structured query language for the preset time period, and the table structure of the database within the corresponding question scenario.
[0176] Specifically, the matching information template, user questions, and / or the context information of the structured query language for a preset time period, as well as the table structure of the database within the corresponding question scenario, can be input into the qwen-plus-latest (Qwen3) large model in the text2sql task workflow. The matching information template can guide the Qwen3 large model to execute specific execution steps to obtain the structured query language.
[0177] The classifier model in this application can also use qwen3-30b-a3b-instruct-2507, which reduces the time to milliseconds compared to the original qwen2-plus model, compared to about 5 seconds. In the text2sql part, qwen-plus-latest (Qwen3) is used, which reduces the time to about 5 seconds compared to the qwen2-plus model. Based on architecture optimization and large model adjustment, the current end-to-end time of the query data server is 6-10 seconds, compared to more than 20 seconds, which is a significant improvement.
[0178] For example, the structured query language obtained by adding contextual information can better match the actual application scenario and is more accurate.
[0179] It should be noted that the matching information template may differ for different technical terms.
[0180] This application provides a method for determining a structured query language. In this method, a prompt information template corresponding to a problem scenario is constructed; the matching technical terms are input into the prompt information template to obtain a matching information template; the structured query language is determined based on the matching information template, the user question, and / or the context information of the structured query language within a preset time period, as well as the table structure of the database within the corresponding problem scenario. The matching information template can prompt or guide the specific execution process of the model. For different matching technical terms, the corresponding matching information template may be different, reflecting the uniqueness and exclusivity of the matching information template, enabling the corresponding model to quickly and accurately obtain the structured query language under the guidance of the matching information template.
[0181] like Figure 4 As shown, this application embodiment provides a flowchart for determining access permissions, such as... Figure 4 As shown, the method for determining access permissions provided in this application includes the following steps S41 to S43.
[0182] S41, determine whether the user has the right to access the question data service platform based on the URL parameter.
[0183] For example, if the URL parameter corresponding to the username is lisa and the URL parameter corresponding to the organization address is dept_tech_002, the URL parameter corresponding to the username and / or the URL parameter corresponding to the organization address are input into the conditional branch controller, and the conditional branch controller determines whether the user has the right to access the question data service platform.
[0184] S42, if so, then input the user's question into the input window of the question service platform.
[0185] For example, if the conditional branch controller determines that the user has the right to access the question service platform, it can redirect to the question service platform, where the user can enter their question into the input window of the question service platform.
[0186] S43, if not, then do not enter user questions into the input window of the question service platform.
[0187] For example, if the conditional branch controller determines that the user has permission to access the question and answer service platform, it will output "Hello, user, you have not activated your question and answer permissions. Please contact the administrator" on the corresponding display interface of the conditional branch controller. The user will then be unable to access the question and answer service platform, and will be unable to enter any questions into the platform's input window.
[0188] This application provides a method for determining access permissions. In this method, at least one of the URL parameter, the username, and the organization address is used to determine whether a user has permission to access the question-and-answer service platform. If so, the user enters a question into the input window of the question-and-answer service platform; otherwise, the user does not enter a question into the input window. This method uses a conditional branch controller to determine the user's access permissions. Users without access permissions are denied access by the conditional branch controller, and only users with access permissions are allowed to enter, greatly improving access security. Furthermore, all permission decisions are made by the conditional branch controller, establishing a single source of trust, eliminating logical conflicts, and simplifying the access control process.
[0189] like Figure 5 As shown in the figure, this application embodiment provides a flowchart for determining the drug information corresponding to the user's problem, such as... Figure 5 As shown in the embodiments of this application, the method for determining the drug information corresponding to the user problem includes the following steps S51 to S52.
[0190] S51, If the output result data is abnormal, output the reason for the abnormality corresponding to the result data.
[0191] For example, when an exception occurs in the output of the result data, the DeepSeek-V3 main model is called based on the exception branch, and the corresponding exception reason is returned. The corresponding exception reason could be: the SQL generated by the model contains a field name that does not exist in the database table within the knowledge base (e.g., model illusion).
[0192] It should be noted that the above-mentioned example of calling the DeepSeek-V3 large model when drug information execution is abnormal is only for illustrative purposes. In actual applications, other suitable large models can be selected based on specific application requirements, and this application does not impose any restrictions on this.
[0193] It should be noted that, due to the guidance of the prompt message template, the probability of abnormal output of the result data is low.
[0194] S52, if the result data output is normal, perform template conversion on the result data to obtain the drug information corresponding to the result data, and output the drug information to the user terminal.
[0195] Specifically, when the output of the result data is normal, the template conversion is entered based on the normal branch. The result data, i.e. the dataset, is processed into text using Jinja2 statements. At the same time, the numeric fields are summed and finally output to the user.
[0196] Specifically, it can automatically identify the number of data entries to be output to the user's end and control the output based on the user's wishes, thereby improving the user experience.
[0197] This application provides a method for determining the drug information corresponding to the user's question. This method performs an anomaly check on the result data to ensure the accuracy of the output data. Furthermore, the anomaly branch acts as an "airbag." Even if the model produces unexpected output, the system can capture and handle it promptly, rather than crashing, ensuring the continuous availability of the entire service. The "reason for the anomaly" output by the anomaly branch is extremely valuable debugging information. Developers can quickly locate the problem in the model itself, data preprocessing, or other stages based on these logs, greatly shortening troubleshooting time.
[0198] Figure 6 The diagram shows the overall flowchart of the drug information retrieval method based on multi-scenario information fusion provided in the embodiments of this application. Figure 6 The steps in the above Figure 1 As for Figure 5 The information has already been provided in the previous document, and this application will not repeat it here.
[0199] It should be noted that the drug information retrieval method based on multi-scenario information fusion provided in this application improves the accuracy from 50% to 85% based on knowledge base optimization (multi-path recall and setting a preset matching score, and using vectors with actual matching scores greater than the preset matching scores as matching professional terms in the knowledge base), architecture design optimization (classifying user questions based on conditional branch controllers and scenario question classifiers to determine the problem scenario corresponding to the user questions), and context optimization.
[0200] Furthermore, by embedding the Dify service in the embodiments of this application into the enterprise's work software, Dify's query service can be used through the enterprise's work software entry point.
[0201] The scope of protection of the drug information retrieval method based on multi-scenario information fusion described in this application is not limited to the execution order of the steps listed in this embodiment. Any solution implemented by adding, subtracting, or replacing steps in the prior art based on the principles of this application is included within the scope of protection of this application.
[0202] This application also provides a drug information retrieval device based on multi-scenario information fusion. The drug information retrieval device based on multi-scenario information fusion can implement the drug information retrieval method based on multi-scenario information fusion described in this application. However, the implementation device of the drug information retrieval method based on multi-scenario information fusion described in this application includes, but is not limited to, the structure of the drug information retrieval device based on multi-scenario information fusion listed in this embodiment. All structural modifications and substitutions of the prior art made in accordance with the principles of this application are included within the protection scope of this application.
[0203] like Figure 7 As shown, in one embodiment, the drug information retrieval device 70 based on multi-scenario information fusion of this application includes an access permission acquisition module 71, a user question input module 72, a question scenario determination module 73, a drug information determination module 74, a structured query language generation module 75, a result data acquisition module 76, and a drug information determination module 77.
[0204] Access permission acquisition module 71 is used to acquire user access permissions;
[0205] User question input module 72 is used to input user questions into the input window of the question service platform based on the access permissions;
[0206] Problem scenario determination module 73 is used to classify the user's problem based on the scenario problem classifier in the question data service platform and determine the problem scenario corresponding to the user's problem;
[0207] The drug information determination module 74 is used to retrieve the knowledge base within the problem scenario corresponding to the user's question, and recall matching professional terms that match the user's question in the knowledge base; wherein, the number of problem scenarios corresponding to the user's question is one, and the number of knowledge bases within the problem scenario is multiple; then, the first professional term that matches the user's question is recalled from the matching knowledge bases within the problem scenario; all the first professional terms corresponding to the problem scenario are sorted to obtain the sorted second professional term; based on the actual matching score, a preset number of third professional terms are selected from the second professional terms, and the third professional terms are fused using a variable aggregator to obtain the fused matching professional term;
[0208] The structured query language generation module 75 is used to perform semantic parsing on the matched professional terms and generate a structured query language executable by the knowledge base.
[0209] The result data acquisition module 76 is used to input the structured query language into the knowledge base for query operation and obtain the result data output by the knowledge base corresponding to the user's question;
[0210] The drug information determination module 77 is used to determine the drug information corresponding to the user question based on the result data.
[0211] The structure and principle of the permission acquisition module 71, user question input module 72, question scenario determination module 73, drug information determination module 74, structured query language generation module 75, result data acquisition module 76, and drug information determination module 77 correspond one-to-one with the steps in the drug information retrieval method based on multi-scenario information fusion mentioned above, so they will not be described in detail here.
[0212] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, or methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules / units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or units may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of apparatuses or modules or units may be electrical, mechanical, or other forms.
[0213] The modules / units described as separate components may or may not be physically separate. The components shown as modules / units may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules / units can be selected to achieve the objectives of the embodiments of this application, depending on actual needs. For example, the functional modules / units in the various embodiments of this application may be integrated into one processing module, or each module / unit may exist physically separately, or two or more modules / units may be integrated into one module / unit.
[0214] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0215] This application also provides an electronic device. Figure 8 The diagram shown is a structural schematic of an electronic device 80 in one embodiment of this application. The drug information retrieval method based on multi-scene information fusion provided in this embodiment can be applied to... Figure 8 The electronic devices shown are 80, but not limited to these. For example... Figure 8As shown, the electronic device 80 includes a processor 81, a memory, a system bus 83, and a network interface 85. The memory may include a non-volatile storage medium 82 and internal memory 84.
[0216] The non-volatile storage medium 82 can store an operating system and a computer program. The computer program includes program instructions that, when executed, cause the processor to perform any of the drug information retrieval methods based on multi-scenario information fusion provided in the embodiments of this application.
[0217] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0218] The internal memory 84 provides an environment for the execution of computer programs in non-volatile storage media. When the computer program is executed by the processor, it enables the processor to execute any of the drug information retrieval methods based on multi-scenario information fusion provided in the embodiments of this application.
[0219] This network interface 85 is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0220] It should be understood that processor 81 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, the general-purpose processor can be a microprocessor or any conventional processor.
[0221] The electronic device 80 in this application embodiment may include terminal devices such as tablet computers, laptop computers, mobile phones, supercomputers, and smart wearable devices. It can also be applied to databases, servers, and service response systems based on terminal artificial intelligence. This application embodiment does not impose any restrictions on the specific type of electronic device.
[0222] For example, electronic devices can be stations (STAION, ST) in WLANs, cellular phones, cordless phones, Session Initiation Protocol (SIP) phones, Wireless Local Loop (WLL) stations, handheld devices with wireless communication capabilities, computing devices or other processing devices connected to a wireless modem, computers, laptops, handheld communication devices, handheld computing devices, and / or other devices for communicating over wireless systems, as well as next-generation communication systems, such as mobile terminals in 5G networks, mobile terminals in future evolved Public Land Mobile Networks (PLMNs), or mobile terminals in future evolved Non-terrestrial Networks (NTNs).
[0223] This application also provides a computer-readable storage medium. Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing a processor. The program can be stored in a computer-readable storage medium, which is a non-transitory medium, such as random access memory, read-only memory, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disk, and any combination thereof. The storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state drive (SSD)).
[0224] This application embodiment may also provide a computer program product comprising one or more computer instructions. When the computer instructions are loaded and executed on a computing device, all or part of the processes or functions described in this application embodiment are generated. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0225] When the computer program product is executed by a computer, the computer performs the method described in the foregoing method embodiments. The computer program product can be a software installation package; when the foregoing method is required, the computer program product can be downloaded and executed on the computer.
[0226] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.
[0227] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. A drug information retrieval method based on multi-scenario information fusion, characterized in that, The method includes: Obtain user access permissions; Based on the access permissions, the user enters their question into the input window of the question service platform; The user questions are classified based on the scenario question classifier in the question data service platform to determine the question scenario corresponding to the user question; The system retrieves knowledge bases within the problem scenario corresponding to the user's question, and recalls matching professional terms from these knowledge bases. The user's question corresponds to one problem scenario, and there are multiple knowledge bases within each problem scenario. First professional terms matching the user's question are then recalled from the matching knowledge bases within the problem scenario. All first professional terms corresponding to the problem scenario are sorted to obtain sorted second professional terms. Based on actual matching scores, a preset number of third professional terms are selected from the second professional terms, and a variable aggregator is used to fuse these third professional terms to obtain fused matching professional terms. Semantic parsing is performed on the matched technical terms to generate an executable structured query language for the knowledge base; The structured query language is input into the knowledge base to perform a query operation, and the result data corresponding to the user's question is obtained from the output of the knowledge base; Based on the results data, the corresponding drug information for the user's problem is determined.
2. The method according to claim 1, characterized in that, The method further includes: Perform word segmentation on the user question to generate multiple word segmentation sequences; The word segmentation sequence is mapped into a high-dimensional vector through an embedding layer, forming a vector sequence that can be processed by the modified model. By utilizing the bidirectional contextual understanding capability of the modified model, global context modeling is performed on each vector sequence through a self-attention mechanism. The semantic associations of the word segmentation sequences corresponding to each vector sequence are analyzed, the contextual confidence of the word segmentation sequences is calculated, and potential erroneous candidate words of the word segmentation sequences are identified. The correct word segmentation of the erroneous candidate word is predicted based on the masked language model, so as to automatically correct the erroneous candidate word and obtain the corrected user question.
3. The method according to claim 1, characterized in that, Retrieve matching technical terms from the knowledge base that correspond to the user's question, including: Determine the preset matching score; The professional terms in the knowledge base are vectorized based on the vectorization model to obtain name vectors; In the knowledge base, determine the actual matching score of the name vector that matches the user's question; The professional terms corresponding to the name vectors whose actual matching scores are greater than the preset matching scores are used as the matching professional terms for recall.
4. The method according to claim 3, characterized in that, In the knowledge base, the expression corresponding to the actual matching score of the name vector that matches the user's question is: Where a represents the user question vector, and b represents the name vector. Represents the first element in the user question vector. vector elements, Represents the first in the name vector There are 1 vector elements, where n represents the number of vector elements in the user question vector and the name vector.
5. The method according to claim 1, characterized in that, Based on the matching technical terms, the structured query language corresponding to the user's question is determined, including: Develop prompt message templates for corresponding problem scenarios; Input the matching technical terms into the prompt information template to obtain the matching information template; The structured query language is determined based on the matching information template, the user question, and / or the context information of the structured query language for the preset time period, as well as the table structure of the database within the corresponding question scenario.
6. The method according to claim 1, characterized in that, The acquisition of user access permissions includes: Obtain a URL parameter representing user identity information, wherein the URL parameter includes at least one of username and organization address; The access permissions corresponding to the URL parameter are determined based on the conditional branch controller.
7. The method according to claim 6, characterized in that, The step of determining the access permissions corresponding to the URL parameter based on the conditional branch controller includes: Determine whether the user has permission to access the question data service platform based on the URL parameters; If so, then enter the user's question into the input window of the question service platform; If not, then no user question will be entered into the input window of the question service platform.
8. The method according to claim 1, characterized in that, Based on the result data, the drug information corresponding to the user's problem is determined, including: If the output result data is abnormal, output the reason for the abnormality corresponding to the result data; If the output of the result data is normal, the result data is converted into a template to obtain the drug information corresponding to the result data, and the drug information is output to the user terminal.
9. A drug information retrieval device based on multi-scenario information fusion, characterized in that, The device includes: The access permission acquisition module is used to acquire a user's access permissions; The user question input module is used to input user questions into the input window of the question service platform based on the access permissions. The problem scenario determination module is used to classify the user's problem based on the scenario problem classifier in the question data service platform, and determine the problem scenario corresponding to the user's problem; The drug information determination module is used to retrieve a knowledge base within the problem scenario corresponding to the user's question, and recall matching professional terms that match the user's question from the knowledge base; wherein, the number of problem scenarios corresponding to the user's question is one, and the number of knowledge bases within the problem scenario is multiple; then, a first professional term that matches the user's question is recalled from the matching knowledge bases within the problem scenario; all the first professional terms corresponding to the problem scenario are sorted to obtain a sorted second professional term; based on the actual matching score, a preset number of third professional terms are selected from the second professional terms, and a variable aggregator is used to fuse the third professional terms to obtain a fused matching professional term; The structured query language generation module is used to perform semantic parsing on the matched professional terms and generate an executable structured query language for the knowledge base. The result data acquisition module is used to input the structured query language into the knowledge base for query operation and obtain the result data output by the knowledge base corresponding to the user's question; The drug information determination module is used to determine the drug information corresponding to the user's question based on the result data.
10. An electronic device, characterized in that, The electronic device includes: A memory that stores a computer program; The processor, which is communicatively connected to the memory, executes the method of any one of claims 1 to 8 when the computer program is invoked.