Question and answer method, device and equipment suitable for traditional Chinese medicine compound molecular mechanism scene, medium and product
By combining a large language model and a domain knowledge graph, a question-answering method was developed to address the problem of unclear molecular mechanisms in traditional Chinese medicine (TCM) compound preparations. This enabled the application of a professional question-answering model in the context of TCM compound preparation molecular mechanisms, supporting the modernization and international recognition of TCM.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies cannot systematically elucidate the component-target relationships and pathways of action of traditional Chinese medicine (TCM) compound prescriptions, resulting in unclear molecular mechanisms and hindering the modernization and international recognition of TCM.
By employing a large language model combined with a multi-step thinking paradigm and domain knowledge graph, a question-answering model is used to answer questions in the context of molecular mechanisms of traditional Chinese medicine compound prescriptions. This includes domain entity recognition, path recognition, and pattern recognition, generating professional response information.
This study has enabled a systematic analysis of the molecular mechanisms of traditional Chinese medicine compound formulas, improved the accuracy and professionalism of question-and-answer models in this field, addressed the shortcomings of traditional manual experimental methods, and supported the modernization and international recognition of traditional Chinese medicine.
Smart Images

Figure CN121808019A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of large-scale modeling and knowledge graph technology in the context of molecular mechanisms of traditional Chinese medicine compound preparations, and particularly to a question-answering method, apparatus, equipment, medium, and product adapted to the context of molecular mechanisms of traditional Chinese medicine compound preparations. Background Technology
[0002] As the core form of clinical application of traditional Chinese medicine, the efficacy of compound Chinese medicine relies on the synergistic effects of multiple components, multiple targets, and multiple pathways. Clarifying its molecular mechanism is the key to realizing the modernization of traditional Chinese medicine and an important prerequisite for promoting traditional Chinese medicine to the international stage and gaining widespread recognition.
[0003] Currently, given that each traditional Chinese medicine prescription contains hundreds of chemical components and thousands of active ingredients, most rely on traditional manual experimental methods to screen the target of each component one by one.
[0004] While existing technologies identify components and targets through omics methods, they suffer from problems such as fragmented component-target associations and unclear pathways of action, making it impossible to systematically explain the compatibility logic and overall regulatory mechanism of compound preparations. Summary of the Invention
[0005] This disclosure provides a question-and-answer method, apparatus, equipment, medium, and product adapted to the molecular mechanism scenario of traditional Chinese medicine compound prescriptions, so as to realize the application of question-and-answer models in the vertical field of molecular mechanism scenario of traditional Chinese medicine compound prescriptions, and solve the problem that traditional artificial experimental methods cannot systematically analyze the regulatory effects of traditional Chinese medicine prescription components.
[0006] According to one aspect of this disclosure, a question-and-answer method adapted to the molecular mechanism scenario of traditional Chinese medicine compound prescriptions is provided, including:
[0007] Problems in the context of understanding the molecular mechanisms of traditional Chinese medicine compound formulas;
[0008] Based on the question and the first prompt template, a first prompt message is generated, and the first prompt message is input into the trained question-answering model to obtain the answer message for the question; wherein, the first prompt template includes a multi-step thinking paradigm, and the multi-step thinking paradigm includes multiple processing steps;
[0009] The question-answering model generates multiple sub-tasks corresponding to the question based on multiple processing steps in the multi-step thinking paradigm. Based on the multiple sub-tasks, it searches the domain knowledge graph of the molecular mechanism scenario of the traditional Chinese medicine compound to obtain recall knowledge. Based on the recall knowledge, it generates the answer information for the question.
[0010] According to another aspect of this disclosure, a question-and-answer device adapted to the molecular mechanism scenario of traditional Chinese medicine compound prescriptions is provided, the device comprising:
[0011] The problem acquisition module is used to acquire problems related to the molecular mechanism of traditional Chinese medicine compound prescriptions.
[0012] The response information acquisition module is used to generate first prompt information based on the question and the first prompt template, and input the first prompt information into the trained question-answering model to obtain the response information of the question; wherein, the first prompt template includes a multi-step thinking paradigm, and the multi-step thinking paradigm includes multiple processing steps;
[0013] The subtask generation module is used by the question-answering model to generate multiple subtasks corresponding to the question based on multiple processing steps in the multi-step thinking paradigm, and to search the domain knowledge graph of the molecular mechanism scenario of the traditional Chinese medicine compound based on the multiple subtasks to obtain recall knowledge, and to generate the answer information of the question based on the recall knowledge.
[0014] According to another aspect of this disclosure, an electronic device is provided, the electronic device comprising:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the question-and-answer method adapted to the molecular mechanism scenario of traditional Chinese medicine compound prescriptions as described in any embodiment of this disclosure.
[0018] According to another aspect of this disclosure, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the question-and-answer method adapted to the molecular mechanism scenario of traditional Chinese medicine compound prescriptions as described in any embodiment of this disclosure.
[0019] According to another aspect of this disclosure, a computer program product is provided, which, when executed by a processor, implements a question-and-answer method adapted to the molecular mechanism scenario of traditional Chinese medicine compound prescriptions as described in any of the embodiments of this disclosure.
[0020] This embodiment of the disclosure obtains a question in the context of the molecular mechanism of traditional Chinese medicine compound; based on the question and a first prompt template, a first prompt information is generated, and the first prompt information is input into a trained question-answering model to obtain the answer information for the question; wherein, the first prompt template includes a multi-step thinking paradigm, and the multi-step thinking paradigm includes multiple processing steps; the question-answering model generates multiple sub-tasks corresponding to the question based on the multiple processing steps in the multi-step thinking paradigm, and performs retrieval in the domain knowledge graph of the molecular mechanism of traditional Chinese medicine compound scenario based on the multiple sub-tasks to obtain recall knowledge, and generates the answer information for the question based on the recall knowledge. By generating multiple sub-tasks through the multi-step thinking paradigm and combining them with the domain knowledge graph of the molecular mechanism of traditional Chinese medicine compound scenario, the problem of low accuracy of general question-answering models in the analysis of traditional Chinese medicine mechanisms is solved, and the application of question-answering models in the vertical field of the molecular mechanism of traditional Chinese medicine compound scenario is realized.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this disclosure and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a reasoning flowchart of a question-and-answer method adapted to the molecular mechanism scenario of traditional Chinese medicine compound in this embodiment of the present disclosure;
[0024] Figure 2 This is a training flowchart of a question-and-answer method adapted to the molecular mechanism scenario of traditional Chinese medicine compound in this embodiment of the present disclosure;
[0025] Figure 3 This is a flowchart of the question-and-answer method for the molecular mechanism scenario of traditional Chinese medicine compound in this embodiment of the present disclosure;
[0026] Figure 4 This is a schematic diagram of the structure of a question-and-answer device adapted to the molecular mechanism scenario of traditional Chinese medicine compound in this embodiment of the present disclosure;
[0027] Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present disclosure.
[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0031] Figure 1 This disclosure provides a reasoning flowchart for a question-answering method adapted to the molecular mechanism scenario of traditional Chinese medicine (TCM) compound prescriptions. This embodiment is applicable to scenarios where a question-answering model is used to answer questions related to the molecular mechanism of TCM compound prescriptions. The method can be executed by a question-answering device for the TCM compound molecular mechanism scenario in this disclosure. This device can be implemented in software and / or hardware and can be integrated into electronic devices such as computer equipment, servers, mobile terminals, or processors. Figure 1 As shown, the method specifically includes the following steps:
[0032] S110, Problems in the Context of Obtaining the Molecular Mechanism of Traditional Chinese Medicine Compound Formulas.
[0033] In this embodiment, the question in the context of the molecular mechanism of traditional Chinese medicine compound can be specifically understood as a question concerning the molecular-level mechanism related to the therapeutic effect of the traditional Chinese medicine compound. For example, the question in the context of the molecular mechanism of traditional Chinese medicine compound could be, "What is the mechanism of action of Musk Heart-Protecting Pill in treating coronary heart disease?"
[0034] Specifically, users, based on their actual needs in the field of molecular mechanisms of traditional Chinese medicine (TCM) compound preparations, raise questions in the context of TCM compound preparation molecular mechanisms. These questions fall within the scope of TCM compound preparation molecular mechanisms, providing precise input for subsequent analysis and solutions based on domain knowledge graphs.
[0035] S120, based on the question and the first prompt template, a first prompt message is formed, and the first prompt message is input into the trained question-answering model to obtain the answer message for the question; wherein, the first prompt template includes a multi-step thinking paradigm, and the multi-step thinking paradigm includes multiple processing steps.
[0036] In this embodiment, the first prompt template can be specifically understood as a pre-set structured template that guides the trained question-answering model to decompose the problem and reason step by step according to the logical rules of the molecular mechanism of traditional Chinese medicine compound prescriptions. The first prompt template includes a multi-step thinking paradigm. The first prompt information can be specifically understood as a complete prompt information that can be directly input into the trained question-answering model after combining the specific question in the context of the molecular mechanism of traditional Chinese medicine compound prescriptions with the content of the pre-set first prompt template. The first prompt information contains the multi-step thinking paradigm that guides the model to reason step by step in the first prompt template, and is also filled with the specific question raised by the user. Based on this complete prompt information, answer information that meets the requirements of the molecular mechanism of traditional Chinese medicine compound prescriptions is generated.
[0037] The trained question-answering model can be understood as a question-answering model formed by training an existing large language model with a dataset from the field of molecular mechanisms of traditional Chinese medicine (TCM) compound formulas and a multi-step thinking paradigm. For example, existing large language models can use LLaMA (Large Language Model Meta AI) models, GPT (Chat Generative Pre-trained Transformer) series models, etc. These models have strong general language understanding and generation capabilities. Through dataset fine-tuning and multi-step thinking paradigm guidance, they can quickly adapt to the scenario of TCM compound formula molecular mechanisms, transforming general reasoning capabilities into domain-specific professional question-answering capabilities to meet the needs of answering professional questions in this scenario. The answer information can be understood as the response information given by the trained question-answering model to user questions related to the molecular mechanisms of TCM compound formulas. The multi-step thinking paradigm can be understood as a paradigm adapted to the field of TCM compound formula molecular mechanisms, guiding the question-answering model to decompose the question and reason step by step according to a fixed logical order. The multi-step thinking paradigm consists of multiple progressive processing steps designed specifically for this field. The processing steps can be understood as a progressive operational process that constitutes a multi-step thinking paradigm in the field of molecular mechanisms of traditional Chinese medicine compound prescriptions. It is the specific execution link that guides the question-answering model to complete professional reasoning. Each step is interconnected and goes deeper and deeper, ensuring that the entire process of the model from problem decomposition to logical integration conforms to the professional rules of the field.
[0038] Specifically, based on the user's question regarding the molecular mechanism of traditional Chinese medicine compound formulas and the initial prompt template, a first prompt message is generated. The user's specific question is input into the question input field of the first prompt template, which contains guidance instructions for a multi-step thinking paradigm, forming a complete prompt text, i.e., the first prompt message. This first prompt message is then input into a trained question-answering model. The model will follow the multi-step thinking paradigm in the prompt message to perform multiple processing steps and obtain the answer to the question. The multi-step thinking paradigm provides a clear reasoning framework for the question-answering model, preventing logical jumps or fragmentation when answering complex questions about the molecular mechanism of traditional Chinese medicine compound formulas. This ensures that the derivation process conforms to the conventional logic of field research, making the answer more coherent.
[0039] Optionally, based on the above embodiments, the processing steps in the multi-step thinking paradigm include: domain entity recognition, domain path recognition, and domain pattern recognition.
[0040] In this embodiment, domain entity identification is the first processing step in a multi-step thinking paradigm. Its core is to enable the trained question-answering model to accurately extract and classify domain entities from user-submitted questions related to the molecular mechanisms of traditional Chinese medicine (TCM) compound formulas. The domain entities to be identified belong to the research field of TCM compound molecular mechanisms, rather than general entities, and mainly include, but are not limited to: TCM compound entities, such as classic formulas (Guizhi Tang, Mahuang Tang), clinically experienced formulas (a modified Qingfei formula), and compound preparations (Compound Danshen Dripping Pills); active ingredient entities, i.e., key chemical substances in the compound formula that exert its medicinal effect, such as cinnamaldehyde in Guizhi Tang and ephedrine in Mahuang Tang; and molecular target entities, referring to the biomolecules in the body that the active ingredients act upon, such as proteins, genes, and enzymes. Through this step, the question-answering model can extract domain entities from complex questions, laying the foundation for subsequent analysis of the relationships between entities and construction of complete mechanistic logic, avoiding subsequent reasoning deviations due to entity identification omissions or misjudgments.
[0041] Domain path identification is the second processing step in a multi-step thinking paradigm. Its core is to allow the trained question-answering model to mine and clarify the specific relationships between entities identified in the previously identified molecular mechanism domain of traditional Chinese medicine (TCM) compound formulas, and to construct a complete action chain. The identified relationships must closely match the research logic of the TCM compound molecular mechanism, rather than general relationships. These relationships mainly include, but are not limited to: "component-compound" association paths, such as "Guizhi Tang - contains - cinnamaldehyde" and "Ma Huang Tang - contains - ephedrine," clarifying the subordinate relationship between the compound and the active ingredient; and "component-target-pathway" action paths, such as "cinnamaldehyde - inhibits - cyclooxygenase 2," analyzing the mechanism of action of the active ingredient on the target and the physiological and pathological pathways involved by the target. Through this step, the question-answering model can connect isolated domain entities into logical action chains, providing crucial relational support for forming a complete mechanism logic and avoiding incomplete mechanism analysis due to broken relationships between entities.
[0042] Domain pattern recognition, the third step in a multi-step thinking paradigm, is essentially an operation that integrates the trained question-answering model with key objects extracted from domain entity recognition and the relationships identified from domain path recognition, into triplets that conform to the molecular mechanisms of traditional Chinese medicine compound formulas. Its core objective is to transform fragmented "entity-relationship" information into a systematic understanding of mechanisms. Following the standardized format of "<entity 1, relation, entity 2>", the previously identified entities and relationships are combined into domain-specific knowledge triplets, such as "<Guizhi Tang, contains, cinnamaldehyde>" and "<cinnamaldehyde, inhibits, cyclooxygenase 2>". These triplets are the core carriers of domain knowledge, accurately representing the key information of the compound formula's molecular mechanism. This step provides a clear logical basis for the final natural language response and ensures that the response content aligns with the professional cognitive framework of the molecular mechanisms of traditional Chinese medicine compound formulas, avoiding fragmented information or logical breaks.
[0043] S130, the question-answering model generates multiple sub-tasks corresponding to the question based on multiple processing steps in the multi-step thinking paradigm, and performs retrieval in the domain knowledge graph of the molecular mechanism scenario of the traditional Chinese medicine compound based on the multiple sub-tasks to obtain recall knowledge, and generates answer information for the question based on the recall knowledge.
[0044] In this embodiment, a subtask can be understood as a specific task formed by decomposing a user question into a processing step within a multi-step thinking paradigm, used for precise information retrieval in the domain knowledge graph. Each subtask corresponds to a retrieval target of a processing step, and there is a progressive dependency between tasks. Knowledge retrieval can be understood as the structured domain knowledge directly related to the user question obtained by the question-answering model after targeted retrieval in the domain knowledge graph of the molecular mechanisms of traditional Chinese medicine compound formulas, based on multiple decomposed subtasks. The domain knowledge graph for the molecular mechanisms of traditional Chinese medicine compound formulas can be understood as a knowledge graph pre-constructed based on publicly available database formula mechanism data and experimental data, with the research on the molecular mechanisms of traditional Chinese medicine compound formulas as its core, and which stores knowledge in this domain in a structured manner. This knowledge graph not only achieves the systematic and structured storage of knowledge on the molecular mechanisms of traditional Chinese medicine compound formulas, but also, through the association of entities and relationships, can quickly respond to the retrieval needs of subtasks, providing accurate and efficient data source support for the acquisition of knowledge retrieval.
[0045] Specifically, the question-answering model, based on the multi-step thinking paradigm within the initial prompt template, breaks down the user's question about the molecular mechanism of traditional Chinese medicine (TCM) compound formulas into multiple sub-tasks. The model then searches the domain knowledge graph of the TCM compound molecular mechanism scenario for each sub-task. GraphRAG (Graph-based Retrieval-Augmented Generation) technology can be used for this search. The retrieved knowledge from each sub-task is then integrated according to the logic of the multi-step thinking paradigm to obtain recall knowledge. Based on this recall knowledge, the answer to the question is generated. Breaking down the sub-tasks according to the multi-step thinking paradigm and performing targeted searches allows for more precise access to the knowledge graph, avoiding indiscriminate scraping of redundant information and improving the matching accuracy of the answer.
[0046] Optionally, based on the above embodiments, the plurality of subtasks include a domain entity recognition subtask, a domain path recognition subtask, and a domain pattern recognition subtask; wherein, the domain entity recognition subtask is used to identify key entities in the problem; the domain path recognition subtask is used to identify the association relationships between the key entities in the domain knowledge graph of the molecular mechanism scenario of the traditional Chinese medicine compound, forming relationship paths; the domain pattern recognition subtask is used to match the key entities with the relationship paths, and use the matched key entities and the relationship paths as the recalled knowledge.
[0047] In this embodiment, the domain entity identification subtask can be understood as a basic task among multiple subtasks. Its core objective is to accurately extract and define entities closely related to the domain from the questions raised by users regarding the molecular mechanism of traditional Chinese medicine compound prescriptions. This is used to identify key entities in the questions and provide clear retrieval objects for subsequent subtasks.
[0048] The domain path recognition subtask can be understood as a follow-up task to the domain entity recognition subtask. Its core objective is to accurately mine and connect the specific relationships between the key entities identified in the previous stage within the knowledge graph of the molecular mechanisms of traditional Chinese medicine compound formulas, forming an "entity-relationship" link. Through this subtask, the key entities in the domain entity recognition subtask can be transformed into a logically connected path network, providing a structured relational framework for the matching and integration of subsequent domain pattern recognition subtasks, thus avoiding the lack of logical coherence in recalled knowledge due to broken relationships between entities.
[0049] The domain pattern recognition subtask can be understood as following the domain path recognition subtask. Its core objective is to systematically match and integrate the key entities identified in the previous stage with the established relationship paths, ultimately outputting retrieved knowledge. Through this subtask, fragmented entities and relationships can be transformed into structured retrieved knowledge that can directly support responses. This ensures both the accuracy and logic of the knowledge and provides a reliable content foundation for the question-answering model to generate professional and systematic response information.
[0050] The technical solution of this embodiment obtains a question in the context of the molecular mechanism of traditional Chinese medicine compound; forms a first prompt information based on the question and a first prompt template; inputs the first prompt information into a trained question-answering model to obtain the answer information for the question; wherein, the first prompt template includes a multi-step thinking paradigm, which includes multiple processing steps; the question-answering model generates multiple sub-tasks corresponding to the question based on the multiple processing steps in the multi-step thinking paradigm, and performs retrieval in the domain knowledge graph of the molecular mechanism of traditional Chinese medicine compound scenario based on the multiple sub-tasks to obtain recalled knowledge, and generates the answer information for the question based on the recalled knowledge; by generating multiple sub-tasks through the multi-step thinking paradigm and combining them with the domain knowledge graph of the molecular mechanism of traditional Chinese medicine compound scenario, the problem of low accuracy of general question-answering models in the analysis of traditional Chinese medicine mechanisms is solved, and the application of question-answering models in the vertical field of the molecular mechanism of traditional Chinese medicine compound scenario is realized.
[0051] Figure 2 This embodiment provides a training flowchart for a question-answering method adapted to the molecular mechanism scenario of traditional Chinese medicine compound prescriptions. The embodiment also provides a training method for the question-answering model, such as... Figure 2 As shown, the method specifically includes the following steps:
[0052] S210, Obtain the first question-and-answer sample data under the scenario of the molecular mechanism of the traditional Chinese medicine compound.
[0053] In this embodiment, the first question-and-answer sample data can be specifically understood as initial questions set by experts in the field of molecular mechanisms of traditional Chinese medicine compound formulas, within a scenario based on these molecular mechanisms. These questions are then paired with standardized answers written by experts based on authoritative literature, experimental data, and clinical research conclusions, forming a "question-answer" pair of first question-and-answer sample data. This type of sample data ensures both the domain-specific relevance of the questions and the professionalism and accuracy of the answers through the experts' knowledge, providing a high-quality data foundation for subsequent model training or prompt template optimization.
[0054] Specifically, initial questions are set up by experts in the field of molecular mechanisms of traditional Chinese medicine compound prescriptions, and answers are written based on these initial questions to obtain the first question-and-answer sample data. This first question-and-answer sample data can comprehensively cover the field of molecular mechanisms of traditional Chinese medicine compound prescriptions.
[0055] S220, expand the sample based on the first question and answer sample data to obtain the second question and answer sample data, and form the sample dataset under the molecular mechanism scenario of the traditional Chinese medicine compound based on the second question and answer sample data and the first question and answer sample data.
[0056] In this embodiment, the second question-and-answer sample data can be specifically understood as question-and-answer sample data obtained by expanding upon the first question-and-answer sample data, combining it with the domain knowledge graph of the molecular mechanism scenario of traditional Chinese medicine compound prescriptions, and by changing the questions and appropriately adjusting the answers to those questions. The sample dataset can be specifically understood as a dataset including both the first and second question-and-answer sample data, which can be used for training the subsequent question-and-answer model, providing sufficient and comprehensive data support for the model to learn domain knowledge and reasoning logic.
[0057] Specifically, based on the first question-and-answer sample data and the knowledge graph of the molecular mechanism scenario of traditional Chinese medicine compound prescriptions, the samples were expanded to obtain the second question-and-answer sample data. The second question-and-answer sample data and the first question-and-answer sample data were merged to obtain the sample dataset under the molecular mechanism scenario of traditional Chinese medicine compound prescriptions, ensuring that the sample dataset under the molecular mechanism scenario of traditional Chinese medicine compound prescriptions meets the model training requirements in terms of quantity, diversity of expression, and coverage of professional logic.
[0058] Optionally, based on the above embodiments, sample expansion is performed on the first question-and-answer sample data to obtain second question-and-answer sample data, including: performing at least one of entity-level expansion, relation-level expansion, and sentence-level expansion on the first question-and-answer sample data to obtain second question-and-answer sample data; wherein, the entity-level expansion is used to perform semantic replacement of key entities of sample questions in the first question-and-answer sample data based on the domain knowledge graph of the molecular mechanism scenario of the traditional Chinese medicine compound; the relation-level expansion is used to adjust the relation path of sample questions in the first question-and-answer sample data based on the entity relation network of the key entities in the domain knowledge graph of the molecular mechanism scenario of the traditional Chinese medicine compound; the sentence-level expansion is used to perform at least one of sentence transformation, tone adjustment, and sentence structure reorganization on sample questions in the first question-and-answer sample data.
[0059] In this embodiment, entity-level expansion can be understood as using the domain knowledge graph of the molecular mechanism scenario of traditional Chinese medicine compound as a basis. For the key entities contained in the sample questions in the first question-and-answer sample data, synonymous or near-synonymous entities in the knowledge graph that are semantically consistent with the key entity, point to the same category, or belong to the same category are selected and replaced in the original sample question. Simultaneously, the core logic of the question and the original answer are maintained, ultimately generating a new sample question and the corresponding expanded answer. For example, if the original sample question is "What is the molecular mechanism of Jin Kui Shen Qi Wan in treating kidney-yang deficiency type edema?", and the synonym entity of "Jin Kui Shen Qi Wan" in the knowledge graph includes "Gui Fu Di Huang Wan" (an alternative name in the pharmacopoeia), then after entity-level expansion, the question becomes "What is the molecular mechanism of Gui Fu Di Huang Wan in treating kidney-yang deficiency type edema?", and the answer only replaces "Jin Kui Shen Qi Wan" with "Gui Fu Di Huang Wan", while the rest of the expression remains unchanged. By extending the entity level, the domain knowledge graph based on the molecular mechanism of traditional Chinese medicine compound can perform semantic replacement on the key entities of the sample questions in the first question-and-answer sample data. This can enrich the diversity of the expression of the sample questions without changing the core knowledge of the first question-and-answer sample, and also allow the extended second question-and-answer sample data to cover more domain standard terms, avoiding the comprehension bias caused by the difference in expression, and providing more comprehensive corpus support for subsequent model training.
[0060] Relationship-level expansion can be understood as using the domain knowledge graph of the molecular mechanism scenario of traditional Chinese medicine compound prescriptions as a basis to sort out the relationships between key entities involved in the sample questions in the first question-and-answer sample data. Based on the entity relationship network of key entities in the knowledge graph, that is, the multi-level association paths between the entity and other entities, the core relationship paths of the original sample questions are extended, refined, or supplemented to generate new sample questions. The answer content is then adapted and adjusted to form an expansion method of pairing new questions with answers. For example, one of the questions in the first question-and-answer sample is "How does Musk Heart-Protecting Pill treat coronary heart disease?". Based on the domain knowledge graph of the molecular mechanism scenario of traditional Chinese medicine compound prescriptions, and based on the relationship in the knowledge graph that "Musk Heart-Protecting Pill contains ginsenoside Rg1, which can promote coronary angiogenesis and thus improve myocardial blood supply", a new question can be generated: "How does ginsenoside Rg1 in Musk Heart-Protecting Pill help patients with coronary heart disease grow new coronary blood vessels?". An adapted answer is then given based on the new question. By extending the relationship level, the entity relationship network of key entities in the domain knowledge graph based on the molecular mechanism of traditional Chinese medicine compound can be used to adjust the relationship path of sample questions in the first question and answer sample data. This can quickly enrich the quantity and expression diversity of the second question and answer sample data without increasing the domain knowledge annotation cost.
[0061] Sentence-level expansion can be understood as generating new sample questions with diverse expressions but consistent core logic by optimizing and adjusting the language of the original questions through sentence transformation, tone adjustment, and sentence structure reorganization, without changing the core semantics of the sample questions in the first question-and-answer sample data. The original answers are then directly reused or slightly adjusted, ultimately forming an expansion method that pairs new questions with answers. For example, one of the questions in the first question-and-answer sample is "How does Musk Heart-Protecting Pill treat coronary heart disease?". New questions can be generated through sentence transformation: "How does Musk Heart-Protecting Pill treat coronary heart disease?" and "How is coronary heart disease treated by Musk Heart-Protecting Pill?". Other new questions can be generated through sentence transformation: "Please explain in detail how Musk Heart-Protecting Pill treats coronary heart disease?" and "In what ways might Musk Heart-Protecting Pill treat coronary heart disease?". New questions can also be generated through sentence restructuring: "How does Musk Heart-Protecting Pill exert its therapeutic effect on coronary heart disease? What is its mechanism?" and "How does Musk Heart-Protecting Pill work in treating coronary heart disease?". Through sentence-level expansion, performing at least one of sentence transformation, tone adjustment, and sentence restructuring on the sample questions in the first question-and-answer sample data, while keeping the answer unchanged or fine-tuning the answer according to different expressions, can help the question-and-answer model adapt to different questioning habits and improve the stability of answers to similar questions.
[0062] Optionally, based on the above embodiments, sample expansion is performed on the first question-and-answer sample data to obtain second question-and-answer sample data, including: generating second prompt information from the first question-and-answer sample data and the second prompt template, wherein the second prompt template includes sample expansion conditions, the sample expansion conditions including semantic similarity and different expressions; and inputting the second prompt information into a pre-trained text generation model to obtain second question-and-answer sample data that satisfies the sample expansion conditions.
[0063] In this embodiment, the second prompt template can be specifically understood as a standardized instruction framework template designed to guide the pre-trained text generation model in generating second question-and-answer sample data. The second prompt template includes sample expansion conditions, which can be specifically understood as the conditions that the second question-and-answer sample data generated by the pre-trained text generation model must meet. The sample expansion conditions include semantic similarity and expression difference, which together constitute the quality standard of the expanded samples.
[0064] Semantic similarity means that the generated second question-and-answer sample data must be consistent with the core information of the corresponding first question-and-answer sample data. This means the direction of the second question-and-answer sample data must be consistent with the corresponding first question-and-answer sample data, key entities must not be missing or replaced, core relationships must not be changed, and the new response must fully retain the core mechanism information of the original response, only adjusting details to adapt to the new question's wording, without adding or deleting key mechanisms. Different wording means that the generated second question-and-answer sample data must differ from the corresponding first question-and-answer sample data in its language expression. This means the question in the second question-and-answer sample data needs to be presented differently by changing sentence structure, adjusting word order, or using domain-specific synonyms. Its response must correspond to the rhythm of the new question's wording, adjusting the wording without changing the core mechanism, and avoiding repetition with the original response.
[0065] The second prompt information can be understood as filling the first question-and-answer sample data into the second prompt template, forming complete information that can be directly input into the pre-trained text generation model. It integrates the task objective and sample expansion conditions in the second prompt template with the specific content of the first question-and-answer sample, clearly informing the text generation model which original sample to expand upon. The second prompt information is the direct basis for the pre-trained text generation model to generate the second question-and-answer sample data, ensuring that the new samples generated by the pre-trained text generation model not only meet the expansion conditions but also fit the needs of the molecular mechanism scenario of traditional Chinese medicine compound prescriptions.
[0066] A pre-trained text generation model can be understood as a model that has been trained in advance using large-scale general text data and text data related to the molecular mechanisms of traditional Chinese medicine compound prescriptions. This model possesses the ability to understand text semantics, follow instruction logic, and generate text that conforms to domain norms. Based on the input second prompt information, the pre-trained text generation model can automatically identify the core semantics of the original sample and, under the condition of sample expansion with "semantic similarity but different expression," generate diverse new questions and appropriate new answers, ultimately outputting second question-and-answer sample data that meets the requirements.
[0067] Specifically, the framework of the second prompt template is defined, and the first question-and-answer sample data is filled into the second prompt template to form complete second prompt information. The text generation model needs to be jointly trained in advance with general text and text in the field of molecular mechanism of traditional Chinese medicine compound, so as to obtain a well-trained text generation model with the ability to recognize instruction constraints, reproduce core semantics, and generate diverse texts. After inputting the constructed second prompt information into the text generation model, the text generation model first parses the sample expansion conditions in the prompt, and then locates the core information of the original sample. Under the constraint of "semantic similarity", the text generation model performs operations such as sentence transformation and word order adjustment on the original question, and at the same time adjusts the wording of the original answer accordingly to ensure that the new answer is both suitable for the expression of the new question and does not lose the core mechanism. The text generation model outputs second question-and-answer samples that meet the conditions, which can efficiently generate a large number of second question-and-answer samples that are semantically consistent with the first question-and-answer samples and have diverse expressions, quickly expanding the sample dataset in the scenario of molecular mechanism of traditional Chinese medicine compound, and providing richer corpus support for subsequent model training.
[0068] S230, a question-answering model for the molecular mechanism scenario of traditional Chinese medicine compound is trained based on the sample dataset in the scenario of molecular mechanism of traditional Chinese medicine compound.
[0069] Specifically, based on the sample dataset of traditional Chinese medicine compound molecular mechanism scenarios, a question-answering model for traditional Chinese medicine compound molecular mechanism scenarios was trained.
[0070] Optionally, based on the above embodiments, training a question-answering model for the molecular mechanism scenario of traditional Chinese medicine compound based on a sample dataset includes: forming a third prompt information from sample questions and a first prompt template in the question-answering sample data of the sample dataset; inputting the third prompt information into the question-answering model to be trained to obtain the task results of each sub-task in the process of the question-answering model to be trained processing the sample questions; obtaining the task labels of the question-answering sample data for each sub-task; obtaining the loss function corresponding to each sub-task based on the task result and the task label of each sub-task; and adjusting the model parameters of the question-answering model to be trained based on the loss function corresponding to each sub-task until a trained question-answering model for the molecular mechanism scenario of traditional Chinese medicine compound is obtained.
[0071] In this embodiment, the third prompt information can be specifically understood as a prompt message constructed to guide the question-answering model to be trained to decompose the sample question according to the preset logic and generate the results of each sub-task. Its core is to combine the first prompt template with the specific sample question in the sample dataset, clearly telling the model how to decompose the question and which sub-task results need to be output, providing a standardized input basis for subsequent calculation of loss by sub-task and optimization of model parameters. The question-answering model to be trained can be specifically understood as a general question-answering model that has not yet been adapted to the molecular mechanism scenario of traditional Chinese medicine compound. Its core function is to receive the third prompt information, process the sample question step by step according to the preset sub-task logic in the prompt, output the preliminary results of each sub-task, and finally, through the optimization of the sub-task loss function, gradually master the professional knowledge and reasoning logic of the molecular mechanism scenario of traditional Chinese medicine compound, and become a mature question-answering model adapted to this scenario.
[0072] The question-answering sample data, specifically the task labels for each subtask, can be understood as standard answers formulated by experts in the field of molecular mechanisms of traditional Chinese medicine compound formulas. These answers are based on "sample questions + standard answers" from the sample dataset, combined with a domain knowledge graph, for each subtask preset in the third-party prompt information. Their core function is to serve as a benchmark for measuring the accuracy of the subtask results output by the question-answering model under training, providing a clear basis for calculating the subtask loss function and optimizing model parameters. The loss function for each subtask can be understood as a mathematical function used to quantify the difference between the subtask result output by the question-answering model under training and the task label for that subtask. Its core function is to reflect the model's performance on that subtask through numerical error values, providing a clear direction for model parameter adjustment. Different subtasks have different output formats, and the calculation method of the loss function will adapt to their characteristics, but the core logic is to measure the degree of inconsistency between the predicted result and the target label.
[0073] Specifically, the specific sample questions in the sample dataset are combined with the first prompt template to form the third prompt information. This third prompt information is then input into the question-answering model to be trained. The model processes the sample questions step by step according to the sub-task logic, outputting preliminary results for each sub-task. Domain experts assign task labels to each sub-task based on the standard answers to the sample questions and the domain knowledge graph. For each sub-task, the difference between the output result and the task label is calculated using a loss function. The loss function for each sub-task is adjusted according to preset weights to reduce the total loss of the question-answering model. Optionally, a threshold for the stable decrease of the total loss on the sample dataset and a threshold for the matching degree between the sub-task results and the labels can be set. When the stable decrease of the total loss is less than or equal to the threshold, and the matching degree between the sub-task results and the labels is greater than or equal to the matching degree threshold, a well-trained question-answering model for the molecular mechanism scenario of traditional Chinese medicine compound formulas is obtained.
[0074] The technical solution of this embodiment involves acquiring first question-and-answer sample data under the scenario of the molecular mechanism of traditional Chinese medicine compound; expanding the first question-and-answer sample data to obtain second question-and-answer sample data; forming a sample dataset under the scenario of the molecular mechanism of traditional Chinese medicine compound based on the second question-and-answer sample data and the first question-and-answer sample data; training a question-and-answer model under the scenario of the molecular mechanism of traditional Chinese medicine compound based on the sample dataset. By expanding the first question-and-answer sample data to obtain the second question-and-answer sample data, and then training the question-and-answer model, the question-and-answer model is trained. Based on a limited initial sample, by expanding and enriching the number of samples and the diversity of expressions, while retaining the core professional logic of the molecular mechanism scenario of traditional Chinese medicine compound, the final trained question-and-answer model can more accurately and comprehensively understand and answer questions in this field, improving the adaptability to different question formats and the professionalism and completeness of the answers.
[0075] Based on the above embodiments, an optional example is provided, which can be used in a question-and-answer scenario adapted to the molecular mechanism of traditional Chinese medicine compound formulas. The workflow diagram of this question-and-answer method adapted to the molecular mechanism of traditional Chinese medicine compound formulas is shown below. Figure 3 As shown.
[0076] First, the system obtains the question raised by the user regarding the molecular mechanism of traditional Chinese medicine compound, which is "What is the main mechanism of Musk Heart-Protecting Pill in treating coronary heart disease?" Based on this question and the first prompt template, a first prompt message is generated. The first prompt message is then input into the trained question-answering model to obtain the answer to the question.
[0077] The question-answering model generates multiple sub-tasks corresponding to the question based on multiple processing steps in a multi-step thinking paradigm. These include: a domain entity recognition sub-task, enabling the model to accurately identify entities related to the field of traditional Chinese medicine (TCM), namely "Musk Heart-Protecting Pill" and "Coronary Heart Disease"; a domain path recognition sub-task, enabling the model to accurately identify the relationships formed by entities related to TCM, namely "Musk Heart-Protecting Pill → Core Component → Target Point → Regulatory Pathway → Improvement of Coronary Heart Disease Pathology"; and a domain pattern recognition sub-task, enabling the model to accurately identify the triples formed by entities and relationships related to TCM, namely "Musk Heart-Protecting Pill - Contains - Core Component", "Core Component - Acts on - Target Point", "Target Point - Regulates - Regulatory Pathway", and "Regulatory Pathway - Improves - Coronary Heart Disease Pathology". Based on these sub-tasks, the model uses GraphRAG technology in the knowledge graph of the molecular mechanism of TCM compound formulas to first retrieve the core components associated with "Musk Heart-Protecting Pill", then locate the targets and pathways corresponding to the components, and finally associate the pathways with the improvement effect on the pathology of coronary heart disease, recalling key knowledge and ultimately integrating it into the answer information.
[0078] Figure 4This is a schematic diagram of a question-and-answer device adapted to the molecular mechanism scenario of traditional Chinese medicine compound prescriptions, provided in an embodiment of this disclosure. This embodiment is applicable to scenarios where a question-and-answer model is used to answer questions related to the molecular mechanism of traditional Chinese medicine compound prescriptions. The device can be implemented using software and / or hardware, and can be integrated into any device that provides question-and-answer functionality adapted to the molecular mechanism scenario of traditional Chinese medicine compound prescriptions, such as… Figure 4 As shown, the question-and-answer device adapted to the molecular mechanism scenario of traditional Chinese medicine compound includes: a question acquisition module 310, a response information acquisition module 320, and a subtask generation module 330.
[0079] Problem acquisition module 310 is used to acquire problems in the context of molecular mechanisms of traditional Chinese medicine compound prescriptions;
[0080] The response information acquisition module 320 is used to generate first prompt information based on the question and the first prompt template, and input the first prompt information into the trained question-answering model to obtain the response information of the question; wherein, the first prompt template includes a multi-step thinking paradigm, and the multi-step thinking paradigm includes multiple processing steps;
[0081] The subtask generation module 330 is used to generate multiple subtasks corresponding to the question based on multiple processing steps in the multi-step thinking paradigm, and to retrieve recall knowledge based on the multiple subtasks in the domain knowledge graph of the molecular mechanism scenario of the traditional Chinese medicine compound, and to generate answer information for the question based on the recall knowledge.
[0082] The technical solution of this embodiment obtains a question in the context of the molecular mechanism of traditional Chinese medicine compound; forms a first prompt information based on the question and a first prompt template; inputs the first prompt information into a trained question-answering model to obtain the answer information for the question; wherein, the first prompt template includes a multi-step thinking paradigm, which includes multiple processing steps; the question-answering model generates multiple sub-tasks corresponding to the question based on the multiple processing steps in the multi-step thinking paradigm; searches the domain knowledge graph of the molecular mechanism of traditional Chinese medicine compound scenario based on the multiple sub-tasks to obtain recalled knowledge; and generates the answer information for the question based on the recalled knowledge. By generating multiple sub-tasks through the multi-step thinking paradigm and combining them with the domain knowledge graph of the molecular mechanism of traditional Chinese medicine compound scenario, the problem of low accuracy of general question-answering models in the analysis of traditional Chinese medicine mechanisms is solved, and the application of question-answering models in the vertical field of the molecular mechanism of traditional Chinese medicine compound scenario is realized.
[0083] Based on the above embodiments, optionally, the processing steps in the multi-step thinking paradigm include: domain entity recognition, domain path recognition, and domain pattern recognition.
[0084] Based on the above embodiments, optionally, the plurality of subtasks include a domain entity recognition subtask, a domain path recognition subtask, and a domain pattern recognition subtask; wherein, the domain entity recognition subtask is used to identify key entities in the problem; the domain path recognition subtask is used to identify the association relationships between the key entities in the domain knowledge graph of the molecular mechanism scenario of the traditional Chinese medicine compound, forming relationship paths; the domain pattern recognition subtask is used to match the key entities with the relationship paths, and use the matched key entities and the relationship paths as the recalled knowledge.
[0085] Optionally, based on the above embodiments, the device further includes a question-answering model training module, used for: acquiring first question-answering sample data under the molecular mechanism scenario of the traditional Chinese medicine compound; expanding the sample data based on the first question-answering sample data to obtain second question-answering sample data; forming a sample dataset under the molecular mechanism scenario of the traditional Chinese medicine compound based on the second question-answering sample data and the first question-answering sample data; and training a question-answering model under the molecular mechanism scenario of the traditional Chinese medicine compound based on the sample dataset under the molecular mechanism scenario of the traditional Chinese medicine compound.
[0086] Based on the above embodiments, optionally, the question-answering model training module is further configured to: perform at least one of entity-level expansion, relation-level expansion, and sentence-level expansion on the first question-answering sample data to obtain second question-answering sample data; wherein, the entity-level expansion is used to perform semantic substitution on key entities of sample questions in the first question-answering sample data based on the domain knowledge graph of the molecular mechanism scenario of the traditional Chinese medicine compound; the relation-level expansion is used to adjust the relation path of sample questions in the first question-answering sample data based on the entity relation network of the key entities in the domain knowledge graph of the molecular mechanism scenario of the traditional Chinese medicine compound; the sentence-level expansion is used to perform at least one of sentence transformation, tone adjustment, and sentence structure reorganization on sample questions in the first question-answering sample data.
[0087] Based on the above embodiments, optionally, the question-answering model training module is further configured to: generate second prompt information from the first question-answering sample data and the second prompt template, wherein the second prompt template includes sample expansion conditions, the sample expansion conditions including semantic similarity and different expressions; input the second prompt information into a pre-trained text generation model to obtain second question-answering sample data that satisfies the sample expansion conditions.
[0088] Based on the above embodiments, optionally, the question-answering model training module is further configured to: form a third prompt information from the sample questions and the first prompt template in the question-answering sample data of the sample dataset; input the third prompt information into the question-answering model to be trained to obtain the task results of each sub-task in the process of the question-answering model to be trained processing the sample questions; obtain the task labels of the question-answering sample data for each sub-task; obtain the loss function corresponding to each sub-task based on the task results and the task labels of each sub-task; and adjust the model parameters of the question-answering model to be trained based on the loss functions corresponding to each sub-task until a well-trained question-answering model in the molecular mechanism scenario of the traditional Chinese medicine compound is obtained.
[0089] The above-described products can perform the methods provided in any embodiment of this disclosure, and have the corresponding functional modules and beneficial effects for performing the methods.
[0090] Figure 5 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0091] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0092] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0093] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as question-answering methods adapted to the molecular mechanism scenario of traditional Chinese medicine compound prescriptions.
[0094] In some embodiments, the question-and-answer method adapted to the molecular mechanism scenario of traditional Chinese medicine compound formulas can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the question-and-answer method adapted to the molecular mechanism scenario of traditional Chinese medicine compound formulas described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the question-and-answer method adapted to the molecular mechanism scenario of traditional Chinese medicine compound formulas by any other suitable means (e.g., by means of firmware).
[0095] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0096] Computer programs used to implement the methods of this disclosure may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0097] In the context of this disclosure, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0098] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0099] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0100] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0101] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0102] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements a question-and-answer method adapted to the molecular mechanism scenario of traditional Chinese medicine compound prescriptions according to any embodiment of this disclosure.
[0103] In implementing a computer program product, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0104] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A question-and-answer method adapted to the molecular mechanism scenario of traditional Chinese medicine compound prescriptions, characterized in that, include: Problems in the context of understanding the molecular mechanisms of traditional Chinese medicine compound formulas; Based on the question and the first prompt template, a first prompt message is generated, and the first prompt message is input into the trained question-answering model to obtain the answer message for the question; wherein, the first prompt template includes a multi-step thinking paradigm, and the multi-step thinking paradigm includes multiple processing steps; The question-answering model generates multiple sub-tasks corresponding to the question based on multiple processing steps in the multi-step thinking paradigm. Based on the multiple sub-tasks, it searches the domain knowledge graph of the molecular mechanism scenario of the traditional Chinese medicine compound to obtain recall knowledge. Based on the recall knowledge, it generates the answer information for the question.
2. The method according to claim 1, characterized in that, The processing steps in the multi-step thinking paradigm include: domain entity recognition, domain path recognition, and domain pattern recognition.
3. The method according to claim 2, characterized in that, The multiple subtasks include a domain entity recognition subtask, a domain path recognition subtask, and a domain pattern recognition subtask; The domain entity identification subtask is used to identify key entities in the problem. The domain path recognition subtask is used to identify the relationships between the key entities in the domain knowledge graph of the molecular mechanism scenario of the traditional Chinese medicine compound, and form relationship paths; The domain pattern recognition subtask is used to match the key entities with the relationship paths, and the matched key entities and the relationship paths are used as the recall knowledge.
4. The method according to claim 1, characterized in that, The training method for the question-answering model includes: Obtain the first question-and-answer sample data under the scenario of the molecular mechanism of the traditional Chinese medicine compound; Based on the first question-and-answer sample data, sample expansion is performed to obtain the second question-and-answer sample data. Based on the second question-and-answer sample data and the first question-and-answer sample data, a sample dataset under the molecular mechanism scenario of the traditional Chinese medicine compound is formed. A question-answering model for the molecular mechanism scenario of traditional Chinese medicine compound prescriptions was trained based on the sample dataset of the scenario.
5. The method according to claim 4, characterized in that, Based on the first question-and-answer sample data, sample expansion is performed to obtain the second question-and-answer sample data, including: The first question-and-answer sample data is expanded by at least one of entity-level expansion, relation-level expansion, and sentence-level expansion to obtain the second question-and-answer sample data; The entity-level expansion is used to perform semantic replacement of key entities in the sample questions in the first question-and-answer sample data based on the domain knowledge graph of the molecular mechanism scenario of the traditional Chinese medicine compound. The relation-level expansion adjusts the relational paths of sample questions in the first question-and-answer sample data based on the entity relation network of key entities in the domain knowledge graph of the molecular mechanism scenario of the traditional Chinese medicine compound. The sentence-level expansion is used to perform at least one of sentence transformation, tone adjustment, and sentence structure reorganization on the sample questions in the first question-and-answer sample data.
6. The method according to claim 4, characterized in that, Based on the first question-and-answer sample data, sample expansion is performed to obtain the second question-and-answer sample data, including: The first question-and-answer sample data and the second prompt template are used to generate a second prompt message. The second prompt template includes sample expansion conditions, which include semantic similarity and different expressions. The second prompt information is input into a pre-trained text generation model to obtain second question-and-answer sample data that meets the sample expansion conditions.
7. The method according to claim 4, characterized in that, A question-answering model for the molecular mechanism scenario of traditional Chinese medicine compound prescriptions is trained based on a sample dataset, including: The sample questions and the first prompt template in the sample dataset are used to form a third prompt information. The third prompt information is then input into the question-answering model to be trained to obtain the task results of each subtask in the process of the question-answering model to process the sample questions. Obtain the task labels for each subtask from the question-and-answer sample data, and obtain the loss function for each subtask based on the task result and the task label of each subtask. The model parameters of the question-answering model to be trained are adjusted based on the loss function corresponding to each sub-task until a well-trained question-answering model under the molecular mechanism scenario of the traditional Chinese medicine compound is obtained.
8. A question-and-answer device adapted to the molecular mechanism scenario of traditional Chinese medicine compound prescriptions, characterized in that, include: The problem acquisition module is used to acquire problems related to the molecular mechanism of traditional Chinese medicine compound prescriptions. The response information acquisition module is used to generate first prompt information based on the question and the first prompt template, and input the first prompt information into the trained question-answering model to obtain the response information of the question; wherein, the first prompt template includes a multi-step thinking paradigm, and the multi-step thinking paradigm includes multiple processing steps; The subtask generation module is used by the question-answering model to generate multiple subtasks corresponding to the question based on multiple processing steps in the multi-step thinking paradigm, and to search the domain knowledge graph of the molecular mechanism scenario of the traditional Chinese medicine compound based on the multiple subtasks to obtain recall knowledge, and to generate the answer information of the question based on the recall knowledge.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the question-and-answer method adapted to the molecular mechanism scenario of traditional Chinese medicine compound prescriptions as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the question-and-answer method according to any one of claims 1-7, adapted to the molecular mechanism scenario of traditional Chinese medicine compound prescriptions.
11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements a question-and-answer method adapted to the molecular mechanism scenario of traditional Chinese medicine compound prescriptions according to any one of claims 1-7.