System for acquiring natural medicinal material domain-specific knowledge

By combining dialogue applications, learning models, search engines and a system of natural medicinal materials special knowledge bases, the problem of difficulty in obtaining accurate natural medicinal materials in the existing technology is solved, and users can easily and efficiently obtain authoritative and accurate natural medicinal materials knowledge, improving the reliability and effectiveness of scientific research.

WO2025123546A1PCT designated stage expired Publication Date: 2025-06-19WESTLAKE UNIV

Patent Information

Application Number
PCT/CN2024/087712
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-04-15
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The existing technology is difficult to obtain accurate, standardized and comprehensive knowledge in the field of natural medicinal materials conveniently, efficiently and intelligently, resulting in incomplete or misleading results in information retrieval, affecting the progress of scientific research on natural medicinal materials.

Method used

It provides a system for obtaining knowledge about natural medicinal materials, combining dialogue applications, learning models, search engines and natural medicinal materials knowledge bases, providing users with a friendly interface through dialogue, and intelligently and accurately obtaining authoritative, accurate, standardized and comprehensive knowledge about natural medicinal materials.

Benefits of technology

It realizes users' convenient and efficient acquisition of special knowledge of natural medicinal materials, ensuring the accuracy and standardization of knowledge, thereby improving the reliability and effectiveness of scientific research in the field of natural medicinal materials, and promoting research progress in this field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024087712_19062025_PF_FP_ABST
    Figure CN2024087712_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to a system for acquiring natural medicinal material domain-specific knowledge. The system comprises a dialogue application program, a first learning model, a search engine, and a natural medicinal material domain-specific knowledge base. The dialogue application program provides a user interaction interface for a user, receives a user question of the user intending to acquire natural medicinal material domain-specific knowledge, and presents to the user an answer generated by the first learning model. The first learning model generates an answer to the latest user question on the basis of a dialogue history of the user, or uses the search engine to perform information retrieval on the natural medicinal material domain-specific knowledge base by employing at least one of coreference-based graph search, vector search and full-text search, to acquire background knowledge associated with the user question, and generates an answer to the user question on the basis of the dialogue history embedded with the background knowledge. According to the system of the present application, the user can intelligently and accurately acquire required authoritative, accurate, standardized and comprehensive natural medicinal material domain-specific knowledge in a friendly dialogue mode.
Need to check novelty before this filing date? Find Prior Art

Description

A system for acquiring domain-specific knowledge of natural medicinal materials Technical Field

[0001] The present application relates to the field of natural medicinal material information processing and application, and in particular to a system for acquiring domain-specific knowledge of natural medicinal materials. Background Art

[0002] Natural medicinal materials (NMMs) have long been recognized as a powerful reservoir of therapeutic agents, their importance reflected in the diversity and biological relevance of the compounds they produce. These compounds play key roles in addressing a wide range of pathological conditions, spanning infectious diseases to cancer, and continue to serve as a rich source of new drug leads. Furthermore, NMMs have an extensive history of clinical application worldwide, particularly in China, India, and the Arab world, demonstrating their enduring relevance in the global healthcare landscape. Despite their significant contributions to healthcare, the inherent complexity of NMMs is such that, for example, even ingredients with the same origin and medicinal site may differ only in their preparation methods, representing different NMMs. Furthermore, since their names are often not rigorously distinguished, searching for information using commonly used internet search tools or existing databases often yields only partial or even erroneous information.

[0003] Taking "Ephedra" as an example, the "Chinese Pharmacopoeia (2020 Edition)" records in detail that the term "Ephedra" refers to natural medicinal materials derived from several different species, specifically including Ephedra sinica (plant grass Ephedra), Ephedra intermedia (plant medium Ephedra) or Ephedra equisetina (plant horsetail Ephedra). However, when trying to search for "Ma Huang" or "Ephedra" through Internet search engines to seek specialized knowledge about "Ephedra", people often only get incomplete or misleading entries, such as "Ephedra is a medicinal preparation from the plant Ephedra sinica". Furthermore, as relevant knowledge about natural medicinal materials has not been stored in an authoritative and rigorous manner, existing dialogue-based platforms based on internet searches are not only unhelpful in acquiring specialized knowledge about natural medicinal materials, but may even exacerbate the aforementioned inaccuracies. For example, these platforms may assert in an affirmative tone that Ephedra sinica is the sole species of Ephedra. The incorrect or imprecise acquisition of specialized knowledge has created significant obstacles to scientific research on natural medicinal materials, and will inevitably undermine the reliability and validity of academic conclusions, and even hinder research progress in this field.

[0004] It can be seen that the current existing technology has not yet found a tool that can enable various users, including professionals in the field, to conveniently, efficiently and intelligently obtain accurate, standardized and comprehensive knowledge related to the field of natural medicinal materials.

[0005] Summary of the Invention

[0006] The present application is provided to solve the above-mentioned problems existing in the prior art.

[0007] There is a need for a system for acquiring domain-specific knowledge of natural medicinal materials, which can enable all types of users, including professionals in the field, to conveniently, efficiently and intelligently acquire accurate, standardized and comprehensive knowledge related to the field of natural medicinal materials.

[0008] According to the first scheme of the present application, a system for acquiring domain-specific knowledge of natural medicinal materials is provided, the system comprising a dialogue application, a first learning model, a search engine and a domain-specific knowledge base of natural medicinal materials, the dialogue application being configured to: provide a user interaction interface for a user, receive dialogue information input by a user on the user interaction interface, the dialogue information including user questions intended to acquire domain-specific knowledge of natural medicinal materials; and present answers to the user questions generated by the first learning model to the user; the first learning model being configured to: based on a dialogue history with the user, determine whether the dialogue history is sufficient to answer the latest user question; and if it is determined that the dialogue history is sufficient to answer the latest user question, generate answers to the user questions based on the dialogue history. the answer to the latest user question; when the first learning model determines that the dialogue history is insufficient to answer the latest user question, the dialogue information containing the latest user question is first processed, and the dialogue information after the first processing is used to interact with the search engine; the search engine is configured to: based on the dialogue information after the first processing, use at least one of a graph search based on co-reference, a vector search and a full-text search to perform information retrieval on the natural medicinal material domain knowledge base to obtain background knowledge associated with the latest user question, embed the background knowledge into the user's dialogue information to generate a dialogue history consisting of each round of dialogue information and the corresponding background knowledge, and return the dialogue history to the first learning model.

[0009] According to the system for acquiring domain-specific knowledge of natural medicinal materials in the embodiment of the present application, by combining and applying dialogue applications, learning models, search engines and domain-specific knowledge bases dedicated to the professional field of natural medicinal materials, users are provided with a dedicated system for acquiring domain-specific knowledge of natural medicinal materials. The system can provide users with a friendly interface in a dialogue manner, and can intelligently and accurately acquire the authoritative, accurate, standardized and comprehensive domain-specific knowledge of natural medicinal materials they need, so that scientific research in the field of natural medicinal materials has a consistent understanding and a unified knowledge base, thereby ensuring the validity and credibility of academic conclusions, and promoting the healthy development of research in this field.

[0010] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In the drawings, which are not necessarily drawn to scale, the same reference numerals may describe similar components in different views. The same reference numerals with letter suffixes or different letter suffixes may represent different instances of similar components. The accompanying drawings illustrate various embodiments generally by way of example and not limitation, and together with the description and claims, serve to illustrate the disclosed embodiments. Such embodiments are illustrative and are not intended to be exhaustive or exclusive of the present apparatus or method.

[0012] FIG1 shows a block diagram of part of the system for acquiring domain-specific knowledge of natural medicinal materials according to an embodiment of the present application.

[0013] FIG2( a ) shows a schematic diagram of a user interaction interface including dialogue guidance information according to an embodiment of the present application.

[0014] FIG2( b ) shows a schematic diagram of example questions and their answers according to an embodiment of the present application.

[0015] FIG3( a ) shows a schematic diagram of dialog information including user customized information according to an embodiment of the present application.

[0016] FIG3( b ) shows another schematic diagram of dialog information including user customized information according to an embodiment of the present application.

[0017] FIG4( a ) shows a schematic diagram of a user interaction interface including irrelevant questions according to an embodiment of the present application.

[0018] FIG4( b ) shows another schematic diagram of a user interaction interface including unrelated questions according to an embodiment of the present application.

[0019] FIG5( a ) shows a schematic diagram of a user interaction interface including multi-round dialogue information according to an embodiment of the present application.

[0020] FIG5( b ) shows another schematic diagram of a user interaction interface including multi-round dialogue information according to an embodiment of the present application.

[0021] FIG6 shows a schematic diagram of a coreference graph according to an embodiment of the present application.

[0022] FIG7 shows a schematic flow chart of a graph search based on coreference according to an embodiment of the present application.

[0023] FIG8 is a schematic diagram showing a process of vector search according to an embodiment of the present application.

[0024] FIG9 shows a schematic diagram of a full-text search process according to an embodiment of the present application. DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the technical solution of the present application, the present application is described in detail below in conjunction with the accompanying drawings and specific embodiments. The embodiments of the present application are further described in detail below in conjunction with the accompanying drawings and specific embodiments, but are not intended to limit the present application.

[0026] The words "first", "second" and similar terms used in this application do not indicate any order, quantity or importance, but are only used to distinguish. Words such as "include" or "comprise" mean that the elements before the word include the elements listed after the word, and do not exclude the possibility of also including other elements. The execution order of the various steps in the method described in this application with reference to the accompanying drawings is not intended to be limiting. As long as the logical relationship between the various steps is not affected, several steps can be integrated into a single step, a single step can be decomposed into multiple steps, and the execution order of the various steps can be changed according to specific needs.

[0027] It should also be understood that the term "and / or" in this application is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Furthermore, the character " / " in this application generally indicates that the associated objects are in an "or" relationship.

[0028] According to an embodiment of the present application, a system for acquiring domain-specific knowledge of natural medicinal materials is provided. FIG1 shows a block diagram of a portion of the system for acquiring domain-specific knowledge of natural medicinal materials according to an embodiment of the present application.

[0029] As shown in FIG1 , a system 100 for acquiring domain-specific knowledge of natural medicinal materials includes at least a dialogue application 101 , a first learning model 102 , a search engine 103 , and a domain-specific knowledge base 104 for natural medicinal materials.

[0030] The conversation application 101 may be configured to provide a user interface 1011 and receive conversation information entered by the user on the user interface 1011. In some embodiments, the conversation information may include, but is not limited to, user questions intended to obtain domain-specific knowledge about natural medicinal materials. Furthermore, the conversation application 101 may be configured to present to the user answers generated by the first learning model 102.

[0031] The first learning model 102 may be configured to determine, based on the conversation history with the user, whether the conversation history is sufficient to answer the latest user question, and then, if it is determined that the conversation history is sufficient to answer the latest user question, generate an answer to the latest user question based on the conversation history. For example, assuming the conversation history is as follows:

[0032] It can be seen that the above conversation history already contains information about "the system Chinese name of Ephedra", namely:

[0033] Therefore, the first model can generate an answer to the latest user question based on the above conversation history:

[0034] In other embodiments, when the first learning model 102 determines that the conversation history is insufficient to answer the latest user question, the conversation information containing the latest user question is first processed, and the conversation information after the first processing is used to interact with the search engine 103. The purpose of performing the first processing on the conversation information is to enable the processed conversation information to be better adapted to the search engine 103, including but not limited to extracting keywords from the conversation information that the user intends to obtain natural medicinal material-specific knowledge, or converting the conversation information into search terms in a specific format that facilitates the search engine 103 to search the natural medicinal material-specific knowledge base 104, etc., and this application does not impose specific restrictions on this.

[0035] The search engine 103 is configured to: based on the first processed conversation information, use at least one of co-reference-based graph search, vector search and full-text search to perform information retrieval on the natural medicinal material domain knowledge base 104 to obtain background knowledge associated with the latest user question, embed the background knowledge into the user's conversation information to generate a conversation history consisting of each round of conversation information and the corresponding background knowledge, and return the conversation history to the first learning model 102.

[0036] In the present application, by combining a dialogue application, a learning model, a search engine, and a domain knowledge base dedicated to the professional field of natural medicinal materials, a dedicated system for obtaining domain knowledge of natural medicinal materials is provided to users, so that when users want to understand and learn professional knowledge of natural medicinal materials, they do not have to use a general search engine to search blindly on the Internet and judge the accuracy of the information by themselves. Instead, they can use the system of the embodiment of the present application to ask questions in a dialogue manner in a friendly interface. Without understanding the process of knowledge search, they can intelligently and accurately obtain the authoritative, accurate, standardized, and comprehensive domain knowledge of natural medicinal materials they need. Therefore, the system for obtaining domain knowledge of natural medicinal materials according to the embodiment of the present application cleverly integrates the knowledge search process into the dialogue process, which not only has a good user experience, but also enables scientific research in the field of natural medicinal materials to have a consistent understanding and a unified knowledge base, thereby ensuring the validity and credibility of academic conclusions, and thus promoting the healthy development of research in this field.

[0037] Unlike general conversation applications, the conversation application according to the embodiments of the present application is specifically used for acquiring knowledge in the field of natural medicinal materials. Therefore, in some embodiments, the conversation application can provide users with conversation guidance information on the user interaction interface, wherein the conversation guidance information includes example questions for users to view the answers to the example questions by clicking.

[0038] Figure 2(a) shows a schematic diagram of a user interaction interface containing dialogue guidance information according to an embodiment of the present application. As shown in Figure 2(a), the dialogue application provides three example questions on the user interaction interface: "What is the systematic name of Ephedra?", "What is the NMM ID of Ephedra?", and "What is the species origin of Ephedra?", and the user can view the answer to the example question by clicking on any one of them. Figure 2(b) shows a schematic diagram of example questions and their answers according to an embodiment of the present application. As shown in Figure 2(b), when the user clicks on the first example question "What is the systematic name of Ephedra?", the system will present the answer corresponding to the question on the user interface, namely: the systematic name of Ephedra is Ephedra equisetina or Ephedra intermedia or Ephedra herbaceous stem (Ephedra equisetina vel intermedia vel sinica Stem-herbaceous).

[0039] By providing users with conversation guidance information, users, especially new users or users with little domain experience, can quickly and easily learn typical usage methods for conversational applications. Furthermore, the conversational application according to this application can also build a multi-angle, fully covered sample question library based on the typical natural medicinal material expertise that the system can obtain, and provide users with switching functions such as "changing a batch of sample questions" to enable users to obtain more accurate question assistance as much as possible.

[0040] In some embodiments, the conversation information between the user and the system can be in the form of natural language, such as text information input by the user, voice information, or text information converted from language information. Furthermore, the conversation information can also be in the form of natural language in different languages. In other words, the system according to the embodiments of the present application has the ability to learn and process natural language, especially cross-language natural language, which is of great significance for promoting international communication and standardization in the field of natural medicinal materials.

[0041] The conversation application according to the embodiment of the present application also supports user customization through conversation. As an example only, in the case where the conversation information contains information related to the answer style desired by the user, the answer to the user's question generated by the first learning model that matches the answer style information is presented to the user, wherein the answer style information includes at least one of the language of the answer, the format of the answer, and the length of the answer. Figure 3(a) shows a schematic diagram of conversation information containing user customization information according to an embodiment of the present application. In Figure 3(a), the user specifies the language of the answer by "Please answer in French" in the conversation information. In response to the user's customization of the language of the answer, the conversation application gives an answer to the user's question in French. In some embodiments, the language specified by the user for the answer includes but is not limited to Chinese and English, and may also include English-Chinese and Chinese-English, etc., and the specification of the language of the answer is not limited to the language used by the current user interface.

[0042] In other embodiments, the user can also specify the format of the answer, where the format of the answer includes but is not limited to JSON format. Figure 3(b) shows another schematic diagram of a dialogue message containing user-customized information according to an embodiment of the present application. In Figure 3(b), since the user specified in the dialogue message that the answer should be in JSON format, in response to the user's specification, the application provides a JSON-formatted answer to the question "What is the species origin of ephedra?", namely: {"species origin": ["Ephedra equisetina", "or", "Ephedra intermedia", "or", "Ephedra sinica"]}. In other embodiments, if supported by a natural medicinal material domain knowledge base, the system can also respond to the user's specification of answering in other formats, thereby allowing the user to conveniently integrate natural medicinal material domain knowledge in a specific format into other programs or databases, etc., further enhancing the practical value of the system in the present application.

[0043] In other embodiments, the user can also specify the length of the answer according to his needs, so as to specify a "short answer", "detailed answer", or specify the word range of the answer, etc. This application does not impose any specific restrictions on this.

[0044] In other embodiments, the conversational application according to embodiments of the present application does not provide answers to user questions that are not related to natural medicinal materials, Chinese medicinal materials, or traditional Chinese medicine. Figures 4(a) and 4(b) respectively illustrate schematic diagrams of user interaction interfaces in Chinese and English containing irrelevant questions according to embodiments of the present application.

[0045] As shown in Figures 4(a) and 4(b), the conversational application receives conversational messages in Chinese and English, respectively, through the user interface that are irrelevant to domain-specific knowledge about natural medicinal materials. In these cases, the conversational application will refuse to answer. Alternatively, it can prompt the user to ask questions related to domain-specific knowledge about natural medicinal materials. This enhances the security and professionalism of the conversational application and prevents users from being misled by inaccurate information or misleading them.

[0046] In some embodiments, the conversation information between the user and the conversation application is multi-round, and the conversation application is further configured to receive the multi-round conversation information input by the user on the user interaction interface, and present to the user the answers to the user questions contained in each round of conversation information generated by the first learning model. Figure 5 (a) and Figure 5 (b) respectively show schematic diagrams of the user interaction interface in Chinese and English containing multi-round conversation information according to an embodiment of the present application. The questions in the multi-round conversation information between the user and the conversation application can be related or independent of each other, and this application does not limit this. In addition, combined with the function of providing conversation guidance information to the user on the user interaction interface as described above, in the case of multi-round conversations, it is also possible to generate example questions in a targeted manner based on the conversation information of the previous round with the user during the interactive conversation with the user, and provide the user with updated example question options as the multi-round conversation progresses, thereby helping the user to acquire and learn the natural medicinal material domain knowledge they need more accurately and in-depth.

[0047] In addition, the conversation application can be further configured to provide options for collecting, downloading and quoting the user's previous conversation records on the user interaction interface, so that the user can add previous conversation records to the user's private information, or download original data associated with natural medicinal material knowledge, or quote pages of each conversation record.

[0048] In other embodiments, the conversational application can be further configured to provide a user evaluation option on the user interaction interface, so as to utilize the evaluation information submitted by the user to improve and train the first learning model. Taking Figure 5(a) as an example, it can be seen that a user evaluation option can be provided for each round of responses to user questions, allowing users to evaluate the responses to each question.

[0049] As described above, each of the conversation applications according to the embodiments of the present application shown in Figures 2(a) to 5(b) can be called by the user by accessing a web page. That is, there is no need to install a special application separately. The web page of the conversation application can be called by simply accessing a specific URL. This greatly facilitates the user to use the system of the present application on various terminal devices.

[0050] The natural medicinal material domain-specific knowledge base according to the embodiment of the present application is specially constructed to facilitate the acquisition of natural medicinal material domain-specific knowledge. It is configured to include parts such as systematic naming of natural medicinal materials, structured and standardized natural medicinal material knowledge, natural medicinal material terminology, natural medicinal material relationship sets and natural medicinal material-related texts.

[0051] The relevant texts on natural medicinal materials include at least the "Chinese Pharmacopoeia: 2020 Edition: Volume 1" or an updated and implemented version of the "Chinese Pharmacopoeia".

[0052] The structured and standardized natural medicinal materials knowledge is obtained by structuring and standardizing the natural medicinal materials related information in the "Chinese Pharmacopoeia: 2020 Edition: Volume 1" or the updated and implemented version of the "Chinese Pharmacopoeia", and covers all natural medicinal materials in the "Chinese Pharmacopoeia: 2020 Edition: Volume 1" or the updated and implemented version of the "Chinese Pharmacopoeia".

[0053] It can be seen that the system for obtaining domain-specific knowledge of natural medicinal materials according to the embodiment of the present application integrates the most authoritative information sources in the field of natural medicinal materials, which is also the basis for providing users with professional, standardized and comprehensive domain-specific knowledge of natural medicinal materials.

[0054] In addition, in order to enable the joint application of the search engine and the natural medicinal material domain knowledge base to better obtain natural medicinal material domain knowledge, the natural medicinal material domain knowledge base is further configured to store the natural medicinal material domain knowledge in JSON format, and the basic data structure is as follows:

[0055] Among them, NMM ID is the unique ID number of the natural medicinal material in the natural medicinal material domain knowledge base, and "field": "value" can be any attribute associated with the NMM ID and its corresponding value.

[0056] Furthermore, the JSON of natural medicinal material domain knowledge can also have a hierarchical nested data structure, as shown in the following example:

[0057] In an embodiment of the present application, the nesting depth of the data structure of the natural medicinal material domain knowledge JSON is set to the fewest possible number of layers, for example, the number of nesting layers does not exceed 4 layers, or 8 layers, or 16 layers, or 32 layers, or 64 layers. In addition, for closely related fields, they can also be stored in a parent field, such as "field 3" and "field 4" in the above example, as closely related fields, they are stored together in the parent field "field 2". In this way, when searching for "field 3" and "field 4", it is only necessary to retrieve the value of "field 2", so that when searching using a search engine, the search efficiency is higher. By controlling the depth of nesting, it is effectively avoided that the user cannot navigate well due to too many levels when browsing knowledge, and the performance degradation caused by the need for a deeper structure search when searching for knowledge.

[0058] According to the natural medicinal material domain knowledge base of the embodiment of the present application, the natural medicinal material relationship set includes multiple groups of natural medicinal material relationships, and each group of natural medicinal material relationships can be represented as a manually annotated [source object, relationship, target object] triple, wherein the source object and the target object include at least natural medicinal material terms, and the relationship includes at least a synonym relationship, a containment relationship, a derived / superior relationship, a derived / subordinate relationship, and a same-level relationship. As an example only, a synonym relationship includes, for example: [Traditional Chinese Medicine, Synonym (Synonym), Chinese Medicine]; a containment relationship includes, for example: [natural medicinal material, containment, processed natural medicinal material]; a derived / superior relationship includes, for example: [Artemisia annua, derived / superior, plant Artemisia annua]; a derived / subordinate relationship includes, for example: [Artemisia annua, derived / subordinate, Artemisia annua segment]; a same-level relationship includes, for example: [Ginseng root, same-level, Ginseng leaf]. In addition, the source object and the target object can be nouns and terms in different languages, respectively. For example, the synonym relationship can also include cross-language synonyms, such as [natural medicinal material, cross-language synonym, Natural Medicinal Material], etc., which are not listed here one by one.

[0059] In some embodiments, each natural medicinal material term in the natural medicinal material domain knowledge base has a corresponding unique coreference subject. A "coreference term" refers to a group of words with the same concept, and the subject of the natural medicinal material term corresponding to a representative word in this group is called the coreference subject of this concept. Labeling natural medicinal material terms with their corresponding unique coreference subject can effectively avoid ambiguity and improve the standardization of natural medicinal material domain knowledge.

[0060] Based on the natural medicinal material relationship set, a search engine can be used to search for each group of natural medicinal material relationships with synonymous relationships in the natural medicinal material relationship set, thereby constructing a coreference primary term graph (CPTG). The constructed coreference primary term graph contains all coreference terms and can reflect the synonymous relationships between each natural medicinal material term. Therefore, it can be seen that the CPTG can basically be considered a directed acyclic graph (DAG).

[0061] FIG6 shows a schematic diagram of a coreference graph according to an embodiment of the present application. When searching for each group of natural medicinal material relationships with synonymous relationships in the natural medicinal material relationship set using a search engine, for example, the synonymous relationships associated with the term "ephedra" in the search results may include [Ephedrae Herba, Synonym, Ephedra], [Ephedra, Synonym, Ephedra], [Ma huang, Synonym, Ephedra], as well as [Ephedra, Synonym, nmm-0006], [Ma-huang, Synonym, nmm-0006], [Ephedra equisetina vel intermedia vel sinica Stem-herbaceous, Synonym, nmm-0006], and based on the above synonymous search results, a coreference graph (partial) as shown in FIG6 can be constructed. As can be seen from Figure 6, NMM ID serves as the subject of each natural medicinal material term in the natural medicinal material knowledge base. In the constructed coreference subject graph, it serves as the coreference subject of each natural medicinal material term. All natural medicinal material terms with the same concept, as well as natural medicinal material terms that have synonymous relationships with these natural medicinal material terms (including cross-language synonymous relationships, etc.), can eventually point to the coreference subject through the search engine. Specifically, in Figure 6, Ephedrae Herba, Ephedra, Ma huang, Ma-huang, Ephedra equisetina or Ephedra intermedia or Ephedra herbaceous stem, and Ephedra equisetina vel intermedia vel sinica Stem-herbaceous all point to the coreference subject nmm-0006.

[0062] Given the aforementioned configurations of the natural medicinal material domain-specific knowledge base and search engine according to an embodiment of the present application, the following, in conjunction with Figures 7-9 , details the process of using a search engine to search the natural medicinal material domain-specific knowledge base to obtain the natural medicinal material domain-specific knowledge required by the user. As an example, a simplified NMM knowledge document for "ephedra" is as follows:

[0063] Figure 7 shows a flow chart of a graph search based on co-reference according to an embodiment of the present application. In some embodiments, the first learning model is further configured to have information extraction capabilities, so that the first processing of the conversation information containing the latest user questions by the first learning model may specifically include extracting natural medicinal material entities and search fields from the conversation information. Thus, when the first learning model determines that the conversation history is not sufficient to answer the latest user questions, it is necessary to interact with the search engine. Before the interaction, it is necessary to perform a first processing on the conversation information containing the latest user questions, wherein the first processing can be associated with different search modes of the search engine. For example, the graph search mode based on co-reference can correspond to the above-mentioned first processing of extracting natural medicinal material entities and search fields from the conversation information. As an example only, assume that the conversation history is as follows:

[0064] Because the conversation history doesn't include information about "the systematic name of ephedra," the first learning model needs to interact with the search engine. As shown in Figure 7 , in step 701, when using a coreference-based graph search to retrieve information from the natural medicinal material domain knowledge base, the first learning model can first extract natural medicinal material entities and search fields from the conversation information containing the most recent user question.

[0065] For a user question like "What is the system name for ephedra?", "ephedra" is the natural medicinal material entity, and "system name" is the search field. That is, the first learning model can extract the following NMM entity name and field name:

[0066] Next, in step 702, the search engine searches the coreference graph for the natural medicinal material entity and the search field to obtain a first coreference term corresponding to the natural medicinal material entity and a second coreference term corresponding to the search field. Taking the natural medicinal material entity's coreference term as the NMM ID and the search field's coreference term as the field ID as an example, by searching the coreference graph for "ephedra" as the natural medicinal material entity and "system name" as the search field, the search engine obtains the first coreference term "nmm-0006" and the second coreference term "nmmsn":

[0067] as well as

[0068] It should be noted that the coreference subject of the natural medicinal material entity and the coreference subject of the search field can also be pre-set to other coreference terms as long as they are unique representations in the system.

[0069] In step 703, the search engine constructs a search instruction based on the first coreference subject and the second coreference subject. For the above example, the natural medicinal material entity "ephedra" and the search field "system name" can be concatenated to construct the following search instruction:

[0070] It is worth noting that in the search instruction "queries", nmm and field_list are still retained. These two fields will not be used directly in subsequent searches, but they can be embedded in the conversation history in the answers generated by the subsequent first learning model. This helps to prompt the context of the first learning model's search, thereby improving the accuracy of the first learning model.

[0071] In some embodiments, two copies of the first learning model can be used to extract NMM entities and search fields in parallel and then concatenate them. This can significantly reduce the time required to construct the search instruction. Furthermore, since the extraction of NMM entities and search fields is independent of each other, when training the first learning model, training can be performed separately for the extraction of NMM entities and search fields, thereby further improving the accuracy of the first learning model. Of course, in this manner, when a user's question contains more than one NMM entity and / or search field, all fields must be searched for each NMM entity.

[0072] In other embodiments, a single first learning model may be used to simultaneously extract NMM entities and search fields, and then construct search instructions. In this way, the fields to be searched may be specified for each NMM, greatly reducing the search space and improving the search speed.

[0073] In other embodiments, a serial extraction method of NMM entities and search fields may be employed. For example, NMM entities may be extracted first, then the user question and the extracted NMM entities are concatenated, the search fields are extracted again using the first learning model, and then a search instruction containing each NMM entity and its corresponding search field is generated. Alternatively, the search field may be extracted first, then the user question and the extracted search field are concatenated, the NMM entity is extracted again using the first learning model, and then a search instruction containing each NMM entity and its corresponding search field is generated.

[0074] As described above, the construction methods of various search instructions can all be trained in a supervised manner, which will not be described in detail here.

[0075] In step 704, the search engine searches for information in the natural medicinal material domain knowledge base based on the search instruction, and uses the information search results as background knowledge associated with the latest user question; wherein, the coreference subject graph is constructed by using the search engine in advance or in real time to search for each group of natural medicinal material relationships with synonymous relationships in the natural medicinal material relationship set, and the coreference subject graph contains all coreference subjects and reflects the synonymous relationship between each natural medicinal material term.

[0076] It is worth noting that when a co-referential term graph is pre-constructed using a search engine, the constructed co-referential term graph can be updated together with other collections such as terms in the natural medicinal materials domain knowledge base to ensure that the co-referential term graph is in the latest state.

[0077] The search engine retrieves information from the natural medicinal materials domain knowledge base based on the above search instructions and obtains the following background knowledge:

[0078] The background knowledge obtained in the above manner is embedded into the user's conversation information to generate the following conversation history consisting of each round of conversation information and the corresponding background knowledge:

[0079] Then, the conversation history is returned to the first learning model so that the first learning model generates an answer to the user's question, as follows:

[0080] shennongchat_response_2: The systematic name of ephedra is Ephedra equisetina vel intermedia vel sinica Stem-herbaceous.

[0081] The coreference-based graph search shown in FIG7 is a preferred search method in a system for acquiring domain-specific knowledge of natural medicinal materials according to an embodiment of the present application, and it can maximize the use of the structured knowledge in the domain-specific knowledge base of natural medicinal materials in the present application. By coreference-based subjects, corresponding standard information in other sets of the domain-specific knowledge base of natural medicinal materials is very suitable for application scenarios such as vocabulary and terminology translation between multiple languages, standardized query or disambiguation.

[0082] In addition to the above-mentioned coreference-based graph search, the search engine of the present application can also perform vector search. To support the search engine's vector search, the natural medicinal material domain knowledge base can be further configured to: store the natural medicinal material knowledge as a slice data set consisting of multiple slice data, and store the first vector embedding representation corresponding to each slice data in the vector database, wherein each slice data is associated with a corresponding coreference subject.

[0083] Slicing refers to pre-dividing the natural medicinal material knowledge into multiple small pieces. Each small piece is slice data, and the granularity of the slice data can be varied. For example, it can contain one or more fields about natural medicinal materials, or it can be part of a field, etc. This provides maximum flexibility and is not specifically limited in this application. Still taking the aforementioned nmm-0006 as an example, its knowledge set can be converted into the following slice data set:

[0084] From the above slice data set, we can see that each slice data is associated with a corresponding coreference subject and a corresponding first vector embedding representation. The first vector will be stored in the vector database for retrieval.

[0085] Correspondingly, the first learning model is further configured to have the ability to extract information and generate vector embedding representations of text, and the first processing of the dialogue information containing the latest user question specifically includes: extracting natural medicinal material entities from the dialogue information, using the search engine to search in the co-referential subject graph to obtain the first co-referential subject corresponding to the natural medicinal material entity, constructing a user question containing the first co-referential subject, and generating a second vector embedding representation corresponding to the user question containing the first co-referential subject. It is worth noting that the first vector embedding representation corresponding to the slice data in the slice data set of natural medicinal material knowledge and the second vector embedding representation corresponding to the user question should use the same embedding method, such as using the same word embedding model or any other applicable model that can generate vector embedding representations based on text, and this application does not limit this.

[0086] FIG8 is a flow chart of a vector search according to an embodiment of the present application. As shown in FIG8 , in step 801, when using vector search to retrieve information from the natural medicinal material domain knowledge base, the first learning model can first extract the natural medicinal material entity from the conversation information, use the search engine to search the co-referential subject graph to obtain the first co-referential subject corresponding to the natural medicinal material entity, construct a user question containing the first co-referential subject, and generate a second vector embedding representation corresponding to the user question containing the first co-referential subject. Take the conversation information containing the following user question as an example:

[0087] {

[0088] "user_message":"What is the species origin of Ephedra?"

[0089] }

[0090] The first learning model extracts natural medicinal material entities from the above conversation information:

[0091] Furthermore, the search engine is used to search the coreference graph to obtain the first coreference corresponding to the natural medicinal material entity, that is, NMM ID:

[0092] As shown in FIG7 , the coreference term graph is constructed by searching each group of natural medicinal material relationships having synonymous relationships in the natural medicinal material relationship set in advance or in real time using the search engine. The coreference term graph contains all coreference terms and reflects the synonymous relationship between each natural medicinal material term.

[0093] Then, construct a natural language user question containing the first coreference subject as follows:

[0094] Next, the second vector embedding representation corresponding to the user question containing the first coreference subject is generated as follows:

[0095] {

[0096] "user_message":"What is the species origin of Ephedra?\n{\"nmm_to_nmm_id\":

[0097] {\"Ephedra\":\"nmm-0006\"}}",

[0098] "user_message_embedding":[...]

[0099] }

[0100] As shown in FIG8 , next, in step 802, the search engine can retrieve matching first vector embedding representations from the vector database based on the second vector embedding representation, i.e., user_message_embedding, and use the set of slice data (each matching content) corresponding to each matching first vector embedding representation as background knowledge associated with the latest user question. In this embodiment, the slice data corresponding to the matching first vector embedding representation is as follows:

[0101] Specifically, when comparing the similarity between user_message_embedding and all first vector embedding representations in the vector database, for example, methods including but not limited to calculating the cosine similarity between vectors can be used to obtain all slice data whose relevance to the user's question meets the preset standard. On the basis of quickly and accurately identifying the natural medicinal material related information associated with the user's question, the search engine parameters can be configured to allow it to flexibly and variably return 1 or n slice data.

[0102] On this basis, background knowledge can be further embedded into the user's conversation information to generate the following conversation history consisting of each round of conversation information and the corresponding background knowledge:

[0103] Then, the conversation history can be returned to the first learning model so that the first learning model generates an answer to the user's question, as follows:

[0104] shennongchat_response_2: The origin of the species of Ephedra is Ephedra equisetina, Ephedra intermedia, or Ephedra sinica.

[0105] The vector search method shown in Figure 8 can accept entire sentences or paragraphs as search input and convert these texts into vector space through embedding, so that the converted vector embedding representation contains the semantic essence of the sentence or paragraph. Subsequently, by determining the similarity between the vector embedding representations, it can quickly identify and retrieve knowledge and information that is semantically closely related to the text to be queried, thereby providing users with more relevant and accurate answers.

[0106] In other embodiments, full-text search can also be used to obtain answers to user questions. To this end, the natural medicinal material domain knowledge base can be further configured to predefine index fields for full-text search, and store the natural medicinal material domain knowledge as a searchable document with an inverted index of each index field. Among them, the inverted index is to sort the searchable documents with the same index field according to the relevance / importance of each document, and store each index field as an easy-to-retrieve "dictionary tree" so that the index field can be found very quickly in the "dictionary tree", and then the corresponding document can be found according to the index field. In addition, the first learning model is further configured to have word segmentation capabilities, information extraction capabilities and the ability to generate document summaries, and the first processing of the dialogue information containing the latest user question includes extracting full-text search keywords based on the word segmentation of the dialogue information.

[0107] Figure 9 shows a schematic diagram of the full-text search process according to an embodiment of the present application. As shown in Figure 9, in step 901, the first learning model first extracts full-text search keywords based on word segmentation of the conversation information.

[0108] In step 902, the search engine retrieves matching searchable documents based on the full-text search keywords and the inverted index while performing information retrieval on the natural medicinal material domain knowledge base using a full-text search.

[0109] In step 903, the searchable documents are used as background knowledge associated with the latest user question, or the first learning model is used to extract document summaries and / or key information from the set of matching searchable documents, and the generated document summaries and / or key information are used as background knowledge associated with the latest user question.

[0110] In other embodiments, when full-text search is used to retrieve information from the natural medicinal material domain knowledge base, the search engine can be further configured to: use other external search systems to search the full-text search keywords to generate Internet search results, and fuse the Internet search results with the document summaries and / or key information generated by the information retrieval of the natural medicinal material domain knowledge base to generate background knowledge associated with the latest user questions, wherein the Internet search results have a lower fusion priority.

[0111] As described above, the full-text search mode allows for efficient search and retrieval of relevant information, and is suitable for application scenarios such as approximate matching, phrase variations, spelling errors, typos, or alternative spellings. In addition, by pre-establishing an inverted index for each searchable document in the natural medicinal materials domain knowledge base, the full-text mode has higher system flexibility and user-friendliness.

[0112] In some embodiments, the specific use of coreference-based graph search, vector search, and full-text search, or a combination thereof, may be determined by the user through selection of a search mode, or may be determined by the first learning model, for example, based on features in the user's conversation information. This application does not limit this. By applying various search modes individually or in combination, search results can be both accurate and comprehensive, thereby improving the overall usability of the system in this application.

[0113] In addition, although exemplary embodiments have been described herein, the scope includes any and all embodiments based on the present application with equivalent elements, modifications, omissions, combinations (e.g., solutions that intersect various embodiments), adaptations, or changes. The elements in the claims are to be interpreted broadly based on the language used in the claims and are not limited to the examples described in this specification or during the prosecution of this application, which examples are to be interpreted as non-exclusive. Therefore, this specification and examples are intended to be considered as examples only, with the true scope and spirit being indicated by the following claims and the full scope of their equivalents.

[0114] The above description is intended to be illustrative rather than restrictive. For example, the above examples (or one or more of their solutions) can be used in combination with each other. For example, those of ordinary skill in the art may use other embodiments when reading the above description. In addition, in the above-mentioned specific embodiments, various features can be grouped together to simplify the application. This should not be interpreted as an intention that a disclosed feature that is not required to be protected is necessary for any claim. On the contrary, the subject matter of the present application may be less than all the features of a specific disclosed embodiment. Thus, the following claims are incorporated into the specific embodiments as examples or embodiments, wherein each claim is independently a separate embodiment, and it is considered that these embodiments can be combined with each other in various combinations or arrangements. The scope of the present invention should be determined with reference to the appended claims and the full scope of equivalents to which these claims are entitled.

[0115] The above embodiments are merely exemplary embodiments of the present application and are not intended to limit the present invention. The scope of protection of the present invention is defined by the claims. Those skilled in the art may make various modifications or equivalent substitutions to the present invention within the spirit and scope of protection of the present application, and such modifications or equivalent substitutions shall also be deemed to fall within the scope of protection of the present invention.

Claims

1. A system for acquiring domain-specific knowledge of natural medicinal materials, characterized in that: The system includes a dialogue application, a first learning model, a search engine, and a natural medicinal material domain-specific knowledge base. The dialog application is configured to: Providing a user interaction interface for the user, and receiving dialogue information input by the user on the user interaction interface, wherein the dialogue information includes questions asked by the user with the intention of obtaining domain-specific knowledge of natural medicinal materials; and Presenting to the user an answer to the user's question generated by the first learning model; The first learning model is configured as: Based on the conversation history with the user, determining whether the conversation history is sufficient to answer the latest user question; If it is determined that the conversation history is sufficient to answer the latest user question, generating an answer to the latest user question based on the conversation history; When the first learning model determines that the conversation history is insufficient to answer the latest user question, performing a first processing on the conversation information including the latest user question, and using the first processed conversation information to interact with the search engine; The search engine is configured to: Based on the first processed dialogue information, at least one of co-reference-based graph search, vector search and full-text search is used to perform information retrieval on the natural medicinal material domain knowledge base to obtain background knowledge associated with the latest user question, and the background knowledge is embedded in the user's dialogue information to generate a dialogue history consisting of each round of dialogue information and the corresponding background knowledge, and the dialogue history is returned to the first learning model.

2. The system according to claim 1, characterized in that The dialogue application is further configured to provide dialogue guidance information to the user on the user interaction interface, wherein the dialogue guidance information includes sample questions for the user to click to view answers to the sample questions.

3. The system according to claim 1 or 2, characterized in that: The dialogue information is dialogue information in a natural language form.

4. The system according to claim 3, characterized in that The dialogue information is dialogue information in natural language form in different languages.

5. The system according to claim 1 or 2, characterized in that: The dialog application is further configured to: In the case where the dialogue information includes information related to an answer style desired by the user, an answer to the user's question generated by the first learning model and matching the information related to the answer style is presented to the user, wherein The answer style related information includes at least one of the language of the answer, the format of the answer, and the length of the answer.

6. The system according to claim 5, characterized in that The languages ​​of the answers include at least Chinese and English.

7. The system according to claim 5, characterized in that The format of the answer includes at least JSON.

8. The system according to claim 1 or 2, characterized in that: The dialogue information is multiple rounds, and the dialogue application is further configured as follows: Receiving multiple rounds of dialogue information input by a user on the user interaction interface; Answers to user questions contained in each round of dialogue information generated by the first learning model are presented to the user.

9. The system according to claim 1 or 2, characterized in that: The conversation application is further configured to not provide answers to user questions that are not related to natural medicinal materials, Chinese medicinal materials, or traditional Chinese medicine.

10. The system according to claim 1 or 2, characterized in that: The dialog application is further configured to: The user interaction interface provides options for collecting, downloading and quoting the user's previous conversation records, so that the user can add previous conversation records to the user's private information, or download original data associated with natural medicinal material knowledge, or quote pages of each conversation record.

11. The system according to claim 1 or 2, characterized in that: The dialog application is further configured to: A user evaluation option is set on the user interaction interface so as to improve and train the first learning model using the evaluation information submitted by the user.

12. The system according to claim 1 or 2, characterized in that: The dialog application is further configured such that: the user invokes the dialog application by accessing a web page.

13. The system according to claim 1 or 2, characterized in that: The natural medicinal material domain-specific knowledge base is configured to include natural medicinal material systematic nomenclature, structured and standardized natural medicinal material knowledge, natural medicinal material terminology, natural medicinal material relationship sets and natural medicinal material related texts, wherein: The relevant texts of natural medicinal materials shall at least include the "Chinese Pharmacopoeia: 2020 Edition: Volume 1" or an updated and implemented version of the "Chinese Pharmacopoeia"; The structured and standardized natural medicinal material knowledge is obtained by structuring and standardizing the natural medicinal material related information in the "Chinese Pharmacopoeia: 2020 Edition: Volume 1" or an updated and implemented version of the "Chinese Pharmacopoeia", and covers the "Chinese Pharmacopoeia: 2020 Edition: Volume 1" or an updated and implemented version of the "Chinese Pharmacopoeia". All natural herbs in; Each of the natural medicinal material terms has a unique coreference subject corresponding thereto; The natural medicinal material relationship set includes multiple groups of natural medicinal material relationships, and the natural medicinal material relationships are represented as manually annotated [source object, relationship, target object] triples, wherein the source object and the target object include at least natural medicinal material terms, and the relationships include at least synonym relationships, inclusion relationships, derived / superior relationships, derived / subordinate relationships, and sibling relationships.

14. The system according to claim 13, characterized in that The first learning model is further configured to have information extraction capability, wherein the first processing of the dialogue information including the latest user question specifically includes extracting natural medicinal material entities and search fields from the dialogue information; The search engine is further configured to, when performing information retrieval on the natural medicinal material domain-specific knowledge base using a coreference-based graph search: Searching the natural medicinal material entity and the search field in a coreference term graph to obtain a first coreference term corresponding to the natural medicinal material entity and a second coreference term corresponding to the search field; constructing a search instruction based on the first coreference subject and the second coreference subject; Based on the search instruction, information is retrieved from the natural medicinal material domain knowledge base, and the information retrieval result is used as background knowledge associated with the latest user question; wherein, The co-referential term graph is constructed by searching each group of natural medicinal material relationships with synonymous relationships in the natural medicinal material relationship set in advance or in real time using the search engine. The co-referential term graph contains all co-referential terms and reflects the synonymous relationship between each natural medicinal material term.

15. The system according to claim 13, characterized in that The natural medicinal material domain-specific knowledge base is further configured to: store the natural medicinal material knowledge as a slice data set consisting of a plurality of slice data, and store the first vector embedding representation corresponding to each slice data in a vector database, wherein each slice data is associated with a corresponding coreference subject; The first learning model is further configured to have information extraction capabilities and the ability to generate vector embedding representations of texts, and the first processing of the dialogue information containing the latest user question specifically includes: extracting a natural medicinal material entity from the dialogue information, using the search engine to search in a co-referential subject graph to obtain a first co-referential subject corresponding to the natural medicinal material entity, constructing a user question containing the first co-referential subject, and generating a second vector embedding representation corresponding to the user question containing the first co-referential subject; The search engine is further configured to: search the vector database based on the second vector embedding representation when performing information retrieval on the natural medicinal material domain knowledge base by using vector search. The first matching vectors are embedded into representations, and the set of slice data corresponding to each matching first matching vector is embedded into representations as background knowledge associated with the latest user question; wherein, The co-referential term graph is constructed by searching each group of natural medicinal material relationships with synonymous relationships in the natural medicinal material relationship set in advance or in real time using the search engine. The co-referential term graph contains all co-referential terms and reflects the synonymous relationship between each natural medicinal material term.

16. The system according to claim 13, characterized in that The natural medicinal material domain-specific knowledge base is further configured to: predefine index fields for full-text search, and store the natural medicinal material domain-specific knowledge as a searchable document with an inverted index of each index field; The first learning model is further configured to have word segmentation capabilities, information extraction capabilities, and document summary generation capabilities, and the first processing of the dialogue information containing the latest user question specifically includes: extracting full-text search keywords based on word segmentation of the dialogue information; The search engine is further configured to, when using full-text search to retrieve information from the natural medicinal materials domain knowledge base: based on the full-text search keywords and the inverted index, retrieve matching searchable documents; use the searchable documents as background knowledge associated with the latest user questions, or use the first learning model to perform document summaries and / or key information extraction on a collection of matching searchable documents, and use the generated document summaries and / or key information as background knowledge associated with the latest user questions.

17. The system according to claim 16, characterized in that The search engine is further configured to, when full-text search is used to retrieve information from the natural medicinal material domain knowledge base: Using other external search systems to search the full-text search keywords to generate Internet search results; The Internet search results are integrated with document summaries and / or key information generated by information retrieval of the natural medicinal material domain knowledge base to generate the latest background knowledge associated with the user's question, wherein the Internet search results have a lower integration priority.

Citation Information

Patent Citations

  • Traditional Chinese medicine video retrieval model based on knowledge graph

    CN116805013A

  • Large language model intelligent inquiry dialogue method, system and device and medium

    CN117093679A

  • Traditional Chinese medicine field knowledge graph question-answering method based on semantic analysis

    CN117171329A

  • System for acquiring special domain knowledge of natural medicinal materials

    CN117648424A

  • Interactive questioning method, interactive questioning system, interactive questioning program, and recording medium with the same program recorded thereon

    JP2007322836A

Cited By

  • Electric power knowledge retrieval system based on large language model

    CN120910228A