Method and device for obtaining information of traditional Chinese medicine compounds online and medium
By combining web crawlers and large language models in an online human-computer interaction method, the system automatically acquires and filters noise from multiple TCM internet knowledge sources, solving the problems of lag and reliability in acquiring information on compounds of authentic medicinal materials, and achieving efficient and accurate data updates and storage.
Patent Information
- Application Number
- CN202410751316.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-12
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2044-06-12
AI Technical Summary
Existing technologies suffer from information lag, limitations, reliability and accuracy issues when acquiring information on compounds from authentic medicinal materials. Furthermore, manual data processing carries the risk of resource waste and misinterpretation.
By utilizing web crawlers and large language models for online human-computer interaction, and by constructing retrieval rules and heuristic rules defined by experts in authentic medicinal materials, the system automatically acquires and filters noise from multiple TCM internet knowledge sources, constructs a structured template for extracting compound information from authentic medicinal materials, and saves it to the database.
It enables timely updates of information, diversified sources, and high reliability, avoiding waste of resources and manpower, and improving the accuracy and reliability of data.
Smart Images

Figure CN118467808B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of traditional Chinese medicine compound information extraction, and more particularly to a traditional Chinese medicine compound information online acquisition method, device and medium. BACKGROUND
[0002] Traditional Chinese medicine is a medicinal plant or animal that grows in a specific geographical environment. Due to the uniqueness of its growth environment, it has special quality and efficacy. In the field of medicinal materials, obtaining the compound information of traditional Chinese medicine is of great significance. It not only helps to understand the pharmacological properties of medicinal materials, but also has a profound significance for the research of drug efficacy, quality control, geographical indication certification, development of new drugs, and ecological protection.
[0003] At present, the methods for obtaining traditional Chinese medicine compound information mainly include direct acquisition method and indirect acquisition method. The direct acquisition method mainly includes the use of traditional analysis methods and modern analysis techniques. According to different research purposes, medicinal material properties and available technical equipment, chromatography, mass spectrometry, infrared spectroscopy, molecular biology methods, computational chemistry methods, etc. are used to directly extract compound information from traditional Chinese medicine and record it. The advantage of this method is that it has the opportunity to discover new traditional Chinese medicine compound information. However, its shortcomings are also very obvious, which will result in a large amount of repeated information acquisition, leading to excessive waste of resources.
[0004] The indirect acquisition method mainly includes the use of academic databases, plant databases, online libraries and electronic books, traditional Chinese medicine databases, etc. According to different research purposes, medicinal material properties and other needs of traditional Chinese medicine experts, different query conditions are designed to obtain a large amount of traditional Chinese medicine compound related data, and then structured traditional Chinese medicine compound information is extracted from these data manually. The advantage of this method is that it effectively utilizes the research results of predecessors, avoids part of the experimental work, reduces the waste of resources, and quickly develops downstream research work. However, the disadvantage of this method is that it has problems in information lag, limitation, reliability, accuracy, etc. For example, the lag of information caused by the out-of-date data source, the limitation caused by the single source of database information, and the problems of reliability and accuracy caused by the loose management and maintenance of the database. In addition, manually sorting a large amount of traditional Chinese medicine compound related data has the risk of objectivity, time waste, manpower loss and error propagation. SUMMARY
[0005] The present application is provided to solve the above-mentioned problems in the prior art. Therefore, a traditional Chinese medicine compound information online acquisition method, device and medium are needed.
[0006] According to the first aspect of the present application, a traditional Chinese medicine compound information online acquisition method is provided, which comprises:
[0007] Utilize the large-scale targeted information webpage acquisition capability of the web crawler, through human-computer interaction of search rule definition with experts of native medicinal materials, online acquire the webpage set of native medicinal material compound information from multiple internet knowledge sources of traditional Chinese medicine;
[0008] Construct the knowledge acquisition prompt rule instruction of the native medicinal material compound field, obtain the candidate set of native medicinal material compound structured data with noise through online human-computer interaction with a large language model; utilize the heuristic rule method to construct the native medicinal material compound structured information implication relationship discriminator, which is used to filter the candidate set of native medicinal material compound structured data with noise, and obtain the data set containing structured native medicinal material compound information;
[0009] Establish the structured native medicinal material compound information extraction template, automatically extract the structured native medicinal material compound information from the data set containing structured native medicinal material compound information, and save it to the database.
[0010] Further, utilize the large-scale targeted information webpage acquisition capability of the web crawler, through human-computer interaction of search rule definition with experts of native medicinal materials, online acquire the webpage set of native medicinal material compound information from multiple internet knowledge sources of traditional Chinese medicine, including:
[0011] Based on the search rule set R constructed by experts of native medicinal materials, the native medicinal material compound information directional webpage acquirer automatically fills in the search rule information, and acquires the webpage set containing the native medicinal material compound information from the open source internet knowledge source of traditional Chinese medicine.
[0012] Further, construct the knowledge acquisition prompt rule instruction of the native medicinal material compound field, obtain the candidate set of native medicinal material compound structured data with noise through online human-computer interaction with a large language model; utilize the heuristic rule method to construct the native medicinal material compound structured information implication relationship discriminator, which is used to filter the candidate set of native medicinal material compound structured data with noise, and obtain the data set containing structured native medicinal material compound information, including:
[0013] The collection of web pages containing information of native medicinal material compounds is traversed in a loop, the abstract of each web page is input into a large language model, a prompt instruction of information extraction of native medicinal material compounds is extracted according to the native medicinal material compound information prepared by experts, the template filling position in the prompt instruction of information extraction is filled by using the native medicinal material list constructed in advance, the large language model automatically extracts the compound information text content contained in the abstract of the web page, and the structured information of the native medicinal material compounds is determined by whether the output structure required in the fill-in-the-blank prompt is contained in the compound information text content: chemical composition: chemical composition name, whether the structured information of the native medicinal material compounds is contained in the compound information text content is determined, and then the semi-structured information of the native medicinal material compounds is extracted.
[0014] Further, a structured native medicinal material compound information extraction template is established, structured native medicinal material compound information is automatically extracted from a data set containing structured native medicinal material compound information, and is saved to a database, including:
[0015] The output structure required in the prompt instruction given by experts is used to construct a heuristic information extraction rule template: chemical composition: [X] -> [X]: {x1, x2, ……}, wherein X is the extracted text description of the native medicinal material compound, -> represents a regular expression, xi represents each native medicinal material compound description obtained after segmentation, and finally a list of native medicinal material compound descriptions {x1, x2, ……} is obtained.
[0016] Further, after the list of native medicinal material compound descriptions {x1, x2, ……} is extracted, the method further includes:
[0017] It is judged whether the text content in the list of native medicinal material compound descriptions {x1, x2, ……} is Chinese or English:
[0018] If it is Chinese, it is translated into English by using a machine translation system to obtain the English name of the native medicinal material compound;
[0019] If it is English, it is translated into Chinese by using a machine translation system to obtain the Chinese name of the native medicinal material compound;
[0020] Finally, the structured information of the native medicinal material compound is obtained.
[0021] According to the second technical solution of the present application, a device for online obtaining information of native medicinal material compounds is provided, and the device includes:
[0022] The webpage set acquisition module is configured to utilize the large-scale directional information webpage acquisition capability of the web crawler, to acquire a webpage set of the native medicinal material compound information online from multiple internet knowledge sources of traditional Chinese medicine through human-computer interaction of search rule definition with experts of the native medicinal material, and to acquire the webpage set of the native medicinal material compound information online from multiple internet knowledge sources of traditional Chinese medicine through human-computer interaction of search rule definition with experts of the native medicinal material.
[0023] The data acquisition module is configured to construct a knowledge acquisition prompt rule instruction of the native medicinal material compound field, to obtain a candidate set of the native medicinal material compound structured data with noise through online human-computer interaction with the large language model, and to construct a native medicinal material compound structured information implication relationship discriminator by using a heuristic rule method, to filter the candidate set of the native medicinal material compound structured data with noise, and to obtain a data set containing the structured native medicinal material compound information.
[0024] The information extraction module is configured to establish a structured native medicinal material compound information extraction template, to automatically extract the structured native medicinal material compound information from the data set containing the structured native medicinal material compound information, and to save the structured native medicinal material compound information to a database.
[0025] Further, the webpage set acquisition module is further configured to automatically fill the search rule information based on a search rule set R constructed by the experts of the native medicinal material, to acquire the webpage set containing the native medicinal material compound information from the open-source internet knowledge sources of traditional Chinese medicine.
[0026] Further, the data acquisition module is further configured to traverse the webpage set containing the native medicinal material compound information in a loop, to extract the abstract of each webpage and input the abstract into the large language model, to respectively fill the template filling position in the information extraction prompt instruction by using the native medicinal material list constructed in advance according to the native medicinal material compound information extraction prompt instruction prepared by the experts of the native medicinal material, to automatically extract the compound information text content contained in the webpage abstract by the large language model, and to judge whether the compound information text content contains the required output structure in the fill-in-the-blank prompt, that is, the chemical composition: chemical composition name, to determine whether the compound information text content contains the structured information of the native medicinal material compound, and to further extract the semi-structured native medicinal material compound information.
[0027] Further, the information extraction module is further configured to construct a heuristic information extraction rule template by using the required output structure in the prompt instruction given by the experts: chemical composition: [X] -> [X]: {x1, x2, ……}, wherein X is the extracted native medicinal material compound text description, “->” represents “adopting a regular expression”, xi represents each native medicinal material compound description obtained after segmentation, and finally a list of native medicinal material compound descriptions {x1, x2, ……} is obtained.
[0028] According to a third aspect of the present application, a readable storage medium is provided, which stores one or more programs, and the one or more programs are executable by one or more processors to implement the method as described above.
[0029] The present application has at least the following beneficial effects:
[0030] 1. The data of the present application can be updated in time, avoiding the lag of information.
[0031] 2. The database constructed by the present application has diversified information sources and can be updated at any time, making the information of native medicinal materials compounds more reliable and accurate.
[0032] 3. The present application can automatically organize a large amount of data related to native medicinal materials compounds, effectively avoiding the risks of objectivity, time waste, human resource loss and error propagation. BRIEF DESCRIPTION OF DRAWINGS
[0033] Figure 1 A flow chart of a native medicinal materials compound information online acquisition method according to an embodiment of the present application is shown;
[0034] Figure 2 A schematic diagram of a native medicinal materials compound information directional webpage acquisition process based on heuristic retrieval rules according to an embodiment of the present application is shown;
[0035] Figure 3 A schematic diagram of a native medicinal materials compound structured information rough extraction method based on a large language model according to an embodiment of the present application is shown;
[0036] Figure 4 A schematic diagram of a native medicinal materials compound structured information rough extraction based on a large language model according to an embodiment of the present application is shown;
[0037] Figure 5 A schematic diagram of a native medicinal materials compound structured information extractor and machine translation processing process given a heuristic extraction rule template according to an embodiment of the present application is shown;
[0038] Figure 6 A structural diagram of a native medicinal materials compound information online acquisition device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0039] For those skilled in the art to better understand the technical solutions of the present application, the present application will be described in detail below in combination with the drawings and specific embodiments. The embodiments of the present application will be further described in detail below in combination with the drawings and specific embodiments, but not as a limitation of the present application. The order in which each step is described herein as an example should not be considered as a limitation, and those skilled in the art should know that the order can be adjusted as long as the logic between them is not destroyed and the whole process cannot be realized.
[0040] Figure 1 A flow chart of a method for obtaining information of a native medicinal material compound online according to an embodiment of the present application is shown, as shown in Figure 1 The embodiment of the present application provides a method for obtaining information of a native medicinal material compound online, which utilizes the large-scale directional information web page acquisition capability of a web crawler, and through human-computer interaction for defining retrieval rules with a native medicinal material expert, obtains a latest, comprehensive and reliable web page set of information of a native medicinal material compound online from multiple internet knowledge sources of traditional Chinese medicine; utilizes the wide field knowledge capability of a large language model, and the native medicinal material expert obtains a candidate set of noisy structured data of a native medicinal material compound through online human-computer interaction with the large language model by constructing a prompt rule instruction of field knowledge of a native medicinal material compound; utilizes a heuristic rule method to construct a structured information of a native medicinal material compound relationship discriminator for filtering the candidate set of noisy structured data of a native medicinal material compound, thereby obtaining a data set containing structured information of a native medicinal material compound; and finally, the native medicinal material expert establishes a structured information extraction template of a native medicinal material compound, automatically extracts structured information of a native medicinal material compound from the data set containing structured information of a native medicinal material compound, and saves it to a database. Specifically, the method comprises the following steps S1-S4, which are described in detail as follows.
[0041] Step S1. Based on the retrieval rule set R constructed by the native medicinal material expert, the native medicinal material compound information directional web page acquisition device automatically fills in the retrieval rule information, and obtains a web page set containing information of a native medicinal material compound from an open-source internet knowledge source of traditional Chinese medicine (such as CNKI, PubMed), and the process is shown in Figure 2 .
[0042] Step S2. Loop through the web page set containing information of a native medicinal material compound, input the abstract of each web page into a large language model, and according to the pre-prepared information extraction prompt instruction of a native medicinal material compound (as shown in Figure 3 ), fill in the template filling position (i.e. “{name}”) in the information extraction prompt instruction with the previously constructed native medicinal material list, and the large language model automatically extracts the compound information text content contained in the web page abstract (the process is shown inFigure 4 The local medicinal material compound structured information implication relationship discriminator judges whether the compound information text content has the structured information implication relationship of the local medicinal material compound by whether the output structure "chemical component: chemical component name" required in the Cloze Prompt is contained in the compound information text content, and then extracts the semi-structured local medicinal material compound information.
[0043] Step S3. The output structure required in the prompt instruction given by the expert (such as "chemical component: chemical component name" required in the Cloze Prompt) is used to construct a heuristic information extraction rule template (such as "chemical component: [X] ->[X]:{x1, x2, ……}", wherein X is the extracted local medicinal material compound text description, "->" represents "adopting a regular expression, and each local medicinal material compound description separated by a comma "," is divided", and xi represents each local medicinal material compound description obtained after division), and finally a list of local medicinal material compound descriptions {x1, x2, ……} is extracted. Figure 3
[0044] Step S4. It is judged whether the text content in each local medicinal material compound description list {x1, x2, ……} is Chinese or English. If it is Chinese, the system will use a machine translation system to translate it into English to obtain the English name of the local medicinal material compound. On the contrary, if it is English, the system will use a machine translation system to translate it into Chinese to obtain the Chinese name of the local medicinal material compound. Finally, the structured local medicinal material compound information (the extraction process is as shown in Figure 5 ).
[0045] Figure 6 A structural diagram of a local medicinal material compound information online acquisition device according to an embodiment of the present application is shown. The embodiment of the present application provides a local medicinal material compound information online acquisition device, as shown in Figure 6 , the device 600 comprises:
[0046] The web page set acquisition module 601 is configured to use the large-scale directional information web page acquisition capability of the web crawler, to acquire a web page set of local medicinal material compound information online from multiple internet knowledge sources of traditional Chinese medicine by human-computer interaction of search rule definition with local medicinal material experts.
[0047] The data acquisition module 602 is configured to construct a native medicinal material compound field knowledge acquisition prompt rule instruction, obtain a candidate set of native medicinal material compound structured data with noise through online human-computer interaction with a large language model, and construct a native medicinal material compound structured information implication relationship discriminator by using a heuristic rule method, to filter the candidate set of native medicinal material compound structured data with noise, and obtain a data set containing structured native medicinal material compound information.
[0048] The information pre-acquisition module 603 is configured to establish a structured native medicinal material compound information extraction template, automatically extract structured native medicinal material compound information from the data set containing structured native medicinal material compound information, and save the information to a database.
[0049] In some embodiments, the web page set acquisition module is further configured to automatically fill the retrieval rule information based on a retrieval rule set R constructed by a native medicinal material expert, and obtain a web page set containing native medicinal material compound information from an open-source traditional Chinese medicine internet knowledge source by using a native medicinal material compound information directional web page acquisition device.
[0050] In some embodiments, the data acquisition module is further configured to cyclically traverse the web page set containing native medicinal material compound information, extract an abstract of each web page, input the abstract into a large language model, fill a template filling position in the information extraction prompt instruction by using a native medicinal material list constructed in advance, automatically extract compound information text content contained in the abstract of the web page by the large language model according to a native medicinal material compound information extraction prompt instruction prepared by a native medicinal material expert, and determine whether the compound information text content contains an output structure required in the fill-in-the-blank prompt by the native medicinal material compound structured information implication relationship discriminator, to determine whether the compound information text content contains the native medicinal material compound structured information implication relationship, and further extract semi-structured native medicinal material compound information.
[0051] In some embodiments, the information pre-acquisition module is further configured to construct a heuristic information extraction rule template by using the output structure required in the prompt instruction given by an expert: chemical composition: [X] -> [X]: {x1, x2, ……}, wherein X is the extracted native medicinal material compound text description, “->” represents “use a regular expression”, xi represents each native medicinal material compound description obtained after segmentation, and finally a list of native medicinal material compound descriptions {x1, x2, ……} is obtained.
[0052] In some embodiments, the information pre-acquisition module is further configured to:
[0053] determine whether the text content in each native medicinal material compound description list {x1, x2, ……} is Chinese or English:
[0054] If it is Chinese, it will be translated into English by a machine translation system to obtain the English name of the native medicinal material compound;
[0055] If it is English, it will be translated into Chinese by a machine translation system to obtain the Chinese name of the native medicinal material compound;
[0056] Finally, structured native medicinal material compound information is obtained.
[0057] It should be noted that the device described in the embodiment belongs to the same technical idea as the method described above, and can achieve the same technical effect, which will not be described here.
[0058] The embodiment of the application provides a readable storage medium, the readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method described in each of the above embodiments.
[0059] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more aspects thereof) can be used in combination with each other. Other embodiments can be used as would be apparent to one of ordinary skill in the art reading the above description. Also, in the above detailed description, various features can be grouped together in one or more embodiments for simplicity. This should not be interpreted as a requirement that the features be grouped together in one or more embodiments. Rather, the subject matter has broader scope indicated by the appended claims and any equivalents thereof. The scope of the application should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. The disclosures of the patents, patent applications and publications referred to herein are hereby incorporated by reference in their entirety.
Claims
1. A method for obtaining information of a traditional Chinese medicine compound online, characterized in that, The method comprises: Using the large-scale targeted information webpage acquisition capability of the web crawler, through human-computer interaction of search rule definition with local medicinal material experts, a webpage set containing local medicinal material compound information is acquired online from multiple traditional Chinese medicine internet knowledge sources, comprising: Based on the search rule set R constructed by the local medicinal material experts, the local medicinal material compound information targeted webpage acquirer automatically fills in the search rule information, and acquires a webpage set containing local medicinal material compound information from an open-source traditional Chinese medicine internet knowledge source; Constructing a local medicinal material compound field knowledge acquisition prompt rule instruction, through online human-computer interaction with a large language model, a candidate set of local medicinal material compound structured data with noise is obtained; a heuristic rule method is used to construct a local medicinal material compound structured information implication relationship discriminator, which is used to filter the candidate set of local medicinal material compound structured data with noise, and obtain a data set containing structured local medicinal material compound information, comprising: Cyclically traversing the webpage set containing local medicinal material compound information, extracting the abstract of each webpage input into the large language model, according to the local medicinal material expert's prefabricated local medicinal material compound information extraction prompt instruction, respectively filling in the template filling position in the information extraction prompt instruction with the local medicinal material list constructed in advance, the large language model automatically extracts the compound information text content contained in the webpage abstract, and the local medicinal material compound structured information implication relationship discriminator judges whether the compound information text content contains the required output structure in the fill-in-the-blank prompt: chemical composition: chemical composition name, judges whether the compound information text content contains the local medicinal material compound structured information implication relationship, and further extracts the semi-structured local medicinal material compound information; Establishing a structured local medicinal material compound information extraction template, automatically extracting structured local medicinal material compound information from the data set containing structured local medicinal material compound information, and saving it to a database, comprising: Using the required output structure in the expert's prompt instruction, constructing a heuristic information extraction rule template: chemical composition: [X] -> [X]:{x1, x2, ……}, wherein X is the extracted local medicinal material compound text description, -> represents using a regular expression, xi represents each local medicinal material compound description obtained after segmentation, and finally a list of local medicinal material compound descriptions {x1, x2, ……} is extracted; After extracting the list of local medicinal material compound descriptions {x1, x2, ……}, the method further comprises: Judging whether the text content in the list of local medicinal material compound descriptions {x1, x2, ……} is Chinese or English: If it is Chinese, it is translated into English using a machine translation system to obtain the English name of the local medicinal material compound; If it is English, it is translated into Chinese using a machine translation system to obtain the Chinese name of the local medicinal material compound; Finally, the structured local medicinal material compound information is obtained.
2. A device for obtaining information of a native medicinal material compound online, characterized in that, The device comprises: The webpage set acquisition module is configured to utilize the large-scale directional information webpage acquisition capability of the web crawler, to acquire a webpage set of the native medicinal material compound information online from multiple internet knowledge sources of traditional Chinese medicine through human-computer interaction with experts of native medicinal materials according to defined search rules; The data acquisition module is configured to construct a knowledge acquisition prompt rule instruction of the native medicinal material compound field, to obtain a candidate set of the native medicinal material compound structured data with noise through online human-computer interaction with a large language model, and to construct a native medicinal material compound structured information implication relationship discriminator by using a heuristic rule method, to filter the candidate set of the native medicinal material compound structured data with noise, and to obtain a data set containing the structured native medicinal material compound information. The information extraction module is configured to establish a structured native medicinal material compound information extraction template, to automatically extract the structured native medicinal material compound information from the data set containing the structured native medicinal material compound information, and to save the information to a database. The webpage set acquisition module is further configured to automatically fill the search rule information according to a rule set R constructed by the experts of the native medicinal materials, to acquire the webpage set containing the native medicinal material compound information from the open-source internet knowledge sources of traditional Chinese medicine. The data acquisition module is further configured to traverse the webpage set containing the native medicinal material compound information, to extract the abstract of each webpage and input the abstract into the large language model, to fill the template filling position in the information extraction prompt instruction by using the native medicinal material list constructed in advance according to the native medicinal material compound information extraction prompt instruction prepared by the experts of the native medicinal materials, to automatically extract the compound information text content contained in the webpage abstract by the large language model, and to determine whether the compound information text content contains the required output structure in the fill-in-the-blank prompt by the native medicinal material compound structured information implication relationship discriminator, to further extract the semi-structured native medicinal material compound information. The information extraction module is further configured to construct the heuristic information extraction rule template by using the required output structure in the prompt instruction given by the experts: chemical composition: [X] -> [X]: {x1, x2, ……}, wherein X is the extracted native medicinal material compound text description, "->" represents "use regular expression", xi represents each native medicinal material compound description obtained after segmentation, and finally a list of native medicinal material compound descriptions {x1, x2, ……} is obtained. Determine whether the text content in the list of native medicinal material compound descriptions {x1, x2, ……} is Chinese or English: If it is Chinese, it is translated into English by using a machine translation system to obtain the English name of the native medicinal material compound; If it is English, it is translated into Chinese by using a machine translation system to obtain the Chinese name of the native medicinal material compound; Finally, the structured native medicinal material compound information is obtained.
3. A readable storage medium, characterized by, The readable storage medium stores one or more programs, and the one or more programs are executable by one or more processors to implement the method in claim 1.
Citation Information
Patent Citations
Academic core author excavation and related information extraction method and system based on complex network
CN103020302A
Unstructured network threat intelligence extraction method and system and medium
CN118013046A