Method and device for acquiring literature information and automatically constructing literature knowledge graph
By combining large language models and knowledge graph technology, unstructured texts are transformed into structured triples, which solves the problem of insufficient deep semantic understanding of language in the existing technology, and realizes efficient information extraction and heterogeneous data integration of scientific literature.
Patent Information
- Application Number
- CN202510522573.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-19
AI Technical Summary
Existing literature extraction techniques rely on shallow language features, lack understanding of deep semantics and contextual relationships in languages, and cannot efficiently extract entities, relationships and key data from unstructured text and tabular data, resulting in the inability to fully understand professional terms and complex concepts in the field of science, and it is difficult to effectively integrate heterogeneous data and cross-domain knowledge.
Combining large language models and knowledge graph technology, through named entity recognition and relationship extraction, unstructured text is transformed into structured triples and integrated into knowledge graphs to achieve an understanding of professional terms and complex concepts in the field of science, and to integrate external knowledge graph information.
It improves the accuracy and completeness of information extraction, achieves a full understanding of professional terms and complex concepts in the scientific field, integrates and associates a large number of heterogeneous data, breaks through the limitations of simple keyword matching, and forms a computable knowledge network.
Smart Images

Figure CN120509472A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing technology, and in particular to a method and device for automatically acquiring document information and constructing a document knowledge graph. Background Art
[0002] With the exponential growth of scientific literature, researchers face unprecedented opportunities and challenges in advancing scientific discovery. Especially in fields such as chemistry and biochemical engineering, research articles contain vast amounts of valuable data, including text descriptions, experimental results, tables, and figures. This data plays a critical role in advancing knowledge, optimizing processes, and driving innovation. However, traditional manual literature review and data extraction methods are time-consuming and inefficient, making it difficult to keep pace with the rapid pace of research publication.
[0003] Among the related technologies, common literature information extraction technologies include text mining methods and advanced natural language processing technologies, which are used to extract meaningful information from large amounts of scientific texts.
[0004] However, the literature extraction technology in related technologies relies on shallow language features and lacks understanding of the deep semantics and contextual associations of the language. Therefore, it does not have strong natural language understanding and generation capabilities, and cannot efficiently extract entities, relationships and key data from unstructured text and tabular data. In addition, the data is heterogeneous and lacks correlation, which leads to the inability to fully understand the professional terms and complex concepts in the scientific field. It is difficult to effectively deal with problems such as the integration of heterogeneous data and the mining of cross-domain knowledge, and it is in urgent need of improvement. Summary of the Invention
[0005] The present application provides a method and apparatus for automatically constructing document information and document knowledge graphs, in order to address the problems in related technologies such as document extraction technology, which relies on shallow language features and lacks understanding of deep language semantics and contextual associations. Therefore, it does not have strong natural language understanding and generation capabilities, cannot efficiently extract entities, relationships, and key data from unstructured text and tabular data, and lacks data heterogeneity and correlation, which in turn leads to an inability to fully understand professional terms and complex concepts in scientific fields, and makes it difficult to effectively deal with problems such as the integration of heterogeneous data and the mining of cross-domain knowledge.
[0006] The first embodiment of the present application provides a method for automatically constructing a document information acquisition and document knowledge graph, comprising the following steps: using at least one document search term provided by a user to retrieve at least one relevant document, and generating text data of different data types based on the document content of the at least one relevant document; constructing a knowledge graph based on the text data of different data types based on a target large language model;
[0007] Integrate at least one target knowledge graph and the knowledge graph to construct a final knowledge graph.
[0008] Through the above technical solution, the embodiment of the present application can perform two-step extraction of named entity recognition and relationship extraction on the text data in the document based on the large language model (LLM), perform accurate knowledge extraction based on the document domain knowledge, achieve a full understanding of professional terms and complex concepts in the scientific field, and improve the accuracy and completeness of information extraction; and can further integrate the information of the target knowledge graph (KG) to construct the final knowledge graph and enrich the information of the local knowledge graph; better express the complex associations and semantics between entities in the form of a knowledge graph structure, and integrate and associate a large amount of heterogeneous data.
[0009] Optionally, in one embodiment of the present application, constructing a knowledge graph based on the text data of different data types includes: identifying named entities in the text data based on a target large language model; and performing relationship extraction on the text data based on the named entities to construct the knowledge graph based on the literature field.
[0010] Through the above technical solution, the embodiment of the present application can use a large language model to perform named entity recognition and relationship extraction, efficiently extract entities, relationships and key data from unstructured text and tabular data, and convert them into structured and actionable knowledge, thereby achieving effective understanding of professional terms and complex concepts in the scientific field, and improving the accuracy and completeness of information extraction.
[0011] Optionally, in one embodiment of the present application, the relationship extraction of the text data based on the named entities to construct the knowledge graph based on the literature field includes: performing relationship extraction on the key entities of the named entities extracted based on the target large language model, associating the entities according to the text content to form literature triples, so as to construct the knowledge graph based on the local knowledge graph.
[0012] Through the above technical solution, the embodiment of the present application can extract and associate relationships between key entities, convert isolated entities in the text into structured knowledge with semantic associations, automatically refine the core content of the text, identify implicit relationships between entities, break through the limitations of simple keyword matching, convert unstructured text into triples, form a computable knowledge network, integrate and associate large amounts of heterogeneous data, and effectively cope with the integration of heterogeneous data and the mining of cross-domain knowledge.
[0013] Optionally, in one embodiment of the present application, generating text data of different data types based on the document content of the at least one relevant document includes: capturing at least one piece of information from metadata, body text, table titles and content, image titles and content, references and attachments according to the web page format of the document publishing platform of each relevant document; and obtaining text data of different data types based on the at least one piece of information.
[0014] Through the above technical solution, the embodiment of the present application can accurately capture valid information according to the web page format of the document publishing platform and save the valid information into a structured JSON format file, retaining the hierarchical relationship of the document text, avoiding the problem of chaotic text structure of PDF files, and realizing the automatic collection and management of document data.
[0015] The second aspect of the present application provides a device for automatically constructing document information and document knowledge graphs, including: a generation module for retrieving at least one relevant document using at least one document search term provided by a user, and generating text data of different data types based on the document content of the at least one relevant document; a construction module for constructing a knowledge graph based on the text data of different data types based on a target large language model; and an integration module for integrating at least one target knowledge graph and the knowledge graph to construct a final knowledge graph.
[0016] Through the above technical solution, the embodiment of the present application can perform two-step extraction of named entity recognition and relationship extraction on the text data in the document based on a large language model, perform accurate knowledge extraction based on the document domain knowledge, achieve a full understanding of professional terms and complex concepts in the scientific field, and improve the accuracy and completeness of information extraction; and can further integrate the information of the target knowledge graph, construct the final knowledge graph, and enrich the information of the local knowledge graph; better express the complex associations and semantics between entities in the form of a knowledge graph structure, and integrate and associate a large amount of heterogeneous data.
[0017] Optionally, in one embodiment of the present application, the construction module includes: an identification unit for identifying named entities in the text data based on a target large language model; and an extraction unit for performing relationship extraction on the text data according to the named entities to construct the knowledge graph according to the literature field.
[0018] Through the above technical solution, the embodiment of the present application can use a large language model to perform named entity recognition and relationship extraction, efficiently extract entities, relationships and key data from unstructured text and tabular data, and convert them into structured and actionable knowledge, thereby achieving effective understanding of professional terms and complex concepts in the scientific field, and improving the accuracy and completeness of information extraction.
[0019] Optionally, in one embodiment of the present application, the extraction unit is specifically used to perform relationship extraction based on the key entities of the named entity extracted by the target large language model, associate the entities according to the text content, form document triples, and construct the knowledge graph based on the local knowledge graph.
[0020] Through the above technical solution, the embodiment of the present application can extract and associate relationships between key entities, convert isolated entities in the text into structured knowledge with semantic associations, automatically refine the core content of the text, identify implicit relationships between entities, break through the limitations of simple keyword matching, convert unstructured text into triples, form a computable knowledge network, integrate and associate large amounts of heterogeneous data, and effectively cope with the integration of heterogeneous data and the mining of cross-domain knowledge.
[0021] Optionally, in one embodiment of the present application, the generation module includes: capturing at least one piece of information including metadata, body text, table titles and content, picture titles and content, references and attachments according to the web page format of the document publishing platform of each relevant document; and obtaining text data of different data types based on the at least one piece of information.
[0022] Through the above technical solution, the embodiment of the present application can accurately capture valid information according to the web page format of the document publishing platform and save the valid information into a structured JSON format file, retaining the hierarchical relationship of the document text, avoiding the problem of chaotic text structure of PDF files, and realizing the automatic collection and management of document data.
[0023] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the method for automatically constructing document information and document knowledge graphs as described in the above embodiments.
[0024] The fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program. When the program is executed by a processor, it implements the above-mentioned method for obtaining document information and automatically constructing a document knowledge graph.
[0025] The fifth aspect of the present application provides a computer program product, wherein the computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned method for obtaining document information and automatically constructing a document knowledge graph.
[0026] The embodiment of the present application can combine a large language model with knowledge graph technology, perform named entity recognition and relationship extraction on the text data in the document based on the large language model, perform accurate knowledge extraction based on the document domain knowledge, convert unstructured text information into a structured triple format, and import it into the knowledge graph; achieve a full understanding of professional terms and complex concepts in the scientific field, improve the accuracy and completeness of information extraction; and further integrate the information of the target knowledge graph, align the data extracted from the document with the external target knowledge graph, and enrich the information of the local knowledge graph; better express the complex associations and semantics between entities in the form of a knowledge graph structure, and integrate and associate a large amount of heterogeneous data. Thus, it solves the problem that the document extraction technology in the related art relies on shallow language features and lacks understanding of the deep semantics and contextual associations of the language, so it does not have strong natural language understanding and generation capabilities, cannot efficiently extract entities, relationships and key data from unstructured text and tabular data, and the data heterogeneity and correlation are insufficient, which leads to the inability to fully understand the professional terms and complex concepts in the scientific field, and it is difficult to effectively deal with the integration of heterogeneous data and the mining of cross-domain knowledge.
[0027] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0029] Figure 1 A flowchart of a method for automatically constructing a document knowledge graph and acquiring document information according to an embodiment of the present application;
[0030] Figure 2 A flowchart of a method for automatically constructing a document knowledge graph and acquiring document information according to a specific embodiment of the present application is provided.
[0031] Figure 3 A schematic diagram of a knowledge graph composed of triples extracted from a paragraph according to a specific embodiment of the present application;
[0032] Figure 4 A schematic diagram of a knowledge graph entity word cloud and a relationship word cloud according to a specific embodiment of the present application;
[0033] Figure 5 This is a schematic diagram of clustering results of word vectors of knowledge graph nodes according to a specific embodiment of the present application;
[0034] Figure 6Schematic diagram of a block diagram of an apparatus for automatically constructing a document information acquisition and document knowledge graph according to an embodiment of the present application;
[0035] Figure 7 A schematic diagram of the structure of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0036] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0037] The following describes the method and device for automatically constructing document information acquisition and document knowledge graphs according to the embodiment of the present application with reference to the accompanying drawings. The document extraction technology in the related technologies mentioned in the above background technology center relies on shallow language features and lacks understanding of the deep semantics and contextual associations of the language. Therefore, it does not have strong natural language understanding and generation capabilities, and cannot efficiently extract entities, relationships and key data from unstructured text and tabular data. In addition, the data heterogeneity and correlation are insufficient, which leads to the inability to fully understand the professional terms and complex concepts in the scientific field, and it is difficult to effectively deal with the integration of heterogeneous data and the mining of cross-domain knowledge. The present application provides a method for automatically constructing document information acquisition and document knowledge graphs, in which a large language model can be integrated with the knowledge graph. By combining the two technologies, the two-step extraction of named entity recognition and relationship extraction is performed on the text data in the literature based on a large language model. Accurate knowledge extraction is performed based on the literature domain knowledge, and the unstructured text information is converted into a structured triple format and imported into the knowledge graph. This achieves a full understanding of the professional terms and complex concepts in the scientific field, improving the accuracy and completeness of information extraction. It can further integrate the information of the target knowledge graph, align the data extracted from the literature with the external target knowledge graph, and enrich the information of the local knowledge graph. The complex associations and semantics between entities are better represented in the form of the knowledge graph structure, and a large amount of heterogeneous data is integrated and associated. This solves the problem that the literature extraction technology in the related technology relies on shallow language features and lacks understanding of the deep semantics and contextual associations of the language. Therefore, it does not have strong natural language understanding and generation capabilities, cannot efficiently extract entities, relationships and key data from unstructured text and tabular data, and the data heterogeneity and association are insufficient, which leads to the inability to fully understand the professional terms and complex concepts in the scientific field, and it is difficult to effectively deal with the integration of heterogeneous data and the mining of cross-domain knowledge.
[0038] Specifically, Figure 1A flowchart of a method for automatically constructing a document knowledge graph and acquiring document information provided in an embodiment of the present application.
[0039] like Figure 1 As shown, the method for automatically constructing the document information acquisition and document knowledge graph includes the following steps:
[0040] In step S101, at least one relevant document is retrieved using at least one document search term provided by a user, and text data of different data types are generated according to the document content of the at least one relevant document.
[0041] Among them, data types include but are not limited to different data types such as text, tables and images; retrieval tools can be general search engines (such as Google, Baidu), academic and literature retrieval tools (such as HowNet, Wanfang Data), etc., among which retrieval tools can be adaptively selected according to actual application scenarios.
[0042] In some embodiments, users can set search terms and filtering conditions, and perform searches based on the set search terms and filtering conditions. For example, users can set the search term to "artificial intelligence" and the filtering conditions to "top conference papers in the past five years, high citation priority". Using the AI agent, the system automatically retrieves relevant documents based on the document search terms provided by the user, downloads and pre-processes the document content, and classifies it into different data types such as text, tables, and images.
[0043] The embodiments of the present application can generate text data of different data types based on the content of the document. Text data of different data types helps to enrich the data source of the model and provide more useful information for the model.
[0044] In step S102, based on the target large language model, a knowledge graph is constructed according to text data of different data types.
[0045] Among them, large-scale language models can be understood as deep learning models trained based on massive text data, which can understand, generate and reason about natural language; with the Transformer architecture as the core, it captures language rules through self-supervised learning (such as masked language modeling and next word prediction) and demonstrates emergent capabilities (such as complex reasoning and cross-task generalization); it includes but is not limited to GPT-3, GPT-4, DeepSeek-V3, etc., and can be adaptively selected according to actual scenarios; knowledge graphs can be understood as a technology that describes real-world entities (such as people, places, events) and their relationships in a structured form. It is essentially a semantic network that stores knowledge through triples (entity-relationship-entity) to support reasoning, search and intelligent question and answer; it includes but is not limited to general knowledge graphs, domain-specific knowledge graphs, open source knowledge graphs, etc.
[0046] During the actual implementation process, the embodiments of the present application can utilize large language models, such as GPT-4, to perform two-step extraction of named entity recognition and relationship extraction on text data in the document, perform precise knowledge extraction based on the document domain knowledge, convert unstructured text information into a structured triple format, and import it into the knowledge graph.
[0047] The embodiments of the present application can use the powerful natural language understanding and generation capabilities of LLM to efficiently extract entities, relationships and key data from unstructured text and tabular data, and convert them into structured and actionable knowledge. The graph structure of the knowledge graph can better represent the complex associations and semantics between entities. The knowledge graph can not only uniformly manage information from different data sources, but also more intuitively display a variety of relationships and hierarchical structures, which makes it easier for computers to understand and process the semantic associations in this data, thereby performing complex reasoning and decision-making. The combination of LLM and KG can effectively overcome the limitations of traditional text mining technology and achieve in-depth mining and integration of knowledge in scientific literature.
[0048] In step S103, at least one target knowledge graph and the knowledge graph are integrated to construct a final knowledge graph.
[0049] The target knowledge graph can be understood as an external knowledge graph, that is, other knowledge graphs other than the knowledge graph constructed in step S102, such as Google Knowledge Graph, Wikidata, DBpedia, etc.
[0050] During the actual implementation process, the embodiment of the present application can align the data extracted from the document with the external knowledge graph by integrating the information of the external knowledge graph. For example: after obtaining the triples extracted from the document, the embodiment of the present application can align the document triples with the external knowledge graph by searching the external knowledge graph, such as Google Knowledge Graph, Wikidata and DBpedia, and supplement the content of the triples with the descriptive words of the external knowledge graph. In this process, the entity names and relationships can first be standardized to ensure that they are consistent with the standard naming in the external knowledge graph, eliminating matching errors caused by naming differences. For specific entities such as chemical substances, information such as CAS number, molecular weight, etc. is automatically supplemented to enrich the content of the triples and integrate duplicate nodes, eliminate redundant information, and ensure the simplicity and accuracy of the knowledge graph. Finally, the aligned triples can be imported into the local knowledge graph by combining the document text data and the external knowledge graph.
[0051] The embodiments of the present application can integrate the information of the external knowledge graph, align the document triples with the external knowledge graph, realize the standardization of entities and relationships, and supplement the descriptive information to enrich and improve the information in the local knowledge graph.
[0052] Optionally, in one embodiment of the present application, a knowledge graph is constructed based on text data of different data types, including: identifying named entities in the text data based on a target large language model; and extracting relationships from the text data based on the named entities to construct a knowledge graph based on the literature field.
[0053] It is understandable that named entities can be understood as real-world objects in natural language processing and knowledge graphs that have specific meaning and can be uniquely identified in the text. They can serve as a bridge connecting text data and structured knowledge. Their identification and linking are the basis for building knowledge graphs and achieving semantic understanding. Named entity recognition can be understood as the technology for detecting and classifying entities from text. Relation extraction can be understood as a core task in natural language processing, which aims to identify semantic relationships between entities in text and convert unstructured text into structured knowledge (such as triples).
[0054] During the actual implementation process, after obtaining the text information of the document in the document, the embodiment of the present application can use a large language model to perform named entity recognition and relationship extraction tasks to convert the unstructured text information into a structured subject-verb-object triple format.
[0055] The embodiments of the present application can utilize large language models for named entity recognition and relationship extraction, efficiently extract entities, relationships, and key data from unstructured text and tabular data, and convert them into structured, actionable knowledge, thereby achieving effective understanding of professional terms and complex concepts in scientific fields and improving the accuracy and completeness of information extraction.
[0056] Optionally, in one embodiment of the present application, relationship extraction is performed on text data based on named entities to construct a knowledge graph based on the literature field, including: performing relationship extraction on key entities of named entities extracted based on the target large language model, associating entities according to text content to form document triples, and constructing a knowledge graph based on the local knowledge graph.
[0057] Key entities can be understood as entities in the text that directly affect semantic understanding, decision-making, or subsequent tasks (such as question-answering, search, and knowledge graph construction). The extraction of key entities depends on different application scenarios. Different scenarios have different types of key entities. For example: In the medical field, disease names, drugs, and symptoms (such as "diabetes" and "aspirin") are key entities. In the financial field, company names, stock codes, and amounts (such as "Apple," "AAPL," and "US$1 billion") are key entities.
[0058] In certain embodiments, key entities can be extracted by pre-designing prompts that adapt to the document content and manually supplementing domain knowledge, allowing large-scale language models to extract key entities in named entity recognition tasks. The output entities of named entity recognition are then fed back into relation extraction, where entities are associated based on text content and formed into document triples, thereby constructing a knowledge graph.
[0059] The embodiments of the present application can extract and associate relationships between key entities, convert isolated entities in the text into structured knowledge with semantic associations, automatically refine the core content of the text, identify implicit relationships between entities, break through the limitations of simple keyword matching, convert unstructured text into triples, form a computable knowledge network, integrate and associate large amounts of heterogeneous data, and effectively cope with the integration of heterogeneous data and the mining of cross-domain knowledge.
[0060] Optionally, in one embodiment of the present application, text data of different data types are generated based on the document content of at least one relevant document, including: capturing at least one piece of information from metadata, body text, table titles and content, image titles and content, references and attachments based on the web page format of the document publishing platform of each relevant document; and obtaining text data of different data types based on at least one piece of information.
[0061] Among them, the web page format can be understood as a structured document standard used to display and transmit information on the Internet. It defines how the content is rendered by the browser and how the user interacts through specific markup language, style rules and interaction logic.
[0062] In the actual implementation process, the embodiment of the present application can utilize the text mining API (Application Programming Interface) of the Elsevier and Web of Science document publishing platforms to obtain a DOI (Digital Object Identifier) list of relevant documents based on the keyword query of the document search. The DOI list refers to a collection of DOI identifiers of a group of documents, data sets or other digital resources; based on the DOI data statistics of the documents, a special HTML (Hypertext Markup Language) document crawling and data cleaning Python script is designed for the seven most popular document publishing platforms (Springer, Elsevier, ACS, MDPI, RSC, Wiley, Taylor & Francis). The document crawling script obtains the URL (Uniform Resource Locator) link based on the document DOI information and downloads the HTML file of the document; the HTML data cleaning script accurately captures metadata, text, table titles and content, image titles and content, references and attachments, etc. according to the web page format of the document publishing platform and saves it into a structured JSON format file.
[0063] The embodiment of the present application can accurately capture valid information according to the web page format of the document publishing platform and save the valid information into a structured JSON format file, retaining the hierarchical relationship of the document text, avoiding the problem of chaotic text structure of PDF files, and realizing automatic collection and management of document data.
[0064] like Figure 2 As shown, the working principle of the method for automatically constructing the document information acquisition and document knowledge graph of the present application is schematically illustrated with a specific embodiment below.
[0065] Step S201: Determine a document search keyword to facilitate document search based on the keyword;
[0066] Step S202: Using the text mining API of the Web of Science literature search platform to search based on search keywords;
[0067] Step S203: Obtaining a DOI list of relevant documents based on the keyword query of the document search;
[0068] Step S204: Design a special Python script for HTML document crawling and data cleaning;
[0069] Step S205: using a document crawling script to obtain a URL link based on the document DOI information, and downloading the HTML file of the document;
[0070] Step S206: Using an HTML data cleaning script to accurately capture metadata, body text, table titles and content, image titles and content, references and attachments, etc. according to the web page format of the document publishing platform;
[0071] Step S207: Save the result of the information captured by the data cleaning script in step S206 into a structured JSON format file;
[0072] Step S208: After obtaining the text information of the document in step S207, a large language model is used to perform named entity recognition and relation extraction tasks to convert the unstructured text information into a structured subject-verb-object triple format;
[0073] Step S209: Designing prompts that can adapt to the content of the document and manually supplementing domain knowledge to enable the large language model to extract key entities in the named entity recognition task;
[0074] Step S210: performing a named entity recognition task using a large language model;
[0075] Step S211: After the large language model performs the named entity recognition task, an entity is generated;
[0076] Step S212: Utilize the large language model to perform the relationship extraction task, and input the output entities of the named entity recognition into the relationship extraction again to associate the entities according to the text content;
[0077] Step S213: Obtain the relationship extraction result of step S212 and generate a relationship;
[0078] Step S214: Group the obtained entities and relationships into triples to construct a knowledge graph;
[0079] Step S215: Retrieve external knowledge graph;
[0080] Step S216: Determine whether there are relevant search results in the external knowledge graph. If so, execute step S217; otherwise, execute step S218;
[0081] Step S217: Align the document triples with the external knowledge graph, and supplement the triples with descriptive terms from the external knowledge graph. During this process, entity names and relationships are first standardized to ensure consistency with the standard naming in the external knowledge graph, eliminating matching errors caused by naming differences. For specific entities such as chemical substances, information such as CAS number and molecular weight is automatically supplemented to enrich the triples and integrate duplicate nodes to eliminate redundant information, ensuring the simplicity and accuracy of the knowledge graph.
[0082] Step S218: After executing step S217, the aligned triples are imported into the local knowledge graph by combining the document text data and the external knowledge graph; if no relevant results are retrieved, the triples obtained in step S214 are directly imported into the local knowledge graph.
[0083] like Figure 3 As shown, Figure 3 The knowledge graph consisting of triples extracted from a paragraph in the embodiment of the present application is presented, wherein: Figure 4 The entity word cloud (left) and relationship word cloud (right) in the knowledge graph in the embodiment of this application are shown; Figure 5 The clustering results of the node word vectors in the knowledge graph in the embodiment of this application are shown.
[0084] In summary, this embodiment proposes an artificial intelligence-assisted framework based on a large language model for automatically extracting and integrating data from scientific literature and building an efficient knowledge graph. This embodiment automatically performs document retrieval and information extraction through an AI agent to obtain structured triple data from unstructured text. The framework uses a document retrieval API to obtain a list of document DOIs, designs HTML document crawling and data cleaning scripts for mainstream publishing platforms, accurately extracts the metadata and text content of the document, and saves it in a structured JSON format. In the information extraction stage, this embodiment uses a large language model for named entity recognition and relationship extraction to convert the text into a subject-verb-object triple format. Then, by aligning external knowledge graphs (such as Google Knowledge Graph, Wikidata, etc.), standardizes entities and relationships, supplements descriptive information, and enriches the content of the knowledge graph. This embodiment can demonstrate the potential of artificial intelligence in accelerating scientific discovery and knowledge integration. It is applicable to document processing in any academic field. By automatically processing a large amount of scientific literature, it provides researchers with powerful tools to help researchers more intuitively understand and explore relationships in data, thereby promoting the efficiency and innovation of scientific research.
[0085] According to the method for automatically constructing document information and document knowledge graphs proposed in the embodiments of the present application, large-scale language models can be combined with knowledge graph technology, and the text data in the document can be extracted in two steps of named entity recognition and relationship extraction based on the large-scale language model. Accurate knowledge extraction can be performed based on the document domain knowledge, and unstructured text information can be converted into a structured triple format and imported into the knowledge graph; a full understanding of professional terms and complex concepts in the scientific field can be achieved, and the accuracy and completeness of information extraction can be improved; and the information of the target knowledge graph can be further integrated, and the data extracted from the document can be aligned with the external target knowledge graph to enrich the information of the local knowledge graph; the complex associations and semantics between entities can be better represented in the form of a knowledge graph structure, and a large amount of heterogeneous data can be integrated and associated. This solves the problem that literature extraction technology in related technologies relies on shallow language features and lacks understanding of the deep semantics and contextual associations of the language. Therefore, it does not have strong natural language understanding and generation capabilities, and cannot efficiently extract entities, relationships and key data from unstructured text and tabular data. In addition, the data is heterogeneous and lacks correlation, which leads to the inability to fully understand professional terms and complex concepts in the scientific field, and it is difficult to effectively deal with problems such as the integration of heterogeneous data and the mining of cross-domain knowledge.
[0086] Next, refer to the attached Figure 6 Describe the automatic construction device of document information acquisition and document knowledge graph proposed in the embodiment of the present application.
[0087] Figure 6 It is a block diagram of the device for automatically constructing document information acquisition and document knowledge graph in an embodiment of the present application.
[0088] like Figure 6 As shown, the device 10 for automatically constructing document information and document knowledge graph includes: a generation module 100, a construction module 200, and an integration module 300.
[0089] The generating module 100 is configured to retrieve at least one relevant document using at least one document search term provided by a user, and generate text data of different data types based on the document content of the at least one relevant document;
[0090] A construction module 200 is used to construct a knowledge graph based on text data of different data types based on a target large language model;
[0091] The integration module 300 is used to integrate at least one target knowledge graph and the knowledge graph to construct a final knowledge graph.
[0092] Optionally, in one embodiment of the present application, the construction module 200 includes: an identification unit for identifying named entities in text data based on a target large language model; and an extraction unit for performing relationship extraction on text data based on the named entities to construct a knowledge graph based on the literature field.
[0093] Optionally, in one embodiment of the present application, the extraction unit is specifically used to perform relationship extraction based on key entities of named entities extracted by the target large language model, associate entities according to text content, form document triples, and construct a knowledge graph based on the local knowledge graph.
[0094] Optionally, in one embodiment of the present application, the generation module 100 includes: capturing at least one piece of information including metadata, body text, table titles and content, image titles and content, references and attachments according to the web page format of the document publishing platform of each relevant document; and obtaining text data of different data types based on at least one piece of information.
[0095] It should be noted that the above explanation of the embodiment of the method for obtaining document information and automatically constructing a document knowledge graph is also applicable to the device for obtaining document information and automatically constructing a document knowledge graph in this embodiment, and will not be repeated here.
[0096] According to the device for automatically constructing document information acquisition and document knowledge graph proposed in the embodiment of the present application, it is possible to combine large-scale language models with knowledge graph technology, perform two-step extraction of named entity recognition and relationship extraction on text data in documents based on large-scale language models, perform accurate knowledge extraction based on document domain knowledge, convert unstructured text information into a structured triple format, and import it into the knowledge graph; achieve a full understanding of professional terms and complex concepts in the scientific field, and improve the accuracy and completeness of information extraction; and further integrate the information of the target knowledge graph, align the data extracted from the document with the external target knowledge graph, and enrich the information of the local knowledge graph; better represent the complex associations and semantics between entities in the form of knowledge graph structure, and integrate and associate a large amount of heterogeneous data. This solves the problem that literature extraction technology in related technologies relies on shallow language features and lacks understanding of the deep semantics and contextual associations of the language. Therefore, it does not have strong natural language understanding and generation capabilities, and cannot efficiently extract entities, relationships and key data from unstructured text and tabular data. In addition, the data is heterogeneous and lacks correlation, which leads to the inability to fully understand professional terms and complex concepts in the scientific field, and it is difficult to effectively deal with problems such as the integration of heterogeneous data and the mining of cross-domain knowledge.
[0097] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0098] Memory 701 , processor 702 , and computer programs stored in the memory 701 and executable on the processor 702 .
[0099] When the processor 702 executes the program, the method for automatically constructing the document information acquisition and document knowledge graph provided in the above embodiment is implemented.
[0100] Furthermore, the electronic device further includes:
[0101] The communication interface 703 is used for communication between the memory 701 and the processor 702 .
[0102] The memory 701 is used to store computer programs that can be run on the processor 702 .
[0103] The memory 701 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0104] If the memory 701, processor 702, and communication interface 703 are implemented independently, the communication interface 703, memory 701, and processor 702 can be interconnected via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0105] Optionally, in a specific implementation, if the memory 701, the processor 702 and the communication interface 703 are integrated on a chip, the memory 701, the processor 702 and the communication interface 703 can communicate with each other through an internal interface.
[0106] The processor 702 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0107] An embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, it implements the above-mentioned method for obtaining document information and automatically constructing a document knowledge graph.
[0108] An embodiment of the present application also provides a computer program product, which stores a computer program, and when the program is executed by a processor, implements the above-mentioned method for obtaining document information and automatically constructing a document knowledge graph.
[0109] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0110] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0111] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0112] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or N wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically by optically scanning the paper or other medium and then editing, interpreting, or otherwise processing in a suitable manner as necessary, and then storing it in a computer memory.
[0113] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented using hardware, as in another embodiment, it can be implemented using any one or a combination of the following technologies known in the art: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0114] Those skilled in the art will appreciate that all or part of the steps in the method for implementing the above-mentioned embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0115] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0116] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for acquiring document information and automatically constructing a document knowledge graph, characterized in that: The following steps are involved: Retrieving at least one relevant document using at least one document search term provided by a user, and generating text data of different data types based on the document content of the at least one relevant document; Based on the target large language model, construct a knowledge graph according to the text data of different data types; Integrate at least one target knowledge graph and the knowledge graph to construct a final knowledge graph.
2. The method according to claim 1, characterized in that The constructing of a knowledge graph based on the text data of different data types includes: identifying named entities in the text data based on a target large language model; Relationship extraction is performed on the text data according to the named entities to construct the knowledge graph according to the literature field.
3. The method according to claim 2, characterized in that The extracting relationships from the text data based on the named entities to construct the knowledge graph based on the literature field includes: In the relationship extraction of the key entities of the named entities extracted based on the target large-scale language model, the entities are associated according to the text content to form document triples, so as to construct the knowledge graph based on the local knowledge graph.
4. The method according to claim 1, wherein Generating text data of different data types according to the content of the at least one relevant document includes: At least one of metadata, text, table titles and content, image titles and content, references, and attachments is captured based on the webpage format of the document publishing platform of each relevant document; Text data of different data types are acquired according to the at least one information.
5. A device for automatically constructing document information acquisition and document knowledge graph, characterized in that: include: A generating module, configured to retrieve at least one relevant document using at least one document search term provided by a user, and generate text data of different data types based on the document content of the at least one relevant document; A construction module, configured to construct a knowledge graph based on the target large language model and the text data of the different data types; An integration module is used to integrate at least one target knowledge graph and the knowledge graph to construct a final knowledge graph.
6. The device according to claim 5, characterized in that The building blocks include: a recognition unit, configured to recognize named entities in the text data based on a target large language model; An extraction unit is used to extract relationships from the text data based on the named entities to construct the knowledge graph based on the literature field.
7. The device according to claim 6, characterized in that The extraction unit is specifically used to extract relationships between key entities of the named entities extracted based on the target large language model, associate entities according to text content, form document triples, and construct the knowledge graph based on the local knowledge graph.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for automatically constructing document information acquisition and document knowledge graph as described in any one of claims 1 to 4.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the method for automatically constructing document information acquisition and document knowledge graph as described in any one of claims 1 to 4.
10. A computer program product comprising a computer program, characterized in that The computer program is executed to implement the method for automatically constructing document information and document knowledge graph as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Literature retrieval method and device based on knowledge graph, electronic equipment and medium
CN113590845A
Lithium battery material property prediction method based on knowledge graph and LLM joint reasoning
CN118070906A
Method and system for constructing domain multi-modal knowledge graph
CN118568271A
Knowledge graph construction method and device for target culture resources and electronic equipment
CN119539044A
Cited By
File digital governance method and system based on large model
CN121092759A
Knowledge graph and digital object mixing method for scientific discovery
CN121524372A
Literature evaluation method and device, electronic equipment and storage medium
CN121809706A