Knowledge graph construction method and device and computer readable storage medium

By adopting the neural network-based entity extraction method and iterative optimization strategy relationship extraction method in the knowledge graph construction, the problem of insufficient efficiency, accuracy and scalability of knowledge graph construction in the existing technology is solved, and a more efficient and accurate knowledge graph construction is achieved.

CN120163216APending Publication Date: 2025-06-17CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311719095.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing knowledge graph construction methods have limitations in terms of efficiency, accuracy, scalability, etc.

Method used

The neural network-based method is used to extract entities from the input text and introduce context information when the entity is extracted; the entity pairs and input text are input into the relationship extraction model, and the iterative optimization strategy is used to continuously update and optimize the extracted relationship to obtain the corresponding knowledge graph.

Benefits of technology

By introducing context information, iterative optimization strategies improve the accuracy and completeness of relationship extraction, achieving more efficient and accurate knowledge graph construction, supporting richer knowledge representation and having good scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163216A_ABST
    Figure CN120163216A_ABST
Patent Text Reader

Abstract

The invention provides a knowledge graph construction method and device and a computer readable storage medium, and the method comprises the steps: employing a method based on a neural network to extract an entity from an input text, and introducing context information when the entity is extracted; entity pairs are extracted from the extracted entities; and inputting the entity pair and the corresponding input text into a relation extraction model, and continuously updating and optimizing the extracted relation by adopting an iterative optimization strategy to obtain a corresponding knowledge graph. According to the method and device and the computer readable storage medium, the problem that an existing knowledge graph construction method has limitations in the aspects of efficiency, accuracy, expansibility and the like can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graphs, and in particular, to a method, device, and computer-readable storage medium for constructing a knowledge graph. Background Art

[0002] Knowledge graph construction is a method based on artificial intelligence and big data technologies, which is used to transform multi-source heterogeneous data into a structured knowledge representation so that machines can understand and reason about this knowledge.

[0003] In the prior art, there have been various knowledge graph construction methods. Among them, some methods are based on data mining and natural language processing technologies, and extract entity, relationship, and attribute information from large-scale text data, and use statistical and machine learning methods to construct a knowledge graph. Some other methods rely on manually labeled data and achieve it by manually constructing and editing a knowledge graph. There are also some methods that combine traditional database technologies with graph databases to support the storage and query of knowledge graphs.

[0004] However, through comprehensive research and analysis of the prior art, it is found that the existing implementation solutions still have some limitations in some aspects, such as efficiency, accuracy, scalability, etc. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method, device, and computer-readable storage medium for constructing a knowledge graph to solve the problem that the existing knowledge graph construction methods have limitations in terms of efficiency, accuracy, scalability, etc., in view of the above deficiencies of the prior art.

[0006] In a first aspect, the present invention provides a method for constructing a knowledge graph, the method

[0007] includes:

[0008] Extract entities from the input text using a neural network-based method, and introduce context information during entity extraction;

[0009] Extract entity pairs from the extracted entities;

[0010] Input the entity pairs and the corresponding input text into a relationship extraction model, and use an iterative optimization strategy to continuously update and optimize the extracted relationships to obtain the corresponding knowledge graph.

[0011] Further, the extracting entities from the input text using a neural network-based method and introducing context information during entity extraction specifically includes:

[0012] Obtain context information containing entities;

[0013] Obtain the input text according to the context information including entities;

[0014] Input the input text into an entity extraction model based on a neural network to obtain the extracted entities.

[0015] Further, the obtaining of the context information including entities specifically includes:

[0016] Obtain the context information including entities through a fixed window, or sentence boundaries, or a context encoder.

[0017] Further, the entity extraction model includes three parts: entity boundary recognition, entity type determination, and domain knowledge transfer. The inputting of the input text into the entity extraction model based on a neural network to obtain the extracted entities specifically includes:

[0018] Input the input text into the entity extraction model, and identify the starting and ending positions of each entity in the input text through the entity boundary recognition part;

[0019] Determine the type to which each entity belongs through the entity type determination part;

[0020] Perform entity disambiguation on each entity through the domain knowledge transfer part to finally obtain the extracted entities.

[0021] Further, the relationship extraction model includes four parts: feature generation, high-level semantic representation, effective semantic segment extraction, and relationship judgment.

[0022] Further, the feature generation part is used to extract high-level features from the input text;

[0023] The high-level semantic representation part is used to convert the input text into a vector representation;

[0024] The effective semantic segment extraction part is used to extract the semantic information related to the entity pair from the input text;

[0025] The relationship judgment part is used to judge the relationship type of the entity pair according to the extracted semantic information.

[0026] Further, after inputting the entity pair and the input text into the relationship extraction model, adopting an iterative optimization strategy to continuously update and optimize the extracted relationship, and obtaining the corresponding knowledge graph, the method further includes:

[0027] Complete the missing relationships in the knowledge graph through semantic association analysis.

[0028] In a second aspect, the present invention provides a device for constructing a knowledge graph, and the device includes:

[0029] An entity extraction module, which is used to extract entities from the input text by using a neural network-based method and introduce context information during entity extraction;

[0030] An entity pair extraction module, connected to the entity extraction module, which is used to extract entity pairs from the extracted entities;

[0031] A knowledge graph construction module, connected to the entity pair extraction module, which is used to input the entity pairs and the corresponding input text into a relation extraction model, and adopt an iterative optimization strategy to continuously update and optimize the extracted relations to obtain the corresponding knowledge graph.

[0032] In a third aspect, the present invention provides a device for constructing a knowledge graph, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to implement the method for constructing a knowledge graph described in the first aspect above.

[0033] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for constructing a knowledge graph described in the first aspect above.

[0034] The method, device and computer-readable storage medium for constructing a knowledge graph provided by the present invention first extract entities from the input text by using a neural network-based method and introduce context information during entity extraction; then extract entity pairs from the extracted entities; and then input the entity pairs and the corresponding input text into a relation extraction model, and adopt an iterative optimization strategy to continuously update and optimize the extracted relations to obtain the corresponding knowledge graph. By introducing context information during entity extraction, the present invention can more accurately link entities to the corresponding entities in the knowledge graph, avoiding ambiguous links. At the same time, by adopting an iterative optimization strategy to continuously update and optimize the extracted relations, the accuracy and integrity of relation extraction can be improved. The present invention has higher efficiency and accuracy when processing large-scale data, can support richer knowledge representation, and has good scalability. It solves the problems that the existing methods for constructing knowledge graphs have limitations in terms of efficiency, accuracy, scalability, etc. Description of the Drawings

[0035] Figure 1 It is a flowchart of a method for constructing a knowledge graph according to Embodiment 1 of the present invention;

[0036] Figure 2 It is a structural schematic diagram of an entity extraction model according to an embodiment of the present invention;

[0037] Figure 3 It is a structural schematic diagram of a relation extraction framework according to an embodiment of the present invention;

[0038] Figure 4 This is a schematic structural diagram of a knowledge graph construction device according to Embodiment 2 of the present invention;

[0039] Figure 5 This is a schematic structural diagram of a knowledge graph construction device according to Embodiment 3 of the present invention. Detailed implementation manners

[0040] To enable those skilled in the art to better understand the technical solutions of the present invention, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0041] It can be understood that the specific embodiments and drawings described herein are only used to explain the present invention, rather than limiting the present invention.

[0042] It can be understood that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.

[0043] It can be understood that, for the convenience of description, only the parts related to the present invention are shown in the drawings of the present invention, and the parts unrelated to the present invention are not shown in the drawings.

[0044] It can be understood that each unit and module involved in the embodiments of the present invention may correspond to only one entity structure, or may be composed of multiple entity structures, or multiple units and modules may also be integrated into one entity structure.

[0045] It can be understood that the terms "first", "second", etc. in the embodiments of the present invention are used to distinguish different objects, or to distinguish different processes for the same object, rather than to describe a specific order of the objects.

[0046] It can be understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of the present invention may occur in a different order from that marked in the drawings.

[0047] It can be understood that in the flowcharts and block diagrams of the present invention, the possible architectures, functions, and operations of the systems, devices, equipment, and methods according to the embodiments of the present invention are shown. Among them, each block in the flowchart or block diagram may represent a unit, module, program segment, or code, which contains executable instructions for implementing the specified function. Moreover, each block or combination of blocks in the block diagram and flowchart can be implemented by a hardware-based system for implementing the specified function, or by a combination of hardware and computer instructions.

[0048] It can be understood that the units and modules involved in the embodiments of the present invention can be implemented in software or in hardware. For example, the units and modules can be located in the processor.

[0049] Example 1:

[0050] This example provides a method for constructing a knowledge graph. As Figure 1 shown, the method includes:

[0051] Step S101: Extract entities from the input text using a neural network-based method, and introduce context information during entity extraction.

[0052] In this example, the construction of the knowledge graph includes entity extraction and relationship extraction. Entity extraction is used to extract atomic information elements from unstructured text, and fine-grained and efficient extraction of named entities is achieved through entity boundary recognition, entity type determination, etc. Among them, in order to improve the accuracy of entity recognition, a neural network-based method is used to extract entities from the input text, and context information is introduced during entity extraction.

[0053] Specifically, with the development of hardware capabilities and the emergence of distributed representations of words, neural network-based methods have become models that can effectively handle many NLP (Natural Language Processing) tasks. As an end-to-end overall process, the processing methods of neural networks for sequence labeling tasks (such as CWS, POS, NER) are similar. Tokens are mapped from discrete one-hot representations to a low-dimensional space to become dense embeddings, and then the embedding sequence of the sentence is input into the RNN to extract features. Finally, the Softmax function predicts the label of each token. Neural network-based methods have further improved the accuracy and scalability of named entity recognition. The corresponding entity extraction models include models based on Bert, XLnet, ERNIE, etc.

[0054] Specifically, the main functions of introducing context information during entity extraction include:

[0055] 1) Disambiguation: Context information can help eliminate the ambiguity of entities. The same entity may have different meanings in different contexts. By analyzing the context around the entity, the specific meaning of the entity can be better understood and associated with the correct entity.

[0056] 2) Entity boundary detection: Context information can help determine the boundaries of entities. In some complex sentence structures or contexts with nested relationships, it may not be possible to accurately determine the start and end positions of entities only relying on the features of individual words. By analyzing the context, the boundaries of entities can be better determined.

[0057] 3) Entity Recognition Consistency: By considering context information, the consistency of entity recognition can be maintained. In a text, the same entity may be mentioned in different ways, or referred to using different names or pronouns. Context information can help normalize these variants to the same entity and maintain the consistency of entity recognition.

[0058] 4) Entity Attribute Recognition: Context information can provide clues about entity attributes. In many cases, the context surrounding an entity can provide information about the entity type, features, and attributes. By analyzing the context, the attributes of the entity can be better understood and associated with the correct entity type.

[0059] Optionally, the method of using a neural network-based approach to extract entities from the input text and introducing context information during entity extraction specifically includes:

[0060] Obtain context information containing the entity;

[0061] Obtain the input text based on the context information containing the entity;

[0062] Input the input text into a neural network-based entity extraction model to obtain the extracted entity.

[0063] In this embodiment, in the entity extraction task, context information is usually the text or sentence around the target entity. The text before and after the target entity can be used as the context to provide more context information to help identify and locate the entity.

[0064] Optionally, the obtaining of the context information containing the entity specifically includes:

[0065] Obtain context information containing the entity through a fixed window, or sentence boundary, or context encoder.

[0066] In this embodiment, to obtain context information, multiple methods can be used, specifically depending on the available data and the characteristics of the task. The specific methods for obtaining context information include but are not limited to the following:

[0067] a) Fixed Window: Select a fixed-size window and extract a certain number of texts before and after the target entity.

[0068] b) Sentence Boundary: Take the sentence as a unit and extract the complete sentence containing the target entity as the context.

[0069] c) Context Encoder: Use a pre-trained language model (such as BERT, GPT, etc.) to take the entire document as the input and capture context information through specific markers or attention mechanisms.

[0070] In this embodiment, once the context information is obtained, it can be introduced into the entity extraction model in the following ways:

[0071] (a) Concatenation: Concatenate the context texts of the target entities together to form a longer input sequence.

[0072] (b) Encoding: Use a specific encoder (such as a recurrent neural network, Transformer) to process the context sequence and obtain the representation of the context.

[0073] (c) Attention mechanism: Use the attention mechanism to dynamically focus on the context part of the target entity so that the model can better utilize the context information.

[0074] For example, assume that we want to extract the name of a service from a user comment. In this case, the sentence containing the target entity can be used as the context. First, collect a comment dataset containing the service name and preprocess it. Then, an entity extraction model based on a recurrent neural network can be designed, taking the context of the target entity as the input and outputting the position or category of the service name. By training the model, evaluating, and tuning, we can obtain an entity extraction system that can accurately extract the service name. Through the above steps, context can be introduced in the entity extraction task and corresponding adjustments and optimizations can be made according to specific requirements and data characteristics.

[0075] Optionally, the entity extraction model includes three parts: entity boundary recognition, entity type determination, and domain knowledge transfer. Inputting the input text into the neural network-based entity extraction model to obtain the extracted entities specifically includes:

[0076] Input the input text into the entity extraction model, and identify the start and end positions of each entity in the input text through the entity boundary recognition part;

[0077] Determine the type to which each entity belongs through the entity type determination part;

[0078] Perform entity disambiguation on each entity through the domain knowledge transfer part to finally obtain the extracted entities.

[0079] In this embodiment, the main tasks of entity extraction include named entity recognition and entity disambiguation. Among them, named entity recognition is an important part of entity extraction and a prerequisite for relation extraction. Information atoms in the text are extracted, such as person names, organization names, geographical locations, dates, monetary values, etc. Named entities are important language units carrying information in the text, with characteristics such as a large number, complex composition rules, and combination nesting. There are two keywords for the entity extraction task: location and classification, which can be abstracted as a sequence labeling problem: given a sentence, label each word in the sentence sequence. At the same time, entity extraction models are applied to various sub-domains, and there are a large number of professional terms in different sub-domains. Therefore, entity extraction models need to transfer knowledge to effectively improve the performance of the extraction engine when the labeled data is insufficient.

[0080] Specifically, as Figure 2 shown, the entity extraction model for the entity extraction engine includes three parts: entity boundary recognition, entity type determination (or entity type judgment), and domain knowledge transfer. Among them, the entity recognition work depends on the entity boundary recognition part and the entity type determination part. The entity boundary recognition part determines the starting and ending positions of the entity in the input sentence, and the entity type determination part determines whether the entity belongs to the defined type set (determined by the entity types required by the business, and different types need to be collected and entered according to different industries here). Entity disambiguation transfers the labeled data knowledge between different sub-domains (domains with rich labeled data and domains with scarce labeled data) through the domain knowledge transfer part, and transfers the entity disambiguation knowledge of the source domain to the target domain, so as to improve the accuracy of entity disambiguation in the target domain. Generally speaking, the entity extraction model is essentially a sequence labeling problem, and the above two stages can be combined together.

[0081] Step S102: Extract entity pairs from the extracted entities.

[0082] In this embodiment, an input text may contain one or more entities. The entity extraction module will identify the entities in the input text and combine two of them into an entity pair. If there is only one entity in an input text (such as a sentence), then no entity pair can be formed. For an input text with three or more entities, the common practice is to combine two of the entities into an entity pair and then perform relation extraction.

[0083] Step S103: Input the entity pair and the corresponding input text into the relation extraction model, and adopt an iterative optimization strategy to continuously update and optimize the extracted relation to obtain the corresponding knowledge graph.

[0084] In this embodiment, the relation extraction model includes four parts: feature generation, high-level semantic representation, effective semantic segment extraction, and relation judgment. Among them, the feature generation part is used to extract high-level features from the input text; the high-level semantic representation part is used to convert the input text into a vector representation; the effective semantic segment extraction part is used to extract semantic information related to the entity pair from the input text; the relation judgment part is used to judge the relation type of the entity pair according to the extracted semantic information.

[0085] Specifically, the relation extraction framework corresponding to the relation extraction model is as Figure 3 shown. Among them, feature generation is used to extract high-level features in the text as input, such as part-of-speech information; high-level semantic representation is used to convert unstructured text into a vector representation; effective semantic segment extraction extracts semantic information related to the input entity pair from the entire input text; relation judgment judges the relation type of the input entity pair according to the extracted effective semantic information.

[0086] In an alternative embodiment, relation extraction includes the following steps:

[0087] 1. Entity extraction: First, entities are extracted from the input text, which means that we need to identify specific entities in the text, such as people, locations, or organizations, etc. A sentence may contain one or more entities.

[0088] 2. Relation extraction model: Next, the input text and the extracted entity pair are input into the relation extraction model. The model includes four main parts:

[0089] (a) Feature generation: In this step, we extract high-level features from the input text, such as part-of-speech information, and these features will be used as one of the inputs of the model.

[0090] (b) High-level semantic representation: This step converts unstructured text into a vector representation, and the purpose of doing this is to represent the text in a form that can be understood and processed by a computer.

[0091] (c) Effective semantic segment extraction: Semantic information related to the input entity pair is extracted from the entire input text, and this information can help judge the relationship between the entity pairs.

[0092] (d) Relation judgment: According to the extracted effective semantic information, judge the relation type of the input entity pair. This step determines what the relationship between the entities is, such as "works at", "is located at", etc.

[0093] 3. Iterative optimization: After relation judgment, an iterative optimization strategy is adopted to continuously update and optimize the extracted relations, which means that we can improve the accuracy and performance of relation extraction by repeatedly training and adjusting the model.

[0094] Optionally, after inputting the entity pair and the input text into the relation extraction model and adopting an iterative optimization strategy to continuously update and optimize the extracted relation to obtain the corresponding knowledge graph, the method further includes:

[0095] Completing the missing relations in the knowledge graph through semantic association analysis.

[0096] In this embodiment, after the knowledge graph is established, automatically completing the missing relations in the graph through semantic association analysis can improve the integrity of the graph.

[0097] It should be noted that the method for constructing the knowledge graph provided by the present invention has the following

[0098] Beneficial effects:

[0099] a) High accuracy: Compared with existing methods, the method of the present invention has achieved a significant improvement in accuracy in entity linking and relation extraction tasks, reaching up to 98%.

[0100] b) Fast processing speed: The algorithm has been optimized and has a fast data processing speed, capable of processing millions of data per day.

[0101] c) Flexibility and scalability: The method of the present invention is applicable to different fields and data sources, and can be conveniently extended to new data types.

[0102] In a specific embodiment, the method for constructing the knowledge graph may include the following steps:

[0103] (1) Determine the domain and scope: Determining the domain and specific scope involved in the knowledge graph helps to determine which data should be collected and can avoid the collection of useless information.

[0104] (2) Collect data: The data can be obtained from structured and semi-structured data sources such as databases, text files, web pages, etc., and supports the import of data through multiple free API data interfaces.

[0105] (3) Data preprocessing: Through preprocessing of structured data, semi-structured data, and text data, the fusion and unified processing of cross-source data are realized.

[0106] (4) Entity recognition and linking: Context information is introduced in entity recognition to more precisely link entities to the corresponding entities in the knowledge graph, avoiding ambiguous links.

[0107] (5) Relation extraction: Adopting an iterative optimization strategy to continuously update and optimize the extracted relations, improving the accuracy and integrity of relation extraction.

[0108] (6) Knowledge representation: Unified representation of entities and relationships.

[0109] (7) Knowledge reasoning: The process of using previously known information to infer new information. Through semantic association analysis (the process of data preparation, entity recognition, relationship extraction, graph construction, knowledge reasoning, evaluation, and optimization), the algorithm can automatically complete the missing relationships in the graph, improving the integrity of the graph.

[0110] (8) Application and maintenance.

[0111] In the method for constructing a knowledge graph provided by an embodiment of the present invention, first, a neural network-based method is used to extract entities from the input text, and context information is introduced during entity extraction; then entity pairs are extracted from the extracted entities; and then the entity pairs and the corresponding input text are input into a relationship extraction model, and an iterative optimization strategy is adopted to continuously update and optimize the extracted relationships to obtain the corresponding knowledge graph. By introducing context information during entity extraction, the present invention can more accurately link entities to the corresponding entities in the knowledge graph, avoiding ambiguous links. At the same time, by adopting an iterative optimization strategy to continuously update and optimize the extracted relationships, the accuracy and integrity of relationship extraction can be improved. The present invention has higher efficiency and accuracy when processing large-scale data, can support richer knowledge representation, and has good scalability. It solves the problems of limitations in efficiency, accuracy, scalability, etc. of existing knowledge graph construction methods.

[0112] Embodiment 2:

[0113] As Figure 4 shown, this embodiment provides a device for constructing a knowledge graph for executing the above method for constructing a knowledge graph. The device includes:

[0114] An entity extraction module 11, configured to extract entities from the input text by using a neural network-based method and introduce context information during entity extraction;

[0115] An entity pair extraction module 12, connected to the entity extraction module 11, configured to extract entity pairs from the extracted entities;

[0116] A knowledge graph construction module 13, connected to the entity pair extraction module 12, configured to input the entity pairs and the corresponding input text into a relationship extraction model, adopt an iterative optimization strategy, and continuously update and optimize the extracted relationships to obtain the corresponding knowledge graph.

[0117] Optionally, the entity extraction module 11 includes:

[0118] An acquisition unit, configured to acquire context information containing entities;

[0119] An obtaining unit, configured to obtain the input text according to the context information including entities;

[0120] An extraction unit, configured to input the input text into an entity extraction model based on a neural network to obtain the extracted entities.

[0121] Optionally, the obtaining unit is specifically configured to:

[0122] Obtain the context information including entities through a fixed window, or sentence boundaries, or a context encoder.

[0123] Optionally, the entity extraction model includes three parts: entity boundary recognition, entity type determination, and domain knowledge transfer. The extraction unit specifically includes:

[0124] A boundary recognition unit, configured to input the input text into the entity extraction model, and identify the start and end positions of each entity in the input text through the entity boundary recognition part;

[0125] A type determination unit, configured to determine the type to which each entity belongs through the entity type determination part;

[0126] An entity disambiguation unit, configured to perform entity disambiguation on each entity through the domain knowledge transfer part to finally obtain the extracted entities.

[0127] Optionally, the relationship extraction model includes four parts: feature generation, high-level semantic representation, effective semantic fragment extraction, and relationship judgment.

[0128] Optionally, the feature generation part is configured to extract high-level features from the input text;

[0129] The high-level semantic representation part is configured to convert the input text into a vector representation;

[0130] The effective semantic fragment extraction part is configured to extract the semantic information related to the entity pair from the input text;

[0131] The relationship judgment part is configured to judge the relationship type of the entity pair according to the extracted semantic information.

[0132] Optionally, the apparatus further includes:

[0133] A relationship completion module, configured to complete the missing relationships in the knowledge graph through semantic association analysis.

[0134] Embodiment 3:

[0135] Reference Figure 5, this embodiment provides a knowledge graph construction device, including a memory 21 and a processor 22. A computer program is stored in the memory 21, and the processor 22 is configured to run the computer program to execute the knowledge graph construction method in Embodiment 1.

[0136] Among them, the memory 21 is connected to the processor 22. The memory 21 can be a flash memory, a read-only memory, or other memories, and the processor 22 can be a central processing unit or a single-chip microcomputer.

[0137] Embodiment 4:

[0138] This embodiment provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the knowledge graph construction method in Embodiment 1 above.

[0139] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, computer program modules, or other data. The computer-readable storage medium includes, but is not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory, or other memory technologies, CD-ROM (Compact Disc Read-Only Memory), digital versatile disc (DVD), or other optical disc storage, magnetic cassette, tape, magnetic disk storage, or any other medium that can be used to store the desired information and can be accessed by a computer.

[0140] In summary, the method, apparatus, and computer-readable storage medium for constructing a knowledge graph provided by the embodiments of the present invention first extract entities from the input text using a neural network-based method and introduce context information during entity extraction; then extract entity pairs from the extracted entities; and then input the entity pairs and the corresponding input text into a relation extraction model, and adopt an iterative optimization strategy to continuously update and optimize the extracted relations to obtain the corresponding knowledge graph. By introducing context information during entity extraction, the present invention can more accurately link entities to the corresponding entities in the knowledge graph, avoiding ambiguous links. At the same time, by adopting an iterative optimization strategy to continuously update and optimize the extracted relations, the accuracy and integrity of relation extraction can be improved. The present invention has higher efficiency and accuracy when processing large-scale data, can support richer knowledge representation, and has good scalability. It solves the problems of limitations in efficiency, accuracy, scalability, etc. of the existing knowledge graph construction methods.

[0141] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principle of the present invention, and the present invention is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also regarded as the protection scope of the present invention.

Claims

1. A method for constructing a knowledge graph, characterized in that, The method includes: Adopting a neural network-based method to extract entities from the input text and introducing context information during entity extraction; Extracting entity pairs from the extracted entities; Inputting the entity pairs and the corresponding input text into a relation extraction model, and adopting an iterative optimization strategy to continuously update and optimize the extracted relations to obtain the corresponding knowledge graph.

2. The method according to claim 1, characterized in that, The adopting a neural network-based method to extract entities from the input text and introducing context information during entity extraction specifically includes: Obtaining context information containing entities; Obtaining the input text according to the context information containing entities; Inputting the input text into a neural network-based entity extraction model to obtain the extracted entities.

3. The method according to claim 2, characterized in that, The obtaining context information containing entities specifically includes: Obtaining context information containing entities through a fixed window, or sentence boundaries, or a context encoder.

4. The method according to claim 2, characterized in that, The entity extraction model includes three parts: entity boundary recognition, entity type determination, and domain knowledge transfer. The inputting the input text into a neural network-based entity extraction model to obtain the extracted entities specifically includes: Inputting the input text into the entity extraction model, and identifying the starting and ending positions of each entity in the input text through the entity boundary recognition part; Determining the type to which each entity belongs through the entity type determination part; Performing entity disambiguation on each entity through the domain knowledge transfer part to finally obtain the extracted entities.

5. The method according to claim 1, characterized in that, The relation extraction model includes four parts: feature generation, high-level semantic representation, effective semantic fragment extraction, and relation determination.

6. The method according to claim 5, characterized in that, The feature generation part is used to extract high-level features from the input text; The high-level semantic representation part is used to convert the input text into a vector representation; The effective semantic fragment extraction part is used to extract semantic information related to the entity pair from the input text; The relation determination part is used to judge the relation type of the entity pair according to the extracted semantic information.

7. The method according to claim 1, characterized in that, After inputting the entity pairs and the input text into the relation extraction model, adopting an iterative optimization strategy to continuously update and optimize the extracted relations to obtain the corresponding knowledge graph, the method further includes: Completing the missing relations in the knowledge graph through semantic association analysis.

8. An apparatus for constructing a knowledge graph, characterized in that, The device includes: An entity extraction module, which is used to adopt a neural network-based method to extract entities from the input text and introduce context information during entity extraction; An entity pair extraction module, connected to the entity extraction module, which is used to extract entity pairs from the extracted entities; A knowledge graph construction module, connected to the entity pair extraction module, which is used to input the entity pairs and the corresponding input text into a relation extraction model, adopt an iterative optimization strategy to continuously update and optimize the extracted relations to obtain the corresponding knowledge graph.

9. An apparatus for constructing a knowledge graph, characterized in that, It includes a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to implement the method for constructing a knowledge graph according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the method for constructing the knowledge graph according to any one of claims 1-7 is implemented.