Knowledge graph construction method and apparatus, and storage medium and electronic device
By using pre-trained extraction models and semantic encoders to encode and cluster entities and relationship types in open domain knowledge graph construction, the problem of insufficient accuracy and integrity of knowledge graph construction in the existing technology is solved, and the ability to efficiently build data layers and pattern layers is realized.
Patent Information
- Application Number
- PCT/CN2024/120324
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-09-23
- Publication Date
- 2025-06-19
AI Technical Summary
The existing open domain knowledge graph construction methods have shortcomings in terms of accuracy and integrity, especially the inability to efficiently construct data layer and schema layer at the same time, resulting in limited graph management, update and reasoning capabilities.
By obtaining the natural language text to be analyzed, using the pre-trained extraction model for entity extraction, and coding and clustering entity types and relationship types in combination with a semantic encoder, determining the target entity types and relationship types, and finally building the target knowledge graph.
It realizes improving accuracy and integrity in the construction of open domain knowledge graphs, and can build data layer and pattern layer at the same time, enhancing the management, update and reasoning capabilities of graphs.
Smart Images

Figure CN2024120324_19062025_PF_FP_ABST
Abstract
Description
Knowledge graph construction method, device, storage medium and electronic device
[0001] Related applications
[0002] This application claims priority to Chinese patent application number 2023117274741, filed on December 14, 2023, entitled “Knowledge graph construction method, device, storage medium and electronic device,” the entire text of which is hereby incorporated by reference. Technical Field
[0003] The present application relates to the field of knowledge graph technology, and specifically, to a knowledge graph construction method, device, storage medium, electronic device and computer program product. Background Art
[0004] A knowledge graph is a representation of the knowledge formed by humans' understanding of the objective world. It organizes and stores the knowledge people use to solve real-world problems through specific structures and classifications. Depending on the type of knowledge contained, knowledge graph question and answer can be divided into two scenarios: open-domain and closed-domain. Closed-domain knowledge graph question and answer limits the scope of knowledge interaction to a specific domain or subject, and the context of questions and answers is relatively limited. In open-domain knowledge graph question and answer, questions and knowledge in the knowledge graph can cover any subject, domain, or topic. Therefore, open-domain knowledge graph question and answer has a broader application prospect.
[0005] Summary of the Invention
[0006] Embodiments of the present application provide a knowledge graph construction method, device, storage medium, electronic device, and computer program product.
[0007] According to one aspect of an embodiment of the present application, a knowledge graph construction method is provided, including: obtaining a first natural language text to be analyzed; performing entity extraction on the first natural language text using a pre-trained extraction model to obtain a first entity and a first entity relationship in the first natural language, and determining a first entity type of the first entity and a first relationship type of the first entity relationship; encoding the first entity type and the first relationship type using a pre-trained semantic encoder to obtain a first semantic vector corresponding to the first entity type and a second semantic vector corresponding to the first relationship type, respectively; calculating a first distance between the first semantic vector and a preset plurality of entity type clustering centers, and a second distance between the second semantic vector and a preset plurality of relationship type clustering centers, and determining a target entity type corresponding to the first entity type and a target relationship type corresponding to the first relationship type based on the first distance and the second distance, respectively; and constructing a target knowledge graph based on the first entity, the first entity relationship, the target entity type corresponding to the first entity type, and the target relationship type corresponding to the first relationship type.
[0008] Optionally, the training process of the extraction model includes: constructing an initial language model; obtaining multiple sets of sample data, wherein the multiple sets of sample data include: a second natural language text, and corresponding second entities and second entity relationships; iteratively training the initial language model based on the multiple sets of sample data to obtain an extraction model.
[0009] Optionally, obtaining multiple groups of sample data includes: obtaining multiple groups of first-class triple samples, wherein the first-class triple samples include: a second entity and a second entity relationship; classifying the multiple groups of first-class triple samples according to the entity name to obtain multiple classification groups, wherein each classification group includes at least one group of triple samples; for each classification group, combining at least two first-class triple samples in the classification group to obtain prompt information, and using a general large model to generate a second natural language text corresponding to the prompt information.
[0010] Optionally, the training process of the semantic encoder includes: constructing an initial deep learning network; obtaining multiple groups of second-category triplet samples, wherein the second-category triplet samples include: baseline samples, positive samples with the same type of labels as the baseline samples, and negative samples with different types of labels as the baseline samples, and the baseline samples, positive samples, and negative samples all include the second entity type corresponding to the second entity and the first semantic vector corresponding to the second entity type, or the second relationship type corresponding to the second entity relationship and the second semantic vector corresponding to the second relationship type; constructing a triplet loss function based on the baseline samples, positive samples, and negative samples in each group of second-category triplet samples, and adjusting the network parameters of the initial deep learning network according to the triplet loss function to obtain a semantic encoder.
[0011] Optionally, multiple groups of second-category triple samples are obtained, including: for the second entity in each first-category triple sample, determining the second entity type corresponding to the second entity, and encoding the second entity type to obtain the corresponding first semantic vector, and forming a first sample by the second entity type corresponding to the second entity and the first semantic vector corresponding to the second entity type; marking the second entity type in each first sample, and determining the entity type label of the second entity type in the first sample; selecting a sample from multiple first samples as a benchmark sample, and determining the first entity type label of the benchmark sample; obtaining positive samples from multiple samples corresponding to the first entity type label, and obtaining negative samples from multiple samples with second entity type labels different from the first entity type labels; and forming a second-category triple sample by the benchmark sample, positive sample, and negative sample.
[0012] Optionally, multiple groups of second-category triple samples are obtained, including: for the second entity relationship within each first-category triple sample, determining the second relationship type corresponding to the second entity relationship, and encoding the second relationship type to obtain the corresponding second semantic vector, and forming a second sample by the second relationship type corresponding to the second entity relationship and the second semantic vector corresponding to the second relationship type; marking the second relationship type within each second sample, and determining the relationship type label of the second relationship type within the second sample; selecting a sample from multiple second samples as a benchmark sample, and determining the first relationship type label of the benchmark sample; obtaining positive samples from multiple samples corresponding to the first relationship type label, and obtaining negative samples from multiple samples with second relationship type labels different from the first relationship type labels; and forming a second-category triple sample by the benchmark sample, positive sample, and negative sample.
[0013] Optionally, the process of determining multiple entity type cluster centers includes: obtaining the first semantic vector of the second entity type corresponding to the second entity in multiple groups of sample data, and selecting one from the multiple first semantic vectors as the first initial centroid; calculating the first distance from the other first semantic vectors in the multiple first semantic vectors except the first initial centroid to the first initial centroid, and determining the next initial centroid based on the first distance, and repeating the initial centroid determination process with the next initial centroid as the first initial centroid until multiple initial centroids are determined; calculating the second distance from the other first semantic vectors in the multiple first semantic vectors except the multiple initial centroids to each initial centroid, and determining multiple clusters based on the second distance; updating the multiple initial centroids according to each first semantic vector in each cluster to obtain multiple entity type cluster centers.
[0014] Optionally, the process of determining multiple relationship type cluster centers includes: obtaining the second semantic vector of the second relationship type corresponding to the second entity relationship in multiple groups of sample data, and selecting one from the multiple second semantic vectors as the first initial centroid; calculating the first distance from the other second semantic vectors in the multiple second semantic vectors except the first initial centroid to the first initial centroid, and determining the next initial centroid based on the first distance, and repeating the initial centroid determination process with the next initial centroid as the first initial centroid until multiple initial centroids are determined; calculating the second distance from the other second semantic vectors in the multiple second semantic vectors except the multiple initial centroids to each initial centroid, and determining multiple clusters based on the second distance; updating the multiple initial centroids according to each second semantic vector in each cluster to obtain multiple relationship type cluster centers.
[0015] Optionally, a target knowledge graph is constructed based on the first entity, the first entity relationship, the target entity type corresponding to the first entity type, and the target relationship type corresponding to the first relationship type, including: constructing a data layer of the target knowledge graph based on the first entity and the first entity relationship, and constructing a model layer of the target knowledge graph based on the target entity type corresponding to the first entity type and the target relationship type corresponding to the first relationship type.
[0016] Optionally, the initial language model is a general large model that is trained to have the ability to generate natural language text based on triples.
[0017] Optionally, using the general large model to generate a second natural language text corresponding to the prompt information includes: using the general large model to expand the prompt information to generate a corresponding second natural language text; wherein the expanded content includes all the contents of at least two first-category triplet samples in the classification group.
[0018] Optionally, the second entity type tag is any one of the entity type tags other than the first entity type tag in the multiple types of entity type tags.
[0019] Optionally, the knowledge graph construction method also includes: performing semantic search and / or content generation based on the constructed knowledge graph.
[0020] According to another aspect of an embodiment of the present application, a knowledge graph construction device is also provided, including: an acquisition module for acquiring a first natural language text to be analyzed; an entity extraction module for performing entity extraction on the first natural language text using a pre-trained extraction model, obtaining a first entity and a first entity relationship in the first natural language, and determining a first entity type of the first entity and a first relationship type of the first entity relationship; an encoding module for encoding the first entity type and the first relationship type respectively using a pre-trained semantic encoder, obtaining a first semantic vector corresponding to the first entity type and a second semantic vector corresponding to the first relationship type; a clustering module for calculating a first distance between the first semantic vector and a preset plurality of entity type clustering centers, and a second distance between the second semantic vector and a preset plurality of relationship type clustering centers, and determining a target entity type corresponding to the first entity type and a target relationship type corresponding to the first relationship type according to the first distance and the second distance, respectively; a construction module for constructing a target knowledge graph based on the first entity, the first entity relationship, the target entity type corresponding to the first entity type, and the target relationship type corresponding to the first relationship type.
[0021] According to another aspect of an embodiment of the present application, a non-volatile storage medium is also provided, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the above-mentioned knowledge graph construction method by running the computer program.
[0022] According to another aspect of an embodiment of the present application, an electronic device is also provided, which includes: a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned knowledge graph construction method through the computer program.
[0023] According to another aspect of an embodiment of the present application, a computer program product is also provided, including a computer program, wherein the computer program implements the above-mentioned knowledge graph construction method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.
[0025] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0026] FIG1 is a hardware structure block diagram of a computer terminal for implementing a knowledge graph construction method according to an embodiment of the present application;
[0027] FIG2 is a flow chart of an optional knowledge graph construction method according to an embodiment of the present application;
[0028] FIG3 is an optional flowchart of extracting entity types and relationship types according to an embodiment of the present application;
[0029] FIG4 a is a schematic diagram of an optional simple triplets (Easy Triplets) according to an embodiment of the present application;
[0030] FIG4 b is a schematic diagram of an optional hard triplets according to an embodiment of the present application;
[0031] FIG4 c is a schematic diagram of an optional general triplet (Semi-Hard Triplets) according to an embodiment of the present application;
[0032] FIG5 is a schematic diagram of an optional type of alignment according to an embodiment of the present application;
[0033] FIG6 is a schematic diagram of an optional data layer and mode layer according to an embodiment of the present application;
[0034] Figure 7 is a structural diagram of an optional knowledge graph construction device according to an embodiment of the present application. DETAILED DESCRIPTION
[0035] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0036] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0038] In addition, the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. For example, an interface is set up between this system and the relevant user or organization. Before obtaining relevant information, it is necessary to send an acquisition request to the aforementioned user or organization through the interface, and obtain the relevant information after receiving the consent information fed back by the aforementioned user or organization.
[0039] Currently, open-domain knowledge graph construction approaches can be broadly categorized into two types: syntactic analysis-based and model-based. Syntactic analysis-based approaches first use a parser to break down sentences into components such as subject, predicate, and object, and then use a classifier to classify each segment. Model-based approaches, on the other hand, first train a model using supervised data, then use the model to extract features and construct the graph. However, both approaches suffer from the following drawbacks: First, they suffer from low construction accuracy: Syntactic analysis-based approaches require text to strictly adhere to grammar, resulting in very low accuracy for non-standard texts. Model-based approaches require massive amounts of training data, but currently available open-domain training data is scarce, resulting in low accuracy. Second, the constructed knowledge graph is incomplete: This is because both approaches only construct entities and relationships, effectively building the data layer but lacking the model layer, making them difficult to manage, update, and reason about.
[0040] Therefore, in the construction of open domain knowledge graphs, there is an urgent need for a high-precision construction method that can simultaneously construct the data layer and the model layer.
[0041] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:
[0042] A knowledge graph is a structured semantic knowledge base used to describe concepts and their relationships in the physical world in symbolic form. Its basic components are "entity-relationship-entity" triples, as well as entities and their related attribute-value pairs. Entities are connected to each other through relationships, forming a network-like knowledge structure. A knowledge graph depicts entities and the relationships between them in the form of nodes and edges. Nodes represent entities (such as people, places, and events), while edges represent various relationships between entities (such as "belong to," "located in," and "participate in"). Knowledge graphs not only provide rich information but also help artificial intelligence systems better understand and reason, thereby improving the effectiveness of applications such as natural language processing, recommendation systems, and question-answering systems.
[0043] Among them, the technical framework of knowledge graph can include:
[0044] 1. Expression
[0045] A common representation of knowledge graphs is a triple, G = (E, R, S), where F represents an entity in the knowledge base, R represents a relationship in the knowledge base, and S represents a triple in the knowledge base. The basic forms of a triple are primarily first entity, relationship, second entity, concept, attribute, and attribute value. Entities, as the most basic elements in a knowledge graph, have different relationships between different entities. Concepts primarily refer to sets, categories, object types, and types of things, such as people and geography. Attributes primarily refer to the properties, characteristics, features, and parameters that an object may possess, such as nationality and birthday. Attribute values primarily refer to the values of a specific attribute of an object, such as 1988-09-08 and China.
[0046] 2. Logical Architecture
[0047] The knowledge graph is logically structured into two layers: the data layer and the schema layer. The data layer is a graph database that stores facts, with facts represented in the form of "first entity - relationship - second entity" or "entity - attribute - attribute value." The schema layer stores refined knowledge, leveraging an ontology library to standardize entities, relationships, and the relationships between entity types and attributes.
[0048] 3. System Architecture
[0049] The knowledge graph architecture is divided into three parts: acquiring source data, knowledge fusion, and knowledge computation and application. Knowledge graphs are constructed in two ways: top-down and bottom-up. In the early days of knowledge graph development, knowledge graphs primarily relied on structured data sources like encyclopedias to extract ontology and schema information, adding this information to the database using a top-down approach. Currently, however, most knowledge graphs are constructed using a bottom-up approach, whereby publicly collected data is automatically extracted and then manually reviewed before being added to the knowledge base.
[0050] Large Language Models (LLMs): These are large language models that are trained on large amounts of text data and can perform a wide range of tasks, including text summarization, translation, and sentiment analysis. LLMs are characterized by their large size, containing billions of parameters, which help them learn complex patterns in language data. These models are often based on deep learning architectures such as transformers, which helps them achieve good performance on various natural language processing tasks.
[0051] Triplet loss is a widely used loss function, often used in face recognition tasks, to distinguish between dissimilar and similar samples. Specifically, Triplet loss excels in detail differentiation. Specifically, when two inputs are similar, Triplet loss can better model these details. This effectively minimizes the differences between the two inputs and learns a better representation of the input, resulting in better performance in these tasks.
[0052] In related technologies, there are two main methods for constructing open-domain knowledge graphs: syntactic analysis and model building. Syntactic analysis involves grammatically breaking down sentences to extract components such as subject, predicate, and object, thereby converting unstructured text information into structured triples. The accuracy of this method is easily constrained by sentence complexity; in other words, accuracy is very low when sentence structures are complex or when sentence components are missing. Constructing open-domain graphs using models requires a large amount of <text, triple> data pairs for model training. Due to the lack of massive sample data, the accuracy of open-domain graph construction using model building is also low.
[0053] In addition, both of the above graph construction methods can only produce information such as entities and relationships at the graph data layer, but cannot produce information such as entity types and relationship types at the graph model layer.
[0054] To address this issue, the present application provides an embodiment of a method for constructing a knowledge graph, which is described in detail below. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in a different order than shown here.
[0055] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 shows a hardware structure block diagram of a computer terminal for implementing a knowledge graph construction method. As shown in Figure 1, the computer terminal 10 may include one or more (102a, 102b, ..., 102n are used in the figure to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that the structure shown in Figure 1 is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may also include more or fewer components than shown in Figure 1, or have a configuration different from that shown in Figure 1.
[0056] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0057] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the knowledge graph construction method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizes the knowledge graph construction method of the above-mentioned application. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include but are not limited to the Internet, corporate intranet, local area network, mobile communication network and combinations thereof.
[0058] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0059] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .
[0060] In the above operating environment, FIG2 is a flow chart of an optional knowledge graph construction method according to an embodiment of the present application. As shown in FIG2 , the method includes at least steps S201-S205, wherein:
[0061] Step S201: Obtain a first natural language text to be analyzed.
[0062] The first natural language text may be various forms of document information, including but not limited to books, papers, news reports, blog posts, social media posts, emails, web page content, etc. These texts may be knowledge in any field, covering a wide range of themes and topics.
[0063] Step S202: perform entity extraction on the first natural language text using a pre-trained extraction model to obtain a first entity and a first entity relationship in the first natural language, and determine a first entity type of the first entity and a first relationship type of the first entity relationship.
[0064] The above steps can be understood as extracting first entities and first entity relationships between first entities, such as names of people, places, and organization names, from the first natural language text through the extraction model.
[0065] Specifically, the training process of the above extraction model includes the following steps S11-S13, wherein:
[0066] Step S11: construct an initial language model.
[0067] Among them, the above-mentioned initial language model can be a general large model, and the training method of the above-mentioned general large model includes: step one, model pre-training, using a masked language model, combined with an unsupervised training method, to perform mask completion training to obtain a pre-trained model; step two, instruction fine-tuning training, using manually labeled instruction data pairs, to supervise the pre-trained model obtained in step one to obtain a fine-tuning model; step three, reward model training, manually labeling the output of the fine-tuning model, and obtaining "good" and "bad" results for training the reward model; step four, using a reinforcement learning method, according to the "fine-tuning model output-reward model evaluation-fine-tuning model retraining" mode, multiple rounds of training are repeated, and finally a general large model can be obtained, and the general large model has the ability to produce natural language text based on triples.
[0068] Step S12: Acquire multiple groups of sample data, wherein the sample data includes: a second natural language text, and corresponding second entities and second entity relationships.
[0069] At present, in the process of constructing open domain knowledge graphs, it is usually necessary to convert unstructured text into structured triple data. However, since the scope of text in open domain scenarios is very wide, if a general large model is directly used for entity extraction, the accuracy of the extracted entities and entity relationships will be low; and if the general large model is directly fine-tuned in the form of <text, triple> to obtain an extraction large model, there will be a problem of lack of sample data.
[0070] Therefore, in the technical solution provided in step S12, multiple sets of sample data can be obtained through the following steps:
[0071] Step 1: Obtain multiple groups of first-category triple samples, where the first-category triple samples include: a second entity and a second-entity relationship;
[0072] Step 2: Classify multiple groups of first-category triple samples according to entity names to obtain multiple classification groups, where each classification group includes at least one group of triple samples;
[0073] Step 3: For each classification group, combine at least two first-category triplet samples within the classification group to obtain prompt information, and use the universal large model to generate a second natural language text corresponding to the prompt information.
[0074] In the above embodiment, multiple groups of first-class triple samples can be collected from the Internet, where the triple information can come from structured or block-structured data such as encyclopedias, websites, and private domain data. Since triple information alone cannot train the extraction model, the collected multiple groups of first-class triple samples are classified and counted according to entity names, and first-class triple samples with the same entity are combined into a classification group; then, within each classification group, multiple first-class triple samples are randomly selected to form prompt information, and the corresponding second natural language text is generated using the universal large model, so that each first-class triple sample and the corresponding second natural language text constitute the sample data.
[0075] For example, multiple triples with the same concept are obtained in the form of <concept, attribute, attribute value>, such as: {"S":"Green tea [tea variety]","P":"Does it contain preservatives?","O":"No"}, {"S":"Green tea [tea variety]","P":"Main nutrients","O":"Tea polyphenols, caffeine peptides"}, {"S":"Green tea [tea variety]","P":"Side effects","O":"Filtering references may lead to reduced sleep"}, {"S":"Green tea [tea variety]","P":"Foreign name","O":"Green Tea"}. These triples are combined and the resulting prompt information is expanded using a general large model to generate corresponding natural language text. The expanded content includes all content from at least two first-category triple samples within the classification group. It is important to note that the expanded content must strictly originate from the triple information and should not generate content other than the triple information, but the entire triple information should also not be omitted. Therefore, the following natural language text can be expanded from the above four triples, namely, "Green tea, also known as Green Tea, is a variety of tea. It is famous for being preservative-free, which is an important consideration for health-conscious people. The main nutrients of green tea include tea polyphenols and caffeine, which give it many popular health benefits. However, it should be noted that excessive consumption of green tea may lead to reduced sleep, so it is necessary to maintain moderation when enjoying it."
[0076] Step S13: iteratively train the initial language model based on multiple groups of sample data to obtain an extraction model.
[0077] The original first-category triplet samples obtained in step S12 and the generated second natural language text are used as training data for the initial language model, and the initial language model is iteratively trained to obtain a fine-tuned extraction model.
[0078] In the training process of the above-mentioned large extraction model, by first expanding the triples into text and then using <text, triples> as supervision for training, compared with the current method of directly obtaining and using <text, triples> as supervision for training, it is possible to obtain training sample data that is more than one thousand times the size of the current largest Chinese training set, thereby ensuring that the extraction accuracy exceeds all the large extraction models currently on the market.
[0079] Furthermore, the first natural language text is subjected to entity extraction using the extraction model trained in steps S11-S13 above, thereby obtaining a first entity and a first entity relationship within the first natural language text. The extraction results from the above steps and the first natural language text are then used as inputs to a general large model, which then outputs the type corresponding to the extraction results, namely, the first entity type of the first entity and the first relationship type of the first entity relationship.
[0080] For example, Figure 3 is an optional entity type and relationship type extraction flow chart according to an embodiment of the present application. As shown in Figure 3, the first natural language text is "Product A is discounted and promoted in a certain month, so that the sales of Product A doubled." The first entity extracted from the first natural language text using the extraction big model is ["Product A discounted and promoted", "sales doubled"], and the first entity relationship is ["so that"]. The first natural language text, the first entity, and the first entity relationship are used as inputs of the general big model, and the general big model outputs the first entity type {"Product A discounted and promoted": "marketing activities", "sales doubled": "marketing results"}, and the first relationship type {"so that": "results"}.
[0081] However, since there are many problems of synonymous multiple expressions in the extraction results of the large extraction model, that is, the same entity or the same entity relationship has multiple expressions, and the large extraction model may only select any one of them when extracting natural language text, resulting in poor integrity of the knowledge graph construction. In this regard, the embodiment of the present application aligns the entity type and relationship type separately by using vector encoding and clustering in the following steps S203-S204.
[0082] Step S203 : Encode the first entity type and the first relationship type using a pre-trained semantic encoder to obtain a first semantic vector corresponding to the first entity type and a second semantic vector corresponding to the first relationship type, respectively.
[0083] Specifically, the training process of the semantic encoder includes the following steps S21-S23, wherein:
[0084] Step S21: construct an initial deep learning network.
[0085] Among them, the above-mentioned initial deep learning network can be BERT (Bidirectional Encoder Representation from Transformers, a pre-trained language representation model), M3E (Moka Massive Mixed Embedding-Bas), etc. Among them, since the relevant content of the initial deep learning network has been described in the existing literature, this application will not go into details about this part.
[0086] Step S22, obtain multiple groups of second-category triplet samples, wherein the second-category triplet samples include: baseline samples, positive samples with the same type of labels as the baseline samples, and negative samples with different types of labels as the baseline samples, and the baseline samples, positive samples, and negative samples all include the second entity type corresponding to the second entity and the first semantic vector corresponding to the second entity type, or the second relationship type corresponding to the second entity relationship and the second semantic vector corresponding to the second relationship type.
[0087] In the technical solution provided in step S22, each sample in the second type of triplet samples may be a sample corresponding to an entity type, or may be a sample corresponding to an entity relationship.
[0088] Therefore, the method for obtaining the second type of triplet samples may include:
[0089] First, for the second entity in each first-category triple sample, determine the second entity type corresponding to the second entity, encode the second entity type to obtain the corresponding first semantic vector, and form a first sample consisting of the second entity type corresponding to the second entity and the first semantic vector corresponding to the second entity type;
[0090] Next, the second entity type in each first sample is labeled to determine an entity type label of the second entity type in the first sample;
[0091] Then, a sample is selected from the multiple first samples as a reference sample, and a first entity type label of the reference sample is determined; a positive sample is obtained from the multiple samples corresponding to the first entity type label, and a negative sample is obtained from the multiple samples with a second entity type label different from the first entity type label. The second entity type label can be understood as any entity type label other than the first entity type label in the multiple entity type labels;
[0092] Finally, the second type of triplet samples is composed of the benchmark samples, positive samples, and negative samples.
[0093] In addition, the second type of triplet samples can be obtained by the following methods:
[0094] First, for each first-category triplet sample, a second relationship type corresponding to the second entity relationship is determined, and the second relationship type is encoded to obtain a corresponding second semantic vector. A second sample is formed by the second relationship type corresponding to the second entity relationship and the second semantic vector corresponding to the second relationship type.
[0095] Next, the second relationship type in each second sample is marked to determine the relationship type label of the second relationship type in the second sample; a sample is selected from the plurality of second samples as a reference sample, and the first relationship type label of the reference sample is determined;
[0096] Then, positive samples are obtained from multiple samples corresponding to the first relationship type label, and negative samples are obtained from multiple samples with a second relationship type label different from the first relationship type label. The second relationship type label is any one of the relationship type labels other than the first relationship type label in the multiple relationship type labels.
[0097] Finally, the second type of triplet samples is composed of the benchmark samples, positive samples, and negative samples.
[0098] The purpose of the second type of triplet samples, which include baseline samples, positive samples, and negative samples, is to improve its ability to understand semantics and learn representations. Training the semantic encoder with the second type of triplet samples, which includes baseline samples, positive samples, and negative samples, allows the semantic encoder to learn how to distinguish samples from different categories and ensures that the distance between the semantic vectors generated by the semantic encoder in the vector space is proportional to the correlation between the semantics.
[0099] Step S23: construct a triplet loss function based on the reference sample, positive sample, and negative sample in each group of second-category triplet samples, and adjust the network parameters of the initial deep learning network according to the triplet loss function to obtain a semantic encoder.
[0100] The basic idea of the triplet loss function is: for a given triplet (Anchor, Positive, Negative), Anchor and Positive are different samples of the same type, while Anchor and Negative are different samples. The triplet loss attempts to learn a feature space in which the reference samples (Anchor) of the same category are closer to the positive samples (Positive), while the reference samples (Anchor) of different categories are farther away from the negative samples (Negative). Therefore, the expression of the triplet loss function is: L = max{d(a,p) - d(a,n) + margin, 0}
[0101] Where a represents anchor, p represents positive, and n represents negative. Margin (distance threshold) is a constant greater than 0. The optimization goal of this loss function is to shorten the distance between anchor and positive and increase the distance between anchor and negative. Therefore, the second type of triple samples mentioned above can be divided into the following three categories:
[0102] Easy Triplets: L = 0, i.e., d(a, p) + margin < d(a, n). In this case, optimization is required, as shown in Figure 4a.
[0103] Hard Triplets: L > margin, that is, d(a, p) > d(a, n). This means that the distance between the anchor and the negative is close, while the distance between the anchor and the positive is far, as shown in Figure 4b. This situation results in the greatest loss and requires optimization.
[0104] Semi-Hard Triplets: L < margin, that is, d(a, p) < d(a, n) < d(a, p) + margin. This means that the distance between the anchor and the positive is closer than the distance between the anchor and the negative, but not close enough to meet the margin requirement, as shown in Figure 4c. In this case, there is a loss. Although the loss is smaller than that of hard triplets, it still needs to be optimized.
[0105] It should be noted that the purpose of setting margin here is: 1. To avoid taking shortcuts in the model. When the embeddings (vector representations in the feature space) of negative and positive are trained to be relatively close, if there is no margin, the triplet loss function becomes L = max{d(a,p) - d(a,n), 0}. Then, as long as d(a,p) = d(a,n), the loss function can be satisfied, that is, the distance between the benchmark sample and the positive sample and the negative sample is the same. In this way, the model will have difficulty in correctly distinguishing between positive and negative samples. Second, by setting the margin, the model can be forced to learn so that the distance between the benchmark sample and the negative sample is larger, and the distance between the benchmark sample and the positive sample is smaller. Third, due to the existence of the margin, the triplet loss function has an additional parameter, and the size of the margin needs to be adjusted. If the margin is large, the model loss will be large, and at the end of learning, the loss result will be difficult to approach 0, and may even cause the network to not converge. However, it can better classify relatively similar samples, that is, distinguish between the benchmark sample and the positive sample. If the margin is small, the loss result is easy to approach 0, the model is easy to train, but it is difficult to distinguish between the benchmark sample and the positive sample.
[0106] In the above embodiment, it can be ensured that the semantic encoder finally trained has good encoding capabilities, and at the same time, it can be ensured that in the vector space, the distance between vectors is proportional to the semantic relevance, that is, the higher the degree of similarity of entities, relationships or types, the closer the vector distance.
[0107] Step S204, calculate the first distance between the first semantic vector and the preset centers of multiple entity type clusters, and the second distance between the second semantic vector and the preset centers of multiple relationship type clusters, and determine the target entity type corresponding to the first entity type and the target relationship type corresponding to the first relationship type based on the first distance and the second distance respectively.
[0108] The entity type cluster centers and relationship type cluster centers are both determined using the K-means++ clustering algorithm. Compared to K-means, this algorithm makes improvements when determining the initial cluster centers. The specific process of this algorithm is as follows:
[0109] The first step is to select a sample point from multiple samples in the data set x as the first initial cluster center c i ;
[0110] The second step is to calculate the shortest distance between each sample in the sample set and the existing cluster center (that is, the distance to the nearest cluster center), which is represented by D(x). The larger this value is, the greater the probability of being selected as the cluster center. Then calculate the probability of each sample being selected as the next cluster center: Select the next cluster center according to the roulette wheel method;
[0111] Step 3: Repeat step 2 until K cluster centers are selected;
[0112] Step 4: For each sample x in the dataset i , calculate its distance to the K cluster centers and divide it into the class corresponding to the cluster center with the smallest distance;
[0113] Step 5: For each category c i , recalculate its cluster center (i.e., the centroid of all samples belonging to that class);
[0114] Step 6: Repeat steps 4 and 5 until the location of the cluster center no longer changes.
[0115] Next, we will determine the cluster centers of multiple entity types and multiple relationship types according to the above-mentioned K-means++ clustering algorithm.
[0116] Specifically, the process of determining the cluster centers of the above-mentioned multiple entity types includes:
[0117] Step 1: Obtain a first semantic vector of a second entity type corresponding to a second entity in multiple sets of sample data, and select one of the multiple first semantic vectors as a first initial centroid;
[0118] Step 2: Calculate the first distances from the first semantic vectors other than the first initial centroid to the first initial centroid in the multiple first semantic vectors, determine the next initial centroid based on the first distances, and repeat the initial centroid determination process using the next initial centroid as the first initial centroid until multiple initial centroids are determined;
[0119] Step 3: Calculate the second distances between the first semantic vectors other than the initial centroids and each initial centroid in the first semantic vectors, and determine the multiple clusters according to the second distances.
[0120] Step 4: Update multiple initial centroids according to the first semantic vectors in each cluster to obtain multiple entity type cluster centers.
[0121] Similarly, the process of determining the cluster centers of the above multiple relationship types includes:
[0122] Step 1: Obtain second semantic vectors of the second relationship type corresponding to the second entity relationship in multiple groups of sample data, and select one of the multiple second semantic vectors as the first initial centroid;
[0123] Step 2: Calculate the first distances from the second semantic vectors other than the first initial centroid to the first initial centroid in the plurality of second semantic vectors, determine the next initial centroid based on the first distances, and repeat the initial centroid determination process using the next initial centroid as the first initial centroid until multiple initial centroids are determined;
[0124] Step 3: Calculate the second distances between the second semantic vectors other than the initial centroids in the plurality of second semantic vectors and each initial centroid, and determine a plurality of clusters according to the second distances;
[0125] Step 4: Update multiple initial centroids according to the second semantic vectors in each cluster to obtain multiple relationship type cluster centers.
[0126] After determining multiple entity type cluster centers and multiple relationship type cluster centers through the above steps, calculate the first distance between the first semantic vector and the multiple entity type cluster centers, and use the category corresponding to the entity type cluster center with the shortest first distance as the target entity type corresponding to the first entity type. Similarly, calculate the second distance between the second semantic vector and the multiple relationship type cluster centers, and use the category corresponding to the relationship type cluster center with the shortest second distance as the target relationship type corresponding to the first relationship type.
[0127] For example, Figure 5 is a schematic diagram of an optional type alignment according to an embodiment of the present application. As shown in Figure 5, universities, colleges, research institutes, higher education institutions, national parks, wildlife reserves, wetland reserves, and nature reserves are respectively encoded by semantic encoders to obtain corresponding semantic vectors, and then the clustering method is used to determine the cluster centers of universities, colleges, research institutes, and higher education institutions as research institutes, and the research institutes are used as the "representative elements" of these types to determine multiple synonymous types of target types. Similarly, the clustering method determines the cluster centers of national parks, wildlife reserves, wetland reserves, and nature reserves as nature reserves, and uses nature reserves as the "representative elements" of these types to determine multiple synonymous types of target types.
[0128] Step S205: construct a target knowledge graph based on the first entity, the first entity relationship, the target entity type corresponding to the first entity type, and the target relationship type corresponding to the first relationship type.
[0129] In the technical solution provided in step S205, a graph construction tool is used to construct a target knowledge graph based on the first entity, the first entity relationship, the target entity type corresponding to the first entity type, and the target relationship type corresponding to the first relationship type, thereby solving the problem that the current open domain knowledge graph only constructs the data layer but not the model layer.
[0130] Specifically, a data layer of the target knowledge graph is constructed based on the first entity and the first entity relationship, and a model layer of the target knowledge graph is constructed based on the target entity type corresponding to the first entity type and the target relationship type corresponding to the first relationship type.
[0131] For example, Figure 6 is a schematic diagram of an optional data layer and model layer according to an embodiment of the present application. As shown in Figure 6, the data layer includes entity relationships between various entities, such as "Actor A" → "Starring" → "Movie B", "Actor E" → "Participating" → "Movie F", and the model layer includes the relationship between the entity types of various entities, such as "Actor" → "Performance" → "Movie".
[0132] Based on the scheme defined by the above steps S201 to S205, it can be known that in an embodiment, a first natural language text to be analyzed is obtained; a pre-trained extraction model is used to perform entity extraction on the first natural language text to obtain a first entity and a first entity relationship in the first natural language, and the first entity type of the first entity and the first relationship type of the first entity relationship are determined; the first entity type and the first relationship type are respectively encoded using a pre-trained semantic encoder to obtain a first semantic vector corresponding to the first entity type and a second semantic vector corresponding to the first relationship type; a first distance between the first semantic vector and a preset plurality of entity type clustering centers and a second distance between the second semantic vector and a preset plurality of relationship type clustering centers are calculated, and a target entity type corresponding to the first entity type and a target relationship type corresponding to the first relationship type are determined based on the first distance and the second distance respectively; a target knowledge graph is constructed based on the first entity, the first entity relationship, the target entity type corresponding to the first entity type, and the target relationship type corresponding to the first relationship type.
[0133] In the above technical solution, the first entity type of the first entity and the first relationship type of the first entity relationship in the first natural language text are clustered and normalized to determine their corresponding target entity type and target relationship type, and the data layer and model layer of the open domain knowledge graph are constructed using the obtained first entity, first entity relationship, target entity type and target relationship type.
[0134] It can be seen that according to the technical solution of the embodiment of the present application, considering the good generation capability of the general large model, it is proposed to first expand the triples into text, and then use <text, triples> as supervision to train the extraction large model, which greatly improves the training samples and ensures the training accuracy of the extraction model; in addition, on the basis of using the extraction large model to extract entities and entity relationships to build a data layer, entity types and relationship types are extracted, and the extracted entity types and relationship types are normalized through semantic coding and clustering algorithms, and the model layer is built through the normalized entity types and relationship types, thereby solving the technical problems of poor integrity and low accuracy of the open domain knowledge graph constructed by related technologies.
[0135] In some embodiments, the knowledge graph construction method according to the present application may also include performing semantic search and / or content generation based on the constructed knowledge graph. For example, intelligent question and answer, automatic content generation, etc. are performed based on the constructed knowledge graph. According to the disclosure of the embodiments of the present application, semantic search and / or content generation can be made more accurate and complete, for example, more relevant results can be automatically returned, and the quality and reliability of the content can be improved. In this regard, the present application can also be understood as the disclosure of a semantic search method or a content generation method.
[0136] According to a further embodiment, the present application further provides a knowledge graph construction device, which, when running, executes the above-mentioned knowledge graph construction method of the above-mentioned embodiment. FIG7 is a schematic structural diagram of an optional knowledge graph construction device according to an embodiment of the present application. As shown in FIG7 , the knowledge graph construction device includes at least an acquisition module 71, an entity extraction module 72, an encoding module 73, a clustering module 74, and a construction module 75, wherein:
[0137] An acquisition module 71 is configured to acquire a first natural language text to be analyzed;
[0138] An entity extraction module 72 is configured to perform entity extraction on the first natural language text using a pre-trained extraction model to obtain a first entity and a first entity relationship in the first natural language, and determine a first entity type of the first entity and a first relationship type of the first entity relationship;
[0139] An encoding module 73 is configured to encode the first entity type and the first relationship type using a pre-trained semantic encoder to obtain a first semantic vector corresponding to the first entity type and a second semantic vector corresponding to the first relationship type;
[0140] a clustering module 74 for calculating a first distance between the first semantic vector and the centers of a plurality of preset entity type clusters, and a second distance between the second semantic vector and the centers of a plurality of preset relationship type clusters, and determining a target entity type corresponding to the first entity type and a target relationship type corresponding to the first relationship type based on the first distance and the second distance, respectively;
[0141] The construction module 75 is used to construct a target knowledge graph based on the first entity, the first entity relationship, the target entity type corresponding to the first entity type, and the target relationship type corresponding to the first relationship type.
[0142] It should be noted that the various modules in the above-mentioned knowledge graph construction device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.
[0143] The features described in the embodiment of the aforementioned knowledge graph construction method are applicable to the embodiment of the knowledge graph construction device. The preferred implementation methods of the embodiment of the knowledge graph construction device can be found in the relevant description in the embodiment of the aforementioned knowledge graph construction method, and will not be repeated here.
[0144] According to a further embodiment, the present application also provides a non-volatile storage medium, in which a program is stored, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the knowledge graph construction method in the above embodiment.
[0145] Optionally, the device where the non-volatile storage medium is located implements the following steps by running the program:
[0146] Step S201, obtaining a first natural language text to be analyzed;
[0147] Step S202: performing entity extraction on the first natural language text using a pre-trained extraction model to obtain a first entity and a first entity relationship in the first natural language, and determining a first entity type of the first entity and a first relationship type of the first entity relationship;
[0148] Step S203: Encode the first entity type and the first relationship type respectively using a pre-trained semantic encoder to obtain a first semantic vector corresponding to the first entity type and a second semantic vector corresponding to the first relationship type;
[0149] Step S204: calculating a first distance between the first semantic vector and the centers of a plurality of preset entity type clusters, and a second distance between the second semantic vector and the centers of a plurality of preset relationship type clusters, and determining a target entity type corresponding to the first entity type and a target relationship type corresponding to the first relationship type based on the first distance and the second distance, respectively;
[0150] Step S205: construct a target knowledge graph based on the first entity, the first entity relationship, the target entity type corresponding to the first entity type, and the target relationship type corresponding to the first relationship type.
[0151] According to an embodiment of the present application, a processor is also provided, which is used to run a program, wherein the knowledge graph construction method in Example 1 is executed when the program is running.
[0152] Optionally, the following steps are performed when the program is running:
[0153] Step S201, obtaining a first natural language text to be analyzed;
[0154] Step S202: performing entity extraction on the first natural language text using a pre-trained extraction model to obtain a first entity and a first entity relationship in the first natural language, and determining a first entity type of the first entity and a first relationship type of the first entity relationship;
[0155] Step S203: Encode the first entity type and the first relationship type respectively using a pre-trained semantic encoder to obtain a first semantic vector corresponding to the first entity type and a second semantic vector corresponding to the first relationship type;
[0156] Step S204: calculating a first distance between the first semantic vector and the centers of a plurality of preset entity type clusters, and a second distance between the second semantic vector and the centers of a plurality of preset relationship type clusters, and determining a target entity type corresponding to the first entity type and a target relationship type corresponding to the first relationship type based on the first distance and the second distance, respectively;
[0157] Step S205: construct a target knowledge graph based on the first entity, the first entity relationship, the target entity type corresponding to the first entity type, and the target relationship type corresponding to the first relationship type.
[0158] The features described in the aforementioned embodiment of the knowledge graph construction method are applicable to the embodiment of the non-volatile storage medium. The preferred implementation method of the embodiment of the non-volatile storage medium can be found in the relevant description in the aforementioned embodiment of the knowledge graph construction method, and will not be repeated here.
[0159] According to an embodiment of the present application, an electronic device is also provided, wherein the electronic device includes one or more processors; a memory for storing one or more programs, which enables the one or more processors to run the programs when the one or more programs are executed by the one or more processors, wherein the program is configured to execute the knowledge graph construction method in the above embodiment when running.
[0160] Optionally, the processor is configured to implement the following steps by executing a computer program:
[0161] Step S201, obtaining a first natural language text to be analyzed;
[0162] Step S202: performing entity extraction on the first natural language text using a pre-trained extraction model to obtain a first entity and a first entity relationship in the first natural language, and determining a first entity type of the first entity and a first relationship type of the first entity relationship;
[0163] Step S203: Encode the first entity type and the first relationship type respectively using a pre-trained semantic encoder to obtain a first semantic vector corresponding to the first entity type and a second semantic vector corresponding to the first relationship type;
[0164] Step S204: calculating a first distance between the first semantic vector and the centers of a plurality of preset entity type clusters, and a second distance between the second semantic vector and the centers of a plurality of preset relationship type clusters, and determining a target entity type corresponding to the first entity type and a target relationship type corresponding to the first relationship type based on the first distance and the second distance, respectively;
[0165] Step S205: construct a target knowledge graph based on the first entity, the first entity relationship, the target entity type corresponding to the first entity type, and the target relationship type corresponding to the first relationship type.
[0166] The features described in the embodiment of the aforementioned knowledge graph construction method are applicable to the embodiment of the electronic device. The preferred implementation methods of the embodiment of the electronic device can be found in the relevant description in the embodiment of the aforementioned knowledge graph construction method, which will not be repeated here.
[0167] According to an embodiment of the present application, a computer program product is further provided, including a computer program, which, when executed by a processor, implements the steps of the knowledge graph construction method according to the above embodiment:
[0168] Step S201, obtaining a first natural language text to be analyzed;
[0169] Step S202: performing entity extraction on the first natural language text using a pre-trained extraction model to obtain a first entity and a first entity relationship in the first natural language, and determining a first entity type of the first entity and a first relationship type of the first entity relationship;
[0170] Step S203: Encode the first entity type and the first relationship type respectively using a pre-trained semantic encoder to obtain a first semantic vector corresponding to the first entity type and a second semantic vector corresponding to the first relationship type;
[0171] Step S204: calculating a first distance between the first semantic vector and the centers of a plurality of preset entity type clusters, and a second distance between the second semantic vector and the centers of a plurality of preset relationship type clusters, and determining a target entity type corresponding to the first entity type and a target relationship type corresponding to the first relationship type based on the first distance and the second distance, respectively;
[0172] Step S205: construct a target knowledge graph based on the first entity, the first entity relationship, the target entity type corresponding to the first entity type, and the target relationship type corresponding to the first relationship type.
[0173] The features described in the embodiment of the aforementioned knowledge graph construction method are applicable to the embodiment of the computer program product. The preferred implementation methods of the embodiment of the computer program product can be found in the relevant description in the embodiment of the aforementioned knowledge graph construction method, and will not be repeated here.
[0174] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0175] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0176] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected to achieve the purpose of the present embodiment according to actual needs.
[0177] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0178] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0179] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
[0180] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0181] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A knowledge graph construction method, comprising: Obtaining a first natural language text to be analyzed; Performing entity extraction on the first natural language text using a pre-trained extraction model to obtain a first entity and a first entity relationship in the first natural language, and determining a first entity type of the first entity and a first relationship type of the first entity relationship; Encode the first entity type and the first relationship type respectively using a pre-trained semantic encoder to obtain a first semantic vector corresponding to the first entity type and a second semantic vector corresponding to the first relationship type; Calculate a first distance between the first semantic vector and a plurality of preset entity type cluster centers, and a second distance between the second semantic vector and a plurality of preset relationship type cluster centers, and determine a target entity type corresponding to the first entity type and a target relationship type corresponding to the first relationship type according to the first distance and the second distance, respectively; A target knowledge graph is constructed based on the first entity, the first entity relationship, the target entity type corresponding to the first entity type, and the target relationship type corresponding to the first relationship type.
2. The method according to claim 1, wherein: The training process of the extraction model includes: Build an initial language model; Acquire multiple sets of sample data, wherein the multiple sets of sample data include: a second natural language text, and a corresponding second entity and a second entity relationship; The initial language model is iteratively trained according to the multiple groups of sample data to obtain the extraction model.
3. The method according to claim 2, wherein: Get multiple sets of sample data, including: Acquire multiple groups of first-category triple samples, wherein the first-category triple samples include: the second entity and the second-entity relationship; Classify multiple groups of the first-category triplet samples according to entity names to obtain multiple classification groups, wherein each of the classification groups includes at least one group of the triplet samples; For each of the classification groups, at least two of the first-category triplet samples in the classification group are combined to obtain prompt information, and a second natural language text corresponding to the prompt information is generated using a universal large model.
4. The method according to claim 3, wherein: The training process of the semantic encoder includes: Build the initial deep learning network; Acquire multiple groups of second-category triple samples, wherein the second-category triple samples include: a benchmark sample, a positive sample with the same type of label as the benchmark sample, and a negative sample with a different type of label from the benchmark sample, and the benchmark sample, the positive sample, and the negative sample all include a second entity type corresponding to the second entity and a first semantic vector corresponding to the second entity type, or a second relationship type corresponding to the second entity relationship and a second semantic vector corresponding to the second relationship type; A triplet loss function is constructed based on the reference sample, the positive sample and the negative sample in each group of the second-category triplet samples, and the network parameters of the initial deep learning network are adjusted according to the triplet loss function to obtain the semantic encoder.
5. The method according to claim 4, wherein: Get multiple sets of second-category triplet samples, including: For each second entity in the first type triple sample, determine the second entity class corresponding to the second entity type, and encode the second entity type to obtain the corresponding first semantic vector, and form a first sample by the second entity type corresponding to the second entity and the first semantic vector corresponding to the second entity type; Marking the second entity type in each of the first samples to determine an entity type label of the second entity type in the first sample; Selecting a sample from the plurality of first samples as the reference sample, and determining a first entity type label for the reference sample; Acquire the positive sample from a plurality of samples corresponding to the first entity type label, and acquire the negative sample from a plurality of samples in a second entity type label different from the first entity type label; The second type of triplet samples is composed of the reference sample, the positive sample, and the negative sample.
6. The method according to claim 4, wherein: Get multiple sets of second-category triplet samples, including: For each second entity relationship in the first-type triple sample, determine a second relationship type corresponding to the second entity relationship, encode the second relationship type to obtain the corresponding second semantic vector, and form a second sample by the second relationship type corresponding to the second entity relationship and the second semantic vector corresponding to the second relationship type; Marking the second relationship type in each of the second samples to determine a relationship type label of the second relationship type in the second sample; Selecting a sample from the plurality of second samples as the reference sample, and determining a first relationship type label for the reference sample; Acquire the positive sample from a plurality of samples corresponding to the first relationship type label, and acquire the negative sample from a plurality of samples in a second relationship type label different from the first relationship type label; The second type of triplet samples is composed of the reference sample, the positive sample, and the negative sample.
7. The method according to claim 3, wherein: The process of determining the cluster centers of the plurality of entity types includes: Obtaining first semantic vectors of a second entity type corresponding to a second entity in a plurality of groups of the sample data, and selecting one of the plurality of first semantic vectors as a first initial centroid; Calculate a first distance from other first semantic vectors except the first initial centroid among the plurality of first semantic vectors to the first initial centroid, determine a next initial centroid according to the first distance, and repeat the initial centroid determination process using the next initial centroid as the first initial centroid until a plurality of the initial centroids are determined; Calculating second distances from the other first semantic vectors except the multiple initial centroids in the multiple first semantic vectors to each of the initial centroids, and determining multiple clusters according to the second distances; The multiple initial centroids are updated according to each of the first semantic vectors in each of the clusters to obtain a plurality of entity type cluster centers.
8. The method according to claim 3, wherein: The process of determining the multiple relationship type cluster centers includes: Obtaining second semantic vectors of a second relationship type corresponding to second entity relationships in a plurality of groups of the sample data, and selecting one of the plurality of second semantic vectors as a first initial centroid; Calculate a first distance from the other second semantic vectors except the first initial centroid among the plurality of the second semantic vectors to the first initial centroid, determine a next initial centroid according to the first distance, and repeat the initial centroid determination process using the next initial centroid as the first initial centroid until a plurality of the initial centroids are determined; Calculating second distances from the other second semantic vectors except the plurality of initial centroids in the plurality of second semantic vectors to each of the initial centroids, and determining a plurality of clusters according to the second distances; The multiple initial centroids are updated according to each of the second semantic vectors in each of the clusters to obtain multiple The relationship type cluster center.
9. The method according to claim 1, wherein: Building a target knowledge graph based on the first entity, the first entity relationship, a target entity type corresponding to the first entity type, and a target relationship type corresponding to the first relationship type, including: A data layer of the target knowledge graph is constructed based on the first entity and the first entity relationship, and a model layer of the target knowledge graph is constructed based on the target entity type corresponding to the first entity type and the target relationship type corresponding to the first relationship type.
10. The method according to claim 2, wherein: The initial language model is a general large model, which is trained to have the ability to generate natural language text based on triples.
11. The method according to claim 3, wherein: Using the general large model to generate a second natural language text corresponding to the prompt information includes: The prompt information is expanded using the general large model to generate a corresponding second natural language text; wherein the expanded content includes all the content of at least two of the first-category triple samples in the classification group.
12. The method according to claim 5 or 6, wherein: The second entity type tag is any one of the entity type tags other than the first entity type tag in the multiple entity type tags.
13. The method according to claim 1, further comprising: Perform semantic search and / or content generation based on the constructed knowledge graph.
14. A knowledge graph construction device, comprising: An acquisition module, used for acquiring a first natural language text to be analyzed; an entity extraction module, configured to perform entity extraction on the first natural language text using a pre-trained extraction model, obtain a first entity and a first entity relationship in the first natural language, and determine a first entity type of the first entity and a first relationship type of the first entity relationship; an encoding module, configured to encode the first entity type and the first relationship type respectively using a pre-trained semantic encoder to obtain a first semantic vector corresponding to the first entity type and a second semantic vector corresponding to the first relationship type; A clustering module, configured to calculate a first distance between the first semantic vector and a plurality of entity type cluster centers preset, and a second distance between the second semantic vector and a plurality of relationship type cluster centers preset, and to determine a target entity type corresponding to the first entity type and a target relationship type corresponding to the first relationship type according to the first distance and the second distance, respectively; A construction module is used to construct a target knowledge graph based on the first entity, the first entity relationship, the target entity type corresponding to the first entity type, and the target relationship type corresponding to the first relationship type.
15. A non-volatile storage medium, wherein: The non-volatile storage medium stores a computer program, wherein the device where the non-volatile storage medium is located executes the knowledge graph construction method described in any one of claims 1 to 13 by running the computer program.
16. An electronic device, comprising: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program, when running, executes the knowledge graph construction method described in any one of claims 1 to 13.
17. A computer program product comprising a computer program, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.
Citation Information
Patent Citations
Buddha knowledge graph construction method, device and apparatus and storage medium
CN113486187A
Method and device for constructing knowledge graph, electronic equipment and storage medium
CN114756690A
Knowledge graph construction method for intestine-brain axis and knowledge graph system
CN116226404A
Knowledge graph construction method and device, storage medium and electronic equipment
CN118093887A
Evaluating techniques for clustering geographic entities
US8676799B1
Cited By
Project decision optimization control method, device and equipment based on knowledge graph
CN120509686A
Large model application construction method based on configurable workflow and domain knowledge base
CN120631869A
Dynamic knowledge graph construction method for three-dimensional security situation analysis
CN120706524A
Construction method and system of power distribution network visualization platform based on SG-CIM
CN120873260A
Product and education collaborative resource intelligent matching system based on knowledge graph
CN120875001A