Knowledge Graph Construction Method, Device, and Electronic Device Based on Audit Information

By performing text extraction and entity relationship extraction of audited text information, a graph network is built and optimized, which solves the problems of low information storage efficiency and reduced semantic correlation, and realizes efficient storage and accurate knowledge graph construction.

CN115481260BActive Publication Date: 2025-06-27STATE GRID CORPORATION OF CHINA +4
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211227641.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-09
Publication Date
2025-06-27
Estimated Expiration
2042-10-09

AI Technical Summary

Technical Problem

In the information storage process, the relationship between the entities contained in the information is not refined, resulting in inefficient use of storage space, and the conversion of information formats will weaken the semantic correlation between contents and reduce the accuracy of the construction of knowledge graphs.

Method used

By obtaining the text information to be audited, text extraction is performed to generate a collection of text information, entity and entity relationship extraction is performed to generate entity-pair information groups, build an initial graph network, and perform graph network optimization to generate a target knowledge graph.

Benefits of technology

Improve the efficiency of the storage space and the accuracy of building knowledge graphs, and reduce invalid storage and retain semantic correlation by refining entities and relationships in audit text information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115481260B_ABST
    Figure CN115481260B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method, apparatus, and electronic device for constructing a knowledge graph based on audit information. A specific implementation of the method includes: obtaining text information to be audited; performing text extraction processing on the text information to be audited to generate a text information set, wherein the text information in the text information set includes a target number of adjacent text segments; performing entity and entity relationship extraction on each text information in the text information set to generate an entity pair information group, obtaining an entity pair information group set, wherein the entity pair information in the entity pair information group in the entity pair information group set includes: an entity information set and relationship information; constructing an initial graph network according to the entity information set and relationship information included in the entity pair information in the entity pair information group set; and performing graph network optimization on the initial graph network to generate a target knowledge graph. This implementation improves the utilization efficiency of storage space and the accuracy of knowledge graph construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technologies, and particularly to a method, an apparatus, and an electronic device for constructing a knowledge graph based on audit information. Background Art

[0002] With the development of computer-related technologies, more and more users have started to electronically process information to improve the storage efficiency of information and reduce the cost consumption of information storage. Currently, when storing information, it is usually necessary to directly or indirectly convert the format corresponding to the information (for example, convert the information into structured, semi-structured, and unstructured types, etc.).

[0003] However, when adopting the above method, the following technical problems often exist:

[0004] First, when there is a large amount of information to be stored, since the entities included in the information and the relationships between the entities are not refined, a relatively large amount of storage space is required to store the information, resulting in low storage space utilization efficiency;

[0005] Second, converting the format of the information often changes the positional relationship between the contents included in the information, thereby weakening the semantic relevance between the contents, and further resulting in a low construction accuracy of the knowledge graph. Summary of the Invention

[0006] This section of the present disclosure is used to briefly introduce concepts, which will be described in detail in the following detailed implementation section. This section of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0007] Some embodiments of the present disclosure propose a method, an apparatus, and an electronic device for constructing a knowledge graph based on audit information to solve one or more of the technical problems mentioned in the above background art section.

[0008] In a first aspect, some embodiments of the present disclosure provide a method for constructing a knowledge graph based on audit information. The method includes: obtaining the text information to be audited; performing text extraction processing on the text information to be audited to generate a text information set, where the text information in the text information set includes a target number of adjacent text segments; performing entity and entity relationship extraction on each text information in the text information set to generate an entity pair information group, obtaining an entity pair information group set, where the entity pair information in the entity pair information group in the entity pair information group set includes: an entity information set and a relationship information; constructing an initial graph network according to the entity information set and the relationship information included in the entity pair information in the entity pair information group set; and performing graph network optimization on the initial graph network to generate a target knowledge graph.

[0009] Optionally, the performing text extraction processing on the text information to be audited to generate a text information set includes: in response to determining that there is a target symbol in the text information to be audited, removing the target symbol from the text information to be audited to generate candidate text information to be audited, where the target symbol is a symbol in a pre-constructed identification symbol library; performing segmentation processing on the candidate text information to be audited to generate a text segment information sequence; and sequentially selecting the target number of adjacent text segments from the text segment information sequence at a fixed step length as text information to obtain the text information set.

[0010] Optionally, the above method further includes: obtaining a basic knowledge graph; for each first entity node in the above target knowledge graph, performing the following processing steps: determining the entity similarity between the above first entity node and each second entity node in the above basic knowledge graph to generate entity similarity information, obtaining an entity similarity information set; in response to determining that the above first entity node meets a first fusion condition, determining the first relationship edges connected to the above first entity node in the above target knowledge graph, obtaining a relationship edge information set, where the above first fusion condition is that there is entity similarity information with a corresponding similarity greater than a first preset similarity in the entity similarity information set; determining the relationship similarity between each relationship edge information in the above relationship edge information set and each second relationship edge in the above basic knowledge graph to generate relationship similarity information, obtaining a relationship similarity information set; in response to determining that there is at least one target relationship similarity information in the above relationship similarity information set, fusing the relationship edges corresponding to the target relationship similarity information in the above at least one target relationship similarity information and the two first entity nodes connected by the relationship edges into the above basic knowledge graph to generate a fused knowledge graph, where the target relationship similarity information in the above at least one target relationship similarity information is the relationship similarity information that meets a second fusion condition, and the above second fusion condition is that the relationship similarity corresponding to the relationship similarity information is less than a second preset similarity.

[0011] Optionally, the above extraction of entities and entity relationships from each text information in the above text information set to generate an entity pair information group includes: performing semantic extraction on the above text information through a pre-trained semantic extraction model to generate semantic feature information; inputting the above semantic feature information and the above text information into a pre-trained entity and entity relationship extraction model to generate the entity pair information group corresponding to the above text information.

[0012] Optionally, constructing an initial graph network based on the entity information set and relationship information included in the entity pair information in the above-mentioned information group of entities includes: performing the following graph network generation steps based on the above-mentioned information group of entity pairs and an empty graph network: selecting an entity pair information from the above-mentioned information group of entity pairs to generate target entity pair information and a candidate entity pair information group set, where the candidate entity pair information group set does not include the target entity pair information; in response to determining that the empty graph network does not include the two entity nodes and relationship edges corresponding to the target entity pair information, adding the entity nodes and relationship edges corresponding to the above-mentioned target entity pair information to the empty graph network to generate a candidate graph network; in response to determining that the empty graph network includes the two entity nodes and relationship edges corresponding to the target entity pair information, increasing the weight of the relationship edge corresponding to the target entity pair information in the empty graph network to generate a candidate graph network; in response to determining that the candidate entity pair information group set is empty, determining the candidate graph network as the above-mentioned initial graph network and ending the above-mentioned graph network generation steps; in response to determining that the candidate entity pair information group set is not empty, determining the candidate entity pair information group set as the information group of entity pairs and determining the candidate graph network as the empty graph network, and performing the above-mentioned graph network generation steps again.

[0013] Optionally, optimizing the above-mentioned initial graph network to generate a target knowledge graph includes: removing a target entity node pair from the above-mentioned initial graph network to generate the above-mentioned target knowledge graph, where the above-mentioned target entity node pair is two first entity nodes that are connected by a relationship edge in the above-mentioned initial graph network and meet the entity node removal condition.

[0014] Optionally, the above method further includes: in response to determining that there is a target entity node in the above-mentioned fused knowledge graph, adding the entity information corresponding to the above-mentioned target entity node to a pre-constructed entity knowledge base, where the above-mentioned target entity node is an entity node whose corresponding entity information does not exist in the above-mentioned entity knowledge base.

[0015] Second aspect, some embodiments of the present disclosure provide a knowledge graph construction device based on audit information. The device includes: an acquisition unit configured to acquire the text information to be audited; a text extraction and processing unit configured to perform text extraction and processing on the text information to be audited to generate a text information set, wherein the text information in the text information set includes a target number of adjacent text segments; an entity and entity relationship extraction unit configured to perform entity and entity relationship extraction on each text information in the text information set to generate an entity pair information group and obtain an entity pair information group set, wherein the entity pair information in the entity pair information group in the entity pair information group set includes: an entity information set and a relationship information; a construction unit configured to construct an initial graph network according to the entity information set and the relationship information included in the entity pair information in the entity pair information group set; a graph network optimization unit configured to optimize the initial graph network to generate a target knowledge graph.

[0016] Optionally, the text extraction and processing unit is further configured to: in response to determining that there is a target symbol in the text information to be audited, remove the target symbol from the text information to be audited to generate candidate text information to be audited, wherein the target symbol is an identification symbol in a pre-constructed identification symbol library; perform segmentation processing on the candidate text information to be audited to generate a text segment information sequence; sequentially select the target number of adjacent text segments from the text segment information sequence at a fixed step length as text information to obtain the text information set.

[0017] Optionally, the above device further includes: obtaining a basic knowledge graph; for each first entity node in the above target knowledge graph, performing the following processing steps: determining the entity similarity between the above first entity node and each second entity node in the above basic knowledge graph to generate entity similarity information, obtaining an entity similarity information set; in response to determining that the above first entity node meets the first fusion condition, determining the first relationship edges connected to the above first entity node in the above target knowledge graph, obtaining a relationship edge information set, where the above first fusion condition is that there is entity similarity information in the entity similarity information set with a corresponding similarity greater than a first preset similarity; determining the relationship similarity between each relationship edge information in the above relationship edge information set and each second relationship edge in the above basic knowledge graph to generate relationship similarity information, obtaining a relationship similarity information set; in response to determining that there is at least one target relationship similarity information in the above relationship similarity information set, fusing the relationship edges corresponding to the target relationship similarity information in the above at least one target relationship similarity information and the two first entity nodes connected by the relationship edges into the above basic knowledge graph to generate a fused knowledge graph, where the target relationship similarity information in the above at least one target relationship similarity information is the relationship similarity information that meets the second fusion condition, and the above second fusion condition is that the relationship similarity corresponding to the relationship similarity information is less than a second preset similarity.

[0018] Optionally, the above entity and entity relationship extraction unit is further configured to: perform semantic extraction on the above text information through a pre-trained semantic extraction model to generate semantic feature information; input the above semantic feature information and the above text information into a pre-trained entity and entity relationship extraction model to generate an entity pair information group corresponding to the above text information.

[0019] Optionally, the above building block is further configured to perform the following knowledge graph network generation steps based on the above entity pair information group set and the empty knowledge graph network: Select an entity pair information from the above entity pair information group set to generate target entity pair information and a candidate entity pair information group set, where the candidate entity pair information group set does not include the target entity pair information; In response to determining that the two entity nodes and the relationship edge corresponding to the target entity pair information are not included in the empty knowledge graph network, add the entity nodes and the relationship edge corresponding to the above target entity pair information to the empty knowledge graph network to generate a candidate knowledge graph network; In response to determining that the two entity nodes and the relationship edge corresponding to the target entity pair information are included in the empty knowledge graph network, increase the weight of the relationship edge corresponding to the target entity pair information in the empty knowledge graph network to generate a candidate knowledge graph network; In response to determining that the candidate entity pair information group set is empty, determine the candidate knowledge graph network as the above initial knowledge graph network, and end the above knowledge graph network generation steps; In response to determining that the candidate entity pair information group set is not empty, determine the candidate entity pair information group set as the entity pair information group set, and determine the candidate knowledge graph network as the empty knowledge graph network, and perform the above knowledge graph network generation steps again. In actual situations, since the knowledge graph network often contains a large number of entity nodes and relationship edges. And the connection method of entity nodes and relationship edges will affect the network structure of the finally generated knowledge graph network. Therefore, by determining whether the two entity nodes and the relationship edge corresponding to the target entity pair information are included in the empty knowledge graph network, it is determined whether to add the two entity nodes and the relationship edge corresponding to the target entity pair information to the empty knowledge graph network. In addition, when the two entity nodes and the relationship edge corresponding to the target entity pair information are included in the empty knowledge graph network, the importance of the two entity nodes and the relationship edge corresponding to the target entity pair information is highlighted by increasing the weight of the relationship edge.

[0020] Optionally, the above knowledge graph network optimization unit is further configured to: Remove the target entity node pair from the above initial knowledge graph network to generate the above target knowledge graph, where the above target entity node pair is two first entity nodes that are connected by a relationship edge in the above initial knowledge graph network and meet the entity node removal condition.

[0021] Optionally, the above device further includes: In response to determining that there is a target entity node in the above fused knowledge graph, add the entity information corresponding to the above target entity node to a pre-constructed entity knowledge base, where the above target entity node is an entity node whose corresponding entity information does not exist in the above entity knowledge base.

[0022] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device, on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect above.

[0023] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the first aspect above is implemented.

[0024] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: The method for constructing a knowledge graph based on audit information in some embodiments of the present disclosure improves the storage space utilization efficiency and the construction accuracy of the knowledge graph. Specifically, the reasons for the low storage space utilization efficiency and the low construction accuracy of the knowledge graph are as follows: First, when there is a large amount of stored information, since the entities and the relationships between the entities included in the information are not refined, a large amount of storage space is required to store the information. Second, when performing format conversion on the information, the positional relationship between the contents included in the information is often changed, thereby weakening the semantic relevance between the contents, and further resulting in a low construction accuracy of the knowledge graph. Based on this, the method for constructing a knowledge graph based on audit information in some embodiments of the present disclosure first obtains the text information to be audited. Then, the above-mentioned text information to be audited is subjected to text extraction processing to generate a text information set, wherein the text information in the above-mentioned text information set includes a target number of adjacent text segments. The present disclosure does not perform format conversion on the audit text information, so that the semantic relevance between the text segments is retained. It is convenient to identify entities and extract entity relationships from adjacent text segments subsequently. At the same time, by constructing a text information from a target number of adjacent text segments, it is ensured that not too many useless features are learned during the entity recognition stage. The construction accuracy of the knowledge graph is greatly improved. Next, entity and entity relationship extraction is performed on each text information in the above-mentioned text information set to generate an entity pair information group, and an entity pair information group set is obtained, wherein the entity pair information in the entity pair information group in the above-mentioned entity pair information group set includes: an entity information set and a relationship information. Through entity and entity relationship extraction, a plurality of entities included in the text information and the relationships between the entities are determined. In addition, an initial graph network is constructed according to the entity information set and the relationship information included in the entity pair information in the above-mentioned entity pair information group set. By constructing the initial graph network, the entities in the text information to be audited and the relationships between the entities are stored in the form of a graph. Finally, the above-mentioned initial graph network is optimized to generate a target knowledge graph. In actual situations, there are often some useless entities and corresponding relationships among the extracted entities. Therefore, through graph network optimization, such entities and corresponding relationships are removed. In this way, the refinement of the entities and the relationships between the entities included in the audit text information is realized, and at the same time, the entities in the text information to be audited and the relationships between the entities are stored in the form of a graph. The storage space utilization efficiency is greatly improved. Description of the Drawings

[0025] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.

[0026] Figure 1 is a schematic diagram of an application scenario of a method for constructing a knowledge graph based on audit information according to some embodiments of the present disclosure;

[0027] Figure 2 is a flowchart according to some embodiments of the method for constructing a knowledge graph based on audit information according to the present disclosure;

[0028] Figure 3 is a flowchart according to some other embodiments of the method for constructing a knowledge graph based on audit information according to the present disclosure;

[0029] Figure 4 is a schematic structural diagram according to some embodiments of the apparatus for constructing a knowledge graph based on audit information according to the present disclosure;

[0030] Figure 5 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Specific Embodiments

[0031] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0032] In addition, it should be noted that for the sake of convenience of description, only the parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0033] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or mutual dependence relationship of the functions performed by these devices, modules or units.

[0034] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0035] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are for illustrative purposes only and are not used to limit the scope of these messages or information.

[0036] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0037] Figure 1 It is a schematic diagram of an application scenario of a method for constructing a knowledge graph based on audit information according to some embodiments of the present disclosure.

[0038] In Figure 1 the application scenario, first, the computing device 101 can obtain the text information to be audited 102; then, the computing device 101 can perform text extraction processing on the above text information to be audited 102 to generate a text information set 103, where the text information in the above text information set 103 includes a target number of adjacent text segments; then, the computing device 101 can perform entity and entity relationship extraction on each text information in the above text information set 103 to generate an entity pair information group, obtaining an entity pair information group set 104, where the entity pair information in the entity pair information group in the above entity pair information group set 104 includes: an entity information set and a relationship information; further, the computing device 101 can construct an initial graph network 105 according to the entity information set and the relationship information included in the entity pair information in the above entity pair information group set 104; finally, the computing device 101 can optimize the above initial graph network 105 to generate a target knowledge graph 106.

[0039] It should be noted that the above computing device 101 can be hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster composed of multiple servers or terminal devices, or can be implemented as a single server or a single terminal device. When the computing device is embodied as software, it can be installed in the above-listed hardware devices. It can be implemented as, for example, multiple software or software modules for providing distributed services, or can be implemented as a single software or software module. No specific limitation is made here.

[0040] It should be understood that Figure 1 the number of computing devices in

[0041] Continuing to refer to Figure 2 , a process 200 according to some embodiments of the method for constructing a knowledge graph based on audit information of the present disclosure is shown. The method for constructing a knowledge graph based on audit information includes the following steps:

[0042] Step 201, obtain the text information to be audited.

[0043] In some embodiments, the execution subject of the method for constructing a knowledge graph based on audit information (such as Figure 1 the computing device 101 shown) can obtain the above-mentioned audit text information through wired connection or wireless connection. Among them, the above-mentioned audit text information can be text information related to auditing. The text format of the above-mentioned audit text information can be, but is not limited to, any one of the following: .txt text format and.pdf text format.

[0044] Step 202: Perform text extraction processing on the text information to be audited to generate a text information set.

[0045] In some embodiments, the execution subject can perform text extraction processing on the text information to be audited to generate a text information set. Among them, the text information in the above-mentioned text information set includes a target number of adjacent text segments. A text segment can represent a passage in the above-mentioned text information to be audited. For example, the above-mentioned target number can be 3.

[0046] As an example, the above-mentioned text information to be audited can be "Statement A. Statement B. Statement C. Statement D. Statement E. Statement F.", and the obtained text information set can be {[Statement A Statement B Statement C], [Statement B Statement C Statement D], [Statement C Statement D Statement E], [Statement D Statement E Statement F]}.

[0047] Step 203: Extract entities and entity relationships from each text information in the text information set to generate an entity pair information group, and obtain an entity pair information group set.

[0048] In some embodiments, the above-mentioned execution entity may extract entities and entity relationships from each text information in the above-mentioned text information set to generate an entity pair information group and obtain an entity pair information group set. Among them, the entity pair information in the entity pair information group in the above-mentioned entity pair information group set includes: an entity information set and relationship information; among them, the entity information in the entity information set represents the entities included in the text information. The relationship information represents the relationship between the entities corresponding to the entity information in the entity information set. For example, the text information may be "Zhang San needs to audit the XX document". The obtained entity information set may be {[Zhang San], [XX document]}. The relationship information may be {audit}. Among them, the entity pair information may also be formed in the form of a triple, where the triple may include the entity information set and relationship information included in the entity pair information. For example, (entity 1: "Zhang San", relationship: "audit", entity 2: "XX document"). Among them, the above-mentioned execution entity may extract the entities included in the text information through an entity extraction model to generate the entity information set included in the entity pair information. The above-mentioned entity extraction model may be an RNN (Recurrent Neural Network) model. The above-mentioned execution entity may extract the relationship between the entities included in the text information through an entity relationship extraction model to generate the relationship information included in the entity pair information. The above-mentioned entity relationship extraction model may be, but is not limited to, any one of the following: TextCNN model and BERT model.

[0049] Step 204, construct an initial graph network according to the entity information set and relationship information included in the entity pair information in the entity pair information group set.

[0050] In some embodiments, the above-mentioned execution entity may construct an initial graph network according to the entity information set and relationship information included in the entity pair information in the above-mentioned entity pair information group set. Among them, the network structure of the above-mentioned initial graph network is a graph structure. The above-mentioned initial graph network includes first entity nodes and relationship edges. The first entity nodes are used to represent entities. The relationship edges are used to represent the relationship between entities. The above-mentioned execution entity may determine the entity information in the entity information set included in the entity pair information as vertices and the relationship information included in the entity pair information as relationship edges to generate the above-mentioned initial graph network.

[0051] Step 205, optimize the graph network of the initial graph network to generate a target knowledge graph.

[0052] In some embodiments, the above-mentioned execution entity may optimize the graph network of the above-mentioned initial graph network to generate a target knowledge graph.

[0053] As an example, the above-mentioned execution entity optimizes the above-mentioned initial graph network to generate a target knowledge graph, which may include the following steps:

[0054] First, determine candidate entity nodes to obtain a set of candidate entity nodes.

[0055] Among them, the candidate entity nodes in the set of candidate entity nodes are the first entity nodes not related to auditing included in the initial graph network.

[0056] Second, in response to determining that there are candidate entity node pairs in the above-mentioned set of candidate entity nodes, remove the above-mentioned candidate entity node pairs from the above-mentioned initial graph network to generate the above-mentioned target knowledge graph.

[0057] Among them, the above-mentioned candidate entity node pairs are two first entity nodes connected by a relationship edge.

[0058] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: The method for constructing a knowledge graph based on audit information according to some embodiments of the present disclosure improves the storage space utilization efficiency and the construction accuracy of the knowledge graph. Specifically, the reasons for the low storage space utilization efficiency and the low construction accuracy of the knowledge graph are as follows: First, when there is a large amount of stored information, since the entities and the relationships between entities included in the information are not refined, a relatively large amount of storage space is required to store the information. Second, when performing format conversion on the information, the positional relationship between the contents included in the information is often changed, thereby weakening the semantic relevance between the contents, and further resulting in a low construction accuracy of the knowledge graph. Based on this, the method for constructing a knowledge graph based on audit information according to some embodiments of the present disclosure first obtains the text information to be audited. Then, perform text extraction processing on the above-mentioned text information to be audited to generate a text information set, where the text information in the above-mentioned text information set includes a target number of adjacent text segments. The present disclosure does not perform format conversion on the audit text information, so as to retain the semantic relevance between text segments. It is convenient to identify entities and extract entity relationships from adjacent text segments subsequently. At the same time, by constructing a text information from a target number of adjacent text segments, it is ensured that not too many useless features are learned during the entity recognition stage. Greatly improving the construction accuracy of the knowledge graph. Next, perform entity and entity relationship extraction on each text information in the above-mentioned text information set to generate an entity pair information group, and obtain an entity pair information group set, where the entity pair information in the entity pair information group in the above-mentioned entity pair information group set includes: an entity information set and a relationship information. Through entity and entity relationship extraction, to determine multiple entities included in the text information and the relationships between the entities. In addition, according to the entity information set and the relationship information included in the entity pair information in the above-mentioned entity pair information group set, construct an initial graph network. By constructing the initial graph network, store the entities in the text information to be audited and the relationships between the entities in the form of a graph. Finally, perform graph network optimization on the above-mentioned initial graph network to generate a target knowledge graph. In actual situations, there are often some useless entities and corresponding relationships among the extracted entities. Therefore, through graph network optimization, such entities and corresponding relationships are removed. In this way, the refinement of the entities and the relationships between the entities included in the audit text information is realized, and at the same time, the entities in the text information to be audited and the relationships between the entities are stored in the form of a graph. Greatly improving the storage space utilization efficiency.

[0059] For further reference Figure 3 , which shows the process 300 of another embodiment of the method for constructing a knowledge graph based on audit information. The process 300 of the method for constructing a knowledge graph based on audit information includes the following steps:

[0060] Step 301, obtain the text information to be audited.

[0061] In some embodiments, for the specific implementation of step 301 and the technical effects brought thereby, reference may be made to Figure 2 step 201 in the corresponding embodiment, which will not be elaborated herein.

[0062] Step 302, in response to determining that there is a target symbol in the text information to be audited, remove the target symbol from the text information to be audited to generate candidate text information to be audited.

[0063] In some embodiments, the execution subject (such as Figure 1 the computing device 101 shown) of the knowledge graph construction method based on audit information may, in response to determining that there is a target symbol in the text information to be audited, remove the target symbol from the text information to be audited to generate candidate text information to be audited. Wherein, the above-mentioned target symbol is a symbol in a pre-constructed identification symbol library. The above-mentioned identification symbol library may be a library for storing symbols to be removed. The symbols to be removed may be symbols other than commas and ellipses. The above-mentioned candidate text information to be audited is information that does not contain the symbols in the above-mentioned identification symbol library. For example, the above-mentioned target symbol may be a quotation mark, the above-mentioned target symbol may also be a comma, and the above-mentioned target symbol may also be a dash.

[0064] Step 303, perform a segmentation process on the candidate text information to be audited to generate a text segment information sequence.

[0065] In some embodiments, the above-mentioned execution subject may perform a segmentation process on the candidate text information to be audited to generate a text segment information sequence. Wherein, the text segment information in the above-mentioned text segment information sequence may represent a sentence in the above-mentioned candidate text information to be audited.

[0066] As an example, the above-mentioned candidate text information to be audited may be "Audit the XX file. Among them, the XX file contains problem A, problem B... The audit conclusion is that the audit passes after modifying the problems. Time: 20200212". The obtained text segment information sequence may be ["Audit the XX file", "Among them, the XX file contains problem A, problem B", "The audit conclusion is that the audit passes after modifying the problems", "Time: 20200212"].

[0067] Step 304, with a fixed step size, sequentially select a target number of adjacent text segments from the text segment information sequence as text information to obtain a text information set.

[0068] In some embodiments, the above-mentioned execution subject may, with a fixed step size, sequentially select a target number of adjacent text segments from the text segment information sequence as text information to obtain a text information set. Wherein, the above-mentioned fixed step size may be 1. The above-mentioned target number may be 3.

[0069] As an example, the above text segment information sequence can be ["Audit XX file", "where XX file contains problem A and problem B", "The audit conclusion is that the audit passes after modifying the problems", "Time: 20200212"]. The above text information set can be ["Audit XX file where XX file contains problem A and problem B and the audit conclusion is that the audit passes after modifying the problems", "Where XX file contains problem A and problem B and the audit conclusion is that the audit passes after modifying the problems and time: 20200212"].

[0070] Step 305: Extract entities and entity relationships from each text information in the text information set to generate an entity pair information group, and obtain a set of entity pair information groups.

[0071] In some embodiments, the above execution subject extracts entities and entity relationships from each text information in the text information set to generate an entity pair information group, and the steps that the obtained set of entity pair information groups may include are as follows:

[0072] The first step: Perform semantic extraction on the above text information through a pre-trained semantic extraction model to generate semantic feature information.

[0073] Among them, the above semantic feature extraction model can be but is not limited to any one of the following: CNN (Convolutional Neural Networks) model and RNN model. The above semantic feature information can represent the semantic features extracted from the above text information, and can be represented by a 1*N feature vector. Where N is the length of the feature vector.

[0074] The second step: Input the above semantic feature information and the above text information into a pre-trained entity and entity relationship extraction model to generate an entity pair information group corresponding to the above text information.

[0075] Among them, the above pre-trained entity and entity relationship extraction model may include: a word embedding layer, an encoding layer, an entity recognition layer, and a relationship classification layer. Among them, the above encoding layer can be a Bi-LSTM (Bi-directional Long Short-Term Memory) model. The above entity recognition layer can be an LSTM model. The above relationship classification layer can be a CNN model. Compared with the traditional linear processing mode, that is, first performing the entity recognition task and then the relationship extraction task. The above entity and entity relationship extraction model realizes model parameter sharing through the underlying encoding layer, that is, both the entity recognition task and the relationship extraction task will update the parameters of the encoding layer through the backpropagation algorithm, so as to realize the dependence between the two subtasks.

[0076] Step 306: Construct an initial graph network according to the entity information set and relationship information included in the entity pair information in the entity pair information set.

[0077] In some embodiments, the above-mentioned execution subject may construct an initial graph network according to the entity information set and relationship information included in the entity pair information in the entity pair information set. Among them, the above-mentioned execution subject may perform the following graph network generation steps based on the above-mentioned entity pair information set and an empty graph network:

[0078] First step: Select an entity pair information from the above-mentioned entity pair information set to generate target entity pair information and a candidate entity pair information set.

[0079] Among them, the candidate entity pair information set does not include the target entity pair information. The above-mentioned target entity pair information is to select an entity pair information from the above-mentioned entity pair information set.

[0080] Second step: In response to determining that the two entity nodes and relationship edge corresponding to the target entity pair information are not included in the empty graph network, add the entity nodes and relationship edge corresponding to the above-mentioned target entity pair information to the empty graph network to generate a candidate graph network.

[0081] Third step: In response to determining that the two entity nodes and relationship edge corresponding to the target entity pair information are included in the empty graph network, increase the weight of the relationship edge corresponding to the target entity pair information in the empty graph network to generate a candidate graph network.

[0082] Among them, the above-mentioned execution subject may add 1 to the weight of the relationship edge corresponding to the target entity pair information in the empty graph network to achieve weight increase.

[0083] Fourth step: In response to determining that the candidate entity pair information set is empty, determine the candidate graph network as the above-mentioned initial graph network and end the above-mentioned graph network generation steps.

[0084] Fifth step: In response to determining that the candidate entity pair information set is not empty, determine the candidate entity pair information set as the entity pair information set, and determine the candidate graph network as the empty graph network, and execute the above-mentioned graph network generation steps again.

[0085] In actual situations, since the graph network often contains a large number of entity nodes and relationship edges, and the connection methods of entity nodes and relationship edges will affect the network structure of the finally generated graph network. Therefore, by determining whether the empty graph network contains two entity nodes and relationship edges corresponding to the target entity pair information, it is determined whether to add the two entity nodes and relationship edges corresponding to the target entity pair information to the empty graph network. In addition, when the empty graph network contains two entity nodes and relationship edges corresponding to the target entity pair information, the importance of the two entity nodes and relationship edges corresponding to the target entity pair information is highlighted by increasing the weight of the relationship edges.

[0086] Step 307, optimize the initial graph network to generate the target knowledge graph.

[0087] In some embodiments, the above-mentioned execution subject can optimize the initial graph network to generate the target knowledge graph. Among them, the above-mentioned execution subject can remove the target entity node pairs from the above-mentioned initial graph network to generate the above-mentioned target knowledge graph. The above-mentioned target entity node pairs are two first entity nodes that are connected by a relationship edge in the above-mentioned initial graph network and meet the entity node removal conditions. The above-mentioned entity node removal conditions can be that both of the two first entity nodes connected by a relationship edge are entity nodes that are not relevant to the audit.

[0088] Optionally, the above-mentioned execution subject can also perform the following processing steps:

[0089] The first step is to obtain the basic knowledge graph.

[0090] Among them, the above-mentioned basic knowledge graph can be a knowledge graph generated according to at least one other audit text.

[0091] The second step is to perform the following processing steps for each first entity node in the above-mentioned target knowledge graph:

[0092] The first sub-step is to determine the entity similarity between the above-mentioned first entity node and each second entity node in the above-mentioned basic knowledge graph to generate entity similarity information, and obtain an entity similarity information set.

[0093] Among them, the above-mentioned execution subject can generate the corresponding entity similarity by determining the cosine similarity between the above-mentioned first entity node and the second entity node to obtain the entity similarity information. The first entity node is an entity node in the above-mentioned target knowledge graph. The second entity node is an entity node in the basic knowledge graph.

[0094] The second sub-step is to, in response to determining that the above-mentioned first entity node meets the first fusion condition, determine the first relationship edge connected to the above-mentioned first entity node in the above-mentioned target knowledge graph to obtain a relationship edge information set.

[0095] Among them, the above first fusion condition is that there is entity similarity information in the entity similarity information set whose corresponding similarity is greater than the first preset similarity.

[0096] The third sub-step is to determine the relationship similarity between each relationship edge information in the above relationship edge information set and each second relationship edge in the above basic knowledge graph to generate relationship similarity information, and obtain a relationship similarity information set.

[0097] Among them, the above execution subject can generate corresponding relationship similarity information by determining the cosine similarity between the above relationship edge information and the second relationship edge.

[0098] The fourth sub-step is to, in response to determining that there is at least one target relationship similarity information in the above relationship similarity information set, fuse the relationship edge corresponding to the target relationship similarity information in the above at least one target relationship similarity information and the two first entity nodes connected by the relationship edge into the above basic knowledge graph to generate a fused knowledge graph.

[0099] Among them, the target relationship similarity information in the above at least one target relationship similarity information is the relationship similarity information that meets the second fusion condition. The above second fusion condition is that the relationship similarity corresponding to the relationship similarity information is less than the second preset similarity. The above execution subject can add the relationship edge corresponding to the target relationship similarity information in the above at least one target relationship similarity information and the two first entity nodes connected by the relationship edge to the above basic knowledge graph to generate a fused knowledge graph.

[0100] The fifth sub-step is to, in response to determining that there is a target entity node in the above fused knowledge graph, add the entity information corresponding to the above target entity node to the pre-constructed entity knowledge base.

[0101] Among them, the above target entity node is an entity node whose corresponding entity information does not exist in the above entity knowledge base. The above entity knowledge base can be a library for storing entity information corresponding to entities.

[0102] From Figure 3 It can be seen that compared with Figure 2Compared with the descriptions of some corresponding embodiments, the present disclosure first performs entity recognition and relationship extraction on text information through an entity and entity relationship extraction model to generate entity pair information in the entity pair information group corresponding to the text information, including: an entity information set and relationship information. Compared with the traditional linear processing mode, that is, first performing an entity recognition task and then a relationship extraction task. The above entity and entity relationship extraction model realizes model parameter sharing through the underlying encoding layer, that is, both the entity recognition task and the relationship extraction task will update the parameters of the encoding layer through the backpropagation algorithm, thereby realizing the dependence between the two subtasks. In addition, considering that inputting the entire text into the network model can enable the model to learn the semantic information before and after, but for the entity recognition and relationship extraction tasks, the obtained results are only the relationship between two entities and two entities directly. Inputting the entire text into the network model will increase the feature redundancy and lead to an increase in the error rate. Therefore, the present disclosure uses text information containing a target number of adjacent text segments as the input of the entity and entity relationship extraction model, which can not only consider the semantic information before and after, but also will not bring too much redundant feature information, greatly improving the recognition accuracy of the model. Secondly, generating a corresponding knowledge graph only for a single audit text information often has independence, that is, it cannot be reused for other audit text information. Therefore, the present disclosure adds a step of fusing the target knowledge graph into the basic knowledge graph, greatly improving the reusability of the obtained fused knowledge graph.

[0103] Further referring to Figure 4 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a knowledge graph construction device based on audit information. These device embodiments correspond to Figure 2 the method embodiments shown, and the device can be specifically applied to various electronic devices.

[0104] As Figure 4As shown, the knowledge graph construction device 400 based on audit information in some embodiments includes: an acquisition unit 401, a text extraction and processing unit 402, an entity and entity relationship extraction unit 403, a construction unit 404, and a graph network optimization unit 405. Among them, the acquisition unit 401 is configured to acquire the text information to be audited; the text extraction and processing unit 402 is configured to perform text extraction and processing on the above-mentioned text information to be audited to generate a text information set, wherein the text information in the above-mentioned text information set includes a target number of adjacent text segments; the entity and entity relationship extraction unit 403 is configured to perform entity and entity relationship extraction on each text information in the above-mentioned text information set to generate an entity pair information group, and obtain an entity pair information group set, wherein the entity pair information in the entity pair information group in the above-mentioned entity pair information group set includes: an entity information set and a relationship information; the construction unit 404 is configured to construct an initial graph network according to the entity information set and the relationship information included in the entity pair information in the above-mentioned entity pair information group set; the graph network optimization unit 404 is configured to perform graph network optimization on the above-mentioned initial graph network to generate a target knowledge graph.

[0105] In some alternative implementation manners of some embodiments, the above-mentioned text extraction and processing unit 402 is further configured to: in response to determining that there is a target symbol in the above-mentioned text information to be audited, remove the above-mentioned target symbol from the above-mentioned text information to be audited to generate candidate text information to be audited, wherein the above-mentioned target symbol is an identification symbol in a pre-constructed identification symbol library; perform segmentation processing on the above-mentioned candidate text information to be audited to generate a text segment information sequence; sequentially select the above-mentioned target number of adjacent text segments from the above-mentioned text segment information sequence at a fixed step length as text information to obtain the above-mentioned text information set.

[0106] In some alternative implementations of some embodiments, the above-mentioned apparatus 400 further includes: obtaining a basic knowledge graph; for each first entity node in the above-mentioned target knowledge graph, performing the following processing steps: determining the entity similarity between the above-mentioned first entity node and each second entity node in the above-mentioned basic knowledge graph to generate entity similarity information, and obtaining an entity similarity information set; in response to determining that the above-mentioned first entity node meets the first fusion condition, determining the first relationship edges connected to the above-mentioned first entity node in the above-mentioned target knowledge graph to obtain a relationship edge information set, where the above-mentioned first fusion condition is that there is entity similarity information with a corresponding similarity greater than a first preset similarity in the entity similarity information set; determining the relationship similarity between each relationship edge information in the above-mentioned relationship edge information set and each second relationship edge in the above-mentioned basic knowledge graph to generate relationship similarity information, and obtaining a relationship similarity information set; in response to determining that there is at least one target relationship similarity information in the above-mentioned relationship similarity information set, fusing the relationship edges corresponding to the target relationship similarity information in the above-mentioned at least one target relationship similarity information and the two first entity nodes connected by the relationship edges into the above-mentioned basic knowledge graph to generate a fused knowledge graph, where the target relationship similarity information in the above-mentioned at least one target relationship similarity information is the relationship similarity information that meets the second fusion condition, and the above-mentioned second fusion condition is that the relationship similarity corresponding to the relationship similarity information is less than a second preset similarity.

[0107] In some alternative implementations of some embodiments, the above-mentioned entity and entity relationship extraction unit 403 is further configured to: perform semantic extraction on the above-mentioned text information through a pre-trained semantic extraction model to generate semantic feature information; input the above-mentioned semantic feature information and the above-mentioned text information into a pre-trained entity and entity relationship extraction model to generate an entity pair information group corresponding to the above-mentioned text information.

[0108] In some alternative implementations of some embodiments, the above-mentioned construction unit 404 is further configured to perform the following graph network generation steps based on the entity pair information group set and the empty graph network: Select an entity pair information from the entity pair information group set to generate target entity pair information and a candidate entity pair information group set, where the candidate entity pair information group set does not include the target entity pair information; In response to determining that the two entity nodes and the relationship edge corresponding to the target entity pair information are not included in the empty graph network, add the entity nodes and the relationship edge corresponding to the target entity pair information to the empty graph network to generate a candidate graph network; In response to determining that the two entity nodes and the relationship edge corresponding to the target entity pair information are included in the empty graph network, increase the weight of the relationship edge corresponding to the target entity pair information in the empty graph network to generate a candidate graph network; In response to determining that the candidate entity pair information group set is empty, determine the candidate graph network as the above-mentioned initial graph network, and end the above-mentioned graph network generation steps; In response to determining that the candidate entity pair information group set is not empty, determine the candidate entity pair information group set as the entity pair information group set, and determine the candidate graph network as the empty graph network, and perform the above-mentioned graph network generation steps again. In actual situations, since the graph network often contains a large number of entity nodes and relationship edges. And the connection methods of the entity nodes and the relationship edges will affect the network structure of the finally generated graph network. Therefore, by determining whether the two entity nodes and the relationship edge corresponding to the target entity pair information are included in the empty graph network, it is determined whether to add the two entity nodes and the relationship edge corresponding to the target entity pair information to the empty graph network. In addition, when the two entity nodes and the relationship edge corresponding to the target entity pair information are included in the empty graph network, the importance of the two entity nodes and the relationship edge corresponding to the target entity pair information is highlighted by increasing the weight of the relationship edge.

[0109] In some alternative implementations of some embodiments, the above-mentioned graph network optimization unit 405 is further configured to remove the target entity node pair from the above-mentioned initial graph network to generate the above-mentioned target knowledge graph, where the target entity node pair is two first entity nodes that are connected by a relationship edge in the above-mentioned initial graph network and meet the entity node removal condition.

[0110] In some alternative implementations of some embodiments, the above-mentioned apparatus 400 further includes: In response to determining that there is a target entity node in the above-mentioned fused knowledge graph, add the entity information corresponding to the target entity node to a pre-constructed entity knowledge base, where the target entity node is an entity node whose corresponding entity information does not exist in the above-mentioned entity knowledge base.

[0111] It can be understood that the various units described in the apparatus 400 are related to the reference Figure 2corresponds to each step in the described method. Thus, the operations, features, and beneficial effects described above for the method also apply to the apparatus 400 and the units included therein, and will not be elaborated herein.

[0112] Reference is now made to Figure 5 , which shows a schematic structural diagram of an electronic device (such as the computing device 101 shown in Figure 1 ) 500 suitable for use in implementing some embodiments of the present disclosure. Figure 5 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0113] As shown in Figure 5 , the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0114] Generally, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 5 shows the electronic device 500 having various devices, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included. Figure 5 Each block shown in

[0115] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such some embodiments, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the methods of some embodiments of the present disclosure are performed.

[0116] It should be noted that the computer-readable medium described in some embodiments of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0117] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0118] The above computer-readable medium can be included in the above electronic device; or can exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: obtain the text information to be audited; perform text extraction processing on the above text information to be audited to generate a text information set, where the text information in the above text information set includes a target number of adjacent text segments; perform entity and entity relationship extraction on each text information in the above text information set to generate an entity pair information group, obtaining an entity pair information group set, where the entity pair information in the entity pair information group in the above entity pair information group set includes: an entity information set and a relationship information; construct an initial graph network according to the entity information set and the relationship information included in the entity pair information in the above entity pair information group set; perform graph network optimization on the above initial graph network to generate a target knowledge graph.

[0119] Computer program code for performing the operations of some embodiments of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++; and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0120] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0121] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an acquisition unit, a text extraction processing unit, an entity and entity relationship extraction unit, a construction unit, and a graph network optimization unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the acquisition unit can also be described as "the unit for acquiring text information to be audited".

[0122] The functions described above can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0123] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A method for constructing a knowledge graph based on audit information, comprising: Obtaining the text information to be audited; In response to determining that there is a target symbol in the text information to be audited, removing the target symbol from the text information to be audited to generate candidate text information to be audited; Performing segmentation processing on the candidate text information to be audited to generate a sequence of text segment information; Taking a target number of adjacent text segments from the sequence of text segment information in turn with a fixed step size as text information to obtain a set of text information; Performing entity and entity relationship extraction on each text information in the set of text information to generate a set of entity pair information groups, and obtaining a set of entity pair information groups, wherein the entity pair information in the entity pair information group in the set of entity pair information groups includes: a set of entity information and relationship information, and performing entity and entity relationship extraction on each text information in the set of text information to generate a set of entity pair information groups, including: performing semantic extraction on the text information through a pre-trained semantic extraction model to generate semantic feature information; inputting the semantic feature information and the text information into a pre-trained entity and entity relationship extraction model to generate a set of entity pair information groups corresponding to the text information; Based on the set of entity pair information groups and an empty graph network, performing the following graph network generation steps: Selecting an entity pair information from the set of entity pair information groups to generate target entity pair information and a set of candidate entity pair information groups, wherein the set of candidate entity pair information groups does not include the target entity pair information; In response to determining that the empty graph network does not include two entity nodes and a relationship edge corresponding to the target entity pair information, adding the entity nodes and the relationship edge corresponding to the target entity pair information to the empty graph network to generate a candidate graph network; In response to determining that the empty graph network includes two entity nodes and a relationship edge corresponding to the target entity pair information, increasing the weight of the relationship edge corresponding to the target entity pair information in the empty graph network to generate a candidate graph network; In response to determining that the set of candidate entity pair information groups is empty, determining the candidate graph network as an initial graph network and ending the graph network generation steps; In response to determining that the set of candidate entity pair information groups is not empty, determining the set of candidate entity pair information groups as the set of entity pair information groups and determining the candidate graph network as the empty graph network, and performing the graph network generation steps again; Performing graph network optimization on the initial graph network to generate a target knowledge graph.

2. The method according to claim 1, wherein, The target symbol is a symbol in a pre-constructed identification symbol library.

3. The method according to claim 1, wherein The method further includes: Obtaining a basic knowledge graph; For each first entity node in the target knowledge graph, performing the following processing steps: Determining the entity similarity between the first entity node and each second entity node in the basic knowledge graph to generate entity similarity information and obtaining a set of entity similarity information; In response to determining that the first entity node satisfies a first fusion condition, determine the first relationship edges connected to the first entity node in the target knowledge graph to obtain a set of relationship edge information, where the first fusion condition is that there is entity similarity information in the entity similarity information set with a corresponding similarity greater than a first preset similarity; Determine the relationship similarity between each piece of relationship edge information in the set of relationship edge information and each second relationship edge in the basic knowledge graph to generate relationship similarity information, and obtain a set of relationship similarity information; In response to determining that there is at least one target relationship similarity information in the set of relationship similarity information, fuse the relationship edges corresponding to the target relationship similarity information and the two first entity nodes connected by the relationship edges in the at least one target relationship similarity information into the basic knowledge graph to generate a fused knowledge graph, where the target relationship similarity information in the at least one target relationship similarity information is relationship similarity information that satisfies a second fusion condition, and the second fusion condition is that the relationship similarity corresponding to the relationship similarity information is less than a second preset similarity.

4. The method according to claim 1, wherein The optimizing the graph network of the initial graph network to generate a target knowledge graph includes: Removing target entity node pairs from the initial graph network to generate the target knowledge graph, where the target entity node pairs are two first entity nodes that are connected by a relationship edge in the initial graph network and satisfy the entity node removal condition.

5. The method according to claim 3, wherein The method further includes: In response to determining that there is a target entity node in the fused knowledge graph, add the entity information corresponding to the target entity node to a pre-constructed entity knowledge base, where the target entity node is an entity node whose corresponding entity information does not exist in the entity knowledge base.

6. A knowledge graph construction device based on audit information, comprising: An acquisition unit configured to acquire text information to be audited; A text extraction and processing unit configured to, in response to determining that there is a target symbol in the text information to be audited, remove the target symbol from the text information to be audited to generate candidate text information to be audited; perform segmentation processing on the candidate text information to be audited to generate a sequence of text segment information; and sequentially select a target number of adjacent text segments from the sequence of text segment information at a fixed step size as text information to obtain a set of text information; An entity and entity relationship extraction unit, configured to perform entity and entity relationship extraction on each text information in the text information set to generate an entity pair information group, and obtain an entity pair information group set, wherein the entity pair information in the entity pair information group in the entity pair information group set includes: an entity information set and relationship information, and the performing entity and entity relationship extraction on each text information in the text information set to generate an entity pair information group includes: performing semantic extraction on the text information through a pre-trained semantic extraction model to generate semantic feature information; inputting the semantic feature information and the text information into a pre-trained entity and entity relationship extraction model to generate an entity pair information group corresponding to the text information; A construction unit, configured to perform the following graph network generation steps based on the entity pair information group set and an empty graph network: selecting an entity pair information from the entity pair information group set to generate target entity pair information and a candidate entity pair information group set, wherein the candidate entity pair information group set does not include the target entity pair information; in response to determining that the empty graph network does not include two entity nodes and a relationship edge corresponding to the target entity pair information, adding the entity nodes and the relationship edge corresponding to the target entity pair information to the empty graph network to generate a candidate graph network; in response to determining that the empty graph network includes two entity nodes and a relationship edge corresponding to the target entity pair information, increasing the weight of the relationship edge corresponding to the target entity pair information in the empty graph network to generate a candidate graph network; in response to determining that the candidate entity pair information group set is empty, determining the candidate graph network as an initial graph network and ending the graph network generation step; in response to determining that the candidate entity pair information group set is not empty, determining the candidate entity pair information group set as the entity pair information group set and determining the candidate graph network as the empty graph network, and performing the graph network generation step again; A graph network optimization unit, configured to optimize the initial graph network to generate a target knowledge graph.

7. An electronic device, comprising: One or more processors; A storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.

8. A computer-readable medium having a computer program stored thereon, wherein, The program, when executed by the processor, implements the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Electric power knowledge graph construction method and device

    CN112632287A

  • Construction method of auditing system time sequence knowledge graph

    CN113849659A