Figure clue mining method, system and equipment based on knowledge graph, medium and product
By using the BERT-EGRE model to jointly extract entity relationships in massive text data in the information age and building a knowledge graph, the problem of difficult to identify and mine character relationships is solved, and efficient and accurate character clue mining effect is achieved.
Patent Information
- Application Number
- CN202510163712.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
In the information age, it is difficult to identify and mine character relationships in massive text data efficiently and accurately, especially when the data volume is large and information fragmented, traditional methods are difficult to expand, and multi-source heterogeneous data increases the difficulty.
The BERT-EGRE-based entity relationship joint extraction model is used to jointly extract entities and relationships, build a knowledge graph, and tap potential key nodes and correlation clues through multi-dimensional calculation graph analysis method.
It significantly improves the efficiency of character clue mining and the accuracy and diversity of results, can effectively process massive multi-source heterogeneous data, and is suitable for automated character relationship extraction and clue mining in complex scenarios.
Smart Images

Figure CN120107006A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information technology, and in particular to a method, system, device, medium and product for character clue mining based on knowledge graph. Background Art
[0002] In the current information age, the Internet generates massive amounts of text data every day, which contain rich and complex character clues and social relationships. It is difficult to efficiently and accurately identify character relationships in texts by relying solely on manual or traditional information extraction methods. And with the rapid growth of data volume and information fragmentation, traditional methods are also difficult to expand in scale. Limited by surface text information, it is difficult to deeply mine semantic associations. At the same time, data sources are becoming more and more diverse, including multi-source heterogeneous data such as text, images, and videos, which makes information extraction and character clue mining more difficult. Therefore, how to automatically extract valuable character relationship information from these massive data has become an important issue in the fields of information retrieval and data mining. This demand is particularly prominent in the fields of crime investigation, network security, historical research, and public opinion monitoring. Especially in complex scenarios involving multiple characters and multiple relationships, effective and intelligent automated character relationship extraction and clue mining tools are particularly important. Summary of the invention
[0003] The purpose of this application is to provide a method, system, device, medium and product for character clue mining based on knowledge graph to improve mining efficiency and the accuracy and diversity of mining results.
[0004] To achieve the above objectives, this application provides the following solutions.
[0005] In a first aspect, the present application provides a method for character clue mining based on a knowledge graph, comprising:
[0006] Obtain text data to be mined;
[0007] Perform data preprocessing on text data to obtain preprocessed vector data;
[0008] A BERT-EGRE-based entity-relationship joint extraction model is used to extract triple data of vector data; the triple data includes entity pairs and relationships between entity pairs; the entities include characters; the relationships include family relationships, professional relationships, and social relationships;
[0009] The entities in the triple data are taken as nodes, and the relationships between entity pairs are taken as edges to construct a knowledge graph;
[0010] Through multi-dimensional computing graph analysis methods, potential key nodes and related clues in the knowledge graph are mined.
[0011] Optionally, the performing data preprocessing on the text data to obtain preprocessed vector data specifically includes:
[0012] The text data is processed by removing stop words and segmenting words, then adding special tags, and then mapped into vector data consisting of three parallel vectors: input identifier, sentence type identifier, and attention mask.
[0013] Optionally, before extracting triple data of vector data using the BERT-EGRE based entity relation joint extraction model, the method further includes:
[0014] Construct a BERT-EGRE model; the BERT-EGRE model includes an encoding layer and a Global Pointer; the encoding layer includes a pre-trained BERT model and a relation embedding layer; the BERT model is used to encode the input vector data and convert it into a context representation vector containing context information, while capturing the deep-level features of each entity; the relation embedding layer is used to map the predefined relation type ID to a vector and output a relation embedding matrix; the context representation vector and the relation embedding matrix output by the BERT model are concatenated and input into the Global Pointer for relation modeling to generate a score matrix for all possible entity pairs under different relations; based on the score matrix, triple data consisting of entity pairs and relations between entity pairs are screened out;
[0015] The BERT-EGRE model is trained using different hyperparameters, and sparse multi-label cross entropy loss is used as the loss function during training.
[0016] After the training is completed, the BERT-EGRE model with the best performance is encapsulated and used as the entity relationship joint extraction model.
[0017] Optionally, the extracting triple data of vector data using a BERT-EGRE-based entity-relation joint extraction model specifically includes:
[0018] The vector data X = {x 1 ,x 2 ...,x n} is input into the BERT model for encoding, and the BERT model converts it into a context representation vector H = {h 1 ,h 2 ...,h n}; where x n represents the nth vector data, X is the vector data set; h n is the nth context representation vector, H is the context representation vector set;
[0019] The relation embedding layer maps the predefined relation type ID into a vector and outputs the relation embedding matrix E. r =Embedding(R); where R = {0,1,…,N r -1} represents the relationship type ID set, N r is the relationship type number; Embedding() represents the mapping operation;
[0020] Use the formula F = Concat (H, E r ) The context representation vector H and the relationship embedding matrix E r Concatenate to obtain a concatenated vector F; Concat() represents a concatenation operation;
[0021] The concatenated vector F is input to Global Pointer for relationship modeling to generate a score matrix for all possible entity pairs under different relationships. The score matrix calculation formula is: Where Q[i][r] represents the i-th entity token_i as the head entity feature of relation r; K[j][r] represents the j-th entity token_j as the tail entity feature of relation r; d is the projection dimension; Score[r][i][j] is the corresponding score matrix;
[0022] When the score matrix Score[r][i][j] is higher than the score threshold, the corresponding<token_i,r,token_j> Output as triplet data.
[0023] Optionally, after constructing the knowledge graph, the step further includes:
[0024] Add, delete, or modify nodes and edges in the knowledge graph.
[0025] Optionally, the graph analysis method using multi-dimensional computing to mine potential key nodes and related clues in the knowledge graph specifically includes:
[0026] For a specific query where a user enters a keyword, use Cypher query statements to query the knowledge graph for nodes within the specified degree of the keyword, as well as the relationship between the node and other entities; or
[0027] For a specific query where the user specifies the start node and the target node, a complex graph pattern matching is performed in the knowledge graph through Cypher statements to find the path between the start node and the target node as the person relationship link; or
[0028] For a specified node input by the user, use a breadth-first search algorithm to start from the node, expand the network diagram according to the relationship level, give priority to displaying direct contacts with a closer distance, and explore direct relationships between people; or
[0029] For a specified node input by the user, use depth-first search to present one or more paths related to the node as character relationship links; or
[0030] When conducting a specific analysis of nodes in the knowledge graph, use the PageRank algorithm to identify key nodes in the knowledge graph as influential person entities; or
[0031] The cosine similarity method is used to calculate the similarity between two person entities in the knowledge graph for recommending content or finding relationships.
[0032] In a second aspect, the present application provides a character clue mining system based on a knowledge graph, including: a presentation layer, an application layer, and a data storage layer;
[0033] The presentation layer is used to interact with the user on the page and to send and receive business data with the application layer; the business data includes text data and triple data;
[0034] The application layer is used to implement the character relationship extraction and mining process. While receiving text data and operation requests from the presentation layer, it transmits information and instructions to the data storage layer, and can receive data returned by the data storage layer and pass it to the presentation layer for display to the user;
[0035] The application layer includes a data preprocessing module, a character relationship extraction module, a knowledge graph management module and a character clue mining module;
[0036] The data preprocessing module is used to perform data preprocessing on the text data to obtain preprocessed vector data;
[0037] The character relationship extraction module uses a BERT-EGRE-based entity relationship joint extraction model to extract triple data of vector data; the triple data includes entity pairs and relationships between entity pairs; the entities include characters; the relationships include family relationships, professional relationships, and social relationships;
[0038] The knowledge graph management module is used to construct a knowledge graph by taking entities in triple data as nodes and relationships between entity pairs as edges;
[0039] The character clue mining module is used to mine potential key nodes and related clues in the knowledge graph through a multi-dimensional computing graph analysis method;
[0040] The data storage layer is used to store and cache various types of data, including text data, triple data, knowledge graph data and query result data.
[0041] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the knowledge graph-based character clue mining method.
[0042] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for character clue mining based on the knowledge graph.
[0043] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the knowledge graph-based character clue mining method.
[0044] According to the specific embodiments provided in this application, this application discloses the following technical effects.
[0045] The present application provides a method, system, device, medium and product for character clue mining based on knowledge graph, which jointly extracts entities and relationships by introducing a BERT-EGRE-based entity-relationship joint extraction model, and constructs a knowledge graph using the entities in the extracted triple data as nodes and the relationships between entity pairs as edges. It uses a multi-dimensional computing graph analysis method to mine potential key nodes and related clues in the knowledge graph, which can significantly improve the mining efficiency as well as the accuracy and diversity of the mining results. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0047] Figure 1 A flowchart of a method for character clue mining based on knowledge graph is provided for this application;
[0048] Figure 2 This is a schematic diagram of the structure of the entity relationship joint extraction model based on BERT-EGRE in this application;
[0049] Figure 3 A three-layer architecture diagram of a character clue mining system based on knowledge graph for this application;
[0050] Figure 4 This is a schematic diagram of the use process of the character clue mining system based on the knowledge graph in this application. DETAILED DESCRIPTION
[0051] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0052] This application proposes a method, system, device, medium and product for character clue mining based on knowledge graph to improve mining efficiency and the accuracy and diversity of mining results.
[0053] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0054] In an exemplary embodiment, Figure 1 As shown, a method for character clue mining based on knowledge graph is provided, including the following steps 1 to 5.
[0055] Step 1: Obtain the text data to be mined.
[0056] In the current information age, the Internet generates a massive amount of text data every day, which contains rich and complex clues about people and social relationships. The text data to be mined can be extracted from multi-source heterogeneous data such as documents, images, and videos. The extracted text data includes at least one text sentence.
[0057] Step 2: Perform data preprocessing on the text data to obtain preprocessed vector data.
[0058] The data preprocessing of text data includes: removing stop words from text data, segmenting the text data after removing stop words, adding special tags such as [CLS] and [SEP], and then mapping it into vector data consisting of three parallel vectors: input identifier, sentence type identifier, and attention mask. Among them, [CLS] and [SEP] are two special markers in BERT (Bidirectional Encoder Representations from Transformers), which play a special role in the input text of BERT. [CLS] is the abbreviation of "classification". In text classification tasks, it usually indicates the beginning of a sentence or document. In BERT, [CLS] corresponds to the word vector of the first word in the input text, and the first neuron in the output layer is usually used to predict the category of the text. [SEP] is the abbreviation of "separator", which usually indicates the end of a sentence or document. In BERT, [SEP] corresponds to the word vector of the last word in the input text, and its function is to segment different sentences. For example, when processing sentence pairs in BERT, a [SEP] is usually inserted between the two sentences to indicate their dividing point.
[0059] Step 3: Use the BERT-EGRE-based entity-relation joint extraction model to extract triple data of vector data; the triple data includes entity pairs and the relationship between entity pairs.
[0060] The character relationship extraction of this application is mainly based on text data, aiming to extract character entities and their relationships from text sentences and automatically identify triples in the text. In order to extract character (usually name) entities from unstructured text data and classify their mutual relationships, such as family relationships, professional relationships, social relationships, etc., this application establishes an entity relationship joint extraction model based on BERT-EGRE (BERT-Embedding Global Relation Extraction) to achieve triple extraction.
[0061] like Figure 2As shown, the BERT-EGRE model constructed in this application specifically includes a coding layer and a Global Pointer; the coding layer also includes a pre-trained BERT model and a relationship embedding layer (Relationship Embeddings). Among them, the BERT model is used to encode the input vector data, convert it into a context representation vector containing context information, and capture the deep-level features of each entity. The relationship embedding layer is used to map the predefined relationship type ID to a vector and output a relationship embedding matrix. After splicing the context representation vector and the relationship embedding matrix output by the BERT model, they are input into the Global Pointer for relationship modeling to generate a score matrix for all possible entity pairs under different relationships. The triple data composed of entity pairs and the relationship between entity pairs is screened out based on the score matrix, which can be represented by <Entity 1, Relationship, Entity 2>. "Entity 1" and "Entity 2" are two different person entities, constituting an entity pair.
[0062] Specifically, firstly, the original input text s 1 ,s 2 ...,s n Perform preprocessing such as removing stop words, segmenting, adding special tags, and vector mapping to obtain the vector x 1 ,x 2 ...,x n . Set the vector data X = {x 1 ,x 2 ...,x n} is input into the pre-trained BERT model for encoding, and the BERT model converts it into a context representation vector H = {h 1 ,h 2 ...,h n}, used to represent the overall semantics of the sentence and capture the deep features of each entity. n Indicates the nth text data; x n represents the nth vector data, X is the vector data set; h n is the nth context representation vector, and H is the set of context representation vectors.
[0063] The relation embedding layer maps the predefined relation type ID into a vector and outputs the relation embedding matrix E. r , the formula is:
[0064] E r =Embedding(R) (1)
[0065] Where R = {0,1,…,N r -1} represents the relationship type ID set, Nr is the number of relationship types, which may include family relationships, professional relationships, social relationships, etc. Embedding() represents a mapping operation.
[0066] Then the output H of the BERT model and the relation embedding matrix E r They are concatenated and their features are integrated to make the model aware of the type of relationship currently being processed. The formula is as follows:
[0067] F=Concat(H,E r ) (2)
[0068] Where Concat() represents the concatenation operation; F is the concatenated vector.
[0069] The concatenated vector F is input to Global Pointer for relationship modeling to generate a score matrix for all possible entity pairs under different relationships. The threshold can be applied to filter out the final relationship from these distributions to achieve joint extraction of entities and relationships. The score matrix calculation formula is:
[0070]
[0071] Where Q[i][r] indicates that the i-th entity token_i is the head entity feature of relation r; K[k][r] indicates that the j-th entity token_j is the tail entity feature of relation r; d is the projection dimension; Score[r][i][j] is the corresponding score matrix. If Score[r][i][j] is very high, it means that under relation r, token_i is the head entity and token_j is the tail entity. In other words, when the score matrix Score[r][i][j] is higher than a certain score threshold, the corresponding<token_i,r,token_j> Output as triple data <entity 1, relation, entity 2>.
[0072] The BERT-EGRE model aims to solve the problems of long-distance dependencies, nested entities, and low-frequency relationship recognition in character relationship extraction in complex contexts. It captures global text representation based on the pre-trained BERT model to obtain the contextual semantics of the data; introduces learnable relationship embedding vectors to enhance the model's semantic discrimination ability for low-frequency relationships; and further uses GlobalPointer to parallelly calculate the score matrices of all possible entity pairs under different relationship types, covering long-distance entity and nested structure recognition, and achieving efficient and accurate joint extraction of entity relationships.
[0073] During the model training process, we will try to apply different hyperparameters (such as learning rate, batch size, regularization parameters, etc.) to train the BERT-EGRE model to find the optimal configuration. Models from different rounds may perform differently, and the best BERT-EGRE model will be selected and packaged as an entity relationship joint extraction model for users to use.
[0074] During model training, the loss function used is sparse multi-label cross entropy loss. This loss function is designed for multi-label classification tasks and is particularly suitable for scenarios with few valid entity pairs, entity nesting, and global modeling. The loss function formula is as follows:
[0075]
[0076] Among them, A represents the entire set, P represents the negative example set (invalid entity pairs); S k is the model's prediction score for the kth label; LOSS is the sparse multi-label cross entropy loss. In traditional multi-label classification, a complete label matrix is usually required to indicate which categories are positive and which are negative. However, in the entity relationship extraction task of this application, the number of positive classes is far less than that of negative classes, and constructing a complete label matrix will bring huge storage and computational overhead. Therefore, the sparse multi-label cross entropy loss function is used to calculate the loss only through the positive class subscript, without knowing the specific location of all negative classes, which significantly reduces memory and computational overhead.
[0077] Step 4: Use the entities in the triple data as nodes and the relationships between entity pairs as edges to construct a knowledge graph.
[0078] As a structured semantic expression technology, knowledge graph draws a complex semantic network by modeling entities and their relationships. It can not only enhance the semantic expression ability of data, but also allow complex queries and reasoning, providing semantic data support for search engines, intelligent question and answer, as well as finance, medical care, intelligence analysis and other fields. Using knowledge graph to manage person relationship data can not only effectively solve the problems of diverse types, large scale, complex attributes, scattered data distribution, low utilization rate, and poor correlation of person relationship information, but also provide a unified model and standard to standardize data, realize scientific management and multi-dimensional presentation of data, and build a complete person relationship knowledge system, thereby greatly improving the comprehensive utilization of person relationship data. At the same time, knowledge graph has the characteristics of information storage, management and transmission, and can provide support for downstream retrieval, mining and visualization and other application tasks.
[0079] This application uses the encapsulated entity-relationship joint extraction model to extract the triple data <entity 1, relationship, entity 2>, and then uses the entities in the triple data as nodes and the relationships between entity pairs as edges to construct the corresponding knowledge graph (hereinafter referred to as graph). In the character clue mining task, the knowledge graph can realize the intuitive expression and analysis of complex character relationship networks by abstracting characters and relationships into nodes and edges, providing a deep-level conversion method from data to knowledge.
[0080] After building the knowledge graph, you can also add, delete, and modify the nodes and edges in the knowledge graph. Users can create a knowledge graph by uploading text data in the form of files, and store the automatically extracted triples in the form of nodes and edges in the corresponding graph of the Neo4j graph database, and display the uploaded file information and graph information. Users can also delete and edit uploaded files and completed graphs. Users can flexibly manage and maintain the knowledge graph by adding, deleting, and modifying the nodes and edges in the graph. Users can click the "Add Node" and "Add Edge" buttons to enter the name and type information of the node and edge, and the background will automatically add new nodes to the graph and create new edges. Users can also upload data files containing multiple nodes and edges (such as CSV, JSON, etc.), and the background will automatically parse the files and add them to the graph in batches. Users can also select one or more nodes or edges in the graph view, and by clicking the "Delete Node" and "Delete Edge" buttons, the background will delete the selected nodes and their related edges. Users can also select nodes or edges in the graph view and directly modify the name and type of the node or edge, and the background will save the updated graph information.
[0081] Step 5: Use multi-dimensional computing graph analysis methods to mine potential key nodes and related clues in the knowledge graph.
[0082] Based on the knowledge graph composed of the results of joint extraction of entity relationships, this application mines potential related clues and key nodes through multi-dimensional computing graph analysis methods to support users in exploring complex relationship networks. Multi-dimensional computing graph analysis methods can include Cypher query statements, breadth-first search (BFS), depth-first search (DFS), PageRank algorithm, similarity calculation method and other clue mining methods for users to choose as needed.
[0083] For example, for a specific query where a user enters a keyword, a Cypher query statement can be used to query the knowledge graph for nodes within the specified degree of the keyword, as well as the relationship between the node and other entities. For a specific query where a user specifies a start node and a target node, a Cypher statement can be used to perform complex graph pattern matching in the knowledge graph stored in the Neo4j graph database to find the path between the start node and the target node, and output it as a person relationship link.
[0084] For a specified node input by a user, a breadth-first search algorithm can be used to start from the node, expand the network diagram according to the relationship level, and give priority to displaying direct contacts with a closer distance to explore the direct relationship between people. For another example, for a specified node input by a user, a depth-first search can also be used to present one or more in-depth exploration paths related to the node, which are output as person relationship links.
[0085] When conducting a specific analysis of the nodes in the knowledge graph, the PageRank algorithm can be used to identify the key nodes in the knowledge graph and find the influential person entities in the knowledge graph. The calculation formula is as follows:
[0086]
[0087] Among them, PR(p) represents the PageRank score of node p; b is the damping factor, usually 0.85; Q(p) represents the set of all nodes pointing to node p; PR(q) represents the PageRank score of node q pointing to node p; L(q) represents the number of nodes pointed to by node q. When the calculated PR(p) is greater than a certain threshold, node p is considered to be a key node in the knowledge graph, representing an influential person entity in the knowledge graph.
[0088] When calculating character clues, you can also use the cosine similarity method to calculate the similarity between two character entities in the knowledge graph for recommending content or finding relationships. The calculation formula is as follows:
[0089]
[0090] Among them, M and N represent the feature vectors of two nodes respectively, M·N is the dot product of the vectors, ||M|| and ||N|| are the respective vector norms; Similarity(M,N) is the cosine similarity of M and N.
[0091] This application considers the performance of knowledge extraction and the calculation of character clues, and designs a character clue mining method based on knowledge graph, aiming to analyze unstructured text data, introduce automated entity extraction and relationship recognition technology, automatically extract name entities and their relationships, and realize display and mining functions. This application improves the extraction accuracy through entity-relationship joint extraction technology, combines the graph structure of the knowledge graph to mine clues for multi-level character relationships, deepens the understanding of character relationships hidden in the data, and provides innovative solutions for the semantic analysis and exploration of large-scale graph data, realizing the intelligent transformation from data to knowledge. Furthermore, this application provides Cypher query statements, breadth-first search, depth-first search, PageRank algorithm, cosine similarity method and other multi-dimensional calculation methods, which can improve the accuracy and diversity of mining results, and complete the comprehensive analysis of graph data and deep mining of clues.
[0092] In an exemplary embodiment, the present application also provides a character clue mining system based on knowledge graph. Figure 3 As shown in the figure, the character clue mining system includes a three-layer system architecture, namely the presentation layer, the application layer and the data storage layer, which together realize the full life cycle management and analysis of the data in the system. Among them, the presentation layer is a unified interactive portal for users, providing a visual operation interface and a real-time feedback mechanism, and sending and receiving data with the application layer, including file upload interface, file preview interface, graph display interface, graph query interface, etc. The data storage layer is a layer that implements structured storage and access to character and relationship data, using databases such as MongoDB, MySQL, Neo4j, and Redis. The application layer processes business logic, and while receiving data and operation requests from the presentation layer, it conveys information and instructions to the data storage layer. At the same time, it receives the data returned by the data storage layer and passes it to the presentation layer for display to users. The application layer includes a knowledge graph management module, a character relationship extraction module, a character clue mining module, and a data preprocessing module. The character relationship extraction module is composed of an entity relationship joint extraction model based on BERT-EGRE. Through entity relationship joint extraction, a direct mapping from original text to entity relationship triples is achieved.
[0093] Specifically, the presentation layer, as the first layer of the three-layer system architecture, is used to interact with users on pages and send and receive business data with the application layer. The business data includes at least text data and triple data. The presentation layer is mainly an interactive interface for character clue extraction and mining. The main functions implemented by this interface include data upload function, knowledge graph visualization display function, graph modification function, node query and calculation function, etc. The knowledge graph visualization display function can obtain the corresponding query results by entering the node relationship to be queried in the front-end interface and selecting the query calculation method. Users can judge the importance and relevance of the characters represented by the nodes based on the query results, so as to conduct data analysis and decision-making more effectively.
[0094] The application layer, as an intermediate transition layer, mainly implements the process of character relationship extraction and mining. While receiving text data and operation requests from the presentation layer, it conveys information and instructions to the data storage layer, and can receive query data returned by the data storage layer and pass it to the presentation layer for display to users.
[0095] The application layer includes three main modules: a character relationship extraction module, a knowledge graph management module and a character clue mining module, and also includes a data preprocessing module. Among them, the data preprocessing module is used to perform data preprocessing on text data to obtain preprocessed vector data. The character relationship extraction module aims at the problem of character relationship extraction. For the unstructured data input by the user (such as text data in the form of files and documents, etc.), the entity relationship joint extraction model based on BERT-EGRE proposed in this application is used to extract triples (such as Zhang San-Friend-Li Si), and the processed triple results are saved in the database to provide basic data for subsequent analysis and mining.
[0096] Specifically, the character relationship extraction module uses a BERT-EGRE-based entity relationship joint extraction model to extract triple data of vector data. The triple data includes entity pairs and relationships between entity pairs; the entities mainly refer to characters; the relationships include social relationships such as family relationships, professional relationships, and social relationships. In the entity recognition and relationship extraction stages, the vector data composed of three vectors is input into the BERT model for encoding, and the context representation vector of each word is obtained. These context representation vectors will be used for subsequent relationship extraction and entity recognition. In order to enhance the model's understanding of different relationships, a relationship embedding layer is introduced, which can use pre-trained embeddings or randomly initialized embeddings. The output of BERT encoding (context representation vector) and the relationship embedding matrix are used as input, combined with Global Pointer for relationship modeling, and the difficulty of extracting nested entities NE is solved through a two-dimensional pointer annotation method. The output of Global Pointer will provide a relationship probability distribution for each pair of entities, indicating the possible relationship between the pair of entities. The application of thresholds can filter out the final relationship from these probability distributions. Finally, all high-confidence relations and their corresponding entities are extracted from the output of Global Pointer to form triple data <entity 1, relation, entity 2>. The extracted triple results are stored in the Neo4j graph database.
[0097] The knowledge graph management module is used to construct a knowledge graph using entities in triple data as nodes and relationships between entity pairs as edges. The graph management module is also used to perform operations such as adding, deleting, and modifying the knowledge graph data therein. Users can use this module to add new nodes and edges, delete unnecessary nodes and edges, modify the properties of existing nodes and edges, and query specific information in the graph. Through these operations, users can flexibly manage and maintain the knowledge graph to ensure its accuracy and completeness.
[0098] The character clue mining module is used to mine potential key nodes and related clues in the knowledge graph through a multi-dimensional computing graph analysis method. The character line clue mining module targets character nodes, combines structured information in the database, and performs matching queries and index calculations on the knowledge graph data at the back end to obtain clue information such as character links, key characters and their relationships.
[0099] The data storage layer is used to store and cache multiple types of data, including text data, triple data, knowledge graph data, and query result data. The data storage layer designs different storage solutions for the multiple data types involved in the system, such as text data, triple data, knowledge graph data, and query data, to achieve efficient data management and performance optimization. In this system, MongoDB is used to store text data, MySQL stores triple data, Neo4j stores knowledge graph data, and Redis cache is used to store query data obtained by mining calculations to optimize query performance.
[0100] Figure 4 This is a schematic diagram of the use process of the character clue mining system based on the knowledge graph in this application. The offline part uses a constructed data set to train the model, and the best model is encapsulated for entity relationship extraction. The data set refers to the character relationship training data set, including cleaned text data and triples. Ontology modeling means first giving all the limited relationship types, and then the relationship type extracted by the model must be one of them. In the online part, the user enters the character clue mining system page and searches for the corresponding text. If it does not exist, the file upload operation is performed. After the upload is successful, the text data will be pre-processed on the back end, and then the encapsulated entity relationship joint extraction model will be used to extract the entity relationship, and the obtained triple results will be stored in the database. Users can see the corresponding knowledge graph results constructed on the front-end interface, and can also edit and modify them.
[0101] According to the actual needs of character clue mining, different analysis methods are selected for the nodes and relationships in the graph. For example, for a specific query entered by a user for a specific node, a Cypher query statement can be used to query the nodes within the specified degree of the keyword in the Neo4j database, as well as the relationships with other entities. For a specific query in which the user specifies the starting node and the target node, a complex BFS / DFS search is performed in the Neo4j graph database through a Cypher query statement to find the paths and relationships between characters. When performing a specific analysis of the nodes in the graph, the PageRank algorithm can be used to calculate the scores of the nodes in the graph, find the influential character entities in the knowledge graph, or compare the PageRank scores of two entities in the graph to intuitively judge their relative importance in the graph. Through the above different analysis methods, the accuracy and diversity of the mining results can be improved.
[0102] When the presentation layer displays data, in addition to displaying it in the form of a graph, it can also display text data and triple data in a structured table format, which also supports modification.
[0103] In an exemplary embodiment, the present application also provides a computer device, which may be a server or a terminal. The computer device includes a processor, a memory, an input / output interface, and a communication interface. The processor, the memory, and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the character clue mining method based on the knowledge graph is implemented.
[0104] In an exemplary embodiment, the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for character clue mining based on knowledge graph.
[0105] In an exemplary embodiment, the present application also provides a computer program product, including a computer program, which implements the knowledge graph-based character clue mining method when executed by a processor.
[0106] It can be understood by a person of ordinary skill in the art that all or part of the processes in the above-mentioned embodiment method can be completed by hardware related to computer program instructions, and the computer program can be stored in a non-volatile computer-readable storage medium, and the computer program can include the process of the embodiment of the above-mentioned method when executed. Among them, any reference to memory or other media in each embodiment provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0107] It should be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0108] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0109] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for character clue mining based on knowledge graph, characterized in that: include: Obtain text data to be mined; Perform data preprocessing on text data to obtain preprocessed vector data; A BERT-EGRE-based entity-relationship joint extraction model is used to extract triple data of vector data; the triple data includes entity pairs and relationships between entity pairs; the entities include characters; the relationships include family relationships, professional relationships, and social relationships; The entities in the triple data are taken as nodes, and the relationships between entity pairs are taken as edges to construct a knowledge graph; Through multi-dimensional computing graph analysis methods, potential key nodes and related clues in the knowledge graph are mined.
2. The method for character clue mining based on knowledge graph according to claim 1 is characterized in that: The data preprocessing of the text data to obtain preprocessed vector data specifically includes: The text data is processed by removing stop words and segmenting words, then adding special tags, and then mapped into vector data consisting of three parallel vectors: input identifier, sentence type identifier, and attention mask.
3. The method for character clue mining based on knowledge graph according to claim 2 is characterized in that: Before extracting triple data of vector data using the BERT-EGRE-based entity relation joint extraction model, the method further includes: Construct a BERT-EGRE model; the BERT-EGRE model includes an encoding layer and a Global Pointer; the encoding layer includes a pre-trained BERT model and a relation embedding layer; the BERT model is used to encode the input vector data and convert it into a context representation vector containing context information, while capturing the deep-level features of each entity; the relation embedding layer is used to map the predefined relation type ID to a vector and output a relation embedding matrix; the context representation vector and the relation embedding matrix output by the BERT model are concatenated and input into the Global Pointer for relation modeling to generate a score matrix for all possible entity pairs under different relations; based on the score matrix, triple data consisting of entity pairs and relations between entity pairs are screened out; The BERT-EGRE model is trained using different hyperparameters, and sparse multi-label cross entropy loss is used as the loss function during training. After the training is completed, the BERT-EGRE model with the best performance is encapsulated and used as the entity relationship joint extraction model.
4. The method for character clue mining based on knowledge graph according to claim 3 is characterized in that: The method of extracting triple data of vector data using the BERT-EGRE-based entity relation joint extraction model specifically includes: The vector data X={x1,x2...,x n } is input into the BERT model for encoding, and the BERT model converts it into a context representation vector H = {h1,h2...,h n }; where x n represents the nth vector data, X is the vector data set; h n is the nth context representation vector, H is the context representation vector set; The relation embedding layer maps the predefined relation type ID into a vector and outputs the relation embedding matrix E. r =Embedding(R); where R = {0,1,…,N r -1} represents the relationship type ID set, N r is the relationship type number; Embedding() represents the mapping operation; Use the formula F = Concat (H, E r ) The context representation vector H and the relationship embedding matrix E r Concatenate to obtain a concatenated vector F; Concat() represents a concatenation operation; The concatenated vector F is input into GlobalPointer for relationship modeling to generate the score matrix of all possible entity pairs under different relationships. The score matrix calculation formula is: Where Q[i][r] represents the i-th entity token_i as the head entity feature of relation r; K[j][r] represents the j-th entity token_j as the tail entity feature of relation r; d is the projection dimension; Score[r][i][j] is the corresponding score matrix; When the score matrix Score[r][i][j] is higher than the score threshold, the corresponding<token_i,r,token_j> Output as triplet data.
5. The method for character clue mining based on knowledge graph according to claim 1 is characterized in that: After constructing the knowledge graph, the following steps are also included: Add, delete, or modify nodes and edges in the knowledge graph.
6. The method for character clue mining based on knowledge graph according to claim 1 is characterized in that: The graph analysis method using multi-dimensional computing to mine potential key nodes and related clues in the knowledge graph specifically includes: For a specific query where a user enters a keyword, use Cypher query statements to query the knowledge graph for nodes within the specified degree of the keyword, as well as the relationship between the node and other entities; or For a specific query where the user specifies the start node and the target node, a complex graph pattern matching is performed in the knowledge graph through Cypher statements to find the path between the start node and the target node as the person relationship link; or For a specified node input by the user, use a breadth-first search algorithm to start from the node, expand the network diagram according to the relationship level, give priority to displaying direct contacts with a closer distance, and explore direct relationships between people; or For a specified node input by the user, use depth-first search to present one or more paths related to the node as character relationship links; or When conducting a specific analysis of nodes in the knowledge graph, use the PageRank algorithm to identify key nodes in the knowledge graph as influential person entities; or The cosine similarity method is used to calculate the similarity between two person entities in the knowledge graph for recommending content or finding relationships.
7. A character clue mining system based on knowledge graph, characterized in that: include: Presentation layer, application layer and data storage layer; The presentation layer is used to interact with the user on the page and to send and receive business data with the application layer; the business data includes text data and triple data; The application layer is used to implement the character relationship extraction and mining process. While receiving text data and operation requests from the presentation layer, it transmits information and instructions to the data storage layer, and can receive data returned by the data storage layer and pass it to the presentation layer for display to the user; The application layer includes a data preprocessing module, a character relationship extraction module, a knowledge graph management module and a character clue mining module; The data preprocessing module is used to perform data preprocessing on the text data to obtain preprocessed vector data; The character relationship extraction module uses a BERT-EGRE-based entity relationship joint extraction model to extract triple data of vector data; the triple data includes entity pairs and relationships between entity pairs; the entities include characters; the relationships include family relationships, professional relationships, and social relationships; The knowledge graph management module is used to construct a knowledge graph by taking entities in triple data as nodes and relationships between entity pairs as edges; The character clue mining module is used to mine potential key nodes and related clues in the knowledge graph through a multi-dimensional computing graph analysis method; The data storage layer is used to store and cache various types of data, including text data, triple data, knowledge graph data and query result data.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the character clue mining method based on the knowledge graph as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for character clue mining based on knowledge graph described in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for character clue mining based on knowledge graph described in any one of claims 1 to 6 is implemented.