Government Affairs Text Data Classification and Processing System and Method Based on Knowledge Graph

By using knowledge graph technology, entity recognition and importance division in the classification and processing of government text data, and building a multi-level transparent knowledge graph, it solves the problem that it is difficult to accurately classify government text data in the existing technology, and achieves more efficient and flexible government text data processing.

CN119807420BActive Publication Date: 2025-06-27YOUQIAN SOFTLINK (BEIJING) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510288377.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The existing government text data classification processing methods are difficult to accurately evaluate and distinguish high and low key entities in the text, resulting in insufficient ability to capture complex relationships and the importance of global entities, affecting the accuracy of classification results.

Method used

A system based on knowledge graph is adopted to identify entities and divide importance through data acquisition and processing units, and divide entities into high-key entities using the entity importance dual weight division algorithm. A multi-level transparent knowledge graph is constructed through a depth map convolutional graph construction algorithm and a cross-graph embedding alignment inference algorithm to realize the precise classification and storage management of government text data.

Benefits of technology

It improves the accuracy of government text data classification, enhances the flexibility and efficiency of the system when processing complex government information, can dynamically process complex entity relationships in government text data, and provides on-demand classification and query functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807420B_ABST
    Figure CN119807420B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data classification and processing, and specifically, to a government affair text data classification and processing system and method based on a knowledge graph, which includes: a data acquisition and processing unit that acquires and cleans government affair text data, and uses an entity importance double-weight division algorithm to divide key entities into high-key entities and low-key entities; a primary knowledge graph unit that constructs a primary government affair knowledge graph using a deep graph convolutional spectrum construction algorithm based on high-key entities and high-key entity relationships; a secondary knowledge graph unit that constructs multiple secondary government affair knowledge graphs using a cross-spectrum embedding alignment and reasoning algorithm based on low-key entities and low-key entity relationships; a horizontal and vertical classification query unit for data query. The government affair text data classification and processing system and method based on the knowledge graph perform dynamic storage processing and cross-spectrum horizontal and vertical queries on complex government affair text data by constructing a multi-level knowledge graph and applying a cross-spectrum embedding alignment and reasoning algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data classification and processing, and more specifically, to a government affair text data classification and processing system and method based on a knowledge graph. Background Art

[0002] The government affair text data classification and processing system and method based on a knowledge graph aims to improve the accuracy of government affair text data classification and enhance the flexibility and efficiency of the system in processing complex government affair information. By combining the dual-weight division algorithm for entity importance and the cross-graph embedding alignment inference algorithm, a multi-level transparent knowledge graph is constructed to control the accurate identification of high and low key entities and the relationship alignment between graphs, so as to achieve precise horizontal and vertical classification and efficient storage management of government affair text data at multiple levels and dimensions.

[0003] Existing government affair text data classification and processing methods usually have difficulty in accurately evaluating and distinguishing high and low key entities in the text. Moreover, due to the objective reasons of complex and dynamic entity relationships and context information in government affair text data, traditional methods are insufficient in capturing complex relationships and the importance of global entities, which in turn affects the accuracy of classification results. Therefore, a government affair text data classification and processing system and method based on a knowledge graph are provided. Summary of the Invention

[0004] The purpose of the present invention is to provide a government affair text data classification and processing system and method based on a knowledge graph to solve the problem that due to the objective reasons of complex and dynamic entity relationships and context information in government affair text data, traditional methods are insufficient in capturing complex relationships and the importance of global entities, which in turn affects the accuracy of classification results as mentioned in the above background art.

[0005] To achieve the above purpose, the present invention aims to provide a government affair text data classification and processing system based on a knowledge graph, including: a data acquisition and processing unit, which acquires and cleans government affair text data, identifies entities in the government affair text data using a named entity recognition model, divides key entities into high key entities and low key entities using a dual-weight division algorithm for entity importance, and extracts entity relationships using Stanford OpenIE technology;

[0006] The dual-weight division algorithm for entity importance is implemented based on TF-IDF technology and PageRank algorithm technology, and is used to evaluate the criticality of entities in government affair text data and divide entities into high key entities and low key entities;

[0007] It further includes a first-level knowledge graph unit, which constructs a first-level government affairs knowledge graph based on high-key entities and high-key entity relationships using a deep graph convolutional graph construction algorithm, and is used to store and manage high-key government affairs data and its data relationships;

[0008] It further includes a second-level knowledge graph unit, which constructs multiple second-level government affairs knowledge graphs centered on the entity nodes of the first-level government affairs knowledge graph based on low-key entities and low-key entity relationships using a cross-graph embedding alignment and reasoning algorithm, and calculates the correlation degree of the two second-level government affairs knowledge graphs, and is used to store and manage low-key government affairs data and its data relationships;

[0009] It further includes a horizontal and vertical classification query unit, which provides a query interface for cross-node and cross-graph horizontal and vertical data queries.

[0010] As a further improvement of this technical solution, the data acquisition and processing unit includes a data acquisition and cleaning module, an entity recognition and division module, and an entity relationship extraction module;

[0011] Among them, the data acquisition and cleaning module is used to collect government affairs text data from the data source and preprocess and clean the government affairs text data;

[0012] The entity recognition and division module uses a named entity recognition model to recognize entities in the government affairs text data, and uses an entity importance double-weight division algorithm to divide the key entities into high-key entities and low-key entities;

[0013] The entity relationship extraction module uses Stanford OpenIE technology to extract entity relationships from the text.

[0014] As a further improvement of this technical solution, the entity recognition and division module uses a named entity recognition model to recognize entities in the government affairs text data, and uses an entity importance double-weight division algorithm to divide the key entities into high-key entities and low-key entities. The specific method is as follows:

[0015] S1.2.1. Use a named entity recognition model to recognize entities in the government affairs text data, and calculate the local importance weight and global importance weight of each entity:

[0016] ;

[0017] ;

[0018] Among them, is the government affairs text data set; is the amount of government affairs text data; is the index of the government affairs text data; is the i-th government text data; is the entity set; is the total number of entities; is the entity index; is the j-th entity;

[0019] Calculate the local weight of the entity:

[0020] ;

[0021] where, is the local weight of the entity ; is the TF-IDF technical operation; is the entity appears in the government text data the number of times; is the government text data quantity containing the entity ;

[0022] Calculate the global weight of the entity:

[0023] ;

[0024] where, is the global weight of the entity ; is the damping coefficient; is the entity index; is the k-th entity; is the entity out-degree; is the entity global weight;

[0025] S1.2.2. For each entity, calculate the comprehensive score of each entity based on the local weight and global weight of the entity:

[0026] ;

[0027] where, is the comprehensive score of the entity ; is the weighted coefficient of the local weight of the entity;

[0028] S1.2.3. According to the comprehensive score of each entity, divide each entity into high-key entities and low-key entities:

[0029] ;

[0030] where, is the average comprehensive score of the entity; is the set of all entities;

[0031] For an entity , if , then the entity is a high - critical entity; otherwise it is a low - critical entity;

[0032] Finally, obtain the high - critical entity set: ;

[0033] Low - critical entity set: ;

[0034] Among them, is the total number of high - critical entities; is the high - critical entity index; is the th high - critical entity; is the total number of low - critical entities; is the low - critical entity index; is the th low - critical entity.

[0035] As a further improvement of this technical solution, the deep graph convolutional spectrum construction algorithm is implemented based on the dynamic graph convolutional network and the Louvain community discovery algorithm, and is used to construct the first - level government affairs knowledge graph;

[0036] The first - level knowledge graph unit constructs the first - level government affairs knowledge graph using the deep graph convolutional spectrum construction algorithm based on high - critical entities and high - critical entity relationships. The specific method steps are as follows:

[0037] S2.1. Construct an initial first - level knowledge graph according to the high - critical entity set and high - critical entity relationships;

[0038] S2.2. On the basis of the initial first - level knowledge graph, introduce dynamic weights and a time decay factor, construct an improved adjacency matrix, update the node embeddings of each layer using the dynamic graph convolutional network, and perform weighted processing on the node relationships through the improved adjacency matrix;

[0039] S2.3. Calculate the node similarity based on the node embeddings of the dynamic graph convolutional network, and perform community partitioning on the nodes based on the Louvain community discovery algorithm, optimize the graph structure, calculate the modularity of the community partitioning, and dynamically merge relevant communities according to the node similarity;

[0040] S2.4. Repeatedly iterate and execute S2.2 and S2.3 until the change in modularity is less than a predetermined threshold, and finally obtain the first - level government affairs knowledge graph.

[0041] As a further improvement of this technical solution, in S2.1, the method for constructing an initial first - level knowledge graph according to the high - critical entity set and high - critical entity relationships is as follows:

[0042] Define a set of high - critical entities Define a set of nodes, and define the set of high - critical entity relationships as a set of edges :

[0043] ;

[0044] Among them, is the set of edges; is the set of high - critical entity relationships; is the th high - critical entity; is the th high - critical entity; is the relationship strength between high - critical entity and high - critical entity ;

[0045] Calculate the edge weight of each edge based on the comprehensive score of high - critical entities:

[0046] ;

[0047] Among them, is the edge weight between high - critical entity and high - critical entity ; is the comprehensive score of high - critical entity ; is the comprehensive score of high - critical entity ;

[0048] The set of edge weights between all high - critical entities is ;

[0049] Generate an initial first - level knowledge graph:

[0050] ;

[0051] Among them, is the initial first - level knowledge graph;

[0052] In the above - mentioned S2.2, based on the initial first - level knowledge graph, introduce a dynamic weight and a time decay factor, construct an improved adjacency matrix, use a dynamic graph convolutional network to update the node embeddings of each layer, and perform weighted processing on the node relationships through the improved adjacency matrix. The specific method steps are as follows:

[0053] Define the adjacency matrix of the initial first - level knowledge graph as , and the element in the adjacency matrix ;

[0054] Introduce a time decay factor Improved Adjacency Matrix :

[0055] ;

[0056] Wherein, is the improved adjacency matrix; is the time decay factor; is the current time; is the high-key entity and the high-key entity relationship strength of the last update time;

[0057] Based on the improved adjacency matrix , use the dynamic graph convolutional network to update the node embeddings of each layer:

[0058] ;

[0059] Wherein, is the node embedding of the th layer; is the ReLU activation function; is the degree matrix of the improved adjacency matrix; is the node embedding of the th layer; is the learnable parameter matrix.

[0060] As a further improvement of this technical solution, in the S2.3, based on the node embeddings of the dynamic graph convolutional network, calculate the node similarity, and based on the Louvain community discovery algorithm, divide the nodes into communities, optimize the graph structure, calculate the modularity of the community division, and dynamically merge relevant communities according to the node similarity. The specific method steps are as follows:

[0061] Calculate the node similarity between the high-key entity and the high-key entity :

[0062] ;

[0063] Wherein, is the node similarity between the high-key entity and the high-key entity ; is the final layer of the convolutional layer of the dynamic graph convolutional network; The node embedding of the final layer L of the high-key entity ; The node embedding of the final layer L of the high-key entity ;

[0064] Community partition of nodes is performed based on the Louvain community discovery algorithm to optimize the graph structure, and the modularity of the community partition is calculated:

[0065] ;

[0066] ;

[0067] Among them, is the modularity; is the sum of the edge weights of the initial first-level knowledge graph; is the expected connection strength; is the community where the high-key entity is located; is the community where the high-key entity is located; is the Kronecker delta function;

[0068] Set the merging threshold , if the similarity of two communities and satisfies the following equation:

[0069] ;

[0070] Then merge the two communities and into a single community;

[0071] Among them, is the merging threshold; is community a; is community b;

[0072] In the above S2.4, repeatedly iterate S2.2 and S2.3 until the change in modularity is less than a predetermined threshold, and finally obtain the first-level government affairs knowledge graph .

[0073] As a further improvement of this technical solution, the cross-graph embedding alignment reasoning algorithm is implemented based on knowledge graph embedding technology and cross-graph alignment technology, and is used to construct multiple second-level government affairs knowledge graphs and align and reason about entities and entity relationships;

[0074] The second-level knowledge graph unit includes a second-level graph construction module and a second-level graph optimization module;

[0075] Among them, the second-level graph construction module constructs multiple second-level government affairs knowledge graphs based on low-key entities and low-key entity relationships, with the entity nodes of the first-level government affairs knowledge graph as the center, using the cross-graph embedding alignment reasoning algorithm;

[0076] The secondary graph optimization module calculates the correlation degree between two secondary government affairs knowledge graphs, which is used to eliminate redundant secondary government affairs knowledge graphs.

[0077] As a further improvement of this technical solution, the secondary graph construction module constructs multiple secondary government affairs knowledge graphs based on low-key entities and low-key entity relationships, with the entity nodes of the primary government affairs knowledge graph as the center, using a cross-graph embedding alignment reasoning algorithm. The specific method steps are as follows:

[0078] S3.1.1. Taking each high-key entity node of the primary government affairs knowledge graph as the center, construct an initial secondary government affairs knowledge graph:

[0079] For the high-key entity node , define the set of low-key entities related to the high-key entity node as the low-key entity node set , and define the set of low-key entity relationships as the low-key entity edge set ;

[0080] Among them, is the high-key entity index; , represents the set of low-key entities related to the high-key entity node ; , represents the relationship between low-key entities;

[0081] Calculate the edge weight of each edge based on the comprehensive score of the low-key entity:

[0082] ;

[0083] Among them, , are both low-key entity indexes related to the high-key entity node ; is the edge weight of the low-key entity and the low-key entity ; is the comprehensive score of the low-key entity ; is the comprehensive score of the low-key entity ;

[0084] The set of edge weights between all low-key entities related to the high-key entity node is ;

[0085] Generate the initial secondary knowledge graph with the high-key entity node as the center:

[0086] ;

[0087] Among them, is the initial secondary knowledge graph;

[0088] S3.1.2. Use the node embeddings of the first-level government affairs knowledge graph to guide the semantic alignment of the initial secondary knowledge graph:

[0089] Obtain the embedding vectors of the highly critical entity nodes from the first-level government affairs knowledge graph and define a projection matrix to calculate the embedding vectors of the highly critical entity nodes after alignment:

[0090] ;

[0091] Among them, is the embedding vector of the highly critical entity node after alignment; is the embedding vector of the highly critical entity node; is the projection matrix;

[0092] Define the initial embedding vector of the low-critical entity in the initial secondary knowledge graph as ;

[0093] In the initial secondary knowledge graph, calculate the attention weights of each low-critical entity and its corresponding highly critical entity node in the first-level government affairs knowledge graph based on the initial embedding vector of the low-critical entity:

[0094] ;

[0095] Among them, is the attention weight of the low-critical entity node and the highly critical entity node ;

[0096] is the cosine similarity;

[0097] Based on the embedding vector of the highly critical entity node after alignment and the attention weights, align the initial embedding vector of the low-critical entity with the highly critical entity node after alignment, and calculate the aligned embedding vector of the low-critical entity:

[0098] ;

[0099] Among them, is the aligned embedding vector of the low-critical entity ;

[0100] Update the aligned embedding vector of the low-critical entity to the initial secondary knowledge graph to obtain the secondary government affairs knowledge graph 。

[0101] As a further improvement of this technical solution, the secondary graph optimization module calculates the correlation degree between two secondary government affairs knowledge graphs to eliminate redundant secondary government affairs knowledge graphs. The specific method steps are as follows:

[0102] Calculate the correlation degree between two secondary government affairs knowledge graphs:

[0103] ;

[0104] Among them, is the secondary government affairs knowledge graph constructed with the high-key entity node as the center; is the high-key entity index; , indicating the set of low-key entities related to the high-key entity node ; is the secondary government affairs knowledge graph constructed with the high-key entity node as the center; is the secondary government affairs knowledge graph and 's correlation degree;

[0105] Set the secondary graph merging threshold , and judge whether two secondary government affairs knowledge graphs need to be merged, as follows:

[0106] If , then the secondary government affairs knowledge graphs and are merged into one secondary government affairs knowledge graph 。

[0107] On the other hand, the present invention provides a method for classifying and processing government affairs text data based on a knowledge graph, which is used for the system for classifying and processing government affairs text data based on a knowledge graph described in any one of the above, and includes the following steps:

[0108] S10.1. Collect and clean government affairs text data, use a named entity recognition model to identify entities in the government affairs text data, use the entity importance double-weight division algorithm to divide key entities into high-key entities and low-key entities, and use Stanford OpenIE technology to extract entity relationships;

[0109] S10.2. Based on high-key entities and high-key entity relationships, use a deep graph convolutional graph construction algorithm to construct a primary government affairs knowledge graph;

[0110] S10.3. Based on the low-key entities and the relationships between low-key entities, centering on the entity nodes of the first-level government affairs knowledge graph, use the cross-graph embedding alignment reasoning algorithm to construct multiple second-level government affairs knowledge graphs, and calculate the correlation degree between the two second-level government affairs knowledge graphs;

[0111] S10.4. Use the query interface to perform horizontal and vertical data queries across nodes and across graphs.

[0112] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0113] 1. In the government affairs text data classification and processing system and method based on the knowledge graph, the combination of the entity importance dual-weight division algorithm and the TF-IDF technology can accurately evaluate and distinguish the importance of entities in government affairs text data. By dividing entities into high and low criticality levels, the accuracy of key entity recognition in government affairs text is effectively improved.

[0114] 2. In the government affairs text data classification and processing system and method based on the knowledge graph, by introducing the construction of multi-level knowledge graphs and the cross-graph embedding alignment reasoning algorithm, hierarchical management of high and low key entities and dynamic relationship adjustment between multiple graphs are realized. Based on the high-key entities of the first-level knowledge graph and the low-key entities of the second-level knowledge graph, the system can dynamically process complex entity relationships in government affairs text data, and at the same time can provide on-demand classification, on-demand horizontal and vertical query and storage of government affairs text data. BRIEF DESCRIPTION OF THE DRAWINGS

[0115] Figure 1 is the overall flow block diagram of the present invention;

[0116] The meanings of the reference numerals in the figure are as follows:

[0117] 1. Data acquisition and processing unit; 11. Data acquisition and cleaning module; 12. Entity recognition and division module; 13. Entity relationship extraction module; 2. First-level knowledge graph unit; 3. Second-level knowledge graph unit; 31. Second-level graph construction module; 32. Second-level graph optimization module; 4. Horizontal and vertical classification query unit. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0118] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0119] Embodiment 1: Please refer to Figure 1 As shown, a government affairs text data classification and processing system based on the knowledge graph is provided, including:

[0120] A data collection and processing unit 1, wherein the data collection and processing unit 1 collects and cleans government text data, uses a named entity recognition model to identify entities in the government text data, and uses an entity importance dual weight partitioning algorithm to divide key entities into high-key entities and low-key entities, and uses Stanford OpenIE technology to extract entity relationships;

[0121] The entity importance dual weight division algorithm is implemented based on TF-IDF technology and PageRank algorithm technology, and is used to evaluate the criticality of entities in government text data and divide entities into high-criticality entities and low-criticality entities;

[0122] In this embodiment, the data acquisition and processing unit 1 includes a data acquisition and cleaning module 11, an entity recognition and division module 12, and an entity relationship extraction module 13;

[0123] The data collection and cleaning module 11 is used to collect government text data from the data source and pre-process and clean the government text data;

[0124] In this embodiment, government text data is collected from data sources, and the data sources include: government websites, government announcements / government news releases, government documents, policy documents, government social media, news media news websites, and government databases;

[0125] The purpose of preprocessing and cleaning government text data is to ensure that the collected government text data can be effectively used in subsequent text analysis, entity recognition, relationship extraction, etc., as follows:

[0126] Use HTML parsers to remove advertisements, irrelevant text, and HTML tags from web pages; standardize the date format in the text; standardize different spelling formats: "government agencies" and "government agencies" into "government agencies"; use Chinese word segmentation tools to segment text and divide long texts into independent words or phrases; use sentence segmentation technology to split long paragraphs into smaller sentences;

[0127] Use text similarity calculation methods to identify repeated text content for deduplication; filter out noise words and stop words, such as "的", "是", "了" and other common words that are irrelevant to the analysis;

[0128] The entity recognition and division module 12 uses a named entity recognition model to identify entities in government text data, and uses an entity importance dual weight division algorithm to divide key entities into high-key entities and low-key entities;

[0129] In this embodiment, the named entity recognition model is based on BERT technology. The named entity recognition model belongs to an information extraction technology in the field of natural language processing and is used to efficiently recognize entities in government affairs texts, accurately extract key entities in the texts. Moreover, due to its strong context understanding ability, it can correctly recognize different types of entities in complex government affairs texts. The collected government affairs text data is used as input and fed into the named entity recognition model. The named entity recognition model will identify the named entities in the text and label the categories of the entities according to the language structure and entity relationships learned during training. The recognized entities will be processed subsequently through the entity importance dual-weight division algorithm to determine which entities belong to high-key entities and which belong to low-key entities; government agencies, policy-related entities, etc. may be classified as high-key entities, while some relatively minor and auxiliary entities may be classified as low-key entities.

[0130] The entity relationship extraction module 13 uses Stanford OpenIE technology to extract entity relationships from the text.

[0131] The entity recognition and division module 12 uses the named entity recognition model to recognize entities in the government affairs text data and uses the entity importance dual-weight division algorithm to divide the key entities into high-key entities and low-key entities. The specific method is as follows:

[0132] S1.2.1. Use the named entity recognition model to recognize entities in the government affairs text data and calculate the local importance weight and global importance weight of each entity:

[0133] ;

[0134] ;

[0135] Among them, is the set of government affairs text data; is the amount of government affairs text data; is the index of the government affairs text data; is the i-th government affairs text data; is the set of entities; is the total number of entities; is the entity index; is the j-th entity;

[0136] Calculate the local weight of the entity:

[0137] ;

[0138] Among them, is the entity local weight; is the TF-IDF technology operation; For the entity The number of occurrences in the government text data ; For the government text data containing the entity Quantity;

[0139] Calculate the global weight of the entity:

[0140] ;

[0141] Among them, For the entity Global weight; Is the damping coefficient; Is the entity index; Is the k-th entity; For the entity Out-degree; For the entity Global weight;

[0142] S1.2.2. For each entity, calculate the comprehensive score of each entity based on the local weight and global weight of the entity:

[0143] ;

[0144] Among them, For the entity Comprehensive score; Is the weighted coefficient of the local weight of the entity;

[0145] S1.2.3. According to the comprehensive score of each entity, divide each entity into high-key entities and low-key entities:

[0146] ;

[0147] Among them, Is the average comprehensive score of the entity; Is the set of all entities;

[0148] For the entity , if , then the entity Is a high-key entity; otherwise it is a low-key entity;

[0149] Finally, obtain the set of high-key entities: ;

[0150] Set of low-key entities: ;

[0151] Among them, Is the total number of high-key entities; Is the index of the high-key entity; is the th high - key entity; is the total number of low - key entities; is the index of the low - key entity; is the th low - key entity.

[0152] In this embodiment, first, a named - entity recognition model is used to identify entities in government - affair text data. Next, an entity - importance dual - weight partitioning algorithm is used to calculate the local weight and global weight of each entity, and key entities are divided into high - key entities and low - key entities;

[0153] TF - IDF is a commonly used weight - calculation method in text - data processing. TF represents the frequency of a word in a text, and IDF represents the rarity of the word in the entire document collection. The local weight of an entity can reflect the importance of the entity in a specific document, and the TF - IDF method is used to calculate the local importance of each entity;

[0154] PageRank calculates its global weight by considering the relationships between entities;

[0155] The entity - importance dual - weight partitioning algorithm combines the calculation of local weight and global weight to fully explore the importance of entities in a single document and the entire knowledge graph. By calculating the comprehensive score of each entity and comparing it with the average score, entities can be divided into high - key entities and low - key entities.

[0156] It also includes a first - level knowledge - graph unit 2. The first - level knowledge - graph unit 2 constructs a first - level government - affair knowledge graph based on high - key entities and high - key - entity relationships using a deep - graph - convolutional - spectrum construction algorithm for storing and managing high - key government - affair data and its data relationships;

[0157] The deep - graph - convolutional - spectrum construction algorithm is implemented based on a dynamic graph - convolutional network and the Louvain community - discovery algorithm for constructing a first - level government - affair knowledge graph;

[0158] In this embodiment, the first - level knowledge - graph unit 2 constructs a first - level government - affair knowledge graph based on high - key entities and high - key - entity relationships using a deep - graph - convolutional - spectrum construction algorithm. The specific method steps are as follows:

[0159] S2.1. Construct an initial first - level knowledge graph according to the high - key - entity set and high - key - entity relationships;

[0160] S2.2. On the basis of the initial first - level knowledge graph, introduce dynamic weights and a time - decay factor to construct an improved adjacency matrix, use a dynamic graph - convolutional network to update the node embeddings of each layer, and perform weighted processing on node relationships through the improved adjacency matrix;

[0161] In this embodiment, the dynamic weight refers to that when constructing a knowledge graph, the weights assigned to different entities or relationships change dynamically according to the actual situation. The dynamic weight will be adjusted with the passage of time, the occurrence of events, the change of policies or other external factors. For example, the importance of certain entities such as government departments and policy themes may change with the change of political decisions. At this time, the relative importance of entities in the graph can be reflected by dynamically adjusting the weights. In a deep graph convolutional network, the dynamic weight may involve real-time adjustment of the weights of each edge in the adjacency matrix, making the connection relationship with high-weight entities closer and helping the model focus on more important entities and their relationships.

[0162] The time decay factor refers to the mechanism by which the influence of certain entities or relationships gradually weakens over time. For government affairs text data, the influence of some policies or events is short-term, while the influence of some basic contents such as long-term policies and legal provisions is lasting. The time decay factor can adjust the weights of entities or relationships in the graph according to their historical information to ensure that the graph can reflect the importance of entities and relationships affected by time changes. The commonly used form of the time decay factor is exponential decay. As time increases, the decay factor makes older entities and relationships contribute less in the calculation.

[0163] Based on the initial first-level knowledge graph, introducing the dynamic weight and the time decay factor enables the government affairs knowledge graph to not only adapt to the changes of entities and relationships, but also automatically adjust the weights of each entity according to time changes, avoiding the distortion of the graph caused by outdated information.

[0164] S2.3. Calculate the node similarity based on the node embedding of the dynamic graph convolutional network, and perform community division on the nodes based on the Louvain community discovery algorithm to optimize the graph structure, calculate the modularity of the community division, and dynamically merge relevant communities according to the node similarity.

[0165] In this embodiment, the dynamic graph convolutional network is a graph convolutional network model applicable to the dynamic graph structure and time series data of dynamic graphs. Different from the traditional static graph convolutional network, the dynamic graph convolutional network can handle the dynamic changes of the graph structure, such as the addition, deletion of nodes and edges, and the change of weights, etc., and is adapted to the method of introducing the dynamic weight and the time decay factor based on the initial first-level knowledge graph.

[0166] The Louvain community detection algorithm is a community detection algorithm for graph structures, aiming to discover tightly connected subgraphs in a graph. It determines the community partition by optimizing the modularity of the graph. Modularity measures the tightness of connections between nodes in the graph. The goal is to assign nodes to communities with high modularity as much as possible, thereby enhancing the similarity of nodes within the community. Modularity is an indicator to measure the quality of graph partitioning, and a higher value indicates a more reasonable community partition. The Louvain community detection algorithm optimizes modularity through a greedy algorithm, continuously merging nodes until the optimal modularity is reached;

[0167] The deep graph convolutional spectrum construction algorithm combines the dynamic graph convolutional network and the Louvain community detection algorithm, enabling the government affairs knowledge graph to adapt to temporal changes and optimizing the graph structure through community detection.

[0168] S2.4. Repeatedly execute S2.2 and S2.3 until the change in modularity is less than a predetermined threshold, and finally obtain the first-level government affairs knowledge graph.

[0169] In this embodiment, the first-level knowledge graph unit 2 constructs a centralized and dynamically updated knowledge graph through high-key entities and their relationships, which is used to store and manage the core entities and their relationships of government affairs data, generating a graph that can represent entity and entity relationships. Subsequently, it is used to support data queries, relationship reasoning, information fusion, and classification processing across entities. By constructing the first-level government affairs knowledge graph, important entities in government affairs texts, such as government departments, policy clauses, key events, etc., and their mutual relationships are structured into a graph form. The graph can not only effectively store data but also be dynamically adjusted and updated according to the relationships between nodes.

[0170] In S2.1 of this embodiment, an initial first-level knowledge graph is constructed according to the high-key entity set and high-key entity relationships. The specific method is as follows:

[0171] The high-key entity set Define the node set, and the high-key entity relationship set Define it as the edge set :

[0172] ;

[0173] Among them, is the edge set; is the high-key entity relationship set; is the th high-key entity; is the th high-key entity; is the high-key entity and the high-key entity 's relationship strength;

[0174] Calculate the edge weight of each edge based on the comprehensive score of high - key entities:

[0175] ;

[0176] Among them, is the high - key entity and the high - key entity 's edge weight; is the comprehensive score of the high - key entity ; is the comprehensive score of the high - key entity ;

[0177] The set of edge weights between all high - key entities is ;

[0178] Generate the initial first - level knowledge graph:

[0179] ;

[0180] Among them, is the initial first - level knowledge graph;

[0181] In this embodiment S2.2, based on the initial first - level knowledge graph, introduce a dynamic weight and a time decay factor, construct an improved adjacency matrix, use a dynamic graph convolutional network to update the node embeddings of each layer, and perform weighted processing on the node relationships through the improved adjacency matrix. The specific method steps are as follows:

[0182] Define the adjacency matrix of the initial first - level knowledge graph as , and the element in the adjacency matrix ;

[0183] Introduce the time decay factor to improve the adjacency matrix :

[0184] ;

[0185] Among them, is the improved adjacency matrix; is the time decay factor; is the current time; is the high - key entity and the high - key entity relationship strength 's last update time;

[0186] Based on the improved adjacency matrix , use a dynamic graph convolutional network to update the node embeddings of each layer:

[0187] ;

[0188] Among them, is the node embedding of the th layer; is the ReLU activation function; is the degree matrix of the improved adjacency matrix; is the node embedding of the th layer; is the learnable parameter matrix.

[0189] In this embodiment S2.3, the node similarity is calculated based on the dynamic graph convolutional network, and the nodes are divided into communities based on the Louvain community discovery algorithm to optimize the graph structure, the modularity of the community division is calculated, and the relevant communities are dynamically merged according to the node similarity. The specific method steps are as follows:

[0190] Calculate the node similarity between the high-key entity and the high-key entity :

[0191] ;

[0192] Among them, is the node similarity between the high-key entity and the high-key entity ; is the final layer of the convolutional layer of the dynamic graph convolutional network; The node embedding of the final layer L of the high-key entity ; The node embedding of the final layer L of the high-key entity ;

[0193] Based on the Louvain community discovery algorithm, divide the nodes into communities to optimize the graph structure, and calculate the modularity of the community division:

[0194] ;

[0195] ;

[0196] Among them, is the modularity; is the sum of the edge weights of the initial first-level knowledge graph; is the expected connection strength; is the community where the high-key entity is located; is the community where the high-key entity is located; is the Kronecker delta function;

[0197] Kronecker delta function The calculation logic is as follows: when and belong to the same community, the value is 1, otherwise it is 0;

[0198] Set the merging threshold , if two communities and satisfy the following equation in terms of similarity:

[0199] ;

[0200] then merge the two communities and into a single community;

[0201] Among them, is the merging threshold; is community a; is community b;

[0202] In this embodiment S2.4, repeatedly iterate S2.2 and S2.3 until the change in modularity is less than a predetermined threshold, and finally obtain the first-level government affairs knowledge graph .

[0203] In this implementation, when the change rate ΔQ of the modularity is less than the set threshold , stop the iteration and obtain the final first-level government affairs knowledge graph .

[0204] It further includes a secondary knowledge graph unit 3. The secondary knowledge graph unit 3 is based on low-key entities and low-key entity relationships, takes the entity nodes of the first-level government affairs knowledge graph as the center, uses a cross-graph embedding alignment reasoning algorithm to construct multiple secondary government affairs knowledge graphs, and calculates the correlation degree of two secondary government affairs knowledge graphs for storing and managing low-key government affairs data and its data relationships;

[0205] The cross-graph embedding alignment reasoning algorithm is implemented based on knowledge graph embedding technology and cross-graph alignment technology, and is used to construct multiple secondary government affairs knowledge graphs and align and reason about entities and entity relationships;

[0206] In this embodiment, the secondary knowledge graph unit 3 includes a secondary graph construction module 31 and a secondary graph optimization module 32;

[0207] Among them, the secondary graph construction module 31 is based on low-key entities and low-key entity relationships, takes the entity nodes of the first-level government affairs knowledge graph as the center, and uses a cross-graph embedding alignment reasoning algorithm to construct multiple secondary government affairs knowledge graphs;

[0208] In this embodiment, low-critical entities usually have less direct influence in government texts. By focusing on low-critical entities and their relationships, the secondary knowledge graph can mine potential information from a broader context, thus supplementing the main relationships between high-critical entities and ensuring the comprehensiveness of the knowledge graph; the relationships and interactions between low-critical entities usually involve fewer conflicts or competitions, so more details can be provided for the graph; taking the entity nodes of the primary knowledge graph as the center for constructing the secondary knowledge graph can ensure the contextual consistency of the secondary knowledge graph, avoid starting from completely different graph structures, and maintain the integrity of the knowledge structure; the cross-graph embedding alignment inference algorithm can map the data from different graphs into the same space through cross-graph alignment, handle the data differences between different graphs, and establish connections in knowledge graphs at different levels;

[0209] By taking the relationships between low-critical entities and low-critical entities as the core of the secondary knowledge graph while keeping the high-critical entity nodes in the primary graph unchanged, multi-level support can be provided for the classification and processing of government text data; knowledge graphs at different levels, namely the primary and secondary levels, process information classified at different levels respectively; through the construction of multi-level knowledge graphs, more refined classification of government data can be achieved; the primary knowledge graph provides an overview of high-critical entities, while the secondary knowledge graph delves into the specific connections of low-critical entities. The combination of the two can provide richer and more comprehensive feature information for the classification model.

[0210] The secondary knowledge graph optimization module 32 calculates the correlation degree of two secondary government knowledge graphs to eliminate redundant secondary government knowledge graphs.

[0211] In this embodiment, the secondary knowledge graph construction module 31 constructs multiple secondary government knowledge graphs based on low-critical entities and low-critical entity relationships, with the entity nodes of the primary government knowledge graph as the center, using the cross-graph embedding alignment inference algorithm. The specific method steps are as follows:

[0212] S3.1.1. Taking each high-critical entity node of the primary government knowledge graph as the center, construct an initial secondary government knowledge graph:

[0213] For the high-critical entity node , define the set of low-critical entities related to the high-critical entity node as the set of low-critical entity nodes , and define the set of low-critical entity relationships as the set of low-critical entity edges ;

[0214] Among them, is the high-critical entity index; , representing the set of low-critical entities related to the high-critical entity node ; , representing the relationships between low-critical entities;

[0215] Calculate the edge weight of each edge based on the comprehensive scores of low-critical entities:

[0216] ;

[0217] Among them, and are both the indices of low-critical entities related to the high-critical entity node ; is the edge weight between the low-critical entity and the low-critical entity ; is the comprehensive score of the low-critical entity ; is the comprehensive score of the low-critical entity ;

[0218] The set of edge weights between all low-critical entities related to the high-critical entity node is ;

[0219] Generate the initial secondary knowledge graph with the high-critical entity node :

[0220] ;

[0221] Among them, is the initial secondary knowledge graph;

[0222] S3.1.2. Use the node embedding of the primary government affairs knowledge graph to guide the semantic alignment of the initial secondary knowledge graph:

[0223] Obtain the embedding vector of the high-critical entity node from the primary government affairs knowledge graph, and define the projection matrix to calculate the embedding vector of the aligned high-critical entity node:

[0224] ;

[0225] Among them, is the embedding vector of the aligned high-critical entity node ; is the embedding vector of the high-critical entity node; is the projection matrix;

[0226] In this embodiment, the projection matrix is used to map the embedding vectors of the high-critical entity nodes in the primary government knowledge graph to the appropriate semantic space in the secondary knowledge graph to achieve the goal of semantic alignment; the projection matrix P is a transformation matrix that adjusts the dimension or structure of the embedding vector of the high-critical entity node by multiplying it with the embedding vector, so that it has a more appropriate representation in the semantic space of the secondary knowledge graph, thereby achieving semantic consistency between the primary and secondary knowledge graphs. Through the operation of the projection matrix P, the semantic features of the high-critical entity nodes in the primary knowledge graph can be corresponding and mapped in the secondary knowledge graph, making the node embeddings of the two knowledge graphs have better semantic compatibility;

[0227] Define the initial embedding vectors of the low-critical entities in the initial secondary knowledge graph as ;

[0228] In the initial secondary knowledge graph, calculate the attention weights of each low-critical entity and its corresponding high-critical entity node in the primary government knowledge graph based on the initial embedding vectors of the low-critical entities:

[0229] ;

[0230] where, is the attention weight of the low-critical entity node and the high-critical entity node ;

[0231] is the cosine similarity;

[0232] Based on the embedding vectors of the aligned high-critical entity nodes and the attention weights, align the initial embedding vectors of the low-critical entities with the aligned high-critical entity nodes, and calculate the aligned embedding vectors of the low-critical entities:

[0233] ;

[0234] where, is the aligned embedding vector of the low-critical entity ;

[0235] Update the aligned embedding vectors of the low-critical entities to the initial secondary knowledge graph to obtain the secondary government knowledge graph .

[0236] In this embodiment, according to the operation steps of the secondary graph construction module 31, all secondary government knowledge graphs centered on all nodes of the primary government knowledge graph can be obtained in the same way.

[0237] The secondary graph optimization module 32 calculates the correlation degree between two secondary government affairs knowledge graphs, which is used to eliminate redundant secondary government affairs knowledge graphs. The specific method steps are as follows:

[0238] Calculate the correlation degree between two secondary government affairs knowledge graphs:

[0239] ;

[0240] Among them, is the secondary government affairs knowledge graph constructed with the high-key entity node as the center; is the high-key entity index; , representing the set of low-key entities related to the high-key entity node ; is the secondary government affairs knowledge graph constructed with the high-key entity node as the center; is the correlation degree between the secondary government affairs knowledge graphs and ;

[0241] Set the secondary graph merging threshold , and determine whether two secondary government affairs knowledge graphs need to be merged, specifically as follows:

[0242] If , then the secondary government affairs knowledge graphs and are merged into a secondary government affairs knowledge graph .

[0243] It also includes a horizontal and vertical classification query unit 4. The horizontal and vertical classification query unit 4 provides a query interface for cross-node and cross-graph horizontal and vertical data queries. Embodiment 2

[0244] A government affair text data classification processing method based on a knowledge graph, which is used for the government affair text data classification processing system in any one of the above, includes the following steps:

[0245] S10.1. Collect and clean government affair text data, use a named entity recognition model to identify entities in the government affair text data, use the entity importance double-weight division algorithm to divide key entities into high-key entities and low-key entities, and use the Stanford OpenIE technology to extract entity relationships;

[0246] S10.2. Based on the high-key entities and high-key entity relationships, use the deep graph convolutional graph construction algorithm to construct a primary government affair knowledge graph;

[0247] S10.3. Based on the low-key entities and low-key entity relationships, centered around the entity nodes of the first-level government affairs knowledge graph, use the cross-graph embedding alignment inference algorithm to construct multiple second-level government affairs knowledge graphs, and calculate the correlation degree between the two second-level government affairs knowledge graphs;

[0248] S10.4. Use the query interface to perform horizontal and vertical data queries across nodes and across graphs. The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.

Claims

1. The government text data classification processing system based on knowledge graph is characterized by: include: A data collection and processing unit (1), wherein the data collection and processing unit (1) collects and cleans government text data, identifies entities in the government text data using a named entity recognition model, divides key entities into high-key entities and low-key entities using an entity importance dual-weight division algorithm, and extracts entity relationships using Stanford OpenIE technology; The entity importance dual weight division algorithm is implemented based on TF-IDF technology and PageRank algorithm technology, and is used to evaluate the criticality of entities in government text data and divide entities into high-criticality entities and low-criticality entities; The first-level knowledge graph unit (2) is based on high-key entities and high-key entity relationships, and uses a deep graph convolutional graph construction algorithm to construct a first-level government affairs knowledge graph for storing and managing high-key government affairs data and their data relationships. The specific method steps are as follows: S2.

1. Construct an initial first-level knowledge graph based on high-key entity sets and high-key entity relationships; S2.

2. Based on the initial level 1 knowledge graph, dynamic weights and time decay factors are introduced to construct an improved adjacency matrix. The dynamic graph convolutional network is used to update the node embedding of each layer, and the node relationship is weighted through the improved adjacency matrix. S2.3, calculate node similarity based on node embedding of dynamic graph convolutional network, divide nodes into communities based on Louvain community discovery algorithm, optimize graph structure, calculate modularity of community division, and dynamically merge related communities based on node similarity; S2.4, repeatedly iterate and execute S2.2 and S2.3 until the modularity change is less than the predetermined threshold, and finally obtain the first-level government affairs knowledge graph; A secondary knowledge graph unit (3), wherein the secondary knowledge graph unit (3) is based on low-key entities and low-key entity relationships, takes the entity nodes of the primary government knowledge graph as the center, uses a cross-graph embedding alignment reasoning algorithm to construct multiple secondary government knowledge graphs, and calculates the correlation between two secondary government knowledge graphs, so as to store and manage low-key government data and their data relationships; A horizontal and vertical classification query unit (4), wherein the horizontal and vertical classification query unit (4) provides a query interface for performing horizontal and vertical data queries across nodes and graphs.

2. The government text data classification processing system based on knowledge graph according to claim 1 is characterized by: The data acquisition and processing unit (1) comprises a data acquisition and cleaning module (11), an entity recognition and division module (12) and an entity relationship extraction module (13); The data collection and cleaning module (11) is used to collect government text data from a data source, and to pre-process and clean the government text data; The entity recognition and classification module (12) uses a named entity recognition model to recognize entities in government text data, and uses an entity importance dual weight classification algorithm to classify key entities into high-key entities and low-key entities; The entity relationship extraction module (13) uses Stanford OpenIE technology to extract entity relationships from text.

3. The government text data classification processing system based on knowledge graph according to claim 2 is characterized by: The entity recognition and classification module (12) uses a named entity recognition model to recognize entities in government text data, and uses an entity importance dual weight classification algorithm to classify key entities into high-key entities and low-key entities. The specific method is as follows: S1.2.

1. Use the named entity recognition model to identify entities in government text data and calculate the local importance weight and global importance weight of each entity: ; ; in, It is a collection of government text data; The amount of government text data; Indexing of government text data; For the Government text data; is a collection of entities; is the total number of entities; Index for the entity; is the jth entity; Calculate entity local weight: ; in, For Entity Local weight; Operate for TF-IDF technology; For Entity In government text data The number of times it appears in To contain entities The amount of government text data; Calculate the entity global weight: ; in, For Entity Global weight; is the damping coefficient; Index for the entity; is the kth entity; For Entity The out-degree of For Entity Global weight; S1.2.

2. For each entity, calculate the comprehensive score of each entity based on the entity local weight and entity global weight: ; in, For Entity The comprehensive score of is the weighting coefficient of the entity local weight; S1.2.

3. Based on the comprehensive score of each entity, each entity is divided into high-key entity and low-key entity: ; in, is the average comprehensive score of the entity; is the set of all entities; For entities ,like , then the entity is a high-key entity; otherwise, it is a low-key entity; Finally, we get a set of high-key entities: ; Low key entity collection: ; in, is the total number of high-key entities; Index for high-key entities; For the High-key entities; is the total number of low-key entities; Index for low-key entities; For the A low-key entity.

4. The government text data classification processing system based on knowledge graph according to claim 3 is characterized by: The deep graph convolutional graph construction algorithm is implemented based on a dynamic graph convolutional network and the Louvain community discovery algorithm, and is used to construct a first-level government knowledge graph.

5. The government text data classification processing system based on knowledge graph according to claim 4 is characterized by: In S2.1, an initial first-level knowledge graph is constructed based on a set of high-key entities and high-key entity relationships. The specific method is as follows: Group high-key entities Define a node set to set high-key entity relationships Defined as an edge set : ; in, is the edge set; is a set of high-key entity relationships; For the High-key entities; For the High-key entities; High key entity and high-key entities The strength of the relationship; Calculate the edge weight of each edge based on the comprehensive score of high-key entities: ; in, High key entity and high-key entities The edge weight of High key entity The comprehensive score of High key entity The comprehensive score of The set of edge weights between all high-key entities is ; Generate the initial first-level knowledge graph: ; in, is the initial level one knowledge graph; In S2.2, based on the initial level 1 knowledge graph, dynamic weights and time decay factors are introduced to construct an improved adjacency matrix, and each layer of node embedding is updated using a dynamic graph convolutional network. The node relationship is weighted by the improved adjacency matrix. The specific steps are as follows: The adjacency matrix of the initial level knowledge graph is defined as , and the adjacency matrix Elements in ; Introducing the time decay factor Improved adjacency matrix : ; in, is the improved adjacency matrix; is the time decay factor; is the current time; High key entity and high-key entities Relationship Strength The last update time of Based on the improved adjacency matrix , using a dynamic graph convolutional network to update each layer of node embedding: ; in, For the Node embedding of layers; is the ReLU activation function; is the degree matrix of the improved adjacency matrix; For the Node embedding of layers; is the learnable parameter matrix.

6. The government text data classification processing system based on knowledge graph according to claim 5 is characterized by: In S2.3, node similarity is calculated based on node embedding of a dynamic graph convolutional network, and nodes are divided into communities based on the Louvain community discovery algorithm, the graph structure is optimized, the modularity of the community division is calculated, and related communities are dynamically merged according to the node similarity. The specific method steps are as follows: Calculate high-key entities and high-key entities Node similarity: ; in, High key entity and high-key entities Node similarity of It is the final layer of the convolutional layer of the dynamic graph convolutional network; High key entities The node embedding of the final layer L; High key entities The node embedding of the final layer L; Based on the Louvain community discovery algorithm, the nodes are divided into communities, the graph structure is optimized, and the modularity of the community division is calculated: ; ; in, is modularity; is the sum of the edge weights of the initial first-level knowledge graph; is the expected connection strength; High key entity The community where you live; High key entity The community where you live; is the Kronecker delta function; Setting the merge threshold If two communities and The similarity satisfies the following equation: ; Merge the two communities and for a separate community; in, is the merge threshold; For community a; for community b; In S2.4, S2.2 and S2.3 are repeatedly iterated until the modularity change is less than a predetermined threshold, and finally a first-level government affairs knowledge graph is obtained. .

7. The government text data classification processing system based on knowledge graph according to claim 6 is characterized by: The secondary knowledge graph unit (3) includes a secondary graph construction module (31) and a secondary graph optimization module (32); The secondary graph construction module (31) is based on low-key entities and low-key entity relationships, takes the entity nodes of the primary government knowledge graph as the center, and uses a cross-graph embedding alignment reasoning algorithm to construct multiple secondary government knowledge graphs; The secondary graph optimization module (32) calculates the correlation between two secondary government knowledge graphs to eliminate redundant secondary government knowledge graphs.

8. The government text data classification processing system based on knowledge graph according to claim 7 is characterized by: The secondary graph construction module (31) is based on low-key entities and low-key entity relationships, takes the entity nodes of the primary government knowledge graph as the center, and uses a cross-graph embedding alignment reasoning algorithm to construct multiple secondary government knowledge graphs. The specific method steps are as follows: S3.1.

1. First-level government knowledge graph With each high-key entity node as the center, construct the initial secondary government knowledge graph: For high-critical entity nodes , will be associated with high-key entity nodes The related low-key entity definition set is the low-key entity node set , the low-key entity relationship set Defined as a set of low-key entity edges ; in, Index for high-key entities; , indicating that the entity nodes with high key A collection of related low-key entities; , which represents the relationship between low-key entities; Calculate the edge weight of each edge based on the combined score of the low-key entities: ; in, , Both are related to high-key entity nodes Related low-key entity indexes; For low critical entities and low critical entities The edge weight of For low critical entities The comprehensive score of For low critical entities The comprehensive score of All entity nodes with high key The set of edge weights between related low-key entities is ; Generate high-key entity nodes The initial secondary knowledge graph of: ; in, is the initial secondary knowledge graph; S3.1.

2. Use the node embedding of the primary government knowledge graph to guide the semantic alignment of the initial secondary knowledge graph: Get high-key entity nodes from the first-level government knowledge graph The embedding vector of the aligned high-key entity node is calculated by defining the projection matrix: ; in, High key entity nodes after alignment The embedding vector of is the embedding vector of high-key entity nodes; is the projection matrix; Defining low-key entities in the initial secondary knowledge graph The initial embedding vector of ; In the initial secondary knowledge graph, the attention weight of each low-key entity and its corresponding high-key entity node in the primary government knowledge graph is calculated based on the initial embedding vector of the low-key entity: ; in, Low-key entity node Nodes with high criticality The attention weight of is the cosine similarity; Based on the embedding vector and attention weight of the aligned high-key entity node, the initial embedding vector of the low-key entity is aligned with the aligned high-key entity node, and the aligned embedding vector of the low-key entity is calculated: ; in, For low critical entities The aligned embedding vector of Lower key entities The aligned embedding vector is updated to the initial secondary knowledge graph to obtain the secondary government knowledge graph .

9. The government text data classification processing system based on knowledge graph according to claim 8 is characterized by: The secondary graph optimization module (32) calculates the correlation between two secondary government knowledge graphs to eliminate redundant secondary government knowledge graphs. The specific method steps are as follows: Calculate the correlation between two secondary government knowledge graphs: ; in, For high-key entity nodes The secondary government affairs knowledge graph built by the center; Index for high-key entities; , indicating that the entity nodes with high key A collection of related low-key entities; For high-key entity nodes The secondary government affairs knowledge graph built by the center; Secondary government knowledge graph and degree of relevance; Set the secondary graph merging threshold , determine whether the two secondary government knowledge graphs need to be merged, as follows: like , then the secondary government knowledge graph and Execute and merge into a secondary government knowledge graph .

10. A method for classifying and processing government text data based on a knowledge graph, used in a system for classifying and processing government text data based on a knowledge graph as claimed in any one of claims 1 to 9, characterized in that: The steps include: S10.

1. Collect and clean government text data, use named entity recognition model to identify entities in government text data, and use entity importance dual weight partitioning algorithm to divide key entities into high key entities and low key entities, and use Stanford OpenIE technology to extract entity relationships; S10.

2. Based on high-key entities and high-key entity relationships, a deep graph convolutional graph construction algorithm is used to construct a first-level government knowledge graph; S10.

3. Based on low-key entities and low-key entity relationships, with the entity nodes of the first-level government knowledge graph as the center, use the cross-graph embedding alignment reasoning algorithm to construct multiple second-level government knowledge graphs, and calculate the correlation between two second-level government knowledge graphs; S10.

4. Use the query interface to perform horizontal and vertical data queries across nodes and graphs.

Citation Information

Patent Citations

  • Data processing method and device based on knowledge graph and computer equipment

    CN111061859A

  • Domain knowledge graph updating method and system based on power grid equipment

    CN117574898A