Modern service industry policy quantitative analysis method and system based on knowledge graph
By constructing knowledge graphs and automation graphs, combined with dynamic theme membership allocation algorithms, the systematic shortage and high cost of policy text analysis in the existing technology is solved, multi-dimensional quantitative analysis and full-process automation are realized, and the depth and timeliness of the analysis results are significantly improved.
Patent Information
- Application Number
- CN202510285025.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-10
AI Technical Summary
The existing technology lacks a systematic quantitative framework in policy text analysis, has limited keyword extraction accuracy, poor topic modeling stability and single data source, resulting in one-sided and lack of depth in the analysis results, and requires a lot of manual labeling, which is costly.
The modern service industry policy quantitative analysis method based on knowledge graph is adopted, and the full process automation and multi-dimensional quantitative analysis are achieved by building a knowledge graph, extracting subgraphs and nodes of the theme files, assigning topics, and calculating indicators such as consistency, uniqueness, focus and responsiveness.
It has achieved in-depth mining of policy texts, solved the problem that traditional methods are difficult to deal with cross-themes and multi-source data integration, broke through the limitations of a single indicator, provided multi-dimensional quantitative basis, reduced the need for manual intervention and development costs, and improved the timeliness of analysis results.
Smart Images

Figure CN120124628A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of knowledge graphs, and relates to a method and system for quantitative analysis of modern service industry policies based on knowledge graphs. Background Art
[0002] Policy texts, as important tools for government decision-making, usually carry rich information such as policy goals, implementation paths, and enforcement measures for industrial development. At the same time, with the rapid development of the modern service industry, the number of policy texts is also increasing day by day. Therefore, there is an urgent need for a tool method for effectively analyzing and quantifying the information in modern service industry policy texts. Policy texts (especially modern service industry policies) mainly have three characteristics: ① Unstructured. Policy texts are usually in the form of natural language, involving a large amount of domain knowledge and rich semantic content, making it difficult to directly conduct quantitative analysis. ② Multi-source heterogeneity. The sources of modern service industry policy texts are extensive, including government bulletins, planning documents, and news reports. The data from different sources vary significantly in format and content, increasing the difficulty of analysis. ③ Diverse analysis requirements. The analysis requirements for policy texts are becoming increasingly diverse, such as the consistency, attention, and implementation effectiveness of policy content.
[0003] Currently, in policy text analysis, the main technologies used are mainly keyword extraction and topic modeling methods. The keyword extraction method mainly extracts keywords in policy texts using the TF-IDF method on the basis of preprocessing text data, and analyzes the characteristic attributes of policy texts by constructing a co-occurrence matrix and performing clustering analysis. In addition, the method based on keyword extraction usually uses the PMC policy analysis framework to further quantify the policy. The topic modeling method mainly conducts quantitative analysis through the LDA (Latent Dirichlet Allocation) topic modeling method. Similarly, the policy text is processed using TF-IDF, and then the semantic information of the document content is extracted using LDA topic modeling.
[0004] The existing technologies for policy text analysis mainly focus on the uniqueness and similarity of the text: in the field of text similarity, two major categories of methods are mainly adopted: traditional similarity measurement methods based on word frequency and character operations, and semantic embedding models based on deep learning. The former measures similarity by calculating the lexical overlap or character edit distance between texts. Among them, Jaccard similarity depends on the ratio of the intersection and union of word sets and is applicable to short text or keyword matching scenarios; cosine similarity maps texts by constructing a word frequency vector space model (such as TF-IDF weighting) and calculates the cosine value of the angle between vectors, which is suitable for global semantic matching of long texts; edit distance measures the difference in text form by calculating the minimum number of character operations between texts and is often used for spelling correction or short string comparison. Another category of methods relies on deep learning models, such as language models like BERT and SBERT, which generate context-aware semantic vectors by capturing the context-dependent relationships between words. By calculating the cosine similarity or Euclidean distance between vectors, the semantic similarity of texts can be measured more accurately. This method is particularly suitable for solving problems such as synonym replacement and semantic ambiguity and can better handle the similarity analysis of complex texts and long texts.
[0005] In terms of text uniqueness measurement, traditional statistical methods measure the uniqueness of a text by counting the proportion of rare words and the density of domain-specific terms in the text. However, this method is limited to the lexical level and cannot distinguish semantic innovation. Similarity distribution difference analysis measures the uniqueness of a text by constructing a similarity matrix of multiple texts and analyzing the distribution characteristics of similarity values, and uses statistical indicators such as mean, variance, and KL divergence to measure the uniqueness of the text. This method reflects the difference between a single text and a reference set from a group perspective and can indirectly infer the innovation of the text. The existing technologies mainly achieve text similarity and uniqueness measurement through formalized lexical or character matching methods, semantic space modeling, and group distribution analysis. These methods provide basic tools for policy text analysis, but they have not effectively combined multi-dimensional indicators (such as consistency, focus) with structured semantic networks (such as knowledge graphs), nor have they been able to handle the joint analysis problem of multi-source heterogeneous data well.
[0006] The current text similarity and policy quantification analysis technology has achieved basic applications in multiple fields. Based on methods such as word frequency statistics (e.g., TF-IDF), topic modeling (e.g., LDA), and character operations (e.g., edit distance), it can initially achieve keyword extraction, topic clustering, and formal difference analysis of policy texts, meeting the basic policy comparison needs. For example, the Jaccard similarity can quickly screen policy clauses with high repetition, while the LDA model can roughly divide policy topics, providing a directional reference for researchers. In addition, traditional methods still have certain practical value in scenarios such as short text processing and structured data matching. However, with the complication and refinement of policy analysis requirements, the limitations of existing technologies in terms of technical depth, cost-effectiveness, and processing efficiency have gradually emerged, specifically manifested as the following three aspects of defects:
[0007] Firstly, in terms of technology: First, the lack of a systematic quantification framework: Existing methods mostly focus on a single indicator (such as policy consistency), without integrating multi-dimensional analyses such as policy focus and response intensity, resulting in one-sided quantification results. Second, the accuracy of keyword extraction is insufficient: Traditional methods such as TF-IDF only rely on word frequency statistics and cannot distinguish semantic depth. For example, two texts that mention "digital economy" with the same frequency (one details support measures and the other only mentions it generally) will be given the same weight, leading to distorted attention measurement. Third, the stability of topic modeling is poor: The topics extracted by topic modeling technologies such as LDA are highly random and difficult to accurately match the analysis requirements. This is mainly because LDA generates topics iteratively through Dirichlet distribution and Gibbs sampling, but the setting of initial values, fluctuations in term co-occurrence, and cross-semantic interference easily lead to topic drift, thereby making the overall topic modeling less stable. This stability defect not only weakens the clarity of topic boundaries but may also mislead policy quantification analysis and is difficult to support precise governance decisions. Fourth, the single source of data: Existing technologies mainly analyze policy planning texts and ignore public opinion data such as news media, resulting in the lack of evaluation of policy response intensity.
[0008] Secondly, due to the above technical defects, the cost of manual intervention is high: Traditional methods require a large amount of manual annotation (such as topic definition and sensitive word desensitization), especially when dealing with multi-source heterogeneous data (policy documents, news texts), the human and time costs increase significantly. The cost of model adaptation is high: For different policy fields (such as digital economy and green economy), it is necessary to repeatedly train dedicated models, lacking a general framework, resulting in an increase in development and maintenance costs.
[0009] At the same time, the above defects will lead to low efficiency in integrating multi-source data. The formats and structures of policy documents and news texts are very different. Traditional methods need to be processed in stages (such as format conversion and data cleaning), with redundant processes and time-consuming. Secondly, the dynamic update is lagging: Existing technologies rely on static models and it is difficult to incorporate new data (such as public opinion dynamics) in real time, resulting in insufficient response timeliness.
[0010] Although certain progress has been made in the basic applications of current text similarity, uniqueness measurement, and policy quantitative analysis techniques, there are still significant deficiencies in terms of technology, cost, and efficiency. At the technical level, there is a lack of a systematic quantitative framework, the accuracy of keyword extraction is limited, the stability of topic modeling is poor, and the data source is single, resulting in one-sided and shallow analysis results. In terms of cost, traditional methods require a large amount of manual intervention and repeated training of dedicated models, increasing the development and maintenance costs. In terms of efficiency, existing technologies lag in the integration and dynamic update of multi-source data, leading to redundant processing procedures and untimely analysis results. Therefore, improving the accuracy of the technology, reducing costs, and enhancing processing efficiency have become urgent problems to be solved. Summary of the Invention
[0011] The purpose of the present invention is to solve the problems in the prior art, such as the lack of a systematic quantitative framework, limited accuracy of keyword extraction, poor stability of topic modeling, single data source, resulting in one-sided and shallow analysis results, and the need for a large amount of manual annotation with high costs. The present invention provides a method and system for quantitative analysis of modern service industry policies based on a knowledge graph.
[0012] To achieve the above object, the present invention adopts the following technical solutions:
[0013] A method for quantitative analysis of modern service industry policies based on a knowledge graph, comprising the following steps:
[0014] Construct a knowledge graph based on relevant policies of the modern service industry;
[0015] Determine the theme files, extract the sub-graphs of the theme files and the corresponding nodes based on the knowledge graph, assign corresponding themes to the extracted nodes respectively, and obtain the theme assignment results;
[0016] Based on the theme assignment results, calculate indicators from aspects such as consistency, uniqueness, policy focus, and policy responsiveness to obtain the quantitative analysis results.
[0017] A further improvement of the present invention lies in:
[0018] The construction of the knowledge graph based on relevant policies of the modern service industry includes:
[0019] Collect relevant policy documents of the modern service industry and preprocess the collected policy documents;
[0020] Send the preprocessed files to GraphRAG to generate a knowledge graph containing several entity nodes, relationships, communities, and summaries.
[0021] The step of respectively assigning corresponding themes to the extracted nodes to obtain the theme assignment results includes:
[0022] For a certain policy node i, calculate its similarity with all topic text nodes, and obtain the top N nodes closest to node i;
[0023] Statistically analyze the topic distribution of the N nodes, and calculate the membership degree of policy node i to topic j;
[0024] According to the calculation results of the membership degree, distinguish the topics corresponding to these N nodes respectively, and obtain the topic assignment result.
[0025] Calculate the indicators for consistency, including:
[0026] Select the top k nodes with the highest topic membership degrees in the policy document and the planning document respectively. Assume that the policy document nodes are {p 1 , p 2 , p 3 ,..., p k}, and the planning document nodes are {g 1 , g 2 , g 3 ,... g k};
[0027] Calculate the similarity between each node p i in P and each node g j in G to form a similarity matrix:
[0028] Calculate the consistency of the policy through the similarity matrix:
[0029]
[0030] where p i represents the policy document node; gj represents the planning document node.
[0031] Calculate the indicators for uniqueness, including:
[0032] Calculate the similarity matrices of the policy text topics and the planning text topics respectively;
[0033] Utilize the distribution differences of the KL divergence matrices.
[0034] Calculate the indicators for policy focus, including:
[0035] Calculate the eigenvector centrality and PageRank value of all topic-related nodes;
[0036] Combine the topic membership degrees of the nodes to continue weighted and normalized processing of the eigenvector centrality and PageRank value to obtain the attention focus of a specific topic.
[0037] Calculate the indicators for the responsiveness of the policy, including:
[0038] Calculate the eigenvector centrality and PageRank value of the nodes related to the theme;
[0039] Combine the theme membership of the nodes to continue the weighted and normalization processing of the eigenvector centrality and PageRank value, and obtain the calculation result of the policy responsiveness.
[0040] A modern service industry policy quantitative analysis system based on a knowledge graph, including:
[0041] A knowledge graph construction module for constructing a knowledge graph based on modern service industry-related policies;
[0042] A theme assignment module for determining theme files, extracting subgraphs of the theme files and nodes corresponding to the theme files based on the knowledge graph, respectively assigning corresponding themes to the extracted nodes, and obtaining theme assignment results;
[0043] A quantitative analysis module for calculating indicators from aspects of consistency, uniqueness, policy focus, and policy responsiveness based on the theme assignment results, and obtaining quantitative analysis results.
[0044] A terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of any method of the present invention are implemented.
[0045] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any method of the present invention are implemented.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] The present invention discloses a modern service industry policy quantitative analysis method based on a knowledge graph, which combines the knowledge graph with policy text analysis, realizes the in-depth mining of the node association pattern and weight distribution in the policy text, solves the problem that traditional methods are difficult to handle cross-topics and multi-source data integration, and at the same time, calculates indicators from aspects of consistency, uniqueness, policy focus, and policy responsiveness, breaks through the limitation of traditional methods relying on a single indicator, can comprehensively evaluate the scientificity and effectiveness of policy texts, provides a multi-dimensional quantitative basis for policy design and optimization, and avoids the one-sidedness of analysis results.
[0048] Furthermore, in the present invention, a knowledge graph is obtained through GraphRAG, which reduces the need for manual intervention and realizes the full-process automation from unstructured text to a structured knowledge graph.
[0049] Furthermore, in the present invention, a dynamic allocation algorithm is proposed by combining the KNN classification idea and the concept of topic membership, realizing the soft clustering of nodes. This method effectively alleviates the problem of distortion of hard clustering results caused by semantic diversity and significantly improves the accuracy and flexibility of topic classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0051] Figure 1 It is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated in the drawings here can be arranged and designed in various different configurations.
[0053] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0054] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0055] The present invention will be further described in detail below with reference to the drawings:
[0056] See Figure 1, this embodiment discloses a method for quantitative analysis of modern service industry policies based on a knowledge graph. By introducing a four-dimensional quantitative framework, knowledge graph technology, and a dynamic allocation algorithm, this embodiment solves the problems of the lack of multi-dimensional measurement quantification and deep semantic analysis in existing technologies for policy analysis. By analyzing the consistency, focus, responsiveness, and uniqueness of policies, it fills the gap in multi-dimensional quantification of traditional methods. By adopting knowledge graph technology, the present invention can accurately extract entities, relationships, and attributes in policy texts, and through constructing a dynamic knowledge graph, deeply analyze the implicit semantics in the texts, improving the accuracy of policy analysis. In addition, this embodiment also designs a dynamic allocation algorithm for topic membership, solves the problem of unstable topics in the LDA method, significantly improves the accuracy of topic classification, and makes policy analysis more scientific and comprehensive.
[0057] At the same time, by automatically constructing a knowledge graph and adopting a general graph computing framework, the need for manual intervention is effectively reduced, and the applicability of the model is improved, solving the problems of high cost and low generality in processing multi-source heterogeneous data in existing technologies. By automatically constructing a knowledge graph, the need for manual intervention is significantly reduced, making the structured processing process of policy texts more efficient. In addition, by adopting a general graph computing framework, it can adapt to policy analysis in different fields, avoid the problem of repeated model training, and thus greatly reduce the development and maintenance costs; improve the applicability and flexibility of the model, and also achieve the integration of cross-domain data processing, reducing the complexity and cost of long-term maintenance.
[0058] This embodiment introduces graph computing technology. By applying algorithms such as PageRank and eigenvector centrality, it dynamically evaluates the influence of words and concepts in policy texts, improving the efficiency and accuracy of semantic analysis. More importantly, by adopting a distributed computing architecture (such as the combination of Neo4j and Spark), parallel processing of graph data is realized, and the time-consuming of graph traversal and subgraph extraction is greatly shortened from the hour level to the minute level. This technological breakthrough not only improves the processing speed of multi-source data, but also supports real-time public opinion monitoring and policy dynamic updates, thus improving the timeliness of policy analysis results.
[0059] Specifically, this embodiment includes the following steps:
[0060] Step 1: Graph construction
[0061] Graph construction realizes the transformation of unstructured text data into structured data, providing directly computable index data for further topic classification and index calculation, specifically including:
[0062] Step 1.1: Collect policy documents related to the modern service industry. The policy documents mainly include the following four categories:
[0063] First, national modern service industry plans and policy texts.
[0064] Second, local-level modern service industry plans.
[0065] Third, annual local-level industrial policy documents related to the service industry.
[0066] Fourth, relevant news media reports on the service industry by local governments.
[0067] The above different policy text documents can be respectively used for subsequent different indicator calculation tasks, so as to fully quantify and analyze the policy texts from different dimensions and directions.
[0068] Step 1.2: Data preprocessing.
[0069] First, convert all policy documents into a unified file format. In the patent in question, use the txt format to save the policy documents. At the same time, perform cleaning operations on the redundant information in the documents that is irrelevant to the main text content, such as html format tags and redundant time information, etc. In addition, for the highly sensitive political vocabulary in the policy texts, this embodiment also performs desensitization processing on these sensitive words.
[0070] Step 1.3: Atlas construction.
[0071] The txt file after data cleaning and desensitization processing completely represents the policy documents related to the modern service industry. Then, use the corresponding policy documents as input and transfer them to GraphRAG to achieve automated atlas construction. As Figure 1 shown in the process of GraphRAG, for the incoming policy text, the processing of the graph retrieval enhancement generation technology includes:
[0072] First, divide the text into blocks;
[0073] Then, utilize the semantic understanding and perception ability of the LLM to extract the corresponding entity nodes and their relationships in the text segment. At the same time, for each node, use the LLM and the relevant node relationships of the node to generate a corresponding description for it. For example, the description of the digital economy is: "The digital economy is an important trend in current social and economic development. It emphasizes the importance of labor and data factors. This field takes data resources as key elements, modern information networks as the main carriers, and promotes a series of economic activities through the effective use of information and communication technologies. The digital economy is deeply integrated with blockchain technology and is an important field promoting the integrated development of digitalization and greening. In Zhejiang Province, the digital economy is the basis for its development and is also a sustainable development industry in the modern industrial system. It is not only the main economic form following agricultural economy and industrial economy, but also a new economic form that takes data resources as key elements, modern information networks as the main carriers, and the integrated application of information and communication technologies and the full-factor digital transformation as important driving forces to promote a more unified fairness and efficiency. In addition, the digital economy is also one of the supporting measures for implementing the national service industry reform and development deployment, aiming to improve the green and low-carbon level of computing centers."
[0074] Combined with the embedding model, a 1024-dimensional high-dimensional vector is generated for this description.
[0075] At the same time, GraphRAG itself will also achieve node partitioning through algorithms of community detection and hierarchical clustering.
[0076] GraphRAG will finally generate several parquet-type files containing information such as nodes, relationships, communities, summaries, etc. These files contain the basic information of the graph data (nodes and edges) and several important attributes.
[0077] Step 1.4: Neo4j data import and visualization.
[0078] Import the parquet file into the Neo4j file, and the relationships between nodes and edges can be visually displayed. At the same time, combined with the apoc library of Neo4j itself, several process steps such as further calculating the graph data and extracting subgraphs can also be carried out.
[0079] In this step, the automated policy text atlas generation process based on GraphRAG technology includes specific technical steps such as text chunking, entity and relationship extraction driven by LLM, node description embedding generation, community detection, and hierarchical clustering. This technology significantly reduces the need for manual intervention and realizes the full-process automation from unstructured text to structured knowledge graphs.
[0080] By optimizing the graph traversal and subgraph extraction efficiency, the calculation time-consuming is shortened from the hour level to the minute level, significantly improving the processing ability of multi-source heterogeneous data.
[0081] Step 2: Theme Allocation.
[0082] Combined with the specific task requirements of the project, measure and evaluate the policy attention under different themes of the modern service industry. What theme allocation achieves is to generate the relative membership belonging to a certain theme for each node, which specifically includes the following steps:
[0083] Step 2.1: Define Theme Documents.
[0084] In this embodiment, according to the national modern service industry plan and policies, relevant texts of specific themes of the modern service industry such as "digital economy" are first summarized and screened, and used as theme texts to be input into GraphRAG together with other local policy plans and news texts. Then, policy text nodes and theme text nodes are obtained respectively. These nodes usually include attribute data such as Name, ID, Description, and Description embedding.
[0085] Step 2.2: Allocate Themes to Each Node.
[0086] Given a theme, it is found that many nodes may belong to multiple themes at the same time. For example, for the node "digital industrial cluster", it can be found that it may belong to both the "digital economy" and "agglomerated development of the service industry" themes. Therefore, in this case, the concept of theme membership is introduced, which specifically includes:
[0087] Drawing on the classification idea of KNN, for a certain policy node i, first calculate its similarity with all theme text nodes, and then select the top N nodes closest to node i. Then, distinguish the themes corresponding to these N nodes respectively. For example, if there are m nodes belonging to theme j, then the theme membership of node i corresponding to theme j is m / N. The overall calculation process formula is as follows:
[0088] Step 2.2.1: Calculate Node Similarity
[0089] For each node i to be classified and the set of text nodes of each theme, calculate their semantic similarity. The calculation formula is as follows:
[0090]
[0091] where v i is the Description embedding of node i, and v j is the Description embedding of node j.
[0092] Step 2.2.2: Obtain the Top N Closest Nodes. Select the top N nodes according to sim. {n 1 ,n2 ,....,n N}
[0093] Step 2.2.3: Statistically analyze the topic distributions of these N nodes.
[0094] Step 2.2.4: Calculate the membership degree of policy node i to topic j:
[0095]
[0096] where m j represents the number of nodes belonging to topic j among the N nodes.
[0097] Step 3: Index calculation
[0098] Based on the topic classification of the graph nodes, further quantitatively analyze the indicators of the policy text. Specifically, four indicators, namely the consistency, uniqueness, policy focus, and policy responsiveness of the policy text, are calculated here for analysis and quantification.
[0099] Step 3.1: Policy consistency
[0100] Policy consistency is mainly used to measure the consistency between policy documents and planning documents under a specific topic. Specifically, first, select the top k nodes with the highest topic membership degrees in the policy document and the planning document respectively. Then, calculate the corresponding node similarity matrix (k×k) based on the nodes between the policy document and the planning document. Finally, calculate the average value to obtain the policy consistency index.
[0101] The specific process is as follows:
[0102] Select topic-related nodes. Respectively select the top k nodes with the highest topic membership degrees in the policy document and the planning document. Assume the nodes in the policy document are {p 1 , p 2 , p 3 ,..., p k}, and the nodes in the planning document are {g 1 , g 2 , g 3 ,... g k}.
[0103] For each node p in P i and each node g in G j calculate the similarity between the nodes to form a similarity matrix.
[0104]
[0105] Calculate the policy consistency:
[0106]
[0107] Step 3.2: Policy Uniqueness
[0108] Policy uniqueness is mainly used to measure the uniqueness between different policy texts. It is mainly measured by the distribution differences of two similarity matrices, namely the policy text - theme and the planning text - theme, and mainly uses KL divergence for measurement. The calculation process is mainly as follows:
[0109] Step 3.2.1: Calculate the similarity matrices of the policy text - theme text and the planning text - theme text respectively. Assume that the distributions of these two similarity matrices are P and Q respectively.
[0110] Step 3.2.2: Calculate the distribution differences between these two probability distributions, and measure and reflect them with the KL divergence value:
[0111]
[0112] Step 3.3: Policy Focus
[0113] The policy focus mainly measures the relative degree of attention of the policy text to a specific theme. It mainly measures the focus of the policy text on a specific theme by measuring the centrality or PageRank value of the theme - related nodes in the entire graph. The specific process is as follows:
[0114] Step 3.3.1: Calculate the eigenvector centrality and PageRank value of all nodes.
[0115] Eigenvector centrality:
[0116]
[0117] where A ij is the adjacency matrix of the graph.
[0118] Calculate the PageRank value:
[0119]
[0120] where d is the damping factor, L(j): the set of out - edges of node j (the set of nodes pointed to by j), and M(i) is the set of in - edges of node i (the set of nodes pointing to i).
[0121] Step 3.3.2: Use the theme membership degree of the nodes to continue to weight and normalize the eigenvector centrality and PageRank value, and finally obtain the attention focus on a specific theme.
[0122]
[0123] where It is the weighted eigenvector centrality of node i for topic j.
[0124]
[0125] According to the above formula, the policy focus degree based on eigenvector centrality can be calculated, and the PageRank calculation process is similar.
[0126] Step 3.4: Policy response strength
[0127] The policy response strength mainly reflects the focus degree of news reports on specific topics. Its calculation steps and process are similar to those of the policy focus degree. First, calculate the eigenvector centrality and PageRank values of all nodes in the news text graph, and then continue to weight and normalize the eigenvector centrality and PageRank values using the topic membership of the nodes to finally obtain the policy response strength of the news text to specific topics.
[0128] Combined with this overall policy text quantitative analysis process, the more accurate quantitative analysis of policy texts is finally realized, so as to more deeply explore the rich information contained in policy texts.
[0129] The present invention mainly has the following three key points:
[0130] Proposing a systematic policy information quantitative analysis framework:
[0131] In this embodiment, a complete quantitative analysis framework is constructed from multiple dimensions such as policy consistency, uniqueness, focus degree, and response degree. This framework breaks through the limitation of traditional methods relying on a single indicator, can comprehensively evaluate the scientificity and effectiveness of policy texts, and provides multi-dimensional quantitative basis for policy design and optimization.
[0132] Innovative application of knowledge graph and graph computing technology:
[0133] In this embodiment, the knowledge graph technology is introduced into the field of policy text analysis for the first time. Using its structured semantic representation ability, the entities, relationships, and complex semantic networks in policy texts are made explicit. Combined with graph computing technology (such as PageRank, eigenvector centrality), the in-depth mining of the node association patterns and weight distributions in policy texts is realized, and the problems that traditional methods are difficult to handle cross-topics and multi-source data integration are solved.
[0134] Design of the dynamic allocation algorithm for topic membership:
[0135] By combining the KNN classification idea with the concept of topic membership, a dynamic allocation algorithm is proposed to realize the soft clustering of nodes. This method effectively alleviates the problem of distortion of hard clustering results caused by semantic diversity, and significantly improves the accuracy and flexibility of topic classification.
[0136] Compared with the best existing topic modeling (LDA) methods and TF-IDF keyword analysis techniques, the present invention mainly has the following advantages:
[0137] First, breakthrough in deep semantic understanding and cross-topic correlation analysis capabilities:
[0138] Traditional LDA and TF-IDF methods rely on word frequency statistics or probability distribution modeling, and can only capture the shallow semantics of text, unable to analyze entity relationships, context dependencies, and cross-topic correlations in policy texts, and the topic boundaries are blurred. The present invention dynamically extracts entities, attributes, and relationships in policy texts through graph augmentation generation technology (GraphRAG), constructs an explicit semantic network, and accurately reveals the logical link between "policy objectives - implementation paths - implementation measures". For example, when analyzing the "digital economy" policy, it can automatically associate cross-domain nodes such as "data elements" with "blockchain technology" and "green computing power", breaking through the semantic fragmentation problem of traditional methods. At the same time, combined with the dynamic topic membership assignment algorithm, it supports the soft membership assignment of nodes to multiple topics (such as a node belonging to both "digital economy" and "green economy" at the same time), significantly improving the cross-topic analysis ability while enhancing the accuracy of topic classification.
[0139] Second, the scientific and dynamic nature of the multi-dimensional quantitative index system:
[0140] Traditional methods only support a single index (such as keyword weight or coarse-grained topic distribution), lack the quantification ability of deep attributes such as policy consistency and uniqueness, and rely on static data with insufficient timeliness. The present invention proposes a four-dimensional quantitative framework and realizes comprehensive dynamic analysis by combining graph computing technology: (1) Policy consistency: Quantify the degree of policy objective fit through the cross-document node similarity matrix (such as the matching between national plans and local policies); (2) Policy uniqueness: Based on the KL divergence analysis of the similarity distribution differences, accurately identify policy innovation points (such as unique measures of local policies in the "digital transformation of the service industry"); (3) Policy focus: Dynamically evaluate the attention intensity of policies to core topics using PageRank values and eigenvector centrality; (4) Policy response strength: Integrate news and public opinion data in real time, and quantify social feedback through node influence analysis (such as the dissemination breadth and sentiment tendency of news about "green service industry"). Combined with a distributed architecture (Neo4j + Spark), it supports real-time graph updates and index recalculation to ensure the timeliness of the results.
[0141] Third, the efficient integration of multi-source heterogeneous data and full-process automation:
[0142] Traditional methods require manual annotation of multi-source data (such as policy documents and news texts). The processing flow is cumbersome, with the time-consuming ratio exceeding 60%. Moreover, the model adaptability is poor and the cross-domain reuse rate is low. The present invention realizes full-process automation based on the GraphRAG technology:
[0143] (1) Automated graph construction: Utilize the semantic understanding ability of the LLM to automatically extract entities (such as "modern service industry demonstration zones") and generate node description vectors (1024 dimensions), significantly reducing the need for manual annotation;
[0144] (2) Multi-source data fusion: Support unified processing of heterogeneous data such as government bulletins, news, and social media, and automatically divide clusters through community detection algorithms (such as classifying "Yangtze River Delta" and "Beijing-Tianjin-Hebei" policies into different subgraphs);
[0145] (3) General graph computing framework: Based on the Cypher query language of Neo4j and Spark distributed computing, it adapts to different policy fields (such as digital economy and green economy), effectively improving the model reuse rate while shortening the development cycle.
[0146] Fourth, a significant optimization of computing efficiency and real-time performance. Traditional methods need to clean, extract keywords, and perform topic modeling on policy documents one by one, which takes a long time for single analysis and cannot respond to new data in real time.
[0147] Fifth, the advantages of the present invention in policy text analysis:
[0148] (1) Distributed computing speedup: Through Spark parallel optimization of graph traversal and subgraph extraction, the calculation time of similarity matrices for numerous nodes is effectively reduced.
[0149] (2) Real-time dynamic monitoring: Combine streaming computing technology, dynamically incorporate public opinion hotspots and update the graph, and achieve minute-level output of policy response strength indicators, greatly improving the timeliness of analysis results.
[0150] This embodiment discloses a modern service industry policy quantitative analysis system based on a knowledge graph, including:
[0151] A knowledge graph construction module for constructing a knowledge graph based on modern service industry-related policies;
[0152] A theme assignment module for determining theme files, extracting subgraphs of theme files and corresponding nodes based on the knowledge graph, respectively assigning corresponding themes to the extracted nodes, and obtaining theme assignment results;
[0153] A quantitative analysis module for calculating indicators from aspects of consistency, uniqueness, policy focus, and policy responsiveness based on the theme assignment results to obtain quantitative analysis results
[0154] Schematic diagram of a terminal device provided by an embodiment of the present invention. The terminal device of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor executes the computer program, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0155] The computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention.
[0156] The terminal device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory.
[0157] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0158] The memory may be used to store the computer program and / or modules, and the processor realizes various functions of the terminal device by running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory.
[0159] If the modules / units integrated in the terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0160] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A quantitative analysis method for modern service industry policies based on knowledge graph, characterized in that: The following steps are involved: Construct a knowledge graph based on policies related to modern service industry; Determine the subject file, extract the subgraph of the subject file and the nodes corresponding to the subject file based on the knowledge graph, assign corresponding topics to the extracted nodes, and obtain the subject assignment results; Based on the results of topic allocation, indicators are calculated from the aspects of consistency, uniqueness, policy focus and policy responsiveness to obtain quantitative analysis results.
2. According to claim 1, a modern service industry policy quantitative analysis method based on knowledge graph is characterized in that: The construction of a knowledge graph based on policies related to modern service industry includes: Collect relevant policy documents related to modern service industry and pre-process the collected policy documents; The preprocessed files are fed into GraphRAG to generate a knowledge graph containing several entity nodes, relationships, communities, and summaries.
3. According to the method of quantitative analysis of modern service industry policies based on knowledge graph according to claim 1, it is characterized in that: The step of respectively assigning corresponding topics to the extracted nodes and obtaining topic assignment results includes: For a policy node i, calculate its similarity with all topic text nodes and obtain the top N nodes closest to node i; Count the topic distribution of N nodes and calculate the membership between policy node i and topic j; According to the membership calculation results, the topics corresponding to the N nodes are distinguished and the topic allocation results are obtained.
4. According to a method for quantitative analysis of modern service industry policies based on knowledge graph according to claim 1, it is characterized in that: Calculate indicators for consistency, including: Select the first k nodes with the highest subject membership in the policy document and planning document respectively. Assume that the policy document nodes are {p1, p2, p3, ..., p k }, the planning file nodes are {g1,g2,g3,...g k }; Calculate for each node p in P i and every node g in G j The similarity between them forms a similarity matrix: The consistency of the policy is calculated through the similarity matrix: Among them, p i represents a policy document node; gj represents a planning document node.
5. According to claim 1, a modern service industry policy quantitative analysis method based on knowledge graph is characterized in that: Calculate indicators for uniqueness, including: Calculate the similarity matrix of policy text topics and planning text topics respectively; Utilize the distribution difference of KL divergence matrix.
6. According to claim 1, a modern service industry policy quantitative analysis method based on knowledge graph is characterized in that: Calculate indicators for policy focus, including: Calculate the eigenvector centrality and PageRank value of all topic-related nodes; The feature vector centrality and PageRank value are further weighted and normalized in combination with the topic membership of the node to obtain the focus of attention on a specific topic.
7. According to claim 1, a modern service industry policy quantitative analysis method based on knowledge graph is characterized in that: Calculate indicators for policy responsiveness, including: Calculate the eigenvector centrality and PageRank value of topic-related nodes; The feature vector centrality and PageRank value are further weighted and normalized in combination with the topic membership of the node to obtain the policy responsiveness calculation result.
8. A modern service industry policy quantitative analysis system based on knowledge graph, characterized in that: include: Knowledge graph construction module, used to construct a knowledge graph based on policies related to modern service industry; The topic assignment module is used to determine the topic file, extract the subgraph of the topic file and the nodes corresponding to the topic file based on the knowledge graph, assign corresponding topics to the extracted nodes, and obtain the topic assignment results; The quantitative analysis module is used to calculate indicators in terms of consistency, uniqueness, policy focus and policy responsiveness based on the topic allocation results to obtain quantitative analysis results.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Internet of Things industrial development evaluation method and device
CN120725501A
Green watershed policy collaboration analysis system and method
CN120994709A