Text mining and network analysis-based carbon and pollution reduction measure collaboration evaluation method, system and equipment and medium

By constructing a three-level keyword classification system and a co-word network, and combining network analysis indicators, the problem of the failure of existing technologies to fully evaluate the synergy of carbon reduction and pollution reduction was solved, and a precise, dynamic, and quantitative analysis of the synergy of carbon reduction and pollution reduction measures was achieved.

CN121707367APending Publication Date: 2026-03-20BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511492325.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing research has failed to effectively consider the comprehensive and synergistic emission reduction of water, air, soil, and greenhouse gases, and has neglected the quantitative analysis of the internal synergy of carbon reduction and pollution reduction measures. Furthermore, text mining and network analysis methods have not been effectively combined to assess synergistic relationships.

Method used

By constructing a personalized external corpus, using the word frequency-inverse document frequency method to screen keywords, a three-level keyword classification system is built. Based on the co-word network and multiple network analysis indicators, the synergy of carbon reduction and pollution reduction measures is evaluated, including indicators such as network average degree, node feature vector centrality, and edge density.

Benefits of technology

It enables precise assessment of the synergy of carbon reduction and pollution control measures, allows dynamic analysis of their internal synergy, provides a reliable quantitative assessment method, and reveals the evolutionary pattern of measure synergy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121707367A_ABST
    Figure CN121707367A_ABST
Patent Text Reader

Abstract

The invention discloses a carbon reduction and pollution reduction measure collaboration evaluation method, system and equipment based on text mining and network analysis and a medium, and belongs to the technical field of climate and environmental governance measure analysis and evaluation.The method comprises the steps that carbon reduction and pollution reduction related measure files are obtained, and a personalized external corpus is constructed; performing duplicate removal and word segmentation on the text data; initially selecting keywords based on a TF-IDF method, and checking and summarizing the keywords according to a keyword three-level classification system; calculating the co-occurrence frequency between every two keywords, and constructing a word co-occurrence network; and analyzing the structure characteristics of the co-word network from different levels by using indexes such as network average degree, edge density, average shortest path and the like, and summarizing the collaboration of carbon and pollution reduction measures. Based on real and reliable first-hand data and a more targeted personalized corpus, keywords are accurately extracted and classified for extraction and classification, and in combination with a network analysis method, internal collaboration of carbon and pollution reduction measures is analyzed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of climate and environmental analysis and assessment technology, specifically to a method, system, equipment, and medium for assessing the synergy of carbon reduction and pollution reduction measures based on text mining and network analysis. Background Technology

[0002] Mitigating climate change and controlling emissions of various environmental pollutants have become global concerns. Based on this, coordinated emission reduction targets for water, air, soil, and greenhouse gases have been proposed, and the main objectives and tasks for promoting synergy between the two have been clarified. Coordinated governance for carbon reduction and pollution control is a systematic project that requires not only regional and inter-departmental collaboration but also a top-level mechanism to ensure coordinated measures.

[0003] Most existing studies use difference-in-differences (DID) or integrated assessment (IAM) models to analyze the synergistic benefits of carbon reduction and pollution control measures. Regarding the research subjects, current studies primarily focus on the synergistic reduction of carbon dioxide and air pollutants, failing to comprehensively consider the goal of synergistic reduction of water, air, soil, and greenhouse gases. Furthermore, existing studies tend to focus more on the impact of measures involving one aspect of carbon reduction or pollution control on the other, rather than the degree of synergy between the two. At the methodological level, existing studies mainly use empirical analysis or systems modeling to quantitatively analyze the synergistic effects of carbon reduction and pollution control, neglecting the inherent synergistic nature of the measures themselves.

[0004] Some studies have attempted to apply text mining and network analysis to explore the synergy of environmental measures. Text mining methods are mostly combined with econometric models to verify the effectiveness of synergistic implementation of environmental measures; while network analysis primarily targets the external synergy of measures, specifically including two aspects: first, constructing citation networks between measures to explore the diffusion patterns and causes of themes in spatial or temporal dimensions; and second, constructing collaborative networks between issuing departments to measure the degree of cooperation among departments or entities of the same type of measure and to identify the leading actors in measure formulation. Overall, in current research, text analysis can extract key information from measure documents to understand their themes and characteristics, while network analysis is mainly used to understand the development trajectory of measures and the relationships between relevant entities. However, no research has yet combined these two methods to study the internal synergistic relationships of measures. Summary of the Invention

[0005] In view of the above-mentioned problems, the present invention is proposed.

[0006] To address the aforementioned technical problems, this invention provides the following technical solution: a method for evaluating the synergistic effect of carbon reduction and pollution mitigation measures based on text mining and network analysis, comprising, Based on keywords related to carbon reduction and pollution control, search and retrieve documents containing measures to be analyzed within a specified time period; construct a personalized external corpus with a quantity greater than the number of documents to be analyzed; based on the personalized external corpus, use the term frequency-inverse document frequency method to filter measure keywords, construct a three-level keyword classification system of measure elements-measure themes-measure keywords, and manually summarize the measure keywords; construct a co-occurrence network based on the frequency of co-occurrence of the keywords in the documents; divide the co-occurrence network into three levels according to the keyword classification system, and analyze the structural characteristics of the co-occurrence network from different levels based on multiple network analysis indicators to evaluate the synergy of carbon reduction and pollution control measures.

[0007] As a preferred embodiment of the synergy evaluation method for carbon reduction and pollution reduction measures based on text mining and network analysis described in this invention, the measures to be analyzed include all measures to reduce carbon dioxide and all pollutant emissions within a preset time period retrieved from an existing document library, excluding approval and interpretation category documents; The personalized external corpus consists of all published documents retrieved from an existing document library, excluding approval documents and interpretation documents of the measure category.

[0008] As a preferred embodiment of the synergy evaluation method for carbon reduction and pollution reduction measures based on text mining and network analysis described in this invention, the measures to be analyzed further include: data preprocessing of the documents to be analyzed, deleting all duplicate data according to the Uniform Resource Locator (URL) value of the documents, and obtaining a document set; Remove numbers, Uniform Resource Locator values, and garbled characters from the body of the file to be analyzed; The text of the document to be analyzed is segmented using Python. The main text of the document is segmented into words, and stop words are removed from the results to obtain the final word segmentation list.

[0009] As a preferred embodiment of the synergistic evaluation method for carbon reduction and pollution reduction measures based on text mining and network analysis described in this invention, the keyword selection includes extracting and summarizing the keywords based on word frequency-inverse document frequency values ​​and manual methods. All measures documents are merged into one document. The frequency of each word in the word segmentation list is counted. The word frequency-inverse document frequency value is calculated. The words in the word segmentation list are sorted in descending order and keywords are extracted. A three-level keyword classification system is set up: the first level is the measure elements, the second level is the measure theme, and the third level is the measure keywords. The measure keywords are manually screened and summarized according to the keyword classification system.

[0010] As a preferred embodiment of the synergy evaluation method for carbon reduction and pollution reduction measures based on text mining and network analysis described in this invention, the construction of the co-occurrence network includes constructing a word co-occurrence network based on the co-occurrence relationship of the keywords; Calculate the co-occurrence frequency of the keywords in carbon reduction and pollution reduction measure documents. For two measure keywords, the co-occurrence frequency together includes the number of measure documents for both measures, thus obtaining the keyword co-occurrence matrix. Construct a word co-occurrence network based on the keyword co-occurrence matrix, where keywords are network nodes and co-occurrence relationships are edges.

[0011] As a preferred embodiment of the synergistic evaluation method for carbon reduction and pollution reduction measures based on text mining and network analysis described in this invention, the construction of the co-word network further includes, consistent with the keyword classification system, the co-word network being divided into three levels.

[0012] As a preferred embodiment of the synergy evaluation method for carbon reduction and pollution reduction measures based on text mining and network analysis described in this invention, the evaluation of the synergy of carbon reduction and pollution reduction measures includes obtaining the overall and local structural features and evolutionary patterns of the co-word network using network analysis indicators. The metrics used include the network's average degree, node eigenvector centrality, network edge density, and average shortest path length.

[0013] Another objective of this invention is to provide a synergy evaluation system for carbon reduction and pollution control measures based on text mining and network analysis.

[0014] To address the aforementioned technical problems, this invention provides the following technical solution: a synergy evaluation system for carbon reduction and pollution mitigation measures based on text mining and network analysis, comprising: a measure acquisition and preprocessing module, a keyword analysis module, and a co-word network analysis module; The measure collection and preprocessing module searches for and obtains the measure files to be analyzed within a set time period based on keywords related to carbon reduction and pollution reduction, constructs a personalized external corpus with a quantity greater than the number of files to be analyzed, and performs deduplication and word segmentation on the measure files to be analyzed. The keyword analysis module, based on the personalized external corpus, uses the term frequency-inverse document frequency method to filter keywords, constructs a three-level keyword classification system of measure elements-measure themes-measure keywords, and manually summarizes the keywords; The co-occurrence network analysis module constructs a co-occurrence network based on the frequency of co-occurrence of the keywords in the measure document, divides the co-occurrence network into three levels according to the keyword classification system, and analyzes the structural characteristics of the co-occurrence network from different levels based on multiple network analysis indicators to evaluate the synergy of carbon reduction and pollution reduction measures.

[0015] The present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, characterized in that the processor executes the computer program to implement the steps of the aforementioned method for evaluating the synergy of carbon reduction and pollution reduction measures based on text mining and network analysis.

[0016] The present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the steps of the aforementioned method for evaluating the synergy of carbon reduction and pollution reduction measures based on text mining and network analysis.

[0017] The beneficial effects of this invention are as follows: This invention focuses on the content of texts related to carbon reduction and pollution control measures, analyzing the inherent synergy within these measures from three dimensions: target, approach, and tools. By setting search fields to retrieve relevant documents from a document library, reliable first-hand data can be obtained. Furthermore, the time frame covered by the measures and the stages can be arbitrarily selected to analyze the evolution of synergy. Based on a personalized external corpus, TF-IDF method is used to extract key terms for the measures, and a three-level classification system of measure elements, measure themes, and measure keywords is constructed, enabling more accurate keyword extraction and classification. Based on the co-occurrence frequency of key terms, a three-level co-occurrence network is constructed, and multiple network analysis indicators are used to analyze the network structure characteristics at different levels, thereby providing a detailed and accurate assessment of the synergy of carbon reduction and pollution control measures. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 The above is a flowchart of an overall method for evaluating the synergy of carbon reduction and pollution reduction measures based on text mining and network analysis, as provided in one embodiment of the present invention.

[0020] Figure 2 This invention provides a method for evaluating the synergy of carbon reduction and pollution reduction measures based on text mining and network analysis, using a three-level keyword classification system to obtain the measure keyword classification results.

[0021] Figure 3 The following is an evolution diagram of the co-word network of a carbon reduction and pollution reduction measure synergy evaluation method based on text mining and network analysis, provided as an embodiment of the present invention. It includes basic network features such as node classification under a three-level keyword classification system, node and edge count, average network degree, and node degree distribution, taking 2008-2012 as an example.

[0022] Figure 4 This invention provides an embodiment of a method for evaluating the synergy of carbon reduction and pollution control measures based on text mining and network analysis. The evolution trend of the synergy of three types of measure elements—measure objectives, measure paths, and measure tools—at each stage is calculated using the edge density index.

[0023] Figure 5 This invention provides an embodiment of a method for evaluating the synergy of carbon reduction and pollution control measures based on text mining and network analysis, which uses the edge density index to measure the evolution trend of synergy among themes within each stage of the measures.

[0024] Figure 6 This invention provides an embodiment of a method for evaluating the synergy of carbon reduction and pollution control measures based on text mining and network analysis. The direct and indirect synergies among the various measure themes in different measure elements are measured using network edge density and average shortest path, taking 2008-2012 as an example. Detailed Implementation

[0025] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0026] Example 1, referring to Figures 1-6 This is one embodiment of the present invention, which provides a method for evaluating the synergy of carbon reduction and pollution reduction measures based on text mining and network analysis, including: S1. Based on fields such as "carbon reduction", "pollution reduction", "environmental protection" and "carbon neutrality", retrieve detailed information of the measures documents issued by the departments within a specific time period from the publicly available document database. Specifically, this includes: index number, document number, title, issuing agency, publication time, subject, category, and text, while removing documents in the "approval" and "interpretation" categories.

[0027] For example, during the process of obtaining files, the time range of file release can be selected according to actual needs and divided into different stages. In this embodiment, the period from 2008 to 2023 is selected and divided into four stages (2008-2012; 2013-2017; 2018-2022; 2023).

[0028] At the same time, all measures documents that can be retrieved from publicly available document databases were collected to build a personalized external corpus containing more than 25,000 documents.

[0029] S2. Preprocess the collected file data, specifically including the following steps: Step 1: Delete all duplicate measure data based on the file's URL value. The deduplicated measure file is denoted as [file name missing]. N is the number of measures documents; Step 2: For each file, remove numbers, characters, URLs, and garbled text from its body, keeping only Chinese and English characters; Step 3: In this embodiment of the invention, the jieba module in Python is used to perform word segmentation. This module segments Chinese sentences based on a large-scale Chinese dictionary and uses a Hidden Markov Model to identify new words. After segmenting the main text, stop words are removed from the results to obtain the final word segmentation list. Each word is denoted as jieba. .

[0030] In an optional embodiment, the SnowNLP tool is used for word segmentation. Specific steps include: importing the SnowNLP library in Python, which is trained on a Naive Bayes model for Chinese word segmentation; calling SnowNLP's word segmentation function to automatically segment sentences into word sequences; during segmentation, utilizing a built-in model for new word recognition without requiring additional HMM configuration; applying the same stop word list to filter words; and outputting the final word segmentation list for direct use in keyword statistics.

[0031] In another optional embodiment, the THULAC tool is used for word segmentation. Specific steps include: installing and importing the THULAC library, which performs Chinese word segmentation based on sequence labeling models (such as CRF); setting the word segmentation mode to precise mode for the main text to segment words with high accuracy; automatically identifying new words and tagging their parts of speech during segmentation; filtering words using the stop word list from the original method, retaining only keywords such as nouns and verbs; and generating the final word segmentation list, compatible with subsequent TF-IDF analysis.

[0032] In this embodiment, a total of 1804 files were obtained after data preprocessing.

[0033] S3. Filter and summarize keywords. First, merge all the files to be analyzed into one document, and count the frequency of each word in the word segmentation list, denoted as . .

[0034] Use the Text-Inverse Document Frequency (TF-IDF) method to extract keywords from text.

[0035] in, For vocabulary Relative word frequency; For vocabulary Inverse document frequency D represents the total word frequency in the document to be analyzed; D represents the external corpus. For the corpus contains The document.

[0036] The words are sorted in descending order according to their TF-IDF values, and the top few words are selected as candidate keywords. For example, in this embodiment, the top 200 words with the highest TF-IDF values ​​are selected as candidate keywords.

[0037] A three-level classification system is set up. The first level is the elements of the measures, including the measures objectives, the measures pathways, and the measures tools. Among them, the measures objectives represent the specific results or effects that are hoped to be achieved in the formulation of the measures; the measures pathways represent the measures and actions taken to achieve the measures objectives; and the measures tools represent the means to achieve the measures objectives.

[0038] The second level consists of the themes of the measures included in each measure element. Specifically, the measures objectives include four objectives: greenhouse gas emission reduction, air pollution control, water pollution control, and soil pollution control; the measures pathways include four implementation pathways: primary industry, energy structure, industry and construction, and transportation and logistics; and the measures tools include three types of regulatory tools: command and control, market mechanism, and voluntary.

[0039] The third level consists of policy keywords, which are manually selected and summarized based on the meaning of different policy themes.

[0040] For example, this embodiment screened a total of 81 keywords, and the detailed classification results are as follows: Figure 2 As shown.

[0041] It should be noted that this step addresses the issues that relying solely on statistical methods to extract keywords may deviate from the substantive content of the policy documents and that keywords lack structured organization.

[0042] By combining objective statistics from TF-IDF with the domain knowledge framework of the PMC model for manual screening and summarization, we ensured that the final keyword set not only has statistical significance but also accurately reflects the core elements and themes of carbon reduction and pollution control measures. Constructing a structured keyword classification system provides a foundation for subsequent construction of co-word networks with clear meanings (divided by elements and themes) and hierarchical synergy analysis (between elements and themes). This is the key to the method's ability to deeply analyze synergy.

[0043] S4. Construct a co-occurrence network based on the co-occurrence relationships of keywords, as follows: Step 1: Calculate the number of times each keyword co-occurs with the other keyword in a document. To avoid duplicate counting, this is represented as the number of documents that contain both keywords.

[0044] in, and For two keywords, This is a Boolean operation, taking either 1 or 0 based on the condition. After iterating through all keyword pairs, a keyword co-occurrence matrix can be obtained. This matrix is ​​a positive real symmetric matrix.

[0045] Step 2: Construct a word co-occurrence network based on the keyword co-occurrence matrix, denoted as Where V represents a node in the network, i.e., a keyword; and E represents an edge connecting nodes. A set of.

[0046] Step 3: Standardize the word co-occurrence frequency and use it as a connection edge. weight ,make .because It is a positive real symmetric matrix, therefore the edges are connected. These are weighted undirected edges.

[0047] For example, the number of co-word networks depends on the number of stages, so in this embodiment, a co-word network with four stages can be constructed. Figure 3 The following is an example diagram of a co-word network, taking 2008-2012 as an example. From this, we can obtain basic information such as the number of nodes, the number of edges, and the height distribution of the co-word network.

[0048] It should be noted that this step addresses the problem of how to transform the keyword associations in text into a quantifiable and comparable network model.

[0049] By constructing a weighted undirected co-word network, the qualitative content association of measures is transformed into a quantitative network topology; the standardized weight processing makes multiple networks constructed across time periods comparable in terms of structural indicators, thereby reliably revealing the dynamic evolution law of synergy.

[0050] S5. Several network analysis indicators are used to analyze the structural characteristics of co-word networks at each level, thereby evaluating the synergy of carbon reduction and pollution reduction. These indicators include the following: (1) Network average degree. This index is the average degree of all nodes in the network, reflecting the overall connection density of the network. The average degree of the network can reflect the synergy between and within the overall elements of the network.

[0051] in, For nodes In a weighted undirected network, the degree of a node is the sum of the weights of the edges connected to it. For nodes and Edge weights; Represents all nodes A set of connected nodes.

[0052] (2) Network edge density Network density is the ratio of the sum of all edge weights in a network to the sum of the largest possible weights in the network, reflecting the overall connectivity of the network. Using this metric, network edge density is constructed to describe the degree of correlation between network elements and thematic networks.

[0053] in, and Two sub-networks (different elements or themes); and There are two nodes; For nodes and Edge weights; and for and The number of nodes; Let be the maximum possible edge weight. To ensure comparability of results across networks of different sizes, let... .

[0054] (3) Average shortest path length of the network The shortest path between nodes is the path that takes the shortest distance or incurs the least cost from one node to another, used to represent the degree of indirect association between nodes. Dijkstra's algorithm is a commonly used algorithm for finding the single-source shortest path in a network graph, and includes the following steps: Step 1: Distance initialization, set the shortest path estimate of all nodes to infinity; Step 2: Find the node closest to the starting point among the unvisited nodes; Step 3: Update the distance to neighboring nodes. Check all neighbors of this node. If the distance to a neighbor through this node is shorter than the currently recorded distance, update this distance. Step 4: Repeat steps 2 and 3 until all nodes have been visited; Referring to this metric, this invention constructs the average shortest path length of the network to describe the strength of indirect connections between sub-networks at different levels of the co-word network, and is used to measure the indirect synergy between different topics.

[0055] in, and Two sub-networks (different elements or themes); and There are two nodes; for and The shortest path between them is calculated using Dijkstra's algorithm; for and There are node pairs with paths between them when the network is a connected network. Since the edge weights in the co-occurrence network are standardized word co-occurrence frequencies, their reciprocals are used when calculating the shortest path.

[0056] For example, these indicators are used to analyze the structural characteristics of the co-word network, thereby understanding the evolution trend of the synergy between carbon reduction and pollution reduction measures. In this embodiment, indicator 1 is used to analyze the overall evolution trend of the synergy between carbon reduction and pollution reduction measures, see... Figure 3 Indicator 2 was used to analyze the synergistic evolution trend of themes between and within the elements of the measures, see [reference needed]. Figure 4 and Figure 5 Indicators 2 and 3 were used to analyze the direct and indirect synergies among the themes of different measures, taking the period from 2008 to 2012 as an example. Figure 6 .

[0057] It should be noted that this step addresses the problem of how to extract meaningful quantitative indicators of collaboration from the network structure.

[0058] By defining and applying these three highly targeted indicators, this method achieves an objective, multi-dimensional (overall, internal, direct, indirect), multi-level (between elements, between themes), quantifiable, and dynamically comparable assessment of the abstract concept of "synergy in carbon reduction and pollution control measures." This is the core innovation and value of this method, elevating qualitative research on the synergy of measures to a quantitative and refined analytical level, providing strong data support for the assessment and optimization of environmental and climate governance measures. Figure 3-6 The application examples clearly demonstrate how these metrics specifically characterize the evolution and structural features of synergy.

[0059] Example 2 is an embodiment of the present invention. This embodiment provides a system for evaluating the synergy of carbon reduction and pollution reduction measures based on text mining and network analysis, including: a measure document collection and preprocessing module, a keyword analysis module, and a co-word network analysis module; The measure analysis module searches for and obtains the measures to be analyzed within a set time period based on keywords related to carbon reduction and pollution reduction, constructs a personalized external corpus with a quantity greater than the number of measure files to be analyzed, and performs deduplication and word segmentation on the measure files to be analyzed. The keyword analysis module, based on the personalized external corpus, uses the term frequency-inverse document frequency method to filter keywords, constructs a three-level keyword classification system of measure elements-measure themes-measure keywords, and manually summarizes the keywords; The co-occurrence network analysis module constructs a co-occurrence network based on the frequency of co-occurrence of the keywords in the measure document, divides the co-occurrence network into three levels according to the keyword classification system, and analyzes the structural characteristics of the co-occurrence network from different levels based on multiple network analysis indicators to evaluate the synergy of carbon reduction and pollution reduction measures.

[0060] This embodiment also provides an electronic device applicable to a method for evaluating the synergy of carbon reduction and pollution reduction measures based on text mining and network analysis, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method for evaluating the synergy of carbon reduction and pollution reduction measures based on text mining and network analysis as proposed in the above embodiment.

[0061] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements a method for evaluating the synergy of carbon reduction and pollution reduction measures based on text mining and network analysis as proposed in the above embodiment.

[0062] The storage medium proposed in this embodiment and the method for evaluating the synergy of carbon reduction and pollution reduction measures based on text mining and network analysis proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0063] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.

[0064] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for evaluating the synergistic effect of carbon reduction and pollution mitigation measures based on text mining and network analysis, characterized in that: include, Based on keywords related to carbon reduction and pollution control, search and retrieve documents on measures to be analyzed within a specified time period; Build a personalized external corpus, setting its size to be greater than the files to be analyzed; Based on the personalized external corpus, the term frequency-inverse document frequency method was used to filter policy keywords, and a three-level keyword classification system of policy elements-policy themes-policy keywords was constructed. The policy keywords were then manually summarized. Construct a co-occurrence network based on the number of times the keywords co-occur in the file; Based on the keyword classification system, the co-word network is divided into three levels, and the structural characteristics of the co-word network are analyzed from different levels based on multiple network analysis indicators to evaluate the synergy of carbon reduction and pollution reduction measures.

2. The method for evaluating the synergy of carbon reduction and pollution mitigation measures based on text mining and network analysis as described in claim 1, characterized in that: The measures to be analyzed include all measures to reduce carbon dioxide and all pollutant emissions within a preset time period retrieved from an existing document library, excluding approval and interpretation category documents; The personalized external corpus consists of all published documents retrieved from an existing document library, excluding approval documents and interpretation documents of the measure category.

3. The method for evaluating the synergy of carbon reduction and pollution mitigation measures based on text mining and network analysis as described in claim 2, characterized in that: The measures to be analyzed also include data preprocessing of the documents to be analyzed, deleting all duplicate data based on the Uniform Resource Locator (URL) value of the documents, resulting in a document set; Remove numbers, Uniform Resource Locator values, and garbled characters from the body of the file to be analyzed; The text of the document to be analyzed is segmented using Python. The main text of the document is segmented into words, and stop words are removed from the results to obtain the final word segmentation list.

4. The method for evaluating the synergy of carbon reduction and pollution mitigation measures based on text mining and network analysis as described in claim 3, characterized in that: The keyword selection process includes extracting and summarizing keywords based on term frequency-inverse document frequency values ​​and manual methods. All measures documents are merged into one document. The frequency of each word in the word segmentation list is counted. The word frequency-inverse document frequency value is calculated. The words in the word segmentation list are sorted in descending order and keywords are extracted. A three-level keyword classification system is set up: the first level is the measure elements, the second level is the measure theme, and the third level is the measure keywords. The measure keywords are manually screened and summarized according to the keyword classification system.

5. The method for evaluating the synergy of carbon reduction and pollution mitigation measures based on text mining and network analysis as described in claim 4, characterized in that: The construction of the co-occurrence network includes constructing a word co-occurrence network based on the co-occurrence relationships of the keywords; Calculate the co-occurrence frequency of the keywords in carbon reduction and pollution reduction measure documents. For two measure keywords, the co-occurrence frequency together includes the number of measure documents for both measures, thus obtaining the keyword co-occurrence matrix. Construct a word co-occurrence network based on the keyword co-occurrence matrix, where keywords are network nodes and co-occurrence relationships are edges.

6. The method for evaluating the synergy of carbon reduction and pollution mitigation measures based on text mining and network analysis as described in claim 5, characterized in that: The construction of the co-word network also includes, consistent with the keyword classification system, the co-word network is divided into three levels.

7. The method for evaluating the synergy of carbon reduction and pollution mitigation measures based on text mining and network analysis as described in claim 6, characterized in that: The assessment of the synergy of carbon reduction and pollution reduction measures includes using network analysis indicators to obtain the overall and local structural characteristics and evolutionary patterns of the co-word network; The metrics used include the network's average degree, node eigenvector centrality, network edge density, and average shortest path length.

8. A system for evaluating the synergy of carbon reduction and pollution reduction measures based on text mining and network analysis, employing the method for evaluating the synergy of carbon reduction and pollution reduction measures based on text mining and network analysis as described in any one of claims 1 to 7, characterized in that, include: The module includes measures collection and preprocessing, keyword analysis, and co-word network analysis. The measure collection and preprocessing module searches for and obtains the measure files to be analyzed within a set time period based on keywords related to carbon reduction and pollution reduction, constructs a personalized external corpus with a quantity greater than the number of files to be analyzed, and performs deduplication and word segmentation on the measure files to be analyzed. The keyword analysis module, based on the personalized external corpus, uses the term frequency-inverse document frequency method to filter keywords, constructs a three-level keyword classification system of measure elements-measure themes-measure keywords, and manually summarizes the keywords; The co-occurrence network analysis module constructs a co-occurrence network based on the frequency of co-occurrence of the keywords in the measure document, divides the co-occurrence network into three levels according to the keyword classification system, and analyzes the structural characteristics of the co-occurrence network from different levels based on multiple network analysis indicators to evaluate the synergy of carbon reduction and pollution reduction measures.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method for evaluating the synergy of carbon reduction and pollution reduction measures based on text mining and network analysis, as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for evaluating the synergy of carbon reduction and pollution reduction measures based on text mining and network analysis, as described in any one of claims 1 to 7.