Assessment method and system for traffic transportation carbon emission policy
By conducting word segmentation, word frequency analysis and industry vocabulary construction on transportation carbon emission policies, combined with the training of prediction models, the problems of low efficiency and insufficient accuracy of policy analysis in the existing technology are solved, and a more accurate and reliable policy effect evaluation is achieved.
Patent Information
- Application Number
- CN202510076106.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The existing analysis methods for transportation carbon emission policy are insufficient in terms of efficiency and accuracy, and cannot effectively identify the key words and industry rules in the transportation environment policy, resulting in a large deviation from the actual situation.
A method for evaluating transportation carbon emission policy is proposed, including obtaining policy documents for word segmentation and word frequency analysis, constructing industry vocabulary and extracting keywords, using labeled policy documents to build a prediction model and conducting training, and determining whether the effect of the policy documents is obvious.
By specifically obtaining historical environmental policy texts in the transportation field and building an industry vocabulary, text preprocessing and model training are carried out accurately, the predictive model's understanding and processing ability of professional terms in the transportation field is improved, and the accuracy and reliability of policy effect evaluation is enhanced.
Smart Images

Figure CN119990936A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of transportation technology, and in particular to a method and system for evaluating transportation carbon emission policies. Background Art
[0002] Globally, climate change and environmental problems are becoming increasingly severe. Reducing greenhouse gas emissions, mainly carbon dioxide, has become a major issue that humans need to solve urgently, and this has reached a global consensus. The carbon emissions of my country's transportation industry cannot be ignored. It accounts for about 10% of my country's total carbon emissions and continues to grow rapidly. The transportation industry has therefore become a key industry for achieving carbon reduction goals. In the process of moving towards carbon neutrality and carbon peak, energy conservation and emission reduction in the industry are urgent. Solving the problem of greenhouse gas emissions in the transportation industry is undoubtedly the core task of the transportation sector to carry out low-carbon actions. In this context, comprehensive decision-making data is urgently needed to assist energy conservation and emission reduction and promote the industry's transformation to green and low-carbon.
[0003] At present, in terms of policy analysis, traditional research mainly relies on manual interpretation of policy texts word by word, relying on the experience and knowledge of professionals to judge the focus and potential impact of policies. In the data processing stage, the common practice is to use general text analysis software, which uses basic word segmentation algorithms and simple word frequency statistics to perform preliminary processing on the text. In terms of model construction, most of them directly use general machine learning models, such as some traditional classification models, and input simply processed data into them for training and prediction.
[0004] However, these existing technical means have serious deficiencies when it comes to policy analysis and effect evaluation in the field of transportation. Although manual interpretation methods have a certain degree of professionalism, they are extremely inefficient and lack unified standards, making it difficult to process a large number of transportation environment policy texts on a large scale. General text analysis software is not optimized for professional terms and specific language expressions in the field of transportation, and cannot accurately identify key words in the industry, resulting in the loss of important information during word segmentation and word frequency statistics, and cannot provide high-quality data support for subsequent analysis. General machine learning models also do not take into account the uniqueness of the transportation industry. During feature extraction and model training, they cannot effectively capture the industry rules and semantic logic in transportation environment policies, resulting in a large deviation between the prediction results and the actual situation, and cannot provide decision makers with targeted and reliable decision-making recommendations, making it difficult to meet the actual needs of the transportation field for policy effect evaluation. Summary of the invention
[0005] The present invention proposes a method and system for evaluating transportation carbon emission policies, which solves the problem that existing policy analysis methods are insufficiently targeted in the field of transportation carbon emission.
[0006] In order to solve the above technical problems, the present invention provides a method for evaluating transportation carbon emission policies, comprising the following steps:
[0007] Step S1: Obtain several policy documents on carbon emissions from transportation, perform word segmentation and word frequency analysis on the text content in all policy documents, construct an industry vocabulary based on the word segmentation results, extract keywords from the industry vocabulary, and update the word segmenter using high-frequency words and keywords in the word segmentation;
[0008] Step S2: construct a policy evaluation index based on the relationship between the GDP and carbon emissions in the corresponding year of the policy document, and classify the policy document into a policy with obvious effect or a policy with insignificant effect according to the comparison between the value of the policy evaluation index and the set threshold, and mark the policy document according to the classification result;
[0009] Step S3: construct a prediction model, embed the prediction model based on the industry vocabulary, and train the prediction model after word embedding using the annotated policy documents;
[0010] Step S4: Use the updated word segmenter to segment the policy document to be evaluated, input the segmentation results into the trained prediction model, and determine whether the effect of the input policy document is obvious.
[0011] Preferably, the step S1 of extracting keywords from the industry vocabulary comprises the following steps:
[0012] Step S11: Calculate the word frequency TF and inverse document frequency IDF of each word segment in the industry vocabulary, and calculate the first importance score of the word segment according to TF and IDF. The expression for calculating the first importance score is:
[0013] TF-IDF(t,d)=TF(t,d)×IDF(t);
[0014]
[0015]
[0016] In the above formula, TF-IDF(t,d) is the first importance score of the word segmentation; TF(t,d) is the word frequency of the word segmentation; IDF(t) is the inverse document frequency of the word segmentation; f t,d is the number of times the segmentation word t appears in the environmental policy text d; D is the total number of segmentations; N is the total number of environmental policy texts;
[0017] Step S12: Based on the word segmentation network constructed according to the relationship between all word segments, the weight of each word in the word segmentation network is calculated as the second importance score of the word segmentation. The expression for calculating the second importance score is:
[0018]
[0019] Where S(t) is the second importance score of the word; ξ is the damping coefficient; ln(t i ) is the pointing participle t i out(t j ) is the word segmentation network t j The out-degree of
[0020] Step S13: weighted fusion of the first importance score and the second importance score is performed to obtain an importance score for each word segment. The expression for calculating the importance score is:
[0021] Score(t)=α×TF-IDF(t)+(1-α)×S(t);
[0022] In the formula, α is the weight adjustment parameter;
[0023] Step S14: sort all the participles in descending order of the importance scores, and select the first N participles as keywords.
[0024] Preferably, the expression of the policy evaluation index in step S2 is:
[0025]
[0026] In the above formula, DPR is the carbon emission efficiency ratio; PEIR is China PEIR is the domestic policy effect improvement rate; Global is the global policy effect improvement rate; China,i is the ratio of domestic GDP to carbon emissions in year i; Global,i is the ratio of global GDP to carbon emissions in year i; GDP China,i ,CE China,i They are the gross domestic product and carbon emissions; GDP Global,i ,CE Global,i They are the world's gross domestic product and carbon emissions respectively.
[0027] Preferably, the step S3 of embedding the prediction model based on the industry vocabulary comprises the following steps:
[0028] Step S31: define the window size, and for each word in the industry vocabulary, count the co-occurrence frequency of the word with other words in the window to obtain the co-occurrence probability of all word pairs;
[0029] Step S32: randomly sampling a central word from the industry vocabulary and obtaining context word pairs of the central word, and obtaining a predicted context word probability distribution of the sampled central word through a prediction model;
[0030] Step S33: Calculate the loss function of the prediction model based on the predicted context word probability distribution and the actual context words, update the parameters of the prediction model through the optimizer, obtain the word vector of each word, and form an industry embedding matrix;
[0031] Step S34: Use the concatenation method to introduce the industry word embedding matrix into the embedding layer of the prediction model.
[0032] Preferably, the expression for calculating the co-occurrence probability of the word pair in step S31 is:
[0033]
[0034] In the formula, p(w j |w i ) is the word w i and the word w j The co-occurrence probability of ij For word w i and the word w j The number of co-occurrences of industry A glossary of terms for the industry.
[0035] Preferably, the expression for calculating the loss function in step S33 is:
[0036]
[0037] In the formula, p(w j |w i ) is the word w i and the word w j The co-occurrence probability of Context(w i ) is the pair of word w i Analysis of the context; V industry A glossary of terms for the industry.
[0038] Preferably, in step S1, obtaining historical environmental policy texts in the field of transportation through a proxy pool and a crawler includes the following steps:
[0039] Step S101: Collect several proxy IPs, verify the validity of the collected proxy IPs, store the verified proxy IPs in a database, and build a proxy pool;
[0040] Step S102: randomly assigning proxy IPs from a proxy pool, and using crawler software to crawl a large amount of environmental policy texts in the field of transportation from public websites through the proxy IPs;
[0041] Step S103: Regularly verify and update the proxy IP in the proxy pool.
[0042] The present invention also provides a transportation carbon emission policy evaluation system, which is implemented based on the above-mentioned transportation carbon emission policy evaluation method, and includes: a data acquisition module, a text preprocessing module, a policy effect evaluation module, a model training module and a text prediction module;
[0043] The data collection module collects text data such as transportation carbon emission regulations, policies and academic papers;
[0044] The text preprocessing module: constructs a stop word list, performs text cleaning on the collected text data, performs word segmentation on the cleaned text data and counts the word frequency, extracts keywords from all word segments, and uses the top 100 high-frequency words and keywords to update the user dictionary of the word segmenter;
[0045] The policy effect evaluation module: constructs evaluation indicators for evaluating the effects of transportation carbon emission policies, divides policies into two categories: obvious effects and insignificant effects, and annotates policy texts according to policy effects;
[0046] The model training module: constructs a prediction model, adds word embedding based on the industry vocabulary to the embedding layer of the prediction model, and uses the two types of labeled policy text data and the prediction model for training;
[0047] The text prediction module pre-processes the transportation carbon emission policy text to be evaluated, inputs the processed policy text into a trained prediction model, and the prediction model outputs a prediction result of whether the effect of the policy is obvious.
[0048] Preferably, the system also includes a user interaction module, which includes two pages. The first page is used to display the latest traffic and carbon emission related policies and regulations, and the second page accepts policy texts input by users. After the background processes the article, the word cloud diagram and word meaning network analysis diagram of the policy text input by the user are displayed on the second page to help users understand the content and semantic relationship of the policy text in an intuitive way.
[0049] Preferably, the data collection module constructs a proxy pool, crawls text data related to transportation carbon emissions from public websites based on the constructed proxy pool, and regularly updates the proxy IPs in the proxy pool.
[0050] The benefits of the present invention include at least:
[0051] 1. By specifically acquiring historical environmental policy texts in the field of transportation and constructing a stop word list in this field, the text can be accurately preprocessed; in the process of word segmentation, word frequency statistics and keyword extraction, the core content of the transportation environmental policy is focused on, key information is effectively screened out, and a high-quality data foundation is provided for subsequent analysis and model training, ensuring that the data is closely integrated with the actual application scenarios and avoiding interference from irrelevant information;
[0052] 2. Constructing an industry vocabulary and performing word embedding enhances the prediction model’s ability to understand and process professional terms in the transportation field. Using two types of environmental policy texts for training enables the model to fully learn the text features and semantic patterns of policies with different effects, continuously optimize its own parameters, improve the accuracy and reliability of predictions, and better adapt to the task requirements of predicting the effects of transportation environmental policies. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 A schematic diagram of a method flow chart of an embodiment of the present invention;
[0054] Figure 2 A schematic diagram of a proxy pool construction process according to an embodiment of the present invention;
[0055] Figure 3 A schematic diagram of a proxy pool framework according to an embodiment of the present invention;
[0056] Figure 4 A flowchart of word frequency analysis according to an embodiment of the present invention is shown;
[0057] Figure 5 A schematic diagram of a flow chart of word meaning network analysis according to an embodiment of the present invention;
[0058] Figure 6 A flowchart of text prediction according to an embodiment of the present invention;
[0059] Figure 7 Schematic diagram of the system framework of an embodiment of the present invention. DETAILED DESCRIPTION
[0060] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of the present invention.
[0061] like Figure 1 As shown, an embodiment of the present invention provides a method for evaluating a transportation carbon emission policy, comprising the following steps:
[0062] Step S1: Obtain several policy documents on transportation carbon emissions through proxy pools and crawlers, perform word segmentation and word frequency analysis on the text content in all policy documents, build an industry vocabulary based on the word segmentation results, extract keywords from the industry vocabulary, and use the high-frequency words and keywords in the word segmentation to update the word segmenter.
[0063] Specifically, in the embodiments of the present invention, a large number of policy documents in the transportation field are crawled on public websites through a Redis proxy pool and Python software.
[0064] As Figure 2 shown, constructing a Redis proxy pool includes the following steps:
[0065] Step S101: Collect a number of proxy IPs, and verify the validity of the collected proxy IPs. Specifically, verify the connectivity, anonymity, speed, and stability of the proxy IPs.
[0066] Step S102: Store the verified proxy IPs in a database to construct a Redis proxy pool.
[0067] Step S103: Randomly or according to a strategy allocate proxy IPs from the Redis proxy pool. Through the proxy IPs, use crawler software to crawl a large number of environmental policy texts in the transportation field on public websites.
[0068] Step S104: Regularly verify and update the proxy IPs in the Redis proxy pool.
[0069] As Figure 3 shown is a schematic structural diagram of the proxy pool constructed in the embodiments of the present invention. After constructing the Redis proxy pool, through the proxy IPs, use Python to crawl a large number of laws, policies, and standard documents related to transportation carbon emissions on public websites.
[0070] Unify and parse the collected text data and save it as a.txt file for subsequent processing. Clean the collected text data to remove special characters, punctuation marks, numbers, HTML tags, etc. At the same time, unify the text format and filter out blank lines and extra spaces to optimize the text quality.
[0071] Construct a special stop word list for the transportation policy field and剔除 words that do not have discrimination. Include general words such as "is", "of", "and", etc., and words in non-related fields such as "agriculture", "construction", "manufacturing", "electric power", etc.
[0072] As Figure 4 shown is a schematic flowchart of word frequency analysis in the embodiments of the present invention. Segment the cleaned text by calling the Jieba library in Python to generate a list of word segmentation results. Check all word segments to determine whether all word segments are in the stop word character table, and detect whether the length of the word segments is greater than 1. Filter out the word segments that are in the stop word list or have a length less than 1. Convert the filtered word segment list into a long text in string format, and use the functions in the WordCloud library to generate a word cloud image to display the word segments in the historical environmental policy text for users.
[0073] The word frequency in the word segmentation results is counted, and the collections.Counter function is used to identify high-frequency words. After analyzing the high-frequency words in the policy text, it is also necessary to identify the keywords therein. Therefore, the embodiment of the present invention innovatively combines the TF-IDF and TextRank algorithms to extract keywords from the word segmentation results. Among them, the TF-IDF algorithm is a statistical method for evaluating the importance of words in a text. It measures the relative importance of words in a specific document by calculating the word frequency (TF) and the inverse document frequency (IDF), thereby effectively filtering out common words and highlighting characteristic words. The TextRank algorithm is an unsupervised keyword extraction algorithm based on a graph model. It regards the words in the text as nodes in the graph, establishes edges according to the co-occurrence relationship of the words, and uses the PageRank algorithm to calculate the weight of each word, thereby extracting the most representative keywords. The embodiment of the present invention combines the two algorithms to extract keywords from the word segmentation results, including the following steps:
[0074] Step S111: Calculate the term frequency TF and inverse document frequency IDF of each word segment, and calculate the first importance score of the word segment according to TF and IDF. The expression for calculating the first importance score is:
[0075] TF-IDF(t,d)=TF(t,d)×IDF(t);
[0076]
[0077]
[0078] In the above formula, TF-IDF(t,d) is the first importance score of the word segmentation; TF(t,d) is the word frequency of the word segmentation; IDF(t) is the inverse document frequency of the word segmentation; f t,d is the number of times the segmentation word t appears in the environmental policy text d; D is the total number of segmentations; N is the total number of environmental policy texts.
[0079] Step S12: Based on the word segmentation network constructed according to the relationship between all word segments, the weight of each word in the word segmentation network is calculated as the second importance score of the word segmentation. The expression for calculating the second importance score is:
[0080]
[0081] Where, S(t) is the second importance score of the word segment; ξ is the damping coefficient, which is 0.85 in the embodiment of the present invention; ln(t i ) is the pointing participle t i out(t j ) is the word segmentation network t j The out-degree of .
[0082] Step S13: weighted fusion of the first importance score and the second importance score is performed to obtain the importance score of each word segment. The expression for calculating the importance score is:
[0083] Score(t)=α×TF-IDF(t)+(1-α)×S(t);
[0084] Wherein, α is a weight adjustment parameter, and in the embodiment of the present invention, the value is 0.5.
[0085] Step S14: sort all the segmented words in descending order of importance scores, and select the first N segmented words as keywords.
[0086] The first 100 high-frequency words and the extracted keywords are automatically added to the custom word library file, and the user dictionary of the Jieba word segmenter is dynamically updated to achieve adaptive optimization of the word segmentation library. The pickle library is used to persist the generated word segmentation model and word frequency data for subsequent calls.
[0087] like Figure 5 The flowchart of the semantic network analysis of the embodiment of the present invention is shown. By calling the segmentation function in the Jieba library, all environmental policy texts are segmented, and all segmented words are checked to see if they are in the stop word character table, and the length of the segmented words is detected to see if it is greater than 1, and words that do not meet these two conditions are filtered out. The word frequency of all filtered segmented words is calculated, and the top N words with the highest word frequency are selected as keywords. The keywords are used as nodes, and the values in the association matrix of the segmented words are used as the weights of the edges. If two keywords appear in the same sentence, the values of the corresponding positions in the association matrix are increased by a set value, and the graph structure is constructed using the Networkx library to generate a PNG file. The latest policies and regulations related to transportation carbon emissions are displayed through a Web application based on PyWebIO; the relevant libraries used in Python include requests, json, scrapy, time, jieba, redis, BeautifulSoup, and threading.
[0088] Step S2: Construct a policy evaluation index based on the relationship between the GDP and carbon emissions in the corresponding year of the policy document. According to the comparison between the value of the policy evaluation index and the set threshold, classify the policy document into a policy with obvious effect or a policy with insignificant effect, and mark the policy document according to the classification result.
[0089] Specifically, constructing an evaluation index system to judge whether the environmental policy is effective includes the following steps:
[0090] Step S21: crawl all articles in the carbon trading online policy and regulations section through a crawler, and save the release date of each article. The metadata of each article includes policy content, issuing unit, policy category, release time, etc. Calculate the domestic and global policy effect improvement rate PEIR based on the carbon emissions and GDP of the corresponding year of the policy text:
[0091]
[0092] Where PEIR China PEIR is the domestic policy effect improvement rate; Global is the global policy effect improvement rate; China,i is the ratio of domestic GDP to carbon emissions in year i; Global,i is the ratio of global GDP to carbon emissions in year i; GDP China,i ,CE China,i They are the gross domestic product and carbon emissions; GDP Global,i ,CE Global,i They are the world's gross domestic product and carbon emissions respectively.
[0093] Step S22: Calculate the domestic and global carbon emission efficiency ratios DPR based on PEIR:
[0094]
[0095] Step S23: When DPR is greater than 20%, the carbon emission policy effect of the previous year is considered significant; otherwise, it is considered insignificant. The two types of policy texts are sorted into two json files, each with the following structure: {"policy content": "specific text", "effect classification": "significant / insignificant", "release date": "YYYY-MM-DD"}, which is convenient for the subsequent BERT model to call and classify.
[0096] Step S3: Build a prediction model, embed the prediction model based on the industry vocabulary, and use the annotated policy documents to train the prediction model after word embedding.
[0097] Specifically, by calling the contents of Torch, Transformers, and Sklearn libraries and defining the SingelInputBert class to build a prediction model, the BERT model is adopted in the embodiment of the present invention. Considering the diversity of policy texts, in order to further improve the classification accuracy of the BERT model for long texts, the embodiment of the present invention adds pre-trained word embedding based on industry vocabulary to the embedding layer of BERT to enhance the model's recognition ability for professional terms such as transportation and carbon emissions.
[0098] Among them, the Word2Vec method is used to generate embedding vectors, and the BERT model is pre-trained for word embedding based on the industry vocabulary, including the following steps:
[0099] Step S31: Collect a large amount of domain text data containing industry-specific vocabulary from industry documents such as traffic carbon emission regulations, policies and related academic papers, pre-process the corpus by removing stop words, punctuation marks, non-language symbols, and use Jieba to perform text segmentation to obtain the industry vocabulary table V industry .
[0100] Step S32: define the window size, and for each word in the industry vocabulary, count the co-occurrence frequency of the word with other words in the window to obtain the co-occurrence probability of all word pairs:
[0101]
[0102] In the formula, p(w j |w i ) is the word w i and the word w j The co-occurrence probability of ij For word w i and the word w j The number of co-occurrences.
[0103] Step S33: Using the Skip-gram model, randomly sampling a central word from the industry vocabulary and obtaining context word pairs of the central word, and obtaining the predicted context word probability distribution of the sampled central word through the prediction model.
[0104] Step S34: Calculate the loss function of the prediction model based on the predicted context word probability distribution and the actual context words:
[0105]
[0106] Where L is the loss function; Context(w i ) is the pair of word w i Analysis of the context.
[0107] Step S35: Update the parameters of the prediction model through the optimizer to obtain the word vector of each word, form an industry embedding matrix, and introduce the industry word embedding matrix into the embedding layer of the prediction model using the splicing method.
[0108] Step S4: Use the updated word segmenter to segment the policy document to be evaluated, input the segmentation results into the trained prediction model, and determine whether the effect of the input policy document is obvious.
[0109] like Figure 6The figure shows a flow chart of text prediction in an embodiment of the present invention. The trained word segmenter and BERT model are loaded, and the Jieba library is used to segment the policy text input by the user, and specific characters and marks in the text content, such as full-width spaces, line breaks, years, etc., are removed to form a new text. The preprocessed text is input into the BERT model to predict the effect of the policy text and determine whether the policy is a relatively obvious policy or a relatively insignificant policy.
[0110] At the same time, the word frequency matrix of the policy text can be subjected to singular value decomposition to obtain the singular value of each word and the singular vector corresponding to each text, and a keyword-keyword semantic distance table and a text-keyword semantic distance table can be constructed to display the key content of the policy text to users.
[0111] like Figure 7 As shown, an embodiment of the present invention also provides an evaluation system for transportation carbon emission policies, which is implemented based on the above-mentioned evaluation method for transportation carbon emission policies, and includes: a data acquisition module, a text preprocessing module, a policy effect evaluation module, a model training module, a text prediction module and a user interaction module.
[0112] The data collection module is used to build a proxy pool. Based on the constructed proxy pool, text data related to transportation carbon emissions are crawled from public websites, and the proxy IPs in the proxy pool are updated regularly.
[0113] The text preprocessing module is used to build a stop word list, perform text cleaning on the collected text data, segment the cleaned text data and count the word frequency, extract keywords from all segmented words, and use the top 100 high-frequency words and keywords to update the user dictionary of the segmenter.
[0114] The policy effectiveness evaluation module constructs evaluation indicators for evaluating the effectiveness of transportation carbon emission policies, divides policies into two categories: those with obvious effects and those with insignificant effects, and annotates policy texts according to their effects.
[0115] The model training module is used to build a prediction model. Word embedding based on the industry vocabulary is added to the embedding layer of the prediction model, and the prediction model is trained using two types of labeled policy text data.
[0116] The text prediction module is used to pre-process the transportation carbon emission policy text to be evaluated, input the processed policy text into the trained prediction model, and the prediction model outputs the prediction result of whether the policy effect is obvious.
[0117] The user interaction module includes two pages. The first page is used to display the latest traffic and carbon emission related policies and regulations. The second page accepts the policy text entered by the user. After the background processes the article, the word cloud diagram and word meaning network analysis diagram of the policy text entered by the user are displayed on the second page, helping users understand the content and semantic relationship of the policy text in an intuitive way.
[0118] The system of the embodiment of the present invention integrates Python crawler technology and Redis proxy pool, and utilizes the proxy IP pool designed by Redis to achieve efficient and stable management of proxy IP. By collecting, verifying, storing, allocating and updating the proxy IP, the Python crawler task can be ensured to proceed smoothly. In the actual application of transportation, the data structure and operation method of Redis can be adjusted according to the needs to adapt to different carbon emission monitoring scenarios, and the real-time monitoring and data collection of policies and regulations related to transportation carbon emissions are realized. Compared with the traditional periodic data collection method, the system can provide more timely and accurate data support, and provide decision makers with the latest policy dynamics and carbon emissions. A user-friendly Web application interface is designed based on the PyWebIO library, so that non-professional users can also easily use the system for policy text analysis and prediction. This design lowers the threshold for the use of the system and expands the potential user group of the system. At the same time, the system allows easy integration of new data sources and analysis tools, and has good scalability. With the continuous development of the field of transportation carbon emissions, the system can flexibly adapt to new needs and challenges and maintain long-term technical competitiveness.
[0119] The method provided by the present invention can not only perform basic word frequency analysis on the policy text, but also use advanced text analysis technologies such as Jieba word segmentation and BERT model to deeply explore the semantic information of the policy text and generate word cloud diagrams and word meaning network diagrams. This enables decision makers to understand the policy content and potential impact more intuitively and deeply, and improves the depth and breadth of policy analysis. By introducing the text prediction function based on the pre-trained BERT model, it is possible to predict the policy effect and judge whether the policy is relatively effective or not. This prediction mechanism provides decision makers with early warning and response strategies for crisis transformation, improving the practicality of the system and the foresight of decision-making.
[0120] The method of the present invention applies data capture, data analysis and data prediction methods, is interactive and friendly, comprehensively uses relevant methods of artificial intelligence, and has a smooth text processing process. It provides new ideas and reference suggestions for the evaluation of transportation carbon emission policies, reduces the difficulty of decision-making, and has strong feasibility.
[0121] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. Only the preferred embodiments of the present invention are expressed. The description is more specific and detailed, but it cannot be understood as limiting the scope of the present invention. As long as there is no contradiction in the combination of these technical features, they should be considered as within the scope of this specification.
[0122] It should be noted that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these modifications and improvements all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.
Claims
1. A method for evaluating transportation carbon emission policies, characterized in that: The following steps are involved: Step S1: Obtain several policy documents on carbon emissions from transportation, perform word segmentation and word frequency analysis on the text content in all policy documents, construct an industry vocabulary based on the word segmentation results, extract keywords from the industry vocabulary, and update the word segmenter using high-frequency words and keywords in the word segmentation; Step S2: construct a policy evaluation index based on the relationship between the GDP and carbon emissions in the corresponding year of the policy document, and classify the policy document into a policy with obvious effect or a policy with insignificant effect according to the comparison between the value of the policy evaluation index and the set threshold, and mark the policy document according to the classification result; Step S3: construct a prediction model, embed the prediction model based on the industry vocabulary, and train the prediction model after word embedding using the annotated policy documents; Step S4: Use the updated word segmenter to segment the policy document to be evaluated, input the segmentation results into the trained prediction model, and determine whether the effect of the input policy document is obvious.
2. The method for evaluating a transportation carbon emission policy according to claim 1, characterized in that: The step S1 of extracting keywords from the industry vocabulary includes the following steps: Step S11: Calculate the word frequency TF and inverse document frequency IDF of each word segment in the industry vocabulary, and calculate the first importance score of the word segment according to TF and IDF. The expression for calculating the first importance score is: TF-IDF(t,d)=TF(t,d)×IDF(t); In the above formula, TF-IDF(t,d) is the first importance score of the word segmentation; TF(t,d) is the word frequency of the word segmentation; IDF(t) is the inverse document frequency of the word segmentation; f t,d is the number of times the segmentation word t appears in the environmental policy text d; D is the total number of segmentations; N is the total number of environmental policy texts; Step S12: Based on the word segmentation network constructed according to the relationship between all word segments, the weight of each word in the word segmentation network is calculated as the second importance score of the word segmentation. The expression for calculating the second importance score is: Where S(t) is the second importance score of the word; ξ is the damping coefficient; ln(t i ) is the pointing participle t i out(t j ) is the word segmentation network t j The out-degree of Step S13: weighted fusion of the first importance score and the second importance score is performed to obtain an importance score for each word segment. The expression for calculating the importance score is: Score(t)=α×TF-IDF(t)+(1-α)×S(t); In the formula, α is the weight adjustment parameter; Step S14: sort all the participles in descending order of the importance scores, and select the first N participles as keywords.
3. The method for evaluating a transportation carbon emission policy according to claim 1, characterized in that: The expression of the policy evaluation index in step S2 is: In the above formula, DPR is the carbon emission efficiency ratio; PEIR is China PEIR is the domestic policy effect improvement rate; Global is the global policy effect improvement rate; China,i is the ratio of domestic GDP to carbon emissions in year i; Global,i is the ratio of global GDP to carbon emissions in year i; GDP China,i ,CE China,i They are the gross domestic product and carbon emissions; GDP Global,i ,CE Global,i They are the world's gross domestic product and carbon emissions respectively.
4. The method for evaluating a transportation carbon emission policy according to claim 1, characterized in that: The step S3 of embedding the prediction model based on the industry vocabulary includes the following steps: Step S31: define the window size, and for each word in the industry vocabulary, count the co-occurrence frequency of the word with other words in the window to obtain the co-occurrence probability of all word pairs; Step S32: randomly sampling a central word from the industry vocabulary and obtaining context word pairs of the central word, and obtaining a predicted context word probability distribution of the sampled central word through a prediction model; Step S33: Calculate the loss function of the prediction model based on the predicted context word probability distribution and the actual context words, update the parameters of the prediction model through the optimizer, obtain the word vector of each word, and form an industry embedding matrix; Step S34: Use the concatenation method to introduce the industry word embedding matrix into the embedding layer of the prediction model.
5. The method for evaluating transportation carbon emission policy according to claim 4, characterized in that: The expression for calculating the co-occurrence probability of the word pair in step S31 is: In the formula, p(w j |w i ) is the word w i and the word w j The co-occurrence probability of ij For word w i and the word w j The number of co-occurrences of industry A glossary of terms for the industry.
6. A method for evaluating transportation carbon emission policies according to claim 4, characterized in that: The expression for calculating the loss function in step S33 is: Where L is the loss function; p(w j |w i ) is the word w i and the word w j The co-occurrence probability of Context(w i ) is the pair of word w i Analysis of the context; V industry A glossary of terms for the industry.
7. The method for evaluating a transportation carbon emission policy according to claim 1, characterized in that: In step S1, a number of policy documents on transportation carbon emissions are obtained through the proxy pool and crawler, including the following steps: Step S101: Collect several proxy IPs, verify the validity of the collected proxy IPs, store the verified proxy IPs in a database, and build a proxy pool; Step S102: randomly assigning proxy IPs from a proxy pool, and using crawler software to crawl a large amount of environmental policy texts in the field of transportation from public websites through the proxy IPs; Step S103: Regularly verify and update the proxy IP in the proxy pool.
8. A transportation carbon emission policy evaluation system, implemented based on a transportation carbon emission policy evaluation method as claimed in any one of claims 1 to 7, characterized in that: include: Data collection module, text preprocessing module, policy effect evaluation module, model training module and text prediction module; The data collection module collects text content such as transportation carbon emission regulations, policies and academic papers; The text preprocessing module: constructs a stop word list, performs text cleaning on the collected text content, performs word segmentation on the cleaned text data and counts the word frequency, extracts keywords from all word segments, and uses the top 100 high-frequency words and keywords to update the user dictionary of the word segmenter; The policy effect evaluation module: constructs evaluation indicators for evaluating the effects of transportation carbon emission policies, divides policies into two categories: obvious effects and insignificant effects, and annotates policy texts according to policy effects; The model training module: constructs a prediction model, adds word embedding based on the industry vocabulary to the embedding layer of the prediction model, and uses the two types of labeled policy text data and the prediction model for training; The text prediction module pre-processes the transportation carbon emission policy text to be evaluated, inputs the processed policy text into a trained prediction model, and the prediction model outputs a prediction result of whether the effect of the policy is obvious.
9. The transportation carbon emission policy evaluation system according to claim 8, characterized in that: The system also includes a user interaction module, which includes two pages. The first page is used to display the latest traffic and carbon emission related policies and regulations, and the second page accepts policy texts input by users. After the background processes the article, the word cloud diagram and word meaning network analysis diagram of the policy text input by the user are displayed on the second page to help users understand the content and semantic relationship of the policy text in an intuitive way.
10. The transportation carbon emission policy evaluation system according to claim 8, characterized in that: The data collection module constructs a proxy pool, crawls text data related to transportation carbon emissions from public websites based on the constructed proxy pool, and regularly updates the proxy IPs in the proxy pool.
Citation Information
Patent Citations
A new energy policy information extraction method and system
CN109766416A
Policy landing effect evaluation method and system based on policy completion degree analysis
CN114266496A
Public policy participation degree evaluation method and system based on LDA and vector space model
CN114528819A
Resource classification screening method and system based on science and technology policies
CN116644174A
Keyword extraction method, apparatus and server
US20190163690A1