Methods and systems for quantitative analysis of coal mine methane emission reduction and control policies
By analyzing coal mine methane emission reduction and control policy texts using a large language model and aligning them with multi-source monitoring data in time and space, the repeatability and accuracy issues of quantitative assessment in traditional methods are resolved, enabling efficient and reliable assessment of policy effectiveness and optimization of control strategies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INFORMATION RES INST OF EMERGENCY MANAGEMENT DEPT
- Filing Date
- 2026-02-04
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies are insufficient for efficiently quantifying and evaluating coal mine methane emission reduction and control policies. Traditional methods, based on macro-statistics and expert experience, lack repeatability and accuracy, making it difficult to meet the needs of constructing quantitative evaluation technology models.
A large language model is used to parse policy text data. Text blocks are segmented according to the principle of semantic coherence. Combined with multi-dimensional policy evaluation, structured policy elements are generated and spatiotemporally aligned with multi-source monitoring data to construct a policy-monitoring related dataset.
It enables automated and highly accurate quantitative evaluation of coal mine methane emission reduction and control policies, bridges policy texts with monitoring data, and provides objective and dynamic policy effectiveness evaluation and support for optimizing control strategies.
Smart Images

Figure CN122065776B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of data mining and information planning technology, specifically to a method and system for quantitatively analyzing coal mine methane emission reduction and control policies. Background Technology
[0002] Against the backdrop of the "dual carbon" goals and comprehensive greenhouse gas emission reduction, methane, as a pollutant with a greenhouse effect potential far exceeding that of carbon dioxide, has become a key target for global coordinated control of greenhouse gases. Methane released during coal mine production accounts for a significant proportion of my country's total anthropogenic methane emissions, and its emission control has gradually evolved from a traditional ancillary aspect of safe production to an independent and urgent governance task in addressing climate change. In recent years, central and local government departments in my country have successively issued a series of special policies and regulations concerning methane emissions from coal mines. Simultaneously, relying on modern monitoring technologies such as the Internet of Things, satellite remote sensing, and underground sensor networks, relatively complete time-series monitoring data on methane emissions has been accumulated.
[0003] Based on the aforementioned technological foundation, by spatiotemporal coupling and causal correlation analysis of policy implementation nodes and multi-dimensional monitoring data, a quantitative assessment technology model that combines policy capability analysis and emission data verification functions is constructed. By applying this model, data support and quantitative decision-making basis are provided for the precise optimization and differentiated adjustment of coal mine methane emission reduction policies. This has become one of the effective technical paths to connect precise coal mine methane governance with the implementation of the "dual carbon" strategy.
[0004] In implementing this technological path, a quantitative assessment of coal mine methane emission reduction and control policies is required. Existing research in this area includes evaluations of climate, energy, and environmental policies, such as the paper "Low-Carbon Policy Intensity" published by the Peking University team (https: / / www.nature.com / articles / s41597-024-03033-5), which proposes several policy intensity indices or quantitative policy indicators. However, such research primarily relies on macro-level statistics and simplified analyses of policy texts. The results do not meet the requirements for constructing the aforementioned quantitative assessment models. For instance, in traditional econometric frameworks (such as the STIRPAT model), the impact of policies is often simplified to dummy variables (0 or 1 representing policy implementation) or coarse proxies based on policy document counts, compressing their rich connotations, implementation strength, and evolutionary dynamics. On the other hand, policy evaluation practices in the coal mine methane sector mainly rely on expert experience and judgment, and qualitative analysis is conducted by manually interpreting the binding force and importance of policy documents. This approach is difficult to handle large-scale, long-term policy text data, and the evaluation standards are highly subjective and lack repeatability, which also makes it difficult to meet the needs of constructing the aforementioned quantitative evaluation technology model.
[0005] Based on this, in order to effectively implement the above-mentioned technical path, it is urgent to solve the technical problem of constructing a policy quantitative framework for coal mine methane emission reduction and governance, and proposing an automated implementation method closely coupled with the quantitative framework to efficiently achieve quantitative analysis of coal mine methane emission reduction and governance policies. Summary of the Invention
[0006] To at least partially overcome the problems existing in related technologies, this application proposes a method and system for quantitative analysis of coal mine methane emission reduction and control policies. Based on text analysis technology and large language model technology, it efficiently realizes the quantitative analysis of coal mine methane emission reduction and control policies within a constructed policy quantification framework.
[0007] First aspect This application provides a method for quantitatively analyzing coal mine methane emission reduction and control policies, the method comprising: Obtain policy text data on methane emission reduction and control in coal mines; The policy text data is parsed and processed by calling a large language model to extract the structured policy elements and generate corresponding quantitative evaluation information; The quantitative assessment information is spatiotemporally aligned and matched with multi-source monitoring data of methane in coal mines to generate a policy-monitoring correlation dataset. This policy-monitoring correlation dataset is used to drive or validate the quantitative assessment technology model for methane emissions from coal mines.
[0008] In one possible implementation, the large language model parses and processes the policy text data, including: Determine the quantized value of the text length of the policy text data; When the text length quantization value exceeds the first length threshold, the policy text data is divided into multiple continuous text blocks based on the principle of text semantic coherence. For each of the continuous text blocks, the large language model is invoked for parsing to obtain the corresponding local quantitative information. The local quantitative information is then aligned according to a preset policy evaluation dimension. The maximum value of the quantitative score of each text block under the same policy evaluation dimension is selected as the final score of that policy evaluation dimension, thereby generating the quantitative evaluation information of the policy text data.
[0009] In one possible implementation, the process of segmenting the policy text data into multiple consecutive text blocks based on the principle of textual semantic coherence specifically includes: The policy text data is broken down into multiple basic semantic units that maintain semantic integrity; For each of the basic semantic units, a greedy algorithm based on a length threshold is used to sequentially package and combine the basic semantic units. The specific process is as follows: Starting from the first basic semantic unit, subsequent units are added to the current text block in sequence. At the same time, the text length quantization value of the current text block is determined. When adding the next unit will cause the text length quantization value to exceed the second length threshold, the current block is stopped from being packaged and the text block is output. Then, starting with the last N basic semantic units of the previous text block, the aforementioned packing process is repeated to generate the next text block, where N is the preset number of overlapping units, thereby ensuring semantic coherence between adjacent text blocks.
[0010] In one possible implementation, basic semantic units are decomposed based on regular expression matching according to pre-built semantic boundary recognition rules. These semantic boundary recognition rules include: The paragraph boundary recognition rule uses a regular expression with multiple line breaks to match the paragraph separation positions in the text in order to segment independent paragraph units; The sentence boundary recognition rule uses regular expressions combining sentence punctuation marks to match the end positions of Chinese and English sentences, and performs secondary segmentation on independent paragraph units whose length exceeds a preset threshold; The paragraph boundary recognition rule takes precedence over the sentence boundary recognition rule.
[0011] In one possible implementation, a custom token counting function is used to calculate and determine the text length quantification value of the policy text data or text block; The custom token counting function is calculated by configuring differentiated weight coefficients for different character categories, and the weight coefficients are determined based on the lexicalization characteristics of the dedicated large language model.
[0012] In one possible implementation, the step of calling a large language model to parse and process the policy text data further includes: When the text length quantization value does not exceed the first length threshold, the large language model is directly invoked to parse and process the policy text data.
[0013] In one possible implementation, the process of generating the quantitative evaluation information for the policy text data further includes: A source identification is established for the quantitative score of each policy evaluation dimension. The source identification includes the chapter number, paragraph position index and key sentence location information of the original policy text, forming a score-original text mapping relationship table. Based on the score-original text mapping table and the quantitative scores of each policy evaluation dimension, a structured data file is generated for export. The structured data file contains a multidimensional data table, which includes a main table of policy quantitative scoring, a table of original text references, and a table of dimension analysis descriptions.
[0014] In one possible implementation, the process of spatiotemporally aligning and matching the quantitative assessment information with multi-source remote sensing monitoring data of methane in coal mines includes: Based on the effective date, termination date, and revision chain of the coal mine methane emission reduction and control policy documents, the quantitative assessment information of the policies is allocated to the corresponding effective years according to the preset time validity rules; The quantitative assessment information of the same region for each year is accumulated according to the administrative division level to generate a two-dimensional policy intensity matrix of region-year. The row vector of the matrix represents the annual policy intensity index sequence of a specific region, and the column vector represents the regional policy intensity distribution sequence of a specific year. Row and column aggregation operations are performed on the region-year two-dimensional policy intensity matrix to generate a multi-level time series dataset of policy intensity at the national and provincial levels.
[0015] In one possible implementation, obtaining the policy text data on methane emission reduction and control in coal mines includes: The effectiveness of coal mine methane emission reduction and control policies is verified based on the collected location information, and the location information of the verified documents is marked as a valid input source. Extract and parse the file content corresponding to the valid input source to generate the coal mine methane emission reduction and control policy text data; The validity verification includes at least one of format verification, integrity verification, and source credibility verification. Second aspect This application provides a system for quantitatively analyzing coal mine methane emission reduction and control policies, the system comprising: The acquisition and processing module is used to acquire policy text data on methane emission reduction and control in coal mines; The analysis and processing module is used to call the large language model to parse and process the policy text data, extract the structured policy elements and generate corresponding quantitative evaluation information. The matching processing module is used to perform spatiotemporal alignment matching between the quantitative assessment information and the multi-source monitoring data of coal mine methane to generate a policy-monitoring association dataset. The policy-monitoring association dataset is used to drive or verify the quantitative assessment technology model of coal mine methane emissions.
[0016] The technical solution provided in this application utilizes a large language model to perform in-depth analysis and quantification of policy texts. This enables automated and highly accurate extraction of structured policy elements from policy clauses for quantitative evaluation, overcoming the shortcomings of traditional macro-statistical methods, such as insufficient analytical depth and the subjective and inefficient nature of expert interpretation. It achieves standardized and repeatable quantitative representation of policy content. Furthermore, by aligning and matching the generated quantitative evaluation information with multi-source monitoring data of methane in coal mines in time and space, a "policy-monitoring" related dataset is constructed that can directly drive or verify downstream quantitative evaluation technology models. This effectively bridges the gap between the semantics of policy texts and physical monitoring data, providing a reliable data foundation and technical support for objectively and dynamically evaluating policy implementation effects and optimizing emission reduction and control strategies. Consequently, it significantly improves the automation level and decision support capabilities of coal mine methane emission reduction and control policy analysis and evaluation. Attached Figure Description
[0017] Figure 1 A flowchart illustrating a method for quantitatively analyzing coal mine methane emission reduction and control policies, provided as an embodiment of this application; Figure 2 A flowchart illustrating a method for quantitatively analyzing coal mine methane emission reduction and control policies, provided as another embodiment of this application; Figure 3 This is a schematic diagram of the block structure of a system for quantitatively analyzing coal mine methane emission reduction and control policies, provided as an embodiment of this application. Figure 4 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. Detailed Implementation
[0018] To make the purpose, technical solution and advantages of this application clearer, the technical solution of this application will be described in detail below.
[0019] As described in the background section, in current practice, by combining policy implementation nodes with multi-dimensional monitoring data in a spatiotemporal manner and conducting causal correlation analysis, a quantitative assessment technology model that combines policy capability analysis and emission data verification functions is constructed. By applying this model, data support and quantitative decision-making basis are provided for the precise optimization and differentiated adjustment of coal mine methane emission reduction policies. This has become one of the effective technical paths to connect precise coal mine methane governance with the implementation of the "dual carbon" strategy.
[0020] In implementing this technological path, a quantitative assessment of coal mine methane emission reduction and control policies is required. Existing research in this area includes evaluations of climate, energy, and environmental policies, such as the paper "Low-Carbon Policy Intensity" published by the Peking University team (https: / / www.nature.com / articles / s41597-024-03033-5), which proposes several policy intensity indices or quantitative policy indicators. However, such research primarily relies on macro-level statistics and simplified analyses of policy texts. The results do not meet the requirements for constructing the aforementioned quantitative assessment models. For instance, in traditional econometric frameworks (such as the STIRPAT model), the impact of policies is often simplified to dummy variables (0 or 1 representing policy implementation) or coarse proxies based on policy document counts, compressing their rich connotations, implementation strength, and evolutionary dynamics. On the other hand, policy evaluation practices in the coal mine methane sector mainly rely on expert experience and judgment, and qualitative analysis is conducted by manually interpreting the binding force and importance of policy documents. This approach is difficult to handle large-scale, long-term policy text data, and the evaluation standards are highly subjective and lack repeatability, which also makes it difficult to meet the needs of constructing the aforementioned quantitative evaluation technology model.
[0021] Based on this, this application proposes a method for quantitative analysis of coal mine methane emission reduction and control policies. Based on text analysis technology and large language model technology, it can efficiently achieve quantitative analysis of coal mine methane emission reduction and control policies within the constructed policy quantification framework.
[0022] like Figure 1 As shown, in one embodiment, this application proposes a method for quantitatively analyzing coal mine methane emission reduction and control policies. The method can be implemented by a computing device, and includes the following steps: Step S110: Obtain policy text data on methane emission reduction and control in coal mines; It is easy to understand that this step is to obtain the policy text data to be evaluated. In practice, this relevant policy data can be collected in advance through manual or automated collection methods to form the required policy text data.
[0023] For example, the suffix of search websites can be restricted using relevant web search tools (such as gov.cn), and a webpage "traversal" search method can be used to collect original policy data. The search scope covers publicly available documents from relevant departments of the central government and 31 provincial-level administrative regions across the country (excluding Hong Kong, Macao, and Taiwan). The data obtained includes policy and standard texts related to coalbed methane, coal mine gas emission standards, and related control and management measures.
[0024] The collected data may be in the form of PDF, HTML, etc. It is then processed into plain text format using relevant conversion tools to obtain policy text data. For example, a general text reading tool can be used to extract the original data content and convert the text in the web page or PDF into a clean text string to form the required policy text data.
[0025] Then, in step S120, the large language model is called to parse and process the policy text data, extract the structured policy elements, and generate corresponding quantitative assessment information.
[0026] As is known to those skilled in the art, large language models typically employ a Transformer decoder-only architecture, possessing parameters ranging from billions to tens of billions or even trillions. They rely on self-attention mechanisms to efficiently capture long-distance semantic dependencies in text. After large-scale unsupervised pre-training and subsequent alignment training, they can achieve fundamental capabilities such as text understanding, information extraction, and semantic reasoning in general domains. In this step of the application, this characteristic of large language models is utilized to achieve the parsing and processing of policy text data based on relevant prompt word engineering.
[0027] It should be noted that while existing technologies include applications of large language models, these applications differ significantly from the quantitative process specifically developed for the coal mine methane sector in this application. This application constructs a specialized terminology database for coal mine methane and a set of policy scoring quantitative standards. Based on these two core pieces of information, it designs domain-specific prompt templates and repeatedly adjusts the quantitative rules and optimizes the process in practice within the coal mine methane sector, ultimately forming a quantitative methodology chain with precise domain-specific knowledge adaptation capabilities. In contrast, existing tools do not construct a systematic policy element structure and scoring rules, and lack an integrated framework specifically designed for the coal mine methane sector, failing to achieve domain-specific model optimization design.
[0028] Furthermore, to improve the consistency and controllability of the output of the large language model in the quantitative task of my country's coal mine methane emission reduction policy, in a specific implementation, the large language model used in step S120 above is a dedicated large language model. This dedicated large language model is obtained by supervised fine-tuning of the basic large language model.
[0029] Specifically, the training set used for this oversight and fine-tuning includes normative texts in the field of coal mine methane and policy text samples annotated by experts; the expected output of the policy text samples is a predefined structured format, which includes at least one of the following: the policy issuing entity, evidence of policy provisions, and quantitative judgment results generated based on preset rules.
[0030] In other words, during the pre-supervised fine-tuning of the basic large language model to obtain a specialized large language model for practical applications, the supervised training set consists of materials and expert-annotated samples related to coal mine methane, including but not limited to normative documents on coal mine safety production, methane / gas monitoring and measurement specifications, original policy clauses and their corresponding element annotations and expected outputs. The expected output adopts a fixed structured format and should at least include the key fields required for policy quantification (such as the identification results of the issuing entity / level, excerpts of evidence related to coal mine methane clauses, and quantification results or intermediate judgments obtained according to established rules). Through supervised learning of the "input (clause or paragraph) - output (structured result)" sample pairs, the model can more stably produce results that meet the format requirements and are consistent with the rules when faced with similar policy texts.
[0031] Due to the characteristics of large language models, there are inherent constraints on the length of the context. For example, the maximum processing length is usually limited to 4096 to 32768 tokens, making it difficult to directly process complete policy text data. If the policy text is directly input into the model for evaluation, it often leads to the problem of key semantics being ignored and the evaluation accuracy decreasing because it exceeds the model's context window limit or the information density is too high. Based on this, this application specifically sets up relevant judgment and segmentation steps to effectively process such policy texts.
[0032] Specifically, in step S120, the large language model is invoked to parse and process the policy text data, including: Determine the text length quantification value of the policy text data. The text length quantification value can be the number of characters or the number of tokens. If tokens are used to quantify the text length, the text length quantification value can be determined by relevant token segmentation and statistical methods. When the text length quantization value exceeds the first length threshold, such as the first length threshold being 4096 Tokens, the policy text data is divided into multiple continuous text blocks based on the principle of text semantic coherence. Then, for each continuous text block, the large language model is called to parse and process it to obtain the corresponding local quantitative information. The local quantitative information is aligned according to the preset policy evaluation dimension. The maximum value of the quantitative score of each text block under the same policy evaluation dimension is selected as the final score of the policy evaluation dimension, and the quantitative evaluation information of the policy text data is generated. When the text length quantification value does not exceed the first length threshold, the large language model is directly called to parse and process the policy text data, and a similar multi-dimensional policy quantification scoring method is used to generate quantification evaluation information.
[0033] It should be noted that the multi-dimensional policy quantitative scoring method described above is part of a pre-constructed policy quantitative framework, which will be introduced in detail later. It will not be discussed further here.
[0034] As a specific implementation method, policy text data is segmented into multiple continuous text blocks based on the principle of textual semantic coherence, specifically including: The policy text data is broken down into multiple basic semantic units that maintain semantic integrity, namely, the extraction and division of basic semantic units; Then, a greedy algorithm based on a length threshold is used to sequentially package and combine the basic semantic units. The specific process is as follows: Starting from the first basic semantic unit, subsequent units are added to the current text block in sequence. At the same time, the text length quantization value of the current text block is determined. When adding the next unit will cause the text length quantization value to exceed the second length threshold, the current block is stopped from being packaged and the text block is output. Then, starting with the last N basic semantic units of the previous text block, the aforementioned packing process is repeated to generate the next text block, where N is the preset number of overlapping units, thereby ensuring semantic coherence between adjacent text blocks.
[0035] It should be noted that the number of overlapping units N here can be flexibly set according to the requirements of semantic coherence. Through this specific segmentation method, the content overlap between adjacent text blocks can be achieved, effectively avoiding the context breakage caused by block boundaries, and thus ensuring that the split text blocks have semantic continuity and relevance.
[0036] Furthermore, in order to effectively meet the segmentation requirements, the second length threshold should be less than or equal to the first length threshold in the above process, such as 3840 Tokens.
[0037] Furthermore, as a specific implementation method, basic semantic units can be decomposed based on regular expression matching according to pre-built semantic boundary recognition rules. Here, the semantic boundary recognition rules include: paragraph boundary recognition rules, which use regular expressions with multiple line breaks to match the paragraph separation positions in the text to segment independent paragraph units; sentence boundary recognition rules, which use regular expressions with sentence-end punctuation combinations to match the end positions of Chinese and English sentences, and perform secondary segmentation on independent paragraph units whose length exceeds a preset threshold; paragraph boundary recognition rules are executed before sentence boundary recognition rules.
[0038] This involves constructing a rule base for recognizing semantic boundaries in the text. Through regular expression matching and text structure analysis, it decomposes the original long text into basic semantic units. For example, targeting paragraph separation features, multiple line breaks (\\n{2,}) are used as paragraph boundary markers to initially split the text into multiple independent paragraph units. For paragraphs whose length still exceeds a preset threshold after splitting, sentence-level splitting is further performed. This is done by matching combinations of punctuation marks at the end of Chinese and English sentences (such as ".", "!", "?", ".", "!", "?" followed by spaces or line breaks) as sentence boundaries, further subdividing long paragraphs into several sentence units. By anchoring to the inherent boundaries of paragraphs and sentences in natural language, this approach ensures that the split basic units possess minimum semantic integrity, avoiding semantic fragmentation issues.
[0039] Furthermore, as a specific implementation method, the text length quantification value of policy text data or text blocks is calculated and determined by a custom token counting function; the custom token counting function is calculated by configuring differentiated weight coefficients for different character categories, and the weight coefficients are determined according to the lexicalization characteristics of the dedicated large language model.
[0040] For example, the character type differential weighting coefficient rule adopted is as follows: for Chinese characters (including Chinese punctuation), 1 character corresponds to 1 token; for non-Chinese characters (including English, numbers, symbols, etc.), 2 characters correspond to 1 token. This calculation method can better fit the actual model's token consumption logic for different types of characters.
[0041] Regarding the aforementioned long text splitting process, it should be noted that existing related technologies often employ mechanical segmentation methods based on a fixed number of characters or a single paragraph separator. These methods assess text size using a standardized length conversion standard, lacking consideration for the semantic structure of the text and the design of semantic connections between blocks. This can easily lead to semantic breaks and loss of integrity in the split text, and it is difficult to adapt to the actual processing logic of different types of characters in the model. In contrast, the text splitting method described in this application ensures the semantic integrity of the text through a hierarchical semantic splitting mechanism, ensures semantic coherence by combining a dynamic inter-block overlapping connection mechanism, and can flexibly adapt to the needs of specific text structures in the domain based on configurable splitting parameters.
[0042] Then continue back Figure 1 Based on step S120, step S130 is performed, in which the quantitative assessment information obtained in step S120 is spatiotemporally aligned and matched with the multi-source monitoring data of coal mine methane to generate a policy-monitoring association dataset. This policy-monitoring association dataset is used to drive or verify the quantitative assessment technology model of coal mine methane emissions.
[0043] It should be noted that the multi-source monitoring data here includes satellite remote sensing or airborne remote sensing data, which can provide methane column concentration or flux information at the regional or even national scale. The spatial coverage of these data matches the macro scale of policy assessment, which can better achieve spatial consistency mapping between the scope of policy effectiveness and emission monitoring data. The multi-source monitoring data here also includes emission flux datasets at the regional or national scale formed by summarizing, standardizing and spatially interpolating the dispersed sensing data based on local mine-level ground monitoring systems. This realizes the effective aggregation of high-precision micro-viewpoint data into macro-statistical representation, making up for the deficiencies of remote sensing data in vertical distribution, process mechanism and instantaneous anomaly monitoring, and providing a multi-scale, high-confidence data support system for spatiotemporal alignment and matching.
[0044] Specifically, in step S130, the spatiotemporal alignment matching process includes: Based on the effective date, termination date, and revision chain of the coal mine methane emission reduction and control policy documents, the quantitative assessment information of the policies is allocated to the corresponding effective years according to the preset time validity rules; For example, the quantitative assessment information of the "Emission Standard for Coalbed Methane (Coal Mine Gas)" (GB21522—2024) is processed, extracting the standard's publication date as 2024-11-28 and implementation date as 2025-04-01, and identifying its revision relationship chain as "GB21522—2024 replaces GB21522—2008 (complete replacement)". Based on this, according to the preset time validity rules—"after the new document is implemented, the replaced document automatically becomes invalid; documents without a specified termination date remain valid by default until they are replaced or repealed; annual allocation uses the year of implementation as the effective year"—the quantitative score of GB21522—2024 is allocated to 2025 and subsequent effective years; simultaneously, the quantitative score of the replaced GB21522—2008 is allocated to each effective year from 2008 to 2024 (on an annual scale, the new standard takes over from 2025 onwards).
[0045] The quantitative assessment information of the same region for each year is accumulated according to the administrative division level to generate a two-dimensional policy intensity matrix of region-year. The row vectors of the matrix represent the annual policy intensity index sequence of a specific region, and the column vectors represent the regional policy intensity distribution sequence of a specific year.
[0046] Specifically, the method for constructing a two-dimensional policy intensity matrix of "region-year" by accumulating the quantitative assessment information of the same region in each year according to the administrative division level is as follows: Based on the defined analytical dimensions—namely, the set of administrative divisions (e.g., provincial-level administrative regions) and the time range (e.g., 2005-2024)—a two-dimensional matrix is initialized. Each cell in the matrix corresponds to a "region-year" pair, with its cell value being the sum of quantitative scores for all effective policies in that region for that year. The specific calculation rule is as follows: for a cell (region i, year j), the scores of all policies that meet the following two conditions are accumulated: the policy is effective in year j; and the policy's text specifies its applicable scope to cover region i. Following these rules, all cells in the matrix are calculated and filled, completing the construction of the "region-year" two-dimensional policy intensity matrix. The rows and columns of the matrix provide two different organizational views of the data.
[0047] For example, taking Shanxi Province as an example, during the study period (2005-2024), a total of 349 coal mine methane emission reduction / gas utilization related policies were included in the quantitative analysis. After identifying whether each policy was valid in a given year and completing the quantitative scoring, the policy scores for Shanxi Province in the corresponding years were accumulated according to the above rules to obtain the annual intensity sequence of Shanxi Province. This sequence is represented in the matrix as the row vector corresponding to Shanxi Province. For example, the row vector of Shanxi Province in some years can be represented as "Shanxi Province = [9,2,9,6,6,2,9...]", that is, the policy intensity of Shanxi Province in 2005 is 9, the intensity in 2006 is 2, and so on.
[0048] Similarly, in the column vector, taking 2015 as an example, the column vector of 2015 can be represented as "2015=[12,9,9,8,12,18, …]", that is, in 2015, the policy intensity of a certain province A was 12, the policy intensity of a certain province B was 9, and so on, which is used to reflect the relative level and distribution pattern of policy intensity in different regions in that year.
[0049] Row and column aggregation operations are performed on the region-year two-dimensional policy intensity matrix to generate a multi-level time series dataset of policy intensity at the national and provincial levels.
[0050] Furthermore, it should be noted that the construction and application of the quantitative evaluation technology model here pertains to the technical content of other patent applications filed by the applicant, and will only be briefly described here: This quantitative evaluation technology model involves an annual prediction method for coal mine methane emissions. The core is to sequentially complete multi-source annual feature construction, main factor identification, sequence prediction under small sample conditions, and uncertainty quantification matching the model structure in the same set of processes. The entire method corresponds to the aforementioned three technical problems: First, process multi-source information such as policies (the policy intensity index quantified in this application), technologies, and markets into annual feature sequences, corresponding to constructing indices such as P (policy), T (technology), and M (market), and aligning them with the methane emission historical data to form a "small sample, high dimension" annual dataset; Second, construct a relatively stable and interpretable prediction model under small sample and high-dimensional feature conditions; Third, give an uncertainty interval while obtaining the prediction result, facilitating scenario analysis and decision-making applications.
[0051] In the technical solution provided in this application, by invoking a dedicated large language model fine-tuned with parameters in the coal mine methane field to deeply analyze and quantitatively process policy texts, it is possible to automatically and accurately extract structured policy elements from policy clauses for quantitative evaluation, overcoming the deficiencies of insufficient analysis depth in traditional macro-statistical methods and the strong subjectivity and low efficiency of expert manual interpretation, and achieving a standardized and repeatable quantitative representation of policy connotations; furthermore, by performing spatio-temporal alignment and matching of the generated quantitative evaluation information with multi-dimensional monitoring data, a "policy-monitoring" association dataset that can directly drive or verify downstream quantitative evaluation technology models is constructed, effectively bridging the gap between policy text semantics and physical world monitoring data, providing a reliable data basis and technical support for objectively and dynamically evaluating the implementation effect of policies and optimizing emission reduction governance strategies, and thus significantly improving the automation level and decision-making support ability of coal mine methane emission reduction governance policy analysis and evaluation.
[0052] In the technical solution of this application, the quantitative results of coal mine methane emission reduction policies, as described above, in addition to being used as input variables for the PTM framework and emission prediction models, can also support multiple more direct application scenarios. Its core value lies in transforming scattered policy texts into structured information with evidence, reproducibility, and comparability, providing an evidence-based decision-making foundation for coal mine methane governance.
[0053] Specifically, firstly, in terms of policy evaluation and policy gap identification, quantitative results can present the trajectory of policy intensity changes on an annual scale, and further decompose the contribution sources into dimensions such as level, objectives, tools, and MRV (monitoring, reporting, and verification). This allows policy reviews to begin by identifying "when the intensity increased and where the increase came from," further pinpointing the types of shortcomings at a particular stage. Examples include insufficient specific provisions for methane, a lack of quantifiable and time-bound objectives, tools that are more advocacy-oriented and lack assessment and penalties, and requirements that exist but lack MRV binding. This transforms the generalized judgment of "how many policies are there, and how strong are they?" into verifiable and targeted improvement criteria, directly serving the next round of policy optimization and text design.
[0054] Secondly, regarding regional benchmarking and regulatory performance analysis, the unified standards provided in this application make the policy intensity structures of different provinces comparable. The comparison shifts from the number of documents to the strength and weakness of "verifiable objectives" and "executable tools." When conducting special summaries, supervisory evaluations, or differentiated policy recommendations, regulatory departments can use policy intensity and its structure as explanatory variables to analyze the institutional factors behind the differences in the promotion of coal mine gas extraction and utilization in different regions. This shifts the discussion from the level of slogans to the level of institutional arrangements, enhancing the interpretability and operability of the analysis.
[0055] Furthermore, regarding the design of policy mix and optimization of tool structure, this application structurally deconstructs policy tools and identifies MRV (Mean Value Recognition) as a key element closely related to the implementation chain, enabling it to address practical issues such as "insufficient regulation, insufficient incentives, or a missing implementation chain." By comparing the tool structures at different stages or in different regions, more targeted adjustments to the tool mix can be proposed, such as strengthening licensing and assessment constraints, supplementing incentive mechanisms, or more closely linking monitoring and reporting with penalty disbursement. This provides quantitative evidence for the demonstration of policy tool mix and enhances the feasibility of the plan.
[0056] Finally, this application also offers transferable and customizable applications for enterprise-side compliance management and project development. Coal mining enterprises and gas utilization project owners are more concerned with how constraints are implemented, what the acceptance criteria are, how subsidies and electricity prices are aligned, and how MRV data is prepared. The structured extraction and evidence retention capabilities of this method can, without altering the quantitative framework, customize tool-based outputs such as "compliance checklist generation, application path streamlining, and MRV data preparation requirements collection" for specific regions, enterprises, or project types. This provides a methodological foundation for enterprises to reduce policy comprehension costs and implementation deviations, and also provides a clear direction for subsequent customized application development.
[0057] To facilitate understanding of the technical solution of this application, the policy quantification framework involved in this application will be explained below: This application's policy quantification framework first proposes a unified PI index to transform my country's coal mine methane emission reduction and control policies from their original text into policy intensity values with annual and regional dimensions. The PI index is determined based on three dimensions: L (Level), O (Objective), and I (Instrument). Specifically: Hierarchical scoring (L) The scoring criteria for evaluating the level of the issuing authority of policy documents are as follows: 3 points correspond to the central level, 2 points correspond to the ministerial level, 2 points for national standards (GB / GB / T) are awarded based on the issuing authority, 1 point corresponds to the provincial level, and 1 point for local standards (DB / DB / T) are awarded based on the issuing authority. For documents jointly issued by multiple agencies, the highest level will be scored. For forwarded documents or implementation rules, the scoring will be based on the level of the issuing agency. After scoring, the level score L and the scoring reasons including the original text citation will be output.
[0058] Target dimension scoring (O) This assessment is used to evaluate the relevance and objective clarity of policy documents to a specific topic (such as methane control). The scoring criteria are as follows: 3 points corresponds to the content clearly mentioning core keywords such as methane, gas, and extraction, and including rigid elements such as quantifiable objectives, clear timelines, and assessment or licensing constraints; 2 points corresponds to the content clearly mentioning core keywords but lacking rigid elements, or not mentioning core keywords but focusing on safety governance with a clear target / scope; 1 point corresponds to the content being general, indirectly related to the topic, or lacking specific details. After scoring, the target score O and the scoring reasons including original text citations are output.
[0059] Tool Dimension Scoring (I) The tool dimension score is used to assess the type and strength of policy document implementation measures. This dimension score is composed of the command-control tool (CC) score and the market tool (MB) score, and the highest score of CC and MB is taken.
[0060] Command-control tools (CC) are scored based on keywords such as "must," "strictly prohibited," "punishment," "assessment," and "permission" and their context: 3 points are for the presence of words such as "must / strictly prohibited / limited time" and clear accompanying specific consequences such as penalties or production stoppages; 2 points are for procedural requirements but unclear consequences; 1 point is for only advocacy or principle statements; 0 points are for no relevant content.
[0061] Market-based instruments (MB) are scored based on keywords such as “subsidy”, “tax incentive”, “fund”, and “carbon quota” and their context: 3 points indicates that the amount, threshold, and calculation formula of the market-based instrument are clearly stated or that there are pre-operational conditions for its issuance (such as verification); 2 points indicates that subsidies / support are mentioned but the standards are not specific; 1 point indicates that applications are encouraged but there is a lack of specific amounts and thresholds; 0 points indicates that there is no relevant content.
[0062] In addition, an MRV binding adjustment mechanism is introduced: MRV_CC (0 or 1 point) corresponds to situations where MRV requirements are directly associated with command-control tools such as penalties and assessments; MRV_MB (0 or 1 point) corresponds to situations where MRV requirements are directly associated with market-based tools such as subsidies and fund disbursements.
[0063] The tool dimension score calculation formula is as follows: Final CC score CC_final = min(Base CC score + MRV_CC adjustment score, 3); Final MB score MB_final = min(Base MB score + MRV_MB adjustment score, 3); Total tool dimension score I = max(CC_final, MB_final). After scoring, the output includes the score results of CC_final, MB_final, and I, along with detailed scoring reasons.
[0064] Under this policy quantitative framework, the final PI index is determined using the formula PI=L×O×I.
[0065] In other words, if the quantified length of a policy text does not exceed the first length threshold, the PI index can be directly determined using the above calculation formula. However, if the quantified length of a policy text exceeds the first length threshold, the scores of all its text blocks must first be integrated and processed, and the highest score of each text block should be taken as the final score for the corresponding dimension. L_total=max{L1,L2,...,L n} (where L1~L n (These are the hierarchical dimension scores for the 1st to nth text blocks, respectively). O_total=max{O1,O2,...,O n} (where O1~O n (These are the target dimension scores for the 1st to nth text blocks, respectively). I_total=max{I1,I2,...,I n} (where I1~I n (These are the tool dimension scores for the 1st to nth text blocks, respectively). Finally, the corresponding overall policy text's PI index is calculated as PI = L_total × O_total × I_total.
[0066] Based on the embodiments described above, in some embodiments, such as Figure 2 As shown, in the process of acquiring policy text data on coal mine methane emission reduction and control, multiple policy documents to be evaluated can be manually sorted and screened in advance. Then, the obtained location information is stored in a structured table format. The location information can be a Uniform Resource Locator (URL) link to the file or a PDF file path pointing to local and network drives. The table format also facilitates the automatic batch calling and processing of files through technical means in the future. Then, based on the collected location information of coal mine methane emission reduction and control policies, the validity is verified, and the verified documents (document location information) are identified as valid input sources. By extracting and parsing the document content corresponding to the valid input sources, the text data of coal mine methane emission reduction and control policies is generated. The validity verification here may include at least one of format verification, integrity verification and source credibility verification.
[0067] However, if the pre-collected location information has already undergone manual screening in practice, and considering the convenience of actual implementation, for file location information that is a Uniform Resource Locator URL link or a PDF file path pointing to local storage devices and network storage devices, the validity verification here can also include only network reachability verification and file existence verification.
[0068] like Figure 2 As shown, in this embodiment, a general text reading tool, such as LinkReader, is also used to extract text data, converting web page text or text within a PDF into clean text strings to provide basic data for subsequent evaluation.
[0069] Next, the text length of the policy text data is quantified to determine whether the text is long or short. For example, here the character count is used: more than 5,000 characters is considered long text, and less than or equal to 5,000 characters is considered short text.
[0070] Similarly to the above embodiments, the parsing processing of different branches is performed based on the call of the large language model, and the comprehensive influence index PI of the corresponding policy text is obtained under the policy quantification framework.
[0071] The difference is that, in this embodiment, the process of generating quantitative evaluation information for policy text data based on the specific configuration of prompt word constraints during the model invocation process also includes: A source identification is established for the quantitative score of each policy evaluation dimension. The source identification includes the chapter number, paragraph position index and key sentence location information of the original policy text, forming a score-original text mapping relationship table. Based on the score-text mapping table and the quantitative scores of each policy evaluation dimension, a structured data file is generated for export (corresponding to...). Figure 2 In the results output stage, the structured data file contains multidimensional data tables, including a main table for policy quantitative scoring, a table of original text references, and a table of dimensional analysis explanations.
[0072] It should be noted that the main table for policy quantitative scoring records the final scores for each evaluation dimension, the original text reference table stores the original text paragraph identifiers and precise location information corresponding to each score, and the dimension analysis description table records the analysis logic and calculation basis for each evaluation dimension.
[0073] Furthermore, based on specific software function configurations and these data tables, all scoring history can be automatically saved so that past analysis results can be viewed at any time, making it easier to track policy evolution trends.
[0074] Figure 3 This is a schematic diagram of the structure of a system for quantitatively analyzing coal mine methane emission reduction and control policies, provided in one embodiment of this application. Figure 3 As shown, the system 300 for quantitative analysis of coal mine methane emission reduction and control policies includes: The processing module 301 is used to acquire policy text data on methane emission reduction and control in coal mines; The analysis and processing module 302 is used to call the large language model to parse and process the policy text data, extract the structured policy elements, and generate corresponding quantitative evaluation information. The matching processing module 303 is used to perform spatiotemporal alignment matching between the quantitative assessment information and the multi-source monitoring data of coal mine methane to generate a policy-monitoring association dataset. The policy-monitoring association dataset is used to drive or verify the quantitative assessment technology model of coal mine methane emissions.
[0075] Regarding the system 300 for quantitative analysis of coal mine methane emission reduction and control policies in the above-mentioned embodiments, the specific methods by which each module performs its operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0076] Figure 4 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application, as shown below. Figure 4 As shown, the electronic device 400 includes: Memory 401, on which an executable program is stored; Processor 402 is configured to execute an executable program in memory 401 to implement the steps of the method in the above method embodiments.
[0077] Regarding the electronic device 400 in the above embodiments, the specific manner in which its processor 402 executes the program in the memory 401 has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0078] The above description is merely a preferred embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0079] It is understood that the same or similar parts in the above embodiments can be referred to each other, and the contents not described in detail in some embodiments can be referred to the same or similar contents in other embodiments.
[0080] It should be noted that in the description of this application, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means at least two.
[0081] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this application pertain.
[0082] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0083] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A method for quantitatively analyzing coal mine methane emission reduction and control policies, characterized in that, include: Obtain policy text data on methane emission reduction and control in coal mines; The policy text data is parsed and processed by calling a large language model, and the structured policy elements are extracted to generate corresponding quantitative evaluation information. The quantitative assessment information is spatiotemporally aligned and matched with multi-source monitoring data of methane in coal mines to generate a policy-monitoring correlation dataset. This policy-monitoring correlation dataset is used to drive or validate the quantitative assessment technology model for methane emissions from coal mines. The step of calling a large language model to parse and process the policy text data includes: Determine the quantized value of the text length of the policy text data; When the text length quantization value exceeds the first length threshold, the policy text data is divided into multiple continuous text blocks based on the principle of text semantic coherence. For each of the continuous text blocks, the large language model is invoked for parsing to obtain the corresponding local quantitative information. The local quantitative information is aligned according to the preset policy evaluation dimension. The maximum value of the quantitative score of each text block under the same policy evaluation dimension is selected as the final score of the policy evaluation dimension, and the quantitative evaluation information of the policy text data is generated. The process of segmenting the policy text data into multiple consecutive text blocks based on the principle of textual semantic coherence specifically includes: The policy text data is broken down into multiple basic semantic units that maintain semantic integrity; For each of the basic semantic units, a greedy algorithm based on a length threshold is used to sequentially package and combine the basic semantic units. The specific process is as follows: Starting from the first basic semantic unit, subsequent units are added to the current text block in sequence. At the same time, the text length quantization value of the current text block is determined. When adding the next unit will cause the text length quantization value to exceed the second length threshold, the current block is stopped from being packaged and the text block is output. Then, starting with the last N basic semantic units of the previous text block, the aforementioned packing process is repeated to generate the next text block, where N is the preset number of overlapping units, thereby ensuring semantic coherence between adjacent text blocks.
2. The method for quantitatively analyzing coal mine methane emission reduction and control policies according to claim 1, wherein, Based on pre-constructed semantic boundary recognition rules, the basic semantic units are decomposed using regular expression matching. These semantic boundary recognition rules include: The paragraph boundary recognition rule uses a regular expression with multiple line breaks to match the paragraph separation positions in the text in order to segment independent paragraph units; The sentence boundary recognition rule uses regular expressions combining sentence punctuation marks to match the end positions of Chinese and English sentences, and performs secondary segmentation on independent paragraph units whose length exceeds a preset threshold; The paragraph boundary recognition rule takes precedence over the sentence boundary recognition rule.
3. The method for quantitatively analyzing coal mine methane emission reduction and control policies according to claim 1, wherein, The text length quantification value of policy text data or text blocks is determined by calculating a custom token counting function; The custom token counting function is calculated by configuring differentiated weight coefficients for different character categories, and the weight coefficients are determined based on the lexicalization characteristics of the large language model.
4. The method for quantitatively analyzing coal mine methane emission reduction and control policies according to claim 1, wherein, The step of calling a large language model to parse and process the policy text data also includes: When the text length quantization value does not exceed the first length threshold, the large language model is directly invoked to parse and process the policy text data.
5. The method for quantitatively analyzing coal mine methane emission reduction and control policies according to claim 1, wherein, The process of generating quantitative evaluation information for the policy text data also includes: A source identification is established for the quantitative score of each policy evaluation dimension. The source identification includes the chapter number, paragraph position index and key sentence location information of the original policy text, forming a score-original text mapping relationship table. Based on the score-original text mapping table and the quantitative scores of each policy evaluation dimension, a structured data file is generated for export. The structured data file contains a multidimensional data table, which includes a main table of policy quantitative scoring, a table of original text references, and a table of dimension analysis descriptions.
6. The method for quantitatively analyzing coal mine methane emission reduction and control policies according to claim 1, wherein, The process of spatiotemporally aligning and matching the quantitative assessment information with multi-source remote sensing monitoring data of methane in coal mines includes: Based on the effective date, termination date, and revision chain of the coal mine methane emission reduction and control policy documents, the quantitative assessment information of the policies is allocated to the corresponding effective years according to the preset time validity rules; The quantitative assessment information of the same region for each year is accumulated according to the administrative division level to generate a two-dimensional policy intensity matrix of region-year. The row vector of the matrix represents the annual policy intensity index sequence of a specific region, and the column vector represents the regional policy intensity distribution sequence of a specific year. Row and column aggregation operations are performed on the region-year two-dimensional policy intensity matrix to generate a multi-level time series dataset of policy intensity at the national and provincial levels.
7. The method for quantitatively analyzing coal mine methane emission reduction and control policies according to any one of claims 1 to 6, wherein, The acquisition of coal mine methane emission reduction and control policy text data includes: The validity of the policy documents on methane emission reduction and control in coal mines is verified based on the location information collected, and the location information of the verified documents is marked as a valid input source. Extract and parse the file content corresponding to the valid input source to generate the coal mine methane emission reduction and control policy text data; The validity verification includes at least one of format verification, integrity verification, and source credibility verification.
8. A system for quantitatively analyzing coal mine methane emission reduction and control policies, characterized in that, include: The acquisition and processing module is used to acquire policy text data on methane emission reduction and control in coal mines; The analysis and processing module is used to call the large language model to parse and process the policy text data, extract the structured policy elements, and generate corresponding quantitative evaluation information. The matching processing module is used to perform spatiotemporal alignment matching between the quantitative assessment information and the multi-source monitoring data of coal mine methane to generate a policy-monitoring association dataset. The policy-monitoring association dataset is used to drive or verify the quantitative assessment technology model of coal mine methane emissions. The step of calling a large language model to parse and process the policy text data includes: Determine the quantized value of the text length of the policy text data; When the text length quantization value exceeds the first length threshold, the policy text data is divided into multiple continuous text blocks based on the principle of text semantic coherence. For each of the continuous text blocks, the large language model is invoked for parsing to obtain the corresponding local quantitative information. The local quantitative information is aligned according to the preset policy evaluation dimension. The maximum value of the quantitative score of each text block under the same policy evaluation dimension is selected as the final score of the policy evaluation dimension, and the quantitative evaluation information of the policy text data is generated. The process of segmenting the policy text data into multiple consecutive text blocks based on the principle of textual semantic coherence specifically includes: The policy text data is broken down into multiple basic semantic units that maintain semantic integrity; For each of the basic semantic units, a greedy algorithm based on a length threshold is used to sequentially package and combine the basic semantic units. The specific process is as follows: Starting from the first basic semantic unit, subsequent units are added to the current text block in sequence. At the same time, the text length quantization value of the current text block is determined. When adding the next unit will cause the text length quantization value to exceed the second length threshold, the current block is stopped from being packaged and the text block is output. Then, starting with the last N basic semantic units of the previous text block, the aforementioned packing process is repeated to generate the next text block, where N is the preset number of overlapping units, thereby ensuring semantic coherence between adjacent text blocks.