A data missing pollution source identification and governance strategy optimization method and system
Patent Information
- Application Number
- CN202610750385.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-28
- Publication Date
- 2026-08-21
AI Technical Summary
[0003]然而在实际应用中,企业排放数据普遍存在月度、年度缺失问题,传统方法未考虑排放的季节性波动,导致历史排放核算结果失真;重点企业识别仅依赖排放总量排序,忽略企业产值差异,难以精准定位“低效高排”企业;静态物质流分析无法刻画磷在“开采—生产—堆存—治理”全链条的动态反馈与长期积累效应;同时,治理策略评估缺乏对政策、经济、技术等深度不确定性的定量处理,难以输出稳健、可落地的适应性治理方案,无法满足区域精准控磷与长效治理的实际需求
[0047]本发明的数据缺失污染源识别与治理策略优化方法、系统、电子设备及存储介质,通过获取目标区域“三磷”企业多源数据,采用基于时序特征的分级差异化方法处理数据缺失,解决了传统方法未考虑排放季节性规律导致的历史排放核算失真问题,得到完整的多源重构数据,为后续分析提供基础;对多源重构数据进行排放负荷与单位产值排放强度双维度分析,解决了仅依赖排放总量排序无法识别“低效高排”企业的问题,所得重点污染源名单作为后续治理情景设计的靶向对象,使治理资源能够精准投向污染贡献大、治理性价比高的企业;搭建覆盖磷全产业链及水体环境的物质流-系统动力学耦合模型,利用历史监测数据进行参数校准,解决了静态物质流分析无法刻画磷迁移转化动态反馈与长期积累效应的问题,为治理效果推演提供量化工具;识别影响系统演化的关键不确定性参数并确定其概率分布,采用拉丁超立方抽样结合蒙特卡洛模拟生成多组未来世界状态,结合重点污染源名单设计源头减排、末端治理、资源化利用及多措施组合的治理情景集,全面覆盖未来不确定性,保证治理方案的针对性;利用校准后的耦合模型对各治理情景进行推演,通过达标概率与极端风险双重指标筛选出在多数未来情景下能够实现水质目标且极端超标风险较低的稳健候选策略,解决了传统单一确定性情景评估的片面性问题;最后引入水质连续超标与治理成本阈值双重触发机制,对候选策略进行动态调控,当系统状态偏离预期目标时自动切换至强化治理模式,输出重点污染源名单、治理策略稳健性排序、适应性路径图及水质预测曲线等优化结果,形成从数据修复、精准识别、动态模拟到稳健决策、自适应调控的完整技术流程,为区域“三磷”污染精准管控与长效治理提供技术支持。
Smart Images

Figure CN122615221A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial pollution source management and data analysis technology, and more specifically, to a method and system for identifying pollution sources with missing data and optimizing treatment strategies. Background Technology
[0002] The Yichang section of the middle reaches of the Yangtze River is home to a large number of phosphate mining, phosphate chemical, and phosphogypsum storage enterprises (referred to as "three-phosphate" enterprises). Total phosphorus emissions from these industries are the core source of pollution in the region's water environment. Existing pollution source management technologies have been widely applied in the region's water environment governance, mainly relying on enterprise emission ledgers, monitoring data, and static material flow analysis methods to conduct pollution source statistics, screen key enterprises, and formulate governance strategies, providing basic technical support for the control of total phosphorus pollution in the region.
[0003] However, in practical applications, enterprise emission data generally suffers from missing monthly and annual data. Traditional methods do not consider seasonal fluctuations in emissions, leading to distorted historical emission accounting results. The identification of key enterprises relies solely on ranking by total emissions, ignoring differences in enterprise output value, making it difficult to accurately locate "inefficient and high-emission" enterprises. Static material flow analysis cannot depict the dynamic feedback and long-term accumulation effect of phosphorus in the entire chain of "mining-production-storage-treatment". At the same time, the evaluation of treatment strategies lacks quantitative processing of deep uncertainties in policies, economy, and technology, making it difficult to output robust and implementable adaptive treatment solutions, and failing to meet the actual needs of precise phosphorus control and long-term treatment in the region. Summary of the Invention
[0004] The problem addressed by this invention is how to automatically fill in missing data and output robust governance strategies.
[0005] To address the aforementioned issues, this invention provides a method, system, electronic device, and storage medium for identifying and optimizing pollution sources with missing data and for implementing control strategies.
[0006] In a first aspect, the present invention provides a method for identifying and optimizing pollution sources with missing data and for managing pollution strategies, comprising:
[0007] Acquire multi-source data of "three phosphorus" enterprises within the target area, and perform missing data repair processing based on time-series features to obtain multi-source reconstructed data;
[0008] A two-dimensional fusion analysis was performed on the multi-source reconstructed data to obtain a list of key pollution sources;
[0009] A material flow-system dynamics coupled model is constructed, and historical monitoring data is used to calibrate the parameters of the material flow-system dynamics coupled model to obtain the calibrated material flow-system dynamics coupled model.
[0010] Key variable parameters are obtained, and Latin hypercube sampling combined with Monte Carlo simulation is used. Combined with the list of key pollution sources, the key variable parameters are randomly simulated to obtain a set of treatment scenarios.
[0011] Based on the calibrated material flow-system dynamics coupling model, the set of governance scenarios is deduced to obtain candidate governance strategies;
[0012] The candidate governance strategies are dynamically adjusted based on a preset triggering mechanism to obtain identification results and optimization strategies.
[0013] Optionally, the step of acquiring multi-source data of "three phosphorus" enterprises within the target area, and performing time-series feature-based missing data repair processing on the multi-source data to obtain multi-source reconstructed data includes:
[0014] Acquire multi-source data of "three phosphorus" enterprises within the target area, including monthly emission records, annual output value reports, and solid waste storage site composition test reports;
[0015] Based on the blank data points in the monthly emission records, the actual measured historical data of the corresponding enterprise in the same month of previous years are extracted. The arithmetic mean of the actual measured historical data is calculated and then missing data repair is performed to obtain the multi-source reconstructed data.
[0016] Optionally, the step of extracting the corresponding measured historical data for the same month of previous years from the blank data points in the monthly emission records, calculating the arithmetic mean of the measured historical data, and then performing missing data repair to obtain the multi-source reconstructed data further includes:
[0017] If the corresponding enterprise has data missing for more than three consecutive months, the data for the interrupted period will be removed, and the total annual emissions of the corresponding enterprise will be calculated based on the total amount of the remaining valid months.
[0018] If the corresponding company's full-year data is missing, then check the corresponding company's production capacity ranking in the target region.
[0019] When the corresponding enterprise belongs to the key production capacity enterprises within the preset ranking range, the measured historical data of adjacent years or the unit production capacity emission level of enterprises of the same scale is obtained and estimated.
[0020] If the corresponding enterprise is not among the key production capacity enterprises in the preset ranking range, then the data of the corresponding enterprise will be removed.
[0021] Optionally, the step of performing a two-dimensional fusion analysis on the multi-source reconstructed data to obtain a list of key pollution sources includes:
[0022] Based on the multi-source reconstructed data, the annual average pollutant emissions of each of the three phosphorus enterprises are obtained. All the three phosphorus enterprises are sorted from high to low according to the annual average pollutant emissions, and the total emissions of the three phosphorus enterprises are accumulated according to the sorting until the predetermined emission threshold of the total emissions in the target area is reached. Based on all the accumulated three phosphorus enterprises, the set of key load enterprises is obtained.
[0023] Based on the multi-source reconstructed data, the annual output value data of all the "three phosphorus" enterprises are obtained. The corresponding pollutant emission amount is generated based on the annual output value data. The emission intensity per unit output value of the corresponding "three phosphorus" enterprises is obtained based on the pollutant emission amount. The corresponding "three phosphorus" enterprises with emission intensity higher than the preset average intensity threshold are screened out to obtain a set of inefficient key enterprises.
[0024] The list of key pollution sources is obtained by taking the union of the set of key enterprises with high load and the set of key inefficient enterprises.
[0025] Optionally, the step of constructing a material flow-system dynamic coupling model and obtaining historical monitoring data to calibrate the parameters of the material flow-system dynamic coupling model to obtain a calibrated material flow-system dynamic coupling model includes:
[0026] Construct a phosphate mining subsystem, a phosphate chemical production subsystem, a phosphogypsum storage subsystem, a wastewater treatment subsystem, a product consumption subsystem, and a water environment subsystem, and establish the material transfer topology between each subsystem;
[0027] The mass transfer topology is mapped to the mass flow-system dynamics coupled model, and key rate equations are set.
[0028] Based on the historical monitoring data, the material flow-system dynamic coupling model is reverse-tuned to adjust the conversion coefficient and loss parameters within the model. When the output data of the material flow-system dynamic coupling model matches the historical data of the actual monitoring data to a preset standard, the calibrated material flow-system dynamic coupling model is obtained.
[0029] Optionally, the acquisition of key variable parameters involves using Latin hypercube sampling combined with Monte Carlo simulation, and combining this with the list of key pollution sources to perform random simulations on the key variable parameters, resulting in a set of governance scenarios, including:
[0030] Based on the key variable parameters, the corresponding probability distribution characteristics are obtained;
[0031] The probability distribution features are sampled using the Latin hypercube sampling method in a stratified combination to generate several sets of parameter combinations.
[0032] According to the Monte Carlo simulation method, the set of future world states is generated by performing a preset number of random iterations on the several sets of parameter combinations.
[0033] Based on the list of key pollution sources and the set of future world states, several different governance measures are generated. Two or more of these governance measures are selected and combined to form the set of governance scenarios.
[0034] Optionally, the step of deriving candidate governance strategies by extrapolating the set of governance scenarios based on the calibrated material flow-system dynamics coupling model includes:
[0035] The set of treatment scenarios is input into the calibrated material flow-system dynamics coupled model for simulation, and the probability of water quality compliance for each scheme in the set of treatment scenarios is statistically analyzed.
[0036] Obtain the tail risk value of each group of schemes, and obtain candidate treatment strategies based on the water quality compliance probability and the tail risk value.
[0037] Secondly, the data-missing pollution source identification and control strategy optimization system of the present invention includes:
[0038] The data repair unit is used to acquire multi-source data of "three phosphorus" enterprises in the target area, perform missing data repair processing based on time-series features on the multi-source data, and obtain multi-source reconstructed data.
[0039] The data analysis unit is used to perform two-dimensional fusion analysis on the multi-source reconstructed data to obtain a list of key pollution sources.
[0040] The model building unit is used to build a material flow-system dynamics coupled model, acquire historical monitoring data to calibrate the parameters of the material flow-system dynamics coupled model, and obtain the calibrated material flow-system dynamics coupled model.
[0041] The data processing unit is used to obtain key variable parameters, and uses Latin hypercube sampling combined with Monte Carlo simulation, combined with the list of key pollution sources, to perform random simulation on the key variable parameters to obtain a set of treatment scenarios.
[0042] The data extrapolation unit is used to extrapolate the set of governance scenarios based on the calibrated material flow-system dynamics coupling model to obtain candidate governance strategies;
[0043] The decision control unit is used to dynamically control the candidate governance strategies based on a preset triggering mechanism to obtain the identification results and optimization strategies.
[0044] Thirdly, the electronic device of the present invention includes: a processor and a memory, the memory being used to store a computer program;
[0045] When the computer program is loaded by the processor, it causes the processor to execute the aforementioned method for identifying and optimizing the management strategy of data-missing pollution sources.
[0046] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the above-described method for identifying and optimizing the management strategy of data missing pollution sources is implemented.
[0047] This invention relates to a method, system, electronic device, and storage medium for identifying pollution sources with missing data and optimizing treatment strategies. By acquiring multi-source data from phosphorus-related enterprises in a target area, it employs a time-series-based hierarchical differentiation method to handle missing data, solving the problem of historical emission accounting distortion caused by traditional methods failing to consider seasonal emission patterns. This yields complete multi-source reconstructed data, providing a foundation for subsequent analysis. The multi-source reconstructed data undergoes dual-dimensional analysis of emission load and emission intensity per unit output, addressing the issue that relying solely on total emission ranking cannot identify "inefficient, high-emission" enterprises. The resulting list of key pollution sources serves as the target for subsequent treatment scenario design, enabling precise allocation of treatment resources to enterprises with significant pollution contributions and high treatment cost-effectiveness. A material flow-system dynamics coupling model covering the entire phosphorus industry chain and aquatic environment is constructed, using historical monitoring data for parameter calibration. This solves the problem that static material flow analysis cannot characterize the dynamic feedback and long-term accumulation effects of phosphorus migration and transformation, providing a quantitative tool for predicting treatment effects. The invention also identifies key uncertainty parameters affecting system evolution and... The probability distribution of the pollution sources was determined, and multiple sets of future world states were generated using Latin hypercube sampling combined with Monte Carlo simulation. A set of governance scenarios, including source reduction, end-of-pipe treatment, resource utilization, and combinations of multiple measures, was designed based on a list of key pollution sources to comprehensively cover future uncertainties and ensure the targeted nature of the governance solutions. A calibrated coupled model was used to extrapolate each governance scenario. Robust candidate strategies that could achieve water quality targets in most future scenarios and had low extreme exceedance risks were selected using dual indicators of compliance probability and extreme risk, solving the problem of the one-sidedness of traditional single deterministic scenario assessment. Finally, a dual trigger mechanism of continuous water quality exceedance and governance cost threshold was introduced to dynamically adjust the candidate strategies. When the system state deviates from the expected target, it automatically switches to an enhanced governance mode, outputting optimized results such as a list of key pollution sources, a robustness ranking of governance strategies, an adaptive path diagram, and water quality prediction curves. This forms a complete technical process from data repair, accurate identification, dynamic simulation to robust decision-making and adaptive control, providing technical support for the precise control and long-term governance of regional phosphorus pollution. Attached Figure Description
[0048] Figure 1This is a flowchart illustrating the method for identifying and optimizing pollution sources with missing data in an embodiment of the present invention.
[0049] Figure 2 This is a schematic diagram of the dual-dimensional key pollution source identification in an embodiment of the present invention;
[0050] Figure 3 This is a schematic diagram of the dynamic stock-flow rate of the material flow-system dynamic coupling model in an embodiment of the present invention. Detailed Implementation
[0051] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0052] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0053] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; and the term "optionally" means "optional embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first," "second," etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0054] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0055] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0056] To address the problems existing in the aforementioned related technologies, this embodiment provides a method, system, electronic device, and storage medium for identifying and optimizing pollution sources with missing data and for governance strategies.
[0057] Combination Figure 1 As shown in the figure, the present invention provides a method for identifying and optimizing pollution sources with missing data, and includes:
[0058] Acquire multi-source data of "three phosphorus" enterprises within the target area, and perform missing data repair processing based on time-series features to obtain multi-source reconstructed data;
[0059] Specifically, existing technologies commonly suffer from missing monthly and annual emission data for enterprises. Traditional linear interpolation or mean-based data filling fails to consider seasonal emission patterns, leading to distorted historical emission calculations. This embodiment abandons traditional filling methods that ignore temporal characteristics, preserving the seasonal fluctuations in the production and emissions of phosphorus-related enterprises through the monthly mean method. Furthermore, it employs a tiered and differentiated processing scheme to address varying degrees of data loss, such as continuous or year-round gaps, minimizing estimation errors while ensuring data integrity. This provides accurate and reliable foundational data for all subsequent analyses, serving as a prerequisite for precise pollution source identification and scientific decision-making.
[0060] A two-dimensional fusion analysis was performed on the multi-source reconstructed data to obtain a list of key pollution sources;
[0061] Specifically, the current approach to identifying key enterprises relies solely on ranking by total emissions, ignoring differences in enterprise output value. This fails to identify inefficient, high-emission enterprises with high emission intensity per unit of output value but relatively small total emissions. This embodiment constructs an identification system based on two dimensions: emission load and emission intensity per unit of output value. It covers major emitters contributing over 80% of the pollution load while accurately identifying inefficient, high-emission enterprises with emissions per unit of output value significantly higher than the industry average. The list of key pollution sources is not only one of the final results of pollution source identification but also the sole target for all subsequent governance scenario design. This ensures that governance resources are precisely allocated to the enterprises that contribute the most to pollution and offer the highest cost-effectiveness for governance, avoiding the waste of resources from a one-size-fits-all approach.
[0062] A material flow-system dynamics coupled model is constructed, and historical monitoring data is used to calibrate the parameters of the material flow-system dynamics coupled model to obtain the calibrated material flow-system dynamics coupled model.
[0063] Specifically, existing static material flow analysis cannot simulate the dynamic feedback and long-term accumulation effects of phosphorus in the mining, production, storage, and remediation chain. Therefore, this embodiment combines static material flow balance relationships with system dynamics methods that can characterize the dynamic evolution of the system to establish a dynamic simulation model that covers the entire phosphorus industry chain and the aquatic environment. This model accurately reflects the migration, transformation, accumulation delay, and feedback regulation mechanisms of phosphorus at each stage. Parameter calibration of the model using historical monitoring data ensures the consistency between the simulation results and actual conditions, providing a reliable quantitative tool for the scientific extrapolation of subsequent remediation scenarios.
[0064] Key variable parameters are obtained, and Latin hypercube sampling combined with Monte Carlo simulation is used. Combined with the list of key pollution sources, the key variable parameters are randomly simulated to obtain a set of treatment scenarios.
[0065] Specifically, existing governance strategy assessments lack quantitative handling of deep uncertainties surrounding future policies, the economy, and technology. To address this issue, this embodiment first identifies key variable parameters affecting governance effectiveness (such as emission coefficients, infiltration coefficients, economic growth rates, and policy implementation capacity). Then, it employs Latin hypercube sampling combined with Monte Carlo simulation to generate a large set of world scenarios covering various possible future situations. Most importantly, all governance scenarios are designed around the list of key pollution sources obtained in the second step, proposing targeted governance solutions such as source reduction, end-of-pipe treatment, resource utilization, and combinations of multiple measures. The governance scenario set generated in this embodiment comprehensively covers future uncertainties and possesses strong practical relevance.
[0066] Based on the calibrated material flow-system dynamics coupling model, the set of governance scenarios is deduced to obtain candidate governance strategies;
[0067] Specifically, this embodiment utilizes a calibrated material flow-system dynamics coupled model to extrapolate the effects of each governance scenario under all future world conditions, selecting robust governance strategies that can achieve water quality targets in most future scenarios and have a low risk of extreme exceedances. This embodiment incorporates quantitative uncertainty into the decision-making process, avoiding the drawbacks of the singularity of traditional decision-making methods and ensuring that the output governance strategy has broad applicability.
[0068] The candidate governance strategies are dynamically adjusted based on a preset triggering mechanism to obtain identification results and optimization strategies.
[0069] Specifically, this embodiment abandons the traditional static governance strategy and introduces a dynamic adaptive control mechanism. By setting preset triggering mechanisms, such as continuous water quality exceeding standards and governance cost thresholds, the system can automatically switch to stronger governance measures when the system state deviates from the expected target during actual execution. This embodiment enables the governance strategy to have the ability to self-adjust and optimize, ensuring the achievement of long-term water quality goals while effectively controlling governance costs. The final output is a complete governance plan that can be dynamically adjusted according to actual conditions.
[0070] Optionally, the step of acquiring multi-source data of "three phosphorus" enterprises within the target area, and performing time-series feature-based missing data repair processing on the multi-source data to obtain multi-source reconstructed data includes:
[0071] Acquire multi-source data of "three phosphorus" enterprises within the target area, including monthly emission records, annual output value reports, and solid waste storage site composition test reports;
[0072] Based on the blank data points in the monthly emission records, the actual measured historical data of the corresponding enterprise in the same month of previous years are extracted. The arithmetic mean of the actual measured historical data is calculated and then missing data repair is performed to obtain the multi-source reconstructed data.
[0073] Specifically, addressing the issue that existing technologies using traditional linear interpolation or ordinary mean filling neglect the impact of seasonal cycles (such as peak production seasons, dry seasons, and environmental inspection periods) on enterprise emissions, this embodiment employs the same-month mean method. By extracting historical measured data from the same month of previous years and calculating the arithmetic mean, the common patterns of emissions within the same month are accurately preserved (e.g., a chemical company's emissions surge in March due to spring fertilizer preparation and decrease sharply in winter due to maintenance). This avoids smoothing out or distorting seasonal fluctuations caused by cross-month interpolation, ensuring a high degree of matching between the reconstructed data and the actual emission pattern. In scenarios where data for some months is missing (e.g., due to monitoring equipment malfunctions or reporting delays), the mean of historical data for the same month can effectively fill the gaps and reconstruct a relatively complete annual emission sequence (for example, if an enterprise's emission data for March 2021 is missing, the actual emission data for all March months of 2019, 2020, 2022, 2023, and 2024 are extracted, their arithmetic mean is calculated, and the missing data is filled in). Compared to directly removing missing data or simple estimation, this significantly reduces the bias in calculating historical total emissions and emission trends caused by missing data, providing a more reliable input basis for subsequent two-dimensional pollution source identification (annual average emissions and emission intensity per unit output). The repaired multi-source reconstructed data (including monthly emission records, annual output, solid waste composition, etc.) more closely reflects the actual operating status of enterprises, making the screening of "key enterprises with high load" (e.g., a cumulative emission ratio threshold of 80%) and the identification of "key inefficient enterprises" (e.g., emission intensity per unit output exceeding the standard) based on this data more accurate. It also provides a high-quality data source for the historical calibration of the material flow-system dynamics coupling model (e.g., the error between simulated and measured values is ≤20%), ultimately ensuring the scientific rigor and relevance of the governance strategy optimization. This embodiment, through the repair of missing time-series features, fundamentally solves the core defect of "data distortion leading to inaccurate subsequent analysis" in existing technologies, and is a key foundational link in improving the reliability of the entire "three phosphorus" pollution control system.
[0074] In one embodiment, monthly emission data of "three phosphorus" enterprises (including wastewater flow rate, total phosphorus concentration, and total phosphorus emission), influent and effluent data of sewage treatment plants, phosphogypsum component data (including total phosphorus content and pH value), enterprise output value (which can be obtained from statistical yearbooks or enterprise annual reports), and geographic information data are obtained. It is understood that the specific information obtained depends on the actual data correlation in the actual application process. The data mentioned in this embodiment are only an example for this embodiment.
[0075] For missing values in the monthly emissions data of enterprises, the average value of the same month is used for imputation. The imputation formula is as follows:
[0076] ,
[0077] in:
[0078] : Missing value imputation amount for year y, month m, unit is the same as the original data;
[0079] : The actual emissions of this company in month m of other years;
[0080] The total number of years for which actual data is available.
[0081] The method in this embodiment preserves the seasonal fluctuation characteristics of emissions.
[0082] Optionally, the step of extracting the corresponding measured historical data for the same month of previous years from the blank data points in the monthly emission records, calculating the arithmetic mean of the measured historical data, and then performing missing data repair to obtain the multi-source reconstructed data further includes:
[0083] If the corresponding enterprise has data missing for more than three consecutive months, the data for the interrupted period will be removed, and the total annual emissions of the corresponding enterprise will be calculated based on the total amount of the remaining valid months.
[0084] If the corresponding company's full-year data is missing, then check the corresponding company's production capacity ranking in the target region.
[0085] When the corresponding enterprise belongs to the key production capacity enterprises within the preset ranking range, the measured historical data of adjacent years or the unit production capacity emission level of enterprises of the same scale is obtained and estimated.
[0086] If the corresponding enterprise is not among the key production capacity enterprises in the preset ranking range, then the data of the corresponding enterprise will be removed.
[0087] Specifically, in existing technologies, if a company has data gaps for more than three consecutive months, forcibly filling in the gaps (such as through linear interpolation) can easily introduce false fluctuations (e.g., misjudging "zero emissions" during shutdown periods as low emissions during normal production), leading to serious distortion of the annual total emissions. This embodiment avoids interference from invalid data in the annual accounting by removing interrupted periods and calculating the annual total emissions based on the remaining valid months, ensuring that the annual total emissions are closer to the actual operating status of the company (e.g., shutdown periods are not included in the calculation, and the calculation is based only on the valid production months), providing more reliable basic data for subsequent screening of key load enterprises. For companies with missing data for the whole year, directly removing them may miss key production capacity enterprises that contribute significantly to regional emissions (e.g., a large phosphate mine lost its data for the whole year due to a system failure, but its production capacity accounts for 30% of the region); if forcibly estimated, the lack of basis may lead to excessive errors. This embodiment achieves differentiated processing by verifying production capacity rankings. For key production capacity enterprises, estimations are made using extrapolation of trends from adjacent years or by analogy with similar enterprises. This avoids missing key pollution sources due to data gaps and constrains estimation biases through industry commonalities or historical patterns (e.g., referencing the unit production capacity emission intensity of phosphate chemical enterprises of similar scale), ensuring that key enterprises are not overlooked. Non-key production capacity enterprises are directly eliminated to reduce interference from invalid data in subsequent analyses (e.g., missing data from small mining enterprises has a limited impact on total regional emissions), improving overall data quality. Through this layered processing, the multi-source reconstructed data retains the seasonal patterns of historical data for the same month (e.g., the restoration of the monthly average in previous steps) while eliminating long-term missing data and noise from irrelevant enterprises. This provides a "cleaner" and more representative input for subsequent two-dimensional pollution source identification (annual average emissions, emission intensity per unit output value calculation), material flow-system dynamics model calibration, and governance strategy derivation, ultimately enhancing the scientific rigor and decision-making reference value of the entire methodology.
[0088] In one embodiment, for periods with more than three consecutive months missing, these periods are removed from the annual emissions calculation, and the total annual emissions are calculated proportionally based on the remaining months with available data. The calculation formula is as follows:
[0089] ,
[0090] in:
[0091] : Converted annual emissions (kg);
[0092] : The sum of actual emissions (kg) for months with available data;
[0093] The number of months for which data is available.
[0094] For companies with completely missing annual data, this embodiment first determines whether they are key candidates (e.g., top 50%) based on their annual output value or production capacity ranking. If they are key candidates, the annual emissions are estimated using trend extrapolation (linear interpolation) or analogy with similar companies; if they are not key candidates, the data for that year is removed.
[0095] The formula for trend extrapolation is as follows:
[0096]
[0097] in:
[0098] : Estimated annual emissions (kg) for missing years;
[0099] : Annual emissions (kg) for the previous year for which data is available;
[0100] : Annual emissions (kg) for the next year for which data is available;
[0101] : The serial number of the missing year;
[0102] : The serial number of the previous year with data;
[0103] : The serial number of the next year with data; all of the above serial numbers are dimensionless.
[0104] The formula for the analogy method using similar companies is as follows:
[0105] ,
[0106] in:
[0107] Estimated annual emissions (kg) for companies with missing data;
[0108] Average emission intensity per unit capacity (kg / ton of product) of 3 to 5 companies of similar size in the same industry.
[0109] : Production capacity (tons of product / year) of the missing enterprises.
[0110] Optionally, the step of performing a two-dimensional fusion analysis on the multi-source reconstructed data to obtain a list of key pollution sources includes:
[0111] Based on the multi-source reconstructed data, the annual average pollutant emissions of each of the three phosphorus enterprises are obtained. All the three phosphorus enterprises are sorted from high to low according to the annual average pollutant emissions, and the total emissions of the three phosphorus enterprises are accumulated according to the sorting until the predetermined emission threshold of the total emissions in the target area is reached. Based on all the accumulated three phosphorus enterprises, the set of key load enterprises is obtained.
[0112] Based on the multi-source reconstructed data, the annual output value data of all the "three phosphorus" enterprises are obtained. The corresponding pollutant emission amount is generated based on the annual output value data. The emission intensity per unit output value of the corresponding "three phosphorus" enterprises is obtained based on the pollutant emission amount. The corresponding "three phosphorus" enterprises with emission intensity higher than the preset average intensity threshold are screened out to obtain a set of inefficient key enterprises.
[0113] The list of key pollution sources is obtained by taking the union of the set of key enterprises with high load and the set of key inefficient enterprises.
[0114] Specifically, existing technologies only rank and select key enterprises based on total emissions, which easily overlooks small and medium-sized enterprises (SMEs) with "small total emissions but extremely high pollution per unit of output value" (such as some small phosphate chemical enterprises with outdated technology, whose annual emissions are only a few thousand tons, but whose emissions per 10,000 yuan of output value are three times the industry average). The load-bearing key enterprise set: ranked based on average annual total emissions, accumulated to a predetermined threshold (e.g., 80%) for regional total emissions, covering the "major emitters" that contribute the most to the region's total phosphorus (such as large phosphate mines and phosphoric acid plants); the inefficient key enterprise set: screened based on emission intensity per unit of output value, identifying "high-consumption, high-pollution, low-output" "inefficient high-emission" enterprises (even if their annual emissions do not reach the total threshold, they are still included in the regulatory scope because their pollution per unit of output value far exceeds the industry average). The combined list of key pollution sources completely solves the one-sidedness of single-dimensional identification.
[0115] Traditional methods focus solely on environmental benefits (emission reductions) while neglecting economic benefits (output contribution), easily leading to an imbalance between "sacrificing reasonable industrial development for emission reduction" and "protecting outdated production capacity." This embodiment internalizes environmental costs into economic evaluation indicators through the emission intensity per unit output value: for enterprises with "high load + high efficiency" (such as large-scale modern phosphate chemical enterprises, with large emissions but low output per unit), the focus is on optimizing emission reduction technologies rather than restricting production capacity; for enterprises with "low load + low efficiency" (such as small, outdated enterprises, with small emissions but extremely high output per unit), priority is given to promoting transformation and upgrading or shutting down and merging. This differentiated identification logic guides regulatory resources towards enterprises with high emission reduction potential and low economic costs, achieving a win-win situation for both the environment and the economy. Simultaneously, the dual-dimensional attributes of the key pollution source list directly support the targeted design of subsequent governance strategies: for enterprises with high load, the focus is on "total quantity control" (such as setting annual emission reduction quotas); for inefficient key enterprises, the focus is on "efficiency improvement" (such as mandating the promotion of clean production technologies and eliminating outdated processes). For example, by identifying 23 key enterprises (18 high-load enterprises and 5 low-efficiency enterprises) through a dual-dimensional approach, subsequent governance scenarios (such as a 30% reduction in emissions at the source) can be tailored to these enterprises, avoiding the implementation resistance or diminished effectiveness caused by a "one-size-fits-all" approach. Considering the diverse types of "three phosphorus" enterprises (mining, chemical, and stockpiling) and their significant scale differences (from million-ton-per-year phosphate rock to tens of thousands-ton-per-year fine chemical production), a single dimension is easily influenced by enterprise size (large-scale enterprises naturally occupy the top positions in total emissions, masking the pollution problems of small and inefficient enterprises). The dual-dimensional fusion analysis, through cross-validation of total emissions and efficiency, effectively filters out the interference of scale factors, truly reflecting the pollution nature of enterprises. Even under complex circumstances such as fluctuations in enterprise production capacity and industrial restructuring, it maintains stable and reliable identification results. This breaks through the limitations of the single perspective in traditional pollution source identification, providing a more comprehensive, accurate, and operational target list for "three phosphorus" pollution control, and is a core prerequisite for improving the effectiveness of subsequent governance strategies.
[0116] In one embodiment, combined with Figure 2 As shown, the annual average emission Ei of each "three phosphorus" enterprise is calculated. Assuming that the top 18 enterprises, sorted by Ei from largest to smallest, account for 80% of the total emissions, these 18 enterprises are marked as "key load enterprises". The emission intensity per unit output value Ii is calculated based on the enterprise's annual output value. Assuming that the Ii of 5 enterprises exceeds 1.5 times the industry average, they are added as "key inefficient enterprises". Finally, 23 key enterprises are obtained, and the list and reasons for selection are output (e.g., Enterprise A – high load; Enterprise B – inefficient and high emission; Enterprise C – both).
[0117] The formula for calculating the average annual emissions of each enterprise is as follows:
[0118] ,
[0119] in:
[0120] : Annual average emissions of the i-th company (kg / year);
[0121] : Total annual emissions (kg) of the i-th company in year k;
[0122] : The number of years that the i-th company has emissions data.
[0123] Each enterprise Sort by largest to smallest, calculate the cumulative emission percentage using the following formula:
[0124] ,
[0125] in:
[0126] : Cumulative emissions percentage of the top j companies (%) Total number of enterprises.
[0127] Filter out The top j companies with ≥80% load are marked as "key load enterprises";
[0128] For enterprises that have already obtained their annual output value, the emission intensity per unit output value is calculated using the following formula:
[0129] ,
[0130] in:
[0131] : Emission intensity per unit output value of the i-th enterprise (kg / ten thousand yuan);
[0132] : Annual output value (ten thousand yuan) of the i-th enterprise.
[0133] Enterprise annual output value Prioritize obtaining data from statistical yearbooks or corporate annual reports; if unavailable, estimate using the following formula:
[0134] ,
[0135] in:
[0136] The number of employees covered by social insurance in this company (in persons);
[0137] Average output per person in the same industry (ten thousand yuan / person)
[0138] Calculate the average emission intensity of companies in the same industry Filter out , marked as "inefficient key enterprises". It is the strength multiplier, with a value ranging from 1.2 to 2.0, preferably 1.5.
[0139] Optionally, the step of constructing a material flow-system dynamic coupling model and obtaining historical monitoring data to calibrate the parameters of the material flow-system dynamic coupling model to obtain a calibrated material flow-system dynamic coupling model includes:
[0140] Construct a phosphate mining subsystem, a phosphate chemical production subsystem, a phosphogypsum storage subsystem, a wastewater treatment subsystem, a product consumption subsystem, and a water environment subsystem, and establish the material transfer topology between each subsystem;
[0141] The mass transfer topology is mapped to the mass flow-system dynamics coupled model, and key rate equations are set.
[0142] Based on the historical monitoring data, the material flow-system dynamic coupling model is reverse-tuned to adjust the conversion coefficient and loss parameters within the model. When the output data of the material flow-system dynamic coupling model matches the historical data of the actual monitoring data to a preset standard, the calibrated material flow-system dynamic coupling model is obtained.
[0143] Specifically, this embodiment effectively overcomes the shortcomings of existing technologies, such as crude handling of missing data, one-sided identification of pollution sources, static and rigid simulation analysis, and weak decision support. This method is not only applicable to the treatment of phosphorus pollution in the Yichang section of the middle reaches of the Yangtze River, but also provides a replicable and scalable technical solution for the refined environmental management of other complex industrial clusters.
[0144] In one embodiment, a material flow balance model is established for a phosphate mining subsystem, a phosphate chemical production subsystem, a phosphogypsum storage subsystem, a wastewater treatment subsystem, a product consumption subsystem, and a water environment subsystem, wherein each subsystem follows the law of conservation of mass.
[0145] The material transfer topology between subsystems is mapped to a coupled material flow-system dynamics model, combined with... Figure 3 The rectangular stock modules in the dynamic stock-flow diagram of the material flow-system dynamic coupling model shown represent, in turn, the phosphorus accumulation Mmining of the phosphate mining subsystem, the phosphorus accumulation Mprod of the phosphate chemical production subsystem, the phosphorus accumulation Mcons of the product consumption subsystem, the phosphorus accumulation Msludg of the wastewater treatment sludge, the phosphogypsum stockpile Sgyp, and the phosphorus pool in the water body Wp, all in kg. Figure 3Solid arrows in the diagram represent the transfer of phosphorus or pollution load, while dashed arrows represent the feedback relationship between decision intensity, treatment behavior, and water quality response. Wherein, Rmining is the phosphate rock mining rate, Rconcentrate is the phosphate concentrate output rate, Rwaste_mining is the phosphorus loss rate from mining waste rock and tailings; Racid is the phosphorus rate introduced by auxiliary materials, Rproduct is the phosphate chemical product production rate, Rwaste_prod is the phosphorus generation rate entering the treatment system during chemical production, Rwaste_prod_direct is the phosphorus rate directly discharged without treatment; Rconsume is the product consumption rate, Rwaste_cons is the phosphorus rate entering the wastewater or solid waste treatment system after consumption, Rexport is the phosphorus rate output outside the watershed; Ggyp is the phosphogypsum production rate, Rrecyc is the phosphogypsum resource utilization rate, Rleach is the phosphogypsum infiltration rate; Rwwtp is the phosphorus rate removed from the wastewater treatment plant and enters the sludge, Reffluent is the wastewater treatment plant effluent discharge rate, Rsludge_util is the sludge resource utilization rate, Rsludge_landfill is the sludge landfill disposal rate, and Rsed is the natural sedimentation rate of phosphorus in water bodies.
[0146] This embodiment uses 2019 to 2023 as the parameter calibration period, 2024 as the verification period, and 2025 to 2035 as the prediction period. First, the monthly emissions from enterprises, wastewater treatment plant influent and effluent volumes, total phosphorus concentration, total phosphorus content of phosphogypsum, enterprise output, and resource utilization obtained in previous embodiments are uniformly converted to kg / year. Second, the input values for each year are written into... Figure 3 The corresponding flow variables, and according to Figure 3 The arrows update the existing variables in different directions. For example, in the phosphate mining module, Rmining enters Mmining, Rconcentrate outputs to the phosphate chemical production module, and Rwaste_mining enters the phosphate pool Wp in the water body. In the phosphate chemical production module, Rproduct enters the product consumption module, Rwaste_prod enters the wastewater treatment module, and Rwaste_prod_direct directly enters Wp. In the phosphogypsum storage module, Ggyp increases Sgyp, and Rrecyc and Rleach decrease Sgyp, with Rleach entering Wp. In the wastewater treatment module, the phosphorus load entering the treatment system is divided into Rwwtp and Reffluent according to the removal efficiency η. Rwwtp enters Msludg, and Reffluent enters Wp. Finally, the phosphate pool Wp receives Reffluent, Rleach, Rwaste_mining, and Rwaste_prod_direct, and subtracts Rsed to obtain the change in the phosphate pool in the water body for each year.
[0147] The calculation method for the key rate equation is as follows:
[0148] The phosphogypsum production rate Ggyp is determined by the phosphoric acid production Pacid, the phosphogypsum pollution coefficient γ, and the total phosphorus content of phosphogypsum Cp, i.e., Ggyp = Pacid × γ × Cp;
[0149] In this embodiment, the initial γ content is 5.5 tons of phosphogypsum / ton of P2O5, and the initial Cp content is 0.37%.
[0150] The phosphogypsum percolation rate Rleach = Sgyp × ηl
[0151] Wherein, ηl is the annual percolation coefficient, initially taken as 2.0%, and sampled in the uncertainty analysis according to the triangular distribution (0.5%, 2.0%, 5.0%).
[0152] The effluent discharge rate of a wastewater treatment plant is Reffluent = Rin, wwtp × (1-η).
[0153] Where Rin,wwtp is the total phosphorus load entering the wastewater treatment plant, and η is the total phosphorus removal efficiency;
[0154] The phosphorus accumulation rate in sludge is Rwwtp = Rin, wwtp × η.
[0155] The industrial emission rate under the influence of decision-making is calculated as Rind = R0 × (1 - K).
[0156] Where R0 is the baseline emission rate (kg / month) without decision intervention; K is the decision intensity coefficient, with a value range of [0, 0.5].
[0157] The dynamic update equations for each state variable in this embodiment are as follows (t in all equations represents the year number, and t+1 represents the next year):
[0158] Phosphorus accumulation in the phosphate mining subsystem:
[0159] ,
[0160] in:
[0161] Phosphorus accumulation in the phosphate mining subsystem at the end of year t, in kg;
[0162] : Rate of phosphorus introduced by phosphate rock mining in year t, unit: kg / year;
[0163] : Phosphate concentrate output rate in year t, unit: kg / year;
[0164] The rate of phosphorus loss from mining waste rock and tailings in year t, in kg / year.
[0165] Phosphorus accumulation in the phosphorus chemical production subsystem:
[0166] ,
[0167] in:
[0168] Phosphorus accumulation in the phosphorus chemical production subsystem at the end of year t, in kg;
[0169] : Rate of phosphorus introduced by sulfuric acid and other auxiliary materials in year t, unit: kg / year (0 if not specified);
[0170] : Rate of phosphorus removal by phosphorus chemical products in year t, unit: kg / year;
[0171] Phosphorus generation rate of wastewater and solid waste in chemical production process in year t (entering the treatment system), unit: kg / year.
[0172] Phosphorus accumulation in the product consumption subsystem:
[0173] ,
[0174] in:
[0175] Phosphorus accumulation in the product consumption subsystem at the end of year t, in kg;
[0176] Product consumption rate in year t, unit: kg / year;
[0177] Phosphorus rate entering wastewater and solid waste treatment systems after consumption in year t, unit: kg / year;
[0178] Phosphorus output rate outside the basin in year t, unit: kg / year.
[0179] Phosphorus accumulation in wastewater treatment sludge:
[0180] ,
[0181] in:
[0182] Phosphorus accumulation in wastewater treatment sludge at the end of year t, unit: kg;
[0183] : The rate at which a wastewater treatment plant removes phosphorus from wastewater in year t, in kg / year;
[0184] : Sludge resource utilization rate in year t, unit: kg / year;
[0185] : Sludge landfill disposal rate in year t, unit: kg / year.
[0186] Dynamic updates on phosphogypsum stockpiles:
[0187] ,
[0188] in:
[0189] : The stockpile of phosphogypsum at the end of year t, in kg;
[0190] : Phosphogypsum production in year t, unit: kg;
[0191] Resource utilization amount in year t, unit: kg;
[0192] leachate release in year t, in kg.
[0193] Dynamic updates of phosphorus pools in water bodies:
[0194] ,
[0195] in:
[0196] Phosphorus pool in water at the end of year t, unit: kg;
[0197] Wastewater discharge rate from the wastewater treatment plant in year t, unit: kg / year;
[0198] : Percolation rate of phosphogypsum in year t, unit: kg / year;
[0199] Phosphorus rate of mining waste rock and tailings directly entering water bodies in year t, unit: kg / year;
[0200] Phosphorus rate of untreated chemical production wastewater directly discharged in year t, unit: kg / year;
[0201] : Natural sedimentation rate of phosphorus in water in year t, unit: kg / year.
[0202] The parameter calibration process in this embodiment is as follows: Using the measured total phosphorus concentration (Cobs,t) at the cross-section as the calibration target, the phosphorus pool (Wp,t) calculated by the model is converted to the simulated total phosphorus concentration (Csim,t) at the cross-section. During the conversion, the annual average runoff or annual average flow volume (Qsection,t) of the target cross-section is used, and Csim,t = Wp,t / Qsection,t. If there is multi-section monitoring data for the target area, the calculation is weighted according to the cross-section control area or flow rate weight.
[0203] Construct the objective function of average relative error J = (1 / T)Σ|Csim,t-Cobs,t| / Cobs,t,
[0204] Where T represents the number of calibration years.
[0205] By using grid search or least squares iteration, the mineral processing recovery rate, local consumption ratio, wastewater treatment plant removal efficiency η, phosphogypsum infiltration coefficient ηl, water body natural settling coefficient, and decision intensity coefficient K are adjusted to ensure that J does not exceed 20%.
[0206] For example, the exemplary calibration results of a specific embodiment are as follows: the mineral processing recovery rate is 0.86, the local consumption ratio is 0.64, the total phosphorus removal efficiency η of the wastewater treatment plant is 0.78, the annual infiltration coefficient ηl of phosphogypsum is 0.021, the natural settling coefficient of water body is 0.18, and the decision intensity coefficient K is 0 under the baseline scenario. The total phosphorus concentration at the cross-section was calibrated from 2019 to 2023. The measured value in 2019 was 0.22 mg / L, the simulated value was 0.24 mg / L, and the relative error was 9.1%; the measured value in 2020 was 0.21 mg / L, the simulated value was 0.20 mg / L, and the relative error was 4.8%; the measured value in 2021 was 0.24 mg / L, the simulated value was 0.25 mg / L, and the relative error was 4.2%; the measured value in 2022 was 0.23 mg / L, the simulated value was 0.21 mg / L, and the relative error was 8.7%; and the measured value in 2023 was 0.20 mg / L, the simulated value was 0.19 mg / L, and the relative error was 5.0%. The average relative error was 6.36%, which is less than the preset threshold of 20%. Using 2024 as the verification period, the simulated total phosphorus concentration at the cross-section was 0.18 mg / L, while the measured value was 0.20 mg / L, with a relative error of 10%. This demonstrates that the material flow-system dynamics coupled model can be used for subsequent scenario calculations. It is understood that the above values are used to illustrate the implementation process of this embodiment and do not constitute a limitation on the actual monitoring results in the region.
[0207] Optionally, the acquisition of key variable parameters involves using Latin hypercube sampling combined with Monte Carlo simulation, and combining this with the list of key pollution sources to perform random simulations on the key variable parameters, resulting in a set of governance scenarios, including:
[0208] Based on the key variable parameters, the corresponding probability distribution characteristics are obtained;
[0209] The probability distribution features are sampled using the Latin hypercube sampling method in a stratified combination to generate several sets of parameter combinations.
[0210] According to the Monte Carlo simulation method, the set of future world states is generated by performing a preset number of random iterations on the several sets of parameter combinations.
[0211] Based on the list of key pollution sources and the set of future world states, several different governance measures are generated. Two or more of these governance measures are selected and combined to form the set of governance scenarios.
[0212] Specifically, existing technologies often use simple random sampling to generate parameter combinations, which can easily lead to insufficient or repeated sampling of key parameter ranges (such as the high-value area of the phosphogypsum permeability coefficient), requiring numerous iterations to cover the uncertainty space. This embodiment uses Latin hypercube sampling to perform stratified combined sampling of the probability distribution of key variable parameters, ensuring that each parameter range has representative samples. It generates uniformly distributed parameter combinations with fewer iterations (e.g., only 1000 sets are needed to achieve the coverage of 5000 sets in traditional methods), significantly improving computational efficiency while avoiding the omission of high-risk scenarios (such as extreme rainfall causing a sudden increase in the permeability coefficient). To avoid single-scenario bias, this embodiment uses Monte Carlo simulation to perform tens of thousands of random iterations on the parameter combinations, generating thousands of future world states covering "high growth + high permeability" and "low growth + low efficiency," comprehensively covering deep uncertainties such as economic fluctuations, technological changes, and extreme weather. This upgrades the evaluation of governance strategies from "single-point prediction" to "probability distribution analysis," completely eliminating the limitations of single-decision approaches. Furthermore, it closely integrates with the key pollution source list identified from two dimensions to generate targeted governance measures. For example, the focus is on "total emission reduction" for key enterprises with high pollution loads, and on "efficiency improvement" for key inefficient enterprises, avoiding resource waste or implementation resistance caused by a one-size-fits-all approach. In practice, for example, a "mandatory clean production audit" measure is specifically designed for the five inefficient enterprises on the key list, rather than a general requirement for industry-wide emission reduction. In actual governance, single governance measures (such as only source emission reduction or only end-of-pipe treatment) often fail to meet targets due to diminishing marginal returns. This embodiment selects two or more measures in combination (such as "20% source emission reduction + 10% efficiency improvement in end-of-pipe treatment" and "80% resource utilization + strengthened policy enforcement") to fully explore the synergistic effect between measures: source emission reduction reduces the load entering rivers, end-of-pipe treatment covers residual pollution, and resource utilization reduces the risk of stockpiling, achieving a governance effect of 1+1>2. It is understandable that the generated governance scenario set includes multi-tiered solutions ranging from "maintaining the status quo" to "strong combinations," covering different cost-benefit preferences (such as low-cost weak measures vs. high-cost strong measures). By combining subsequent assessments of "probability of compliance + CVaR," robust strategies that "achieve compliance in 80% of the future and have controllable extreme risks" can be selected from the scenario set, avoiding decision-making errors due to excessive optimism or conservatism. For example, if a scenario combination achieves compliance in 90% of the future but has an extremely high CVaR (high extreme risk), it will be excluded, ensuring that the final strategy balances safety and feasibility. Meanwhile, the "three phosphorus" industry involves multiple stages such as mining, chemical processing, and stockpiling, and the industrial characteristics of different regions (such as the Yichang section and the Wujiang River basin) vary significantly. This embodiment uses stratified parameter sampling and customized combination of measures to allow the governance scenario set to flexibly adapt to the industrial structure and pollution characteristics of different regions. For example, for concentrated phosphate mining areas, the focus is on leachate control of mining waste rock; for concentrated chemical processing areas, the focus is on deep treatment of production wastewater; and for concentrated stockpiling areas, the focus is on comprehensive utilization of phosphogypsum, avoiding the incompatibility of general models.
[0213] In one embodiment, for example, key variable parameters are determined to obtain the corresponding probability distribution characteristics:
[0214] Wastewater treatment plant removal efficiency: follows a uniform distribution, with a relative variation range of ±5%.
[0215] The permeability coefficient of phosphogypsum follows a triangular distribution, with a minimum value of 0.5%, a most likely value of 2%, and a maximum value of 5%.
[0216] Regional GDP annual growth rate: follows a triangular distribution, with a minimum of 3%, a most likely value of 5%, and a maximum value of 7%.
[0217] Policy enforcement capacity: follows a uniform distribution, ranging from 0.8 to 1.2.
[0218] Latin hypercube sampling is used to perform stratified sampling of the probability distribution characteristics corresponding to each parameter, generating several sets of parameter combinations. Combined with Monte Carlo simulation, N sets of future world states are constructed, where N≥1000.
[0219] Assume the following governance scenario:
[0220] Baseline scenario: Maintain current emission intensity, removal efficiency, and comprehensive utilization rate of phosphogypsum.
[0221] Source reduction scenario: Reduce the emission intensity of key enterprises by 10%, 20% or 30% respectively.
[0222] End-of-pipe treatment scenario: Increase the total phosphorus removal efficiency of wastewater treatment plants by 5% or 10%, respectively.
[0223] Resource utilization scenario: Increase the comprehensive utilization rate of phosphogypsum to 80% or 90%, respectively.
[0224] Two or more of the aforementioned governance measures are selected and combined to form the governance scenario set.
[0225] In another embodiment, this embodiment assumes that five governance scenarios are set up and applied to respectively. Figure 3 Different flows within.
[0226] BAU is the baseline scenario, with no changes to K, η, and Rrecyc;
[0227] S1 represents the source reduction scenario, which reduces the emission intensity of key enterprises by 20%, corresponding to a reduction in Rwaste_prod and Rwaste_prod_direct.
[0228] S2 represents the end-of-pipe treatment scenario, which increases the total phosphorus removal efficiency of wastewater treatment plants by 10%, corresponding to an increase in Rwwtp and a decrease in Reffluent.
[0229] S3 represents the resource utilization scenario, which increases the comprehensive utilization rate of phosphogypsum to 80%, correspondingly increasing Rrecyc and decreasing Sgyp and subsequent Rleach;
[0230] S4 is a combined scenario, where S1 and S2 are implemented simultaneously.
[0231] After simulation, the success rate of scenario S4 (92%) was significantly higher than that of other single measures (78% for S1 and 71% for S2), verifying the necessity of combined optimization.
[0232] Optionally, the step of deriving candidate governance strategies by extrapolating the set of governance scenarios based on the calibrated material flow-system dynamics coupling model includes:
[0233] The set of treatment scenarios is input into the calibrated material flow-system dynamics coupled model for simulation, and the probability of water quality compliance for each scheme in the set of treatment scenarios is statistically analyzed.
[0234] Obtain the tail risk value of each group of schemes, and obtain candidate treatment strategies based on the water quality compliance probability and the tail risk value.
[0235] Specifically, this embodiment inputs a set of governance scenarios into a calibrated material flow-system dynamics coupled model, simulating thousands of future world conditions and statistically analyzing the probability of water quality compliance for each scheme. This is no longer about predicting what will definitely happen, but rather quantifying the likelihood of what might happen, providing decision-makers with a probability space for strategy success (e.g., "this scheme has a 92% probability of meeting the standard"), greatly reducing the blindness of decision-making. Furthermore, simply looking at the probability of compliance is insufficient; some schemes, while having high average compliance rates, may cause catastrophic exceedances under extremely unfavorable circumstances (such as severe drought leading to loss of water body self-purification capacity, or serious policy ineffectiveness). This embodiment innovatively introduces Conditional Value at Risk (CVaR) as a "tail risk value," specifically measuring the average degree of exceedance under worst-case scenarios. This ensures that the selected candidate strategies are not only highly feasible but also have controllable consequences under low-probability events, aligning with the core principle of "prevention first" in environmental risk management. This is to prevent the pursuit of high compliance probabilities from leading to the selection of "aggressive schemes" with excessively high costs and difficulty in implementation, or the selection of conservative and ineffective "moderate schemes" from simply avoiding risks. This embodiment uses a comprehensive evaluation based on two dimensions: "probability of water quality compliance (efficiency)" and "tail-end risk value (safety)." It automatically eliminates highly efficient but unstable solutions (high compliance rate but extremely high CVaR) and stable but ineffective solutions (low CVaR but insufficient compliance rate), accurately selecting robust strategies that balance environmental benefits and risk resistance. For example, the S4 scenario selected in the above embodiment was chosen as the optimal solution because it meets the standard in 92% of scenarios and has the lowest extreme risk. Furthermore, since this embodiment uses a material flow-system dynamics coupling model calibrated with historical data (with an error set at ≤20% in the experiment), its simulated "treatment measures-water quality response" relationship closely matches physical reality. Therefore, the resulting candidate strategies are not theoretical deductions but quantitative predictions based on the actual laws of regional phosphorus cycling, providing users with a highly operational action guide.
[0236] In one embodiment, a coupled material flow-system dynamics model is run to calculate the probability of achieving the target total phosphorus concentration at the cross-section for each governance scenario under N future world states, using the following formula:
[0237] ,
[0238] in:
[0239] : The probability of achieving the target for the j-th governance scenario, dimensionless;
[0240] Under the j-th governance scenario, the number of future world states in which the simulated total phosphorus concentration at the cross-section is ≤ the water quality standard value;
[0241] Total number of future world states =N≥1000.
[0242] The Conditional Value at Risk (CVaR) for each governance scenario is calculated to measure the average degree of exceedance under extremely adverse conditions, using the following formula:
[0243] ,
[0244] in:
[0245] : Confidence level, indicating that the probability of examining the tail is 1- For the worst-case scenario, we take 0.9;
[0246] K: Number of tail scenarios
[0247] : The concentration value (mg / L) of the kth state after sorting the future world states by total phosphorus concentration in the cross section from high to low.
[0248] Water quality standard concentration (mg / L), take 0.2 mg / L (Class III surface water standard);
[0249] : The out-of-scale value of the k-th tail scenario; if it does not exceed the scale, take 0.
[0250] Using the probability of compliance as the ordinate and CVaR as the abscissa, each governance scenario is mapped onto a two-dimensional plane. Scenarios with a compliance probability greater than 0.8 and a CVaR lower than the median CVaR of all scenarios are selected as candidate governance strategies.
[0251] In another embodiment, this embodiment operates as follows for each governance scenario j and future world state s: Figure 3 The table shows the total phosphorus concentrations (Cj, s, t) at various cross-sections before 2035. Using the Class III surface water total phosphorus standard (Cstd) of 0.2 mg / L as the compliance threshold, a condition is considered compliant if Cj, s, t ≤ Cstd in the target year or evaluation period; otherwise, it is considered non-compliant. Dividing the number of compliance occurrences in 1000 future world states by 1000 yields the compliance probability (Pj) for governance scenario j.
[0252] Conditional Value at Risk The calculation process is as follows: First, sort the total phosphorus concentration of each treatment scenario in 1000 sets of future world states from high to low, and take the 10% with the highest concentration as the tail adverse scenario; then calculate the average excess amount relative to Cstd in these tail scenarios. If a tail scenario does not exceed the standard, the excess amount of that scenario is taken as 0. The smaller the value, the lower the risk of the strategy exceeding the limit under extremely unfavorable conditions.
[0253] The decision calculation results of this embodiment are as follows: the probability of achieving the BAU scenario is 45%. The concentration was 0.17 mg / L; the probability of meeting the target in scenario S1 was 78%. The concentration was 0.09 mg / L; the probability of meeting the target in scenario S2 was 71%. The concentration was 0.11 mg / L; the probability of meeting the target in scenario S3 was 62%. The concentration was 0.13 mg / L; the probability of meeting the target in scenario S4 was 92%. The concentration was 0.05 mg / L. Since the probability of S4 meeting the standard was greater than 0.8, and... Since the CVaR is lower than the median value for each scenario, S4 is selected as the candidate governance strategy, and the strategy robustness ranking is output as S4>S1>S2>S3>BAU.
[0254] The specific switching rules for the dynamic adaptive policy path are as follows: the initial governance path adopts S1; after obtaining the total phosphorus monitoring data of the cross-section each year, the measured concentration of that year is written into... Figure 3 The water quality feedback variables are analyzed, and the Wp and Csim for the following year are recalculated. If the measured or simulated total phosphorus concentration at the cross-section is higher than 0.2 mg / L for two consecutive years, the strategy will be upgraded from the current strategy to S4 at the beginning of the following year. If the cumulative treatment cost accounts for more than 3% of the regional GDP and the water quality target is still not met, the combined strategy with a higher probability of achieving the target per unit cost will be prioritized. For example, if the total phosphorus concentration at the cross-section is 0.223 mg / L in 2027 and 0.231 mg / L in 2028, the strategy upgrade will be triggered at the beginning of 2029, switching from S1 to S4. If the risk of exceeding the target is predicted to be triggered around 2030 under the future world conditions of high growth and high infiltration, the combined measures originally planned to be implemented in 2035 will be brought forward to 2030, that is, the combined treatment will be started 5 years in advance.
[0255] Using the methods described above, Figure 3 The material flow-system dynamics coupled model not only outputs the phosphorus stock changes of each subsystem, but also connects the risk of water quality exceeding standards, treatment costs, and decision intensity K through dashed feedback to generate an adaptive path diagram. The horizontal axis of the adaptive path diagram represents the year, and the vertical axis lists treatment scenarios such as BAU, S1, S2, S3, and S4. The arrows indicate the time points of strategy switching caused by triggering conditions.
[0256] The triggering conditions (preset triggering mechanisms, which can be understood as specific indicators of the triggering mechanism, determined based on actual governance needs; this embodiment is only an example and not a limitation) are set as follows:
[0257] Triggering condition A: The total phosphorus concentration at the monitored section exceeds Cstd for two consecutive years.
[0258] Triggering condition B: The cumulative governance cost accounts for more than 3% of the region's GDP.
[0259] When any trigger condition is met, the system will automatically switch from the current strategy to stronger governance measures (e.g., upgrading from 10% emission reduction at the source to 20% emission reduction at the source and simultaneously implementing end-of-pipe treatment), and output the time point of the strategy switch.
[0260] Finally, this embodiment outputs a list of key enterprises, a strategy robustness ranking (S4>S1>S2>S3>BAU), an adaptive decision-making path diagram, and a future water quality prediction curve under the optimal strategy S4. The prediction shows that after implementing S4, the total phosphorus concentration at the cross-section can be reduced to below 0.15 mg / L by 2035. This embodiment achieves automatic data filling and output of robust governance strategies, and can be dynamically adjusted and updated in a timely manner.
[0261] This invention also provides a system for identifying and optimizing pollution sources with missing data and control strategies, comprising:
[0262] The data repair unit is used to acquire multi-source data of "three phosphorus" enterprises in the target area, perform missing data repair processing based on time-series features on the multi-source data, and obtain multi-source reconstructed data.
[0263] The data analysis unit is used to perform two-dimensional fusion analysis on the multi-source reconstructed data to obtain a list of key pollution sources.
[0264] The model building unit is used to build a material flow-system dynamics coupled model, acquire historical monitoring data to calibrate the parameters of the material flow-system dynamics coupled model, and obtain the calibrated material flow-system dynamics coupled model.
[0265] The data processing unit is used to obtain key variable parameters, and uses Latin hypercube sampling combined with Monte Carlo simulation, combined with the list of key pollution sources, to perform random simulation on the key variable parameters to obtain a set of treatment scenarios.
[0266] The data extrapolation unit is used to extrapolate the set of governance scenarios based on the calibrated material flow-system dynamics coupling model to obtain candidate governance strategies;
[0267] The decision control unit is used to dynamically control the candidate governance strategies based on a preset triggering mechanism to obtain the identification results and optimization strategies.
[0268] The data missing pollution source identification and governance strategy optimization system of the present invention has the same advantages over the prior art as the data missing pollution source identification and governance strategy optimization method described above, and will not be repeated here.
[0269] This invention also provides an electronic device, including: a processor and a memory, wherein the memory is used to store computer programs;
[0270] When the computer program is loaded by the processor, it causes the processor to execute the data missing pollution source identification and control strategy optimization method as described above.
[0271] The electronic device of the present invention has the same advantages over the prior art as the aforementioned method for identifying and optimizing pollution sources with missing data, and will not be repeated here.
[0272] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the data missing pollution source identification and governance strategy optimization method described above.
[0273] The computer-readable storage medium of the present invention has the same advantages over the prior art as the aforementioned method for identifying and optimizing pollution sources and control strategies for data loss, and will not be repeated here.
[0274] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.
Claims
1. A method for identifying pollution sources with missing data and optimizing treatment strategies, characterized in that, The method includes: Acquire multi-source data of "three phosphorus" enterprises within the target area, and perform missing data repair processing based on time-series features to obtain multi-source reconstructed data; A two-dimensional fusion analysis was performed on the multi-source reconstructed data to obtain a list of key pollution sources; A material flow-system dynamics coupled model is constructed, and historical monitoring data is obtained to calibrate the parameters of the material flow-system dynamics coupled model to obtain the calibrated material flow-system dynamics coupled model. Key variable parameters are obtained, and Latin hypercube sampling combined with Monte Carlo simulation is used. Combined with the list of key pollution sources, the key variable parameters are randomly simulated to obtain a set of treatment scenarios. Based on the calibrated material flow-system dynamics coupling model, the set of governance scenarios is deduced to obtain candidate governance strategies; The candidate governance strategies are dynamically adjusted based on a preset triggering mechanism to obtain identification results and optimization strategies.
2. The method for identifying and optimizing pollution sources with missing data according to claim 1, characterized in that, The process involves acquiring multi-source data on "three phosphorus" enterprises within the target area, performing time-series feature-based missing data repair processing on the multi-source data, and obtaining multi-source reconstructed data, including: Acquire multi-source data of "three phosphorus" enterprises within the target area, including monthly emission records, annual output value reports, and solid waste storage site composition test reports; Based on the blank data points in the monthly emission records, the actual measured historical data of the corresponding enterprise in the same month of previous years are extracted. The arithmetic mean of the actual measured historical data is calculated and then missing data repair is performed to obtain the multi-source reconstructed data.
3. The method for identifying and optimizing pollution sources with missing data according to claim 2, characterized in that, The step of extracting historical measured data for the same month of previous years from blank data points in the monthly emission records, calculating the arithmetic mean of the historical measured data, and then performing missing data repair to obtain the multi-source reconstructed data further includes: If the corresponding enterprise has data missing for more than three consecutive months, the data for the interrupted period will be removed, and the total annual emissions of the corresponding enterprise will be calculated based on the total amount of the remaining valid months. If the corresponding company's full-year data is missing, then check the corresponding company's production capacity ranking in the target region. When the corresponding enterprise belongs to the key production capacity enterprises within the preset ranking range, the measured historical data of adjacent years or the unit production capacity emission level of enterprises of the same scale is obtained and estimated. If the corresponding enterprise is not among the key production capacity enterprises in the preset ranking range, then the data of the corresponding enterprise will be removed.
4. The method for identifying and optimizing pollution sources with missing data according to claim 1, characterized in that, The two-dimensional fusion analysis of the multi-source reconstructed data yields a list of key pollution sources, including: Based on the multi-source reconstructed data, the annual average pollutant emissions of each of the three phosphorus enterprises are obtained. All the three phosphorus enterprises are sorted from high to low according to the annual average pollutant emissions, and the total emissions of the three phosphorus enterprises are accumulated according to the sorting until the predetermined emission threshold of the total emissions in the target area is reached. Based on all the accumulated three phosphorus enterprises, the set of key load enterprises is obtained. Based on the multi-source reconstructed data, the annual output value data of all the "three phosphorus" enterprises are obtained. The corresponding pollutant emission amount is generated based on the annual output value data. The emission intensity per unit output value of the corresponding "three phosphorus" enterprises is obtained based on the pollutant emission amount. The corresponding "three phosphorus" enterprises with emission intensity higher than the preset average intensity threshold are screened out to obtain a set of inefficient key enterprises. The list of key pollution sources is obtained by taking the union of the set of key enterprises with high load and the set of key inefficient enterprises.
5. The method for identifying and optimizing pollution sources with missing data according to claim 1, characterized in that, The process of constructing a coupled material flow-system dynamics model, obtaining historical monitoring data to calibrate the parameters of the coupled material flow-system dynamics model, and obtaining the calibrated coupled material flow-system dynamics model includes: Construct a phosphate mining subsystem, a phosphate chemical production subsystem, a phosphogypsum storage subsystem, a wastewater treatment subsystem, a product consumption subsystem, and a water environment subsystem, and establish the material transfer topology between each subsystem; The mass transfer topology is mapped to the mass flow-system dynamics coupled model, and key rate equations are set. Based on the historical monitoring data, the material flow-system dynamic coupling model is reverse-tuned to adjust the conversion coefficient and loss parameters within the model. When the output data of the material flow-system dynamic coupling model matches the historical data of the actual monitoring data to a preset standard, the calibrated material flow-system dynamic coupling model is obtained.
6. The method for identifying and optimizing pollution sources with missing data according to claim 1, characterized in that, The acquisition of key variable parameters involves using Latin hypercube sampling combined with Monte Carlo simulation. Combined with the list of key pollution sources, random simulations are performed on the key variable parameters to obtain a set of governance scenarios, including: Based on the key variable parameters, the corresponding probability distribution characteristics are obtained; The probability distribution features are sampled using the Latin hypercube sampling method in a stratified combination to generate several sets of parameter combinations. According to the Monte Carlo simulation method, the set of future world states is generated by performing a preset number of random iterations on the several sets of parameter combinations. Based on the list of key pollution sources and the set of future world states, several different governance measures are generated. Two or more of these governance measures are selected and combined to form the set of governance scenarios.
7. The method for identifying and optimizing pollution sources with missing data according to claim 1, characterized in that, The process involves extrapolating the set of governance scenarios based on the calibrated material flow-system dynamics coupling model to obtain candidate governance strategies, including: The set of treatment scenarios is input into the calibrated material flow-system dynamics coupled model for simulation, and the probability of water quality compliance for each scheme in the set of treatment scenarios is statistically analyzed. Obtain the tail risk value of each group of schemes, and obtain candidate treatment strategies based on the water quality compliance probability and the tail risk value.
8. A system for identifying pollution sources with missing data and optimizing treatment strategies, characterized in that, include: The data repair unit is used to acquire multi-source data of "three phosphorus" enterprises in the target area, perform missing data repair processing based on time-series features on the multi-source data, and obtain multi-source reconstructed data. The data analysis unit is used to perform two-dimensional fusion analysis on the multi-source reconstructed data to obtain a list of key pollution sources. The model building unit is used to build a material flow-system dynamics coupled model, acquire historical monitoring data to calibrate the parameters of the material flow-system dynamics coupled model, and obtain the calibrated material flow-system dynamics coupled model. The data processing unit is used to obtain key variable parameters, and uses Latin hypercube sampling combined with Monte Carlo simulation, combined with the list of key pollution sources, to perform random simulation on the key variable parameters to obtain a set of treatment scenarios. The data extrapolation unit is used to extrapolate the set of governance scenarios based on the calibrated material flow-system dynamics coupling model to obtain candidate governance strategies; The decision control unit is used to dynamically control the candidate governance strategies based on a preset triggering mechanism to obtain the identification results and optimization strategies.
9. An electronic device, characterized in that, include: Processor and memory, the memory being used to store computer programs; When the computer program is loaded by the processor, it causes the processor to execute the data missing pollution source identification and control strategy optimization method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for identifying and optimizing the management strategy of data missing pollution sources as described in any one of claims 1-7.