Machine Learning-Based Pollution Emission Prediction Method and System

By obtaining template pollution emission data and label guidance knowledge data to generate sample learning data, the data dependence and feature extraction difficulties of pollution emission prediction methods in the existing technology are solved, and more accurate pollution emission prediction and detailed description are achieved, which improves the learning efficiency and prediction capabilities of the model.

CN119940632BActive Publication Date: 2025-07-25CHINESE ACAD OF ENVIRONMENTAL PLANNING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510028500.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-07-25
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The existing pollution emission prediction methods based on machine learning rely on a large amount of historical data and the data quality is uneven, making it difficult to effectively extract complex pollution emission characteristics, resulting in unstable prediction results and lack of detailed description and explanatory information, which limits its value in practical applications.

Method used

By obtaining template pollution emission data, prior pollution emission description labels and label guidance knowledge data of the target machine learning network model, combining the model enable definition information to generate sample learning data, perform model parameter learning, and generate pollution emission prediction models, it can accurately predict any input data and provide detailed description labels.

Benefits of technology

It significantly improves the accuracy and reliability of pollution emission forecasting, optimizes model learning efficiency, enhances the ability to identify and predict complex pollution emission patterns, and supports environmental pollution monitoring, governance and decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940632B_ABST
    Figure CN119940632B_ABST
Patent Text Reader

Abstract

The present invention provides a pollution emission prediction method and system based on machine learning, which significantly improves the accuracy and reliability of pollution emission prediction. The method first obtains target machine learning data containing rich prior knowledge, and combines model enabling definition information to generate example learning data for model update. By using these example data to perform parameter learning on the target machine learning network model, the obtained pollution emission prediction model can more accurately predict any input pollution emission data, and generate corresponding pollution emission description labels and label-guided knowledge data. This process not only optimizes the learning efficiency of the model, but also enhances the model's ability to identify and predict complex pollution emission patterns, thus contributing to environmental pollution monitoring, treatment and decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular, to a pollution emission prediction method and system based on machine learning. Background Art

[0002] With the rapid development of industrialization and urbanization, the problem of environmental pollution has become increasingly severe. Among them, the accurate prediction of pollution emissions is crucial for environmental protection, policy-making, and public health. Traditional pollution emission prediction methods often rely on statistical models or physical models, and these methods have limitations in dealing with complex and variable pollution emission data, making it difficult to accurately reflect the actual situation and future trends of pollution emissions.

[0003] In recent years, machine learning technology has gradually shown great potential in the field of pollution emission prediction due to its powerful data processing and pattern recognition capabilities. However, existing pollution emission prediction methods based on machine learning still face some challenges. On the one hand, these methods usually rely on a large amount of historical pollution emission data for model training, but the acquisition of historical data is often restricted by many factors, and the data quality is uneven, which affects the prediction accuracy of the model. On the other hand, traditional machine learning methods often have difficulty in effectively extracting key information when dealing with pollution emission data with complex features and high non-linearity, resulting in instability of prediction results.

[0004] In addition, when existing machine learning models predict pollution emissions, they usually only output a single pollution emission value or a simple classification label, lacking detailed descriptions and explanatory information about pollution emission characteristics, which limits the value of prediction results in practical applications. Therefore, how to construct a machine learning prediction method that can make full use of limited data resources, effectively extract pollution emission characteristics, and provide rich prediction information has become an urgent problem to be solved in the current field of pollution emission prediction. Summary of the Invention

[0005] In view of the above-mentioned problems, in combination with the first aspect of the present invention, embodiments of the present invention provide a pollution emission prediction method based on machine learning, and the method includes:

[0006] Obtain target machine learning data of a target machine learning network model, where the target machine learning data includes: template pollution emission data, prior pollution emission description labels corresponding to the template pollution emission data, and corresponding label-guided knowledge data, and the prior pollution emission description labels of the template pollution emission data are selected from a plurality of pre-defined pollution emission description labels;

[0007] Obtain the model enabling definition information of the target machine learning network model; the model enabling definition information characterizes that: based on the multiple pollution emission description tags, pollution emission prediction is performed on the input pollution emission data, and the pollution emission description tags and corresponding tag guiding knowledge data of the input pollution emission data are generated;

[0008] Generate example learning data for updating the target machine learning network model according to the model enabling definition information and the target machine learning data, and configure the template pollution emission data in the target machine learning data as the input pollution emission data in the example learning data;

[0009] Perform model parameter learning on the target machine learning network model according to the example learning data, and use the target machine learning network model that has completed model parameter learning as the pollution emission prediction model; the pollution emission prediction model is used to perform pollution emission prediction on any input pollution emission data, and generate corresponding pollution emission description tags and tag guiding knowledge data.

[0010] In a possible implementation manner of the first aspect, the example learning data is a pollution emission data generated by fusing the model enabling definition information and the target machine learning data; the performing model parameter learning on the target machine learning network model according to the example learning data includes:

[0011] Perform feature encoding on the example learning data to generate X encoded feature node data, where X is a positive integer;

[0012] Based on the time series information of the X encoded feature node data, generate the feature to be learned according to the first X - 1 encoded feature node data among the X encoded feature node data; and generate the expected output feature corresponding to the feature to be learned according to the X - 1 encoded feature node data except the first encoded feature node data among the X encoded feature node data;

[0013] Derive the feature to be learned through the target machine learning network model for each encoded feature node data to generate a derivation result; the derivation result includes the derived X - 1 encoded feature node data, and the a-th encoded feature node data in the derivation result is derived based on the first a encoded feature node data in the feature to be learned, where a is a positive integer between 1 and X - 1;

[0014] Based on the expected output features and the derivation result, calculate the error parameters corresponding to each encoded feature node data in the derivation result. The error parameter corresponding to the a-th encoded feature node data in the derivation result represents: the loss result between the a-th encoded feature node data in the derivation result and the a-th encoded feature node data in the expected output features;

[0015] Fuse the error parameters corresponding to each encoded feature node data in the derivation result to generate the training error parameter of the target machine learning network model;

[0016] According to the training objective of minimizing the training error parameter, update the neuron weight information of the target machine learning network model, so as to learn the model parameters of the target machine learning network model.

[0017] In a possible implementation manner of the first aspect, the generating the example learning data for updating the target machine learning network model according to the model enabling definition information and the target machine learning data includes:

[0018] Fuse the model enabling definition information and the template pollution emission data in the target machine learning data and output them as a pollution emission data to generate the first pollution emission data;

[0019] Fuse the prior pollution emission description label in the target machine learning data and the label guiding knowledge data in the target machine learning data and output them as a pollution emission data to generate the second pollution emission data;

[0020] Generate the example learning data for updating the target machine learning network model according to the first pollution emission data and the second pollution emission data.

[0021] In a possible implementation manner of the first aspect, the learning the model parameters of the target machine learning network model according to the example learning data includes:

[0022] Through the target machine learning network model, based on the first pollution emission data in the example learning data, perform pollution emission prediction on the template pollution emission data in the target machine learning data to generate a pollution emission prediction result; the pollution emission prediction result includes: the pollution emission description label of the generated corresponding template pollution emission data, and the corresponding label guiding knowledge data;

[0023] Based on the loss result between the pollution emission prediction result and the second pollution emission data in the example learning data, update the neuron weight information of the target machine learning network model, so as to learn the model parameters of the target machine learning network model.

[0024] In a possible implementation of the first aspect, the target machine learning data is a machine learning data obtained from a training data sequence, and the generation steps of the training data sequence include:

[0025] Obtain a plurality of machine learning data of the target machine learning network model. A machine learning data includes a template pollution emission data and training supervision data for the corresponding template pollution emission data. The training supervision data for any template pollution emission data includes a prior pollution emission description label for the corresponding template pollution emission data and corresponding label-guided knowledge data. The prior pollution emission description label in any machine learning data is selected from the plurality of pollution emission description labels;

[0026] Based on the plurality of machine learning data, determine the number of templates for each pollution emission description label in the plurality of pollution emission description labels. The number of templates for any pollution emission description label is used to represent the number of machine learning data in the plurality of machine learning data that contain the any pollution emission description label;

[0027] Based on the number of templates for each pollution emission description label, determine marginal pollution emission description labels from the plurality of pollution emission description labels. The marginal pollution emission description labels are used to represent pollution emission description labels corresponding to a number of templates less than a set number;

[0028] Extract one or more template pollution emission data from the template pollution emission data included in the machine learning data containing the marginal pollution emission description labels;

[0029] Perform feature derivation on each selected template pollution emission data respectively to generate Y iterative pollution emission data, where Y is a positive integer. An iterative pollution emission data corresponds to a template pollution emission data, and there is a consistent pollution feature path between any iterative pollution emission data and the corresponding template pollution emission data;

[0030] Use the training supervision data of the template pollution emission data corresponding to each iterative pollution emission data as the training supervision data for the corresponding iterative pollution emission data;

[0031] Regard each iterative pollution emission data as an iterative template pollution emission data, and generate Y iterative machine learning data based on the Y iterative pollution emission data and the corresponding training supervision data;

[0032] Generate the training data sequence based on the plurality of machine learning data and the Y iterative machine learning data.

[0033] In a possible implementation of the first aspect, the step of respectively performing feature derivation on the selected template pollution emission data to generate Y iterative pollution emission data includes:

[0034] Obtain a feature derivation request of a feature derivation network, where the feature derivation request represents: based on the input template pollution emission data and the prior pollution emission description label, generate iterative pollution emission data with consistent pollution feature paths for the template pollution emission data;

[0035] Obtain one or more derived example combination data. Any one of the derived example combination data includes an example pollution emission data, the prior pollution emission description label of the corresponding example pollution emission data, and pollution emission data with a consistent pollution feature path as the corresponding example pollution emission data;

[0036] Through the feature derivation network, based on the feature derivation request and the one or more derived example combination data, respectively generate iterative pollution emission data with consistent pollution feature paths for the selected template pollution emission data, generating Y iterative pollution emission data.

[0037] In a possible implementation of the first aspect, the step of through the feature derivation network, based on the feature derivation request and the one or more derived example combination data, respectively generate iterative pollution emission data with consistent pollution feature paths for the selected template pollution emission data, generating Y iterative pollution emission data includes:

[0038] Poll the selected template pollution emission data one by one, and use the currently polled template pollution emission data as the current template pollution emission data;

[0039] Generate an input pollution emission data segment according to the current template pollution emission data and the corresponding prior pollution emission description label;

[0040] Integrate the feature derivation request, the one or more derived example combination data, and the generated input pollution emission data segment to generate derived learning data corresponding to the current template pollution emission data;

[0041] Through the feature derivation network, by learning the derived learning data corresponding to the current template pollution emission data, generate iterative pollution emission data with a consistent pollution feature path as the current template pollution emission data;

[0042] Continue to poll until all the selected template pollution emission data have been polled, generating Y iterative pollution emission data.

[0043] In a possible implementation of the first aspect, the feature derivation network generates iterative pollution emission data with consistent pollution feature paths for each selected template pollution emission data based on the feature derivation request and the one or more derived example combination data, generating Y iterative pollution emission data, including:

[0044] Fuse the feature derivation request and the one or more derived example combination data to generate a training sample for the feature derivation network;

[0045] Respectively use each selected template pollution emission data and the corresponding prior pollution emission description label to generate an input pollution emission data segment corresponding to each template pollution emission data;

[0046] Train the feature derivation network based on the training sample to generate a trained feature derivation network;

[0047] Respectively pass the trained feature derivation network based on the input pollution emission data segments corresponding to each template pollution emission data to generate iterative pollution emission data with consistent pollution feature paths for the corresponding template pollution emission data, generating Y iterative pollution emission data.

[0048] In a possible implementation of the first aspect, the method further includes:

[0049] After receiving a prediction instruction for the target pollution emission data, generate candidate pollution emission data based on the model enabling definition information and the target pollution emission data, and configure the target pollution emission data as input pollution emission data in the candidate pollution emission data;

[0050] Perform pollution emission prediction on the target pollution emission data based on the candidate pollution emission data through the pollution emission prediction model to generate a pollution emission description label and corresponding label guidance knowledge data for the target pollution emission data.

[0051] In yet another aspect, an embodiment of the present invention further provides a pollution emission prediction system based on machine learning, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor, and the machine-readable storage medium is used to store programs, instructions, or codes. The processor is used to execute the programs, instructions, or codes in the machine-readable storage medium to implement the above method.

[0052] Based on the above aspects, the embodiments of the present application significantly improve the accuracy and reliability of pollution emission prediction. The method first obtains target machine learning data containing rich prior knowledge, and combines the model enabling definition information to generate example learning data for model update. By using these example data to perform parameter learning on the target machine learning network model, the obtained pollution emission prediction model can more accurately predict any input pollution emission data, and generate corresponding pollution emission description tags and tag-guided knowledge data. This process not only optimizes the learning efficiency of the model, but also enhances the model's recognition and prediction capabilities for complex pollution emission patterns, thus contributing to environmental pollution monitoring, treatment, and decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a schematic flowchart of the execution process of the pollution emission prediction method based on machine learning provided by an embodiment of the present invention.

[0054] Figure 2 is a schematic diagram of the hardware architecture of the pollution emission prediction system based on machine learning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The present invention will be specifically described below with reference to the accompanying drawings of the specification. Figure 1 is a schematic flowchart of the pollution emission prediction method based on machine learning provided by an embodiment of the present invention. The pollution emission prediction method based on machine learning will be introduced in detail below.

[0056] Step S110: Obtain target machine learning data of the target machine learning network model. The target machine learning data includes: template pollution emission data, prior pollution emission description tags corresponding to the template pollution emission data, and corresponding tag-guided knowledge data. The prior pollution emission description tags of the template pollution emission data are selected from a plurality of pre-defined pollution emission description tags.

[0057] In this embodiment, in the environmental monitoring scenario of an industrial park, the target machine learning network model aims to accurately analyze and predict the air pollution emission situation of enterprises in the park. The target machine learning data is the basic sample for model learning.

[0058] The template pollution emission data can be data collected from the waste gas emission monitoring devices of each enterprise in the industrial park. For example, in the waste gas emitted by the chimney of a chemical enterprise during a specific period, the concentration of sulfur dioxide (SO2) is 50 milligrams per cubic meter, the concentration of nitrogen oxides (NO x ) is 80 milligrams per cubic meter, the concentration of particulate matter (PM) is 30 milligrams per cubic meter, etc. These data form a template pollution emission data sample.

[0059] The prior pollution emission description label is selected from multiple complex pre - defined pollution emission description labels. These pollution emission description labels are not just simple classifications of pollutant concentrations, but also involve various factors such as the nature of the emission source, the emission pattern, and the potential impact on the environment. For example, for the waste gas emission of the above - mentioned chemical enterprise, the prior pollution emission description label may be "High - concentration composite waste gas emission from chemical enterprises - Relatively large potential environmental hazard - Short - term fluctuation type". This label indicates that the emission source is a chemical enterprise, the concentrations of multiple pollutants in the waste gas are high, and due to the types and concentrations of the emitted pollutants, there is a relatively large potential hazard to the environment, and the emission situation fluctuates in the short term, which may be related to the intermittent production process of the enterprise.

[0060] The corresponding label - guiding knowledge data is auxiliary information related to this prior pollution emission description label. For example, it may include the relationship between the production process of chemical enterprises and pollutant generation, historical change data of the surrounding environment under similar emission situations (such as the damage situation of nearby vegetation, the long - term change trend of air quality, etc.), the existing measures taken by the enterprise to control pollution and their effect evaluation. These label - guiding knowledge data help the model to more deeply understand the meaning behind the prior pollution emission description label, so as to better learn and predict.

[0061] By collecting numerous such template pollution emission data, corresponding prior pollution emission description labels, and label - guiding knowledge data, the target machine - learning data of the target machine - learning network model is formed. Each set of data inside is a detailed record and interpretation of different pollution emission situations in the industrial park, providing rich materials for the subsequent training and optimization of the model.

[0062] Step S120, obtain the model enabling definition information of the target machine - learning network model. The model enabling definition information represents: based on the multiple pollution emission description labels, perform pollution emission prediction on the input pollution emission data, and generate the pollution emission description label of the input pollution emission data and the corresponding label - guiding knowledge data.

[0063] For example, in the environmental management system of the industrial park, the model enabling definition information of the target machine - learning network model stipulates how the target machine - learning network model processes the input pollution emission data based on the complex multiple pollution emission description labels.

[0064] For example, in the industrial park, there are various types of enterprises, including chemical, mechanical manufacturing, electronics, etc. The pollution emission characteristics of each enterprise are different, and the forms and meanings of their pollution emission data are also very complex. The model enabling definition information is used to guide the target machine - learning network model on how to perform pollution emission prediction when facing the pollution emission data of different types of enterprises.

[0065] For the waste gas emission data collected from various enterprises (which is a type of input pollution emission data), the model enables the defined information to require the model to conduct a comprehensive analysis based on multiple predefined pollution emission description tags. These tags may cover information in multiple dimensions, ranging from the types of pollutants, concentrations, emission rates, to the temporal patterns of emissions (such as whether it is intermittent emission, whether it is related to production shifts, etc.), and then to the potential impacts of emissions on different areas of the surrounding environment (such as near residential areas, farmlands, water bodies, etc.).

[0066] Taking a chemical enterprise as an example, after the model enables the defined information and receives the waste gas emission data of the enterprise, it needs to consider various complex pollutants generated during the chemical production process (such as volatile organic compounds, heavy metals, etc.), as well as the diffusion laws and environmental impacts of these pollutants under different meteorological conditions (such as wind direction, wind speed, temperature, humidity, etc.). At the same time, for the waste gas emissions of a machinery manufacturing enterprise, the model should pay attention to the emissions of pollutants such as metal dust and lubricant volatiles, and combine factors such as the production scale and equipment operation status of the enterprise to accurately predict pollution emissions based on the pollution emission description tags.

[0067] While conducting pollution emission prediction, the model also generates pollution emission description tags for the input pollution emission data and corresponding tag-guided knowledge data. For example, for the waste gas emission data of a machinery manufacturing enterprise, after analysis, the pollution emission description tags generated by the model may be "Medium-concentration dust and volatile organic compound emissions from a machinery manufacturing enterprise - Local environmental impact - Stable type", and the corresponding tag-guided knowledge data may include the wind direction frequency map of the enterprise's location, historical data comparison of surrounding air quality monitoring stations, maintenance records and effectiveness analysis of environmental protection equipment within the enterprise, etc. These generated tags and knowledge data can provide more detailed and targeted decision-making basis for the environmental management department of the park, helping them better supervise the pollution emission behaviors of enterprises and formulate more reasonable environmental protection policies.

[0068] Step S130: Generate example learning data for updating the target machine learning network model based on the model-enabled defined information and the target machine learning data, and configure the template pollution emission data in the target machine learning data as the input pollution emission data in the example learning data.

[0069] Assume that the target machine learning data of the target machine learning network model has been obtained, where the template pollution emission data comes from the actual monitoring of different enterprises in the park. For example, there is the waste gas emission data of an electronics enterprise, which contains information such as the fluoride concentration of 10 milligrams per cubic meter and the ammonia concentration of 5 milligrams per cubic meter. This is the template pollution emission data part in the target machine learning data. Its corresponding prior pollution emission description label is "Emission of low-concentration special gases from electronics enterprises - Impact on specific areas - Stable type", and the label-guided knowledge data includes the types of chemical agents used in the production process of this enterprise, the distribution of surrounding environmentally sensitive areas (such as the nearby high-precision electronics equipment production workshops being sensitive to these gases), etc.

[0070] The model enabling definition information requires specific processing of the input pollution emission data to generate the information required for pollution emission prediction. According to these requirements, the model enabling definition information is fused with the template pollution emission data in the target machine learning data. For example, the model enabling definition information stipulates the weight allocation method for different types of pollutants (such as the weight of fluoride in a specific environmental impact assessment is 0.3, and that of ammonia is 0.2, etc.). In this way, the waste gas emission data of the electronics enterprise is fused with the model enabling definition information to generate the first pollution emission data. This data integrates the model's requirements for data processing and the actual pollution emission data, and is a form of data after standardized processing.

[0071] Then, the prior pollution emission description label and the label-guided knowledge data in the target machine learning data are fused and output as a pollution emission data to generate the second pollution emission data. For the above example of the electronics enterprise, the "Emission of low-concentration special gases from electronics enterprises - Impact on specific areas - Stable type" in the prior pollution emission description label is integrated with the information such as the types of enterprise production agents and the distribution of surrounding sensitive areas in the label-guided knowledge data to form a second pollution emission data containing more semantic information.

[0072] Finally, based on the first pollution emission data and the second pollution emission data, example learning data for updating the target machine learning network model is generated. The example learning data is a special data structure that integrates the model requirements, actual pollution emission data, prior labels, and relevant knowledge data. It can be understood as a learning sample customized for the target machine learning network model, which not only contains the original pollution emission data characteristics but also incorporates the processing rules required by the model and the prior understanding of the data, enabling the model to better adapt to the complex pollution emission situations in the industrial park during subsequent learning processes and improving the accuracy of pollution emission prediction for different enterprises.

[0073] Step S140: Perform model parameter learning on the target machine learning network model according to the sample learning data, and use the target machine learning network model that has completed model parameter learning as the pollution emission prediction model. The pollution emission prediction model is used to predict pollution emissions for any input pollution emission data, and generate corresponding pollution emission description labels and label-guided knowledge data.

[0074] Taking the previously generated sample learning data as an example, it contains various fused information. First, perform feature encoding on the sample learning data. Suppose the sample learning data is about the pollution emission situation of a certain chemical enterprise. After feature encoding, X encoded feature node data are generated. For example, information such as the concentrations of different pollutants, emission rates, and emission times in the waste gas of the chemical enterprise is encoded into a series of feature node data. If X = 5, then these 5 encoded feature node data may respectively represent the sulfur dioxide concentration feature, nitrogen oxide emission rate feature, emission time period feature, the influence of meteorological conditions (such as wind direction) on emissions feature, and the relationship feature between the enterprise's production load and emissions.

[0075] Based on the time series information of these X encoded feature node data, generate the feature to be learned according to the first X - 1 encoded feature node data. For example, for the above 5 encoded feature node data, construct the feature to be learned based on the first 4 (i.e., the sulfur dioxide concentration feature, nitrogen oxide emission rate feature, emission time period feature, the influence of meteorological conditions on emissions feature). At the same time, generate the expected output feature corresponding to the feature to be learned according to the X - 1 encoded feature node data except the first encoded feature node data. That is, construct the expected output feature according to the latter 4 feature node data (nitrogen oxide emission rate feature, emission time period feature, the influence of meteorological conditions on emissions feature, the relationship feature between the enterprise's production load and emissions).

[0076] Then, use the target machine learning network model to deduce each encoded feature node data of the feature to be learned. For the first encoded feature node data (sulfur dioxide concentration feature), the model makes a deduction according to its own network structure and initial parameters to obtain a preliminary result. Then, based on this preliminary result and the next encoded feature node data (nitrogen oxide emission rate feature), make another deduction to obtain the deduction result of the second encoded feature node data. And so on, finally generating a deduction result, which includes the deduced X - 1 encoded feature node data, and the a-th encoded feature node data is deduced based on the first a encoded feature node data in the feature to be learned (where a is a positive integer between 1 and X - 1).

[0077] Based on the expected output characteristics and the derivation results, the error parameters corresponding to each coded feature node data in the derivation results are calculated. For example, for the second coded feature node data (nitrogen oxide emission rate feature) in the derivation result, it is compared with the corresponding data in the expected output characteristics, and the loss result between the two is calculated. This loss result is the error parameter corresponding to the coded feature node data.

[0078] The error parameters corresponding to each encoded feature node data in the derivation result are fused to generate the training error parameters of the target machine learning network model. This training error parameter comprehensively reflects the overall difference between the model derivation result and the expected output feature.

[0079] According to the training goal of minimizing the training error parameter, the neuron weight information of the target machine learning network model is updated. For example, if a neuron is found to have a large error in the derivation of the nitrogen oxide emission rate characteristics, the weight associated with this neuron is adjusted so that the model can deduce more accurately the next time it processes similar data. By continuously processing the sample learning data in this way, the parameters of the model are gradually adjusted, and finally the target machine learning network model that has completed the model parameter learning is used as the pollution emission prediction model. This pollution emission prediction model can predict pollution emissions for any pollution emission data input in the industrial park, and generate corresponding pollution emission description labels and label-guided knowledge data. For example, for the newly input waste gas emission data of a machinery manufacturing enterprise, the model can accurately predict that its pollution emission description label is "Medium-concentration dust and volatile organic compound emissions of machinery manufacturing enterprises-local environmental impact-stable type", and generate label-guided knowledge data containing the trend of air quality changes around the enterprise and the potential impact assessment on the health of nearby residents, providing strong decision-making support for the environmental management of the industrial park.

[0080] Among them, the pollution emission prediction model plays an important role in the daily environmental management of industrial parks. For example, when the environmental monitoring platform receives a prediction instruction for the target pollution emission data of a newly settled enterprise (assuming it is a small metal processing enterprise), it will generate candidate pollution emission data based on the previous model activation definition information and the target pollution emission data. The target pollution emission data contains information such as the metal dust concentration of 20 mg per cubic meter and the volatile organic compound concentration of 15 mg per cubic meter in the exhaust gas emitted from the chimney of the enterprise. These data are configured in the candidate pollution emission data as input pollution emission data.

[0081] Then, the trained pollution emission prediction model is used to predict the target pollution emission data based on the candidate pollution emission data. After complex calculations and analysis, the model generates a pollution emission description label for the target pollution emission data, which may be "medium-concentration composite pollutant emissions from small metal processing enterprises-local environmental risks-fluctuation type", as well as corresponding label-guided knowledge data, such as analysis of the impact of wind direction and wind speed changes in the area where the enterprise is located on the diffusion of pollutants, cross-contamination risk assessments that other surrounding enterprises may be subject to, and recommendations on environmental protection measures that the enterprise should take.

[0082] These prediction results can help the management departments of industrial parks to timely understand the pollution emissions of newly settled enterprises, formulate corresponding environmental management strategies in advance, ensure the overall environmental quality of industrial parks, and protect the health and safety of surrounding residents and the ecological environment. At the same time, as more enterprises' pollution emission data are continuously input into the model for prediction, the model can also continuously optimize and update itself, improving the accuracy and reliability of pollution emission predictions for various enterprises in the industrial park.

[0083] Based on the above steps, the embodiment of the present application significantly improves the accuracy and reliability of pollution emission prediction. The method first obtains target machine learning data containing rich prior knowledge, and combines the model activation definition information to generate sample learning data for model updating. By using these sample data to learn the parameters of the target machine learning network model, the resulting pollution emission prediction model can more accurately predict any input pollution emission data, and generate corresponding pollution emission description labels and label-guided knowledge data. This process not only optimizes the learning efficiency of the model, but also enhances the model's ability to recognize and predict complex pollution emission patterns, thereby facilitating environmental pollution monitoring, governance and decision-making.

[0084] In a possible implementation manner, the sample learning data is a pollution emission data generated by fusing the model enabling definition information and the target machine learning data. Step S140 includes:

[0085] Step S141, feature encoding is performed on the sample learning data to generate X encoded feature node data, where X is a positive integer.

[0086] In this embodiment, for example, there is a large chemical enterprise in an industrial park, and its pollution emission data contains various complex information. The sample learning data may integrate various aspects of information such as the concentration, emission rate, emission time, and impact of meteorological conditions of various pollutants in the enterprise's waste gas emissions. When performing feature encoding, if X is 5, the data of these 5 encoded feature nodes may correspond to different feature aspects respectively. The data of the first encoded feature node may be the concentration feature of the main pollutant sulfur dioxide (SO2), which accurately records the SO2 concentration value emitted by the enterprise during a specific period; the data of the second encoded feature node can be the emission rate feature of nitrogen oxides (NO x ) that details the emission amount of NO x per unit time; the data of the third encoded feature node is the emission time period feature, which clarifies whether it is the peak production period during the day or the low-load period at night; the data of the fourth encoded feature node is the impact feature of the wind direction in meteorological conditions on emission diffusion, considering that different wind directions will cause great differences in the diffusion path and range of pollutants in the park; the data of the fifth encoded feature node is the relationship feature between the enterprise's production load and emissions, reflecting the correlation between the enterprise's production task volume and the pollutant emissions.

[0087] Step S142: Based on the time series information of the X encoded feature node data, generate a feature to be learned according to the first X - 1 encoded feature node data among the X encoded feature node data. And generate the expected output feature corresponding to the feature to be learned according to the X - 1 encoded feature node data except the first encoded feature node data among the X encoded feature node data.

[0088] For the example of the above chemical enterprise, a feature to be learned is constructed according to the first 4 encoded feature node data (SO2 concentration feature, NOx emission rate feature, emission time period feature, impact feature of wind direction on emission diffusion). This feature to be learned comprehensively covers the main pollution emission-related factors except the relationship between production load and emissions, forming a feature combination with a specific time series logic. At the same time, generate the expected output feature corresponding to the feature to be learned according to the X - 1 encoded feature node data except the first encoded feature node data. That is, construct the expected output feature according to the NO x emission rate feature, emission time period feature, impact feature of wind direction on emission diffusion, and relationship feature between the enterprise's production load and emissions. This expected output feature is the result that the model should output under ideal conditions, and it has a close logical connection with the feature to be learned and is determined based on an in-depth understanding of the pollution emission laws and influencing factors of chemical enterprises.

[0089] Step S143: Derive the encoded feature node data one by one for the feature to be learned through the target machine learning network model, and generate a derivation result. The derivation result includes X - 1 derived encoded feature node data, and the a-th encoded feature node data in the derivation result is derived based on the first a encoded feature node data in the feature to be learned, where a is a positive integer between 1 and X - 1.

[0090] In this process, the target machine learning network model starts from the first encoded feature node data of the feature to be learned. For the first encoded feature node data (SO2 concentration feature), the target machine learning network model processes it according to its existing network structure and initial parameters to obtain a preliminary derivation result. Then, combined with the second encoded feature node data (NO x emission rate feature), the model makes another derivation to obtain the derivation result of the second encoded feature node data based on the first two encoded feature node data. In this way, it successively combines the subsequent encoded feature node data for derivation, and finally generates a derivation result. This derivation result contains X - 1 derived encoded feature node data, and the a-th encoded feature node data in the derivation result is derived based on the first a encoded feature node data in the feature to be learned. For example, when a is 3, the third encoded feature node data in the derivation result is obtained by integrating the first 3 encoded feature node data (SO2 concentration feature, NO x emission rate feature, emission time period feature) in the feature to be learned, reflecting the model derivation result under the combined action of these three factors.

[0091] Step S144: Based on the expected output feature and the derivation result, calculate the error parameters corresponding to each encoded feature node data in the derivation result. The error parameter corresponding to the a-th encoded feature node data in the derivation result represents the loss result between the a-th encoded feature node data in the derivation result and the a-th encoded feature node data in the expected output feature.

[0092] Continuing with the example of a chemical enterprise, for the second encoded feature node data (NO x emission rate feature) in the derivation result, compare it with the corresponding data in the expected output feature. Suppose the NO x emission rate in the derivation result is 100 kilograms per hour, while the NO xIf the emission rate is 120 kilograms per hour, then the difference between the two (20 kilograms per hour) is the error parameter corresponding to the encoded feature node data. This error parameter characterizes the loss result between the encoded feature node data in the derivation result and the corresponding data in the expected output feature, and it intuitively reflects the derivation accuracy of the model at this feature node.

[0093] Step S145: Fuse the error parameters corresponding to the encoded feature node data in the derivation result to generate the training error parameter of the target machine learning network model.

[0094] For the error parameters corresponding to the 5 encoded feature node data of the above chemical enterprise, they are combined through a specific fusion method (such as weighted summation or root mean square calculation, etc.). This fusion method comprehensively considers the error situations of the encoded feature node data, and forms a training error parameter that comprehensively reflects the overall difference between the model derivation result and the expected output feature.

[0095] Step S146: Update the neuron weight information of the target machine learning network model according to the training objective of minimizing the training error parameter, so as to learn the model parameters of the target machine learning network model.

[0096] For example, in the example of the chemical enterprise, if it is found that a certain neuron generates a large error in the derivation of the NO x emission rate feature, this indicates that the weight of this neuron may need to be adjusted. By analyzing the training error parameter, the model can determine which neurons contribute more to the error, and then adjust the weights of these neurons. For example, if the associated weight of a certain neuron with the NO x emission rate feature is too large, resulting in the derivation result of the model on this feature deviating from the expected output feature by a large amount, then the weight of this neuron is appropriately reduced. By continuously adjusting the neuron weight information in this way, the target machine learning network model can gradually improve its processing ability for the sample learning data, and then improve the accuracy of predicting the pollution emissions of enterprises in the industrial park. In the environmental management system of the entire industrial park, this model parameter learning process based on sample learning data helps to build an accurate and reliable pollution emission prediction model, providing strong technical support for the environmental supervision of the park, the pollution control of enterprises, and the overall environmental protection.

[0097] In a possible implementation manner, step S130 includes:

[0098] Step S131: Fuse the model enabling definition information and the template pollution emission data in the target machine learning data and output them as a pollution emission data to generate the first pollution emission data.

[0099] In step S132, the prior pollution emission description label in the target machine learning data and the label-guided knowledge data in the target machine learning data are fused and output as a pollution emission data to generate a second pollution emission data.

[0100] In step S133, based on the first pollution emission data and the second pollution emission data, example learning data for updating the target machine learning network model is generated.

[0101] In a possible implementation manner, step S140 may further include:

[0102] The target machine learning network model performs pollution emission prediction on the template pollution emission data in the target machine learning data based on the first pollution emission data in the example learning data to generate a pollution emission prediction result. The pollution emission prediction result includes: the pollution emission description label of the generated corresponding template pollution emission data, and the corresponding label-guided knowledge data.

[0103] Based on the loss result between the pollution emission prediction result and the second pollution emission data in the example learning data, the neuron weight information of the target machine learning network model is updated, so as to perform model parameter learning on the target machine learning network model.

[0104] In this embodiment, taking a large steel enterprise in an industrial park as an example, the model enabling definition information includes rules and requirements for processing various pollution emission data, such as the weight distribution principle for different pollutants, the importance consideration for different emission time periods, and the calculation method for the impact of meteorological conditions on pollution diffusion, etc. The template pollution emission data of this steel enterprise contains rich information. For example, within a specific time period, the concentration of particulate matter (PM) emitted from its chimney is 80 milligrams per cubic meter, the concentration of sulfur dioxide (SO2) is 150 milligrams per cubic meter, and the concentration of nitrogen oxides (NO x ) is 200 milligrams per cubic meter. At the same time, there are also related data such as the emission flow rate and temperature. During the fusion process, according to the weight distribution principle in the model enabling definition information, for example, the weight of SO2 concentration in the overall pollution assessment is 0.3, the weight of NO x concentration is 0.4, the weight of PM concentration is 0.2, and the weight of other factors is 0.1. These weights are calculated and fused with the actual pollution emission data of the steel enterprise. This fusion is not a simple numerical addition, but is based on pre-set complex calculation rules, fully considering the mutual relationship between various factors and the comprehensive effect on the environmental impact. Finally, the first pollution emission data is generated. This first pollution emission data is a standardized data form that combines model requirements and actual pollution emission characteristics, providing an input data that conforms to the model logical structure for subsequent model learning.

[0105] Next, the prior pollution emission description label in the target machine learning data and the label-guided knowledge data in the target machine learning data are fused and output as a pollution emission data to generate the second pollution emission data. For the above-mentioned steel enterprise, its prior pollution emission description label may be "high-concentration complex pollution emission of large steel enterprises - extensive environmental impact - stable type". This label not only indicates the scale of the enterprise and the type of pollutant emission concentration, but also implies the scope of environmental impact and the stability of emissions. The corresponding label-guided knowledge data contains a lot of information, such as the relationship between the production process of the steel enterprise and pollution generation, that is, which links in the ironmaking and steelmaking processes will generate a large amount of SO2, NO x and PM; the status of the surrounding environment of the enterprise, such as what sensitive areas (such as residential areas, farmlands, water sources, etc.) are around and the relative position relationship between these areas and the enterprise; and the pollution control measures already taken by the enterprise and their effect evaluation, for example, the operating efficiency of the installed desulfurization, denitrification, and dust removal equipment and the actual contribution to pollutant emission reduction. When fusing the prior pollution emission description label and the label-guided knowledge data, the semantic information in the label should be organically integrated with the specific content in the knowledge data. For example, combining "extensive environmental impact" with the distribution of surrounding sensitive areas, and linking "stable type" emissions with the stable operation of the enterprise's existing pollution control equipment. In this way, a second pollution emission data containing rich semantics and actual situations is constructed.

[0106] Then, based on the first pollution emission data and the second pollution emission data, example learning data for updating the target machine learning network model is generated. The example learning data is a special data structure that combines the first pollution emission data and the second pollution emission data. It contains both the actual data characteristics of pollution emissions (the first pollution emission data) processed according to the model requirements and the prior pollution emission description label and related knowledge data (the second pollution emission data). This structure enables the example learning data to provide comprehensive learning information for the target machine learning network model, allowing the model to understand both the actual numerical situation of pollution emissions and the meaning behind these numerical values and related environmental impact factors, thus laying a foundation for the accurate learning and parameter update of the model.

[0107] In terms of learning model parameters for the target machine learning network model based on example learning data, the target machine learning network model performs pollution emission prediction on the template pollution emission data in the target machine learning data based on the first pollution emission data in the example learning data, and generates a pollution emission prediction result. Still taking the steel enterprise as an example, the target machine learning network model takes the first pollution emission data as input, and this first pollution emission data contains the fused pollution emission characteristics of the steel enterprise and relevant information required by the model. The target machine learning network model analyzes and predicts the template pollution emission data of the steel enterprise according to its network structure and existing parameters. For example, when the model processes data such as pollutant concentration, flow rate, and temperature of the steel enterprise, it combines the pollution emission laws and relevant knowledge it has learned to generate a pollution emission prediction result. This pollution emission prediction result includes the pollution emission description label of the corresponding generated template pollution emission data, as well as the corresponding label-guided knowledge data. Suppose the generated pollution emission description label is "High-concentration complex pollution emission of large steel enterprises - Severe environmental impact - Stable type". Here, "Severe environmental impact" is a more accurate environmental impact assessment obtained by the model based on the analysis of pollutant concentration and sensitive areas of the surrounding environment; the corresponding label-guided knowledge data may include more detailed predictions of changes in the surrounding environment, such as the deterioration trend of the surrounding air quality in the future period under the current pollution emission level, the potential increase in the pollution risk of the nearby water source, and the increase in the health risk of the surrounding residents.

[0108] Finally, based on the loss result between the pollution emission prediction result and the second pollution emission data in the sample learning data, update the neuron weight information of the target machine learning network model, so as to learn the model parameters of the target machine learning network model. Compare the pollution emission prediction result of the above steel enterprise with the second pollution emission data, and calculate the loss result between the two. For example, in terms of pollution emission description labels, there are semantic differences between "severe environmental impact" and "extensive environmental impact" in the second pollution emission data, and this difference reflects the deviation between the model prediction and the prior knowledge; in terms of label-guided knowledge data, there may also be deviations of different degrees between the predicted deterioration trend of the surrounding air quality, the potential increase in the pollution risk of the water source quality, etc. and the actual situation in the second pollution emission data (such as the results obtained based on historical data and actual monitoring). Quantify and calculate these deviations through a specific loss function to obtain an overall loss result. According to this loss result, analyze which neurons in the model contribute more to the deviation during the prediction process. For example, if it is found that the neurons related to pollutant concentration analysis play a key role in predicting "severe environmental impact" and their prediction results deviate greatly from the actual situation, then adjust the weight information of these neurons. By continuously updating the neuron weight information according to this loss result, the target machine learning network model can gradually optimize its own parameters, improve the accuracy of predicting the pollution emissions of enterprises in the industrial park, and thus better meet the complex environmental management needs of the industrial park, providing strong technical support for the environmental protection and sustainable development of the park.

[0109] In a possible implementation manner, the target machine learning data is a machine learning data obtained from a training data sequence, and the generation steps of the training data sequence include:

[0110] Step A110, obtain a plurality of machine learning data of the target machine learning network model. A machine learning data includes: a template pollution emission data and training supervision data corresponding to the template pollution emission data. The training supervision data of any template pollution emission data includes the prior pollution emission description label of the corresponding template pollution emission data and the corresponding label-guided knowledge data. The prior pollution emission description label in any machine learning data is selected from the plurality of pollution emission description labels.

[0111] In this embodiment, in the industrial park, different enterprises have different pollution emission situations, and these situations are collected and organized into machine learning data. A machine learning data contains a template pollution emission data and corresponding training supervision data. For example, for a chemical enterprise, its template pollution emission data may be the information of various pollutants emitted from the chimney during a specific period, such as the concentration of sulfur dioxide (SO2) being 50 milligrams per cubic meter, nitrogen oxides (NOx ) with a concentration of 80 milligrams per cubic meter, a particulate matter (PM) concentration of 30 milligrams per cubic meter, etc. The training supervision data for the corresponding template pollution emission data includes prior pollution emission description labels and label-guided knowledge data. The prior pollution emission description label may be "medium-concentration complex pollution emission of chemical enterprises - local environmental impact - fluctuating type", which details information such as the enterprise type, pollution emission concentration level, the scope of environmental impact, and the stability of emissions. The label-guided knowledge data may include the links related to pollution generation in the production process of chemical enterprises, such as specific chemical reactions that produce a large amount of SO2, as well as relevant information about the surrounding environment, such as the presence of farmland nearby and the impact of wind direction on the pollutant diffusion direction. The prior pollution emission description label in each machine learning data is selected from multiple pollution emission description labels, and these labels cover the classification descriptions of various possible pollution emission situations in the industrial park.

[0112] Step A120, based on the multiple machine learning data, determine the template quantity of each pollution emission description label among the multiple pollution emission description labels. The template quantity of any pollution emission description label is used to represent: the quantity of machine learning data containing the any pollution emission description label among the multiple machine learning data.

[0113] After collecting the machine learning data of numerous enterprises in the industrial park, analyze and statistically process these machine learning data. For example, there are various pollution emission description labels such as "high-concentration complex pollution emission of chemical enterprises - extensive environmental impact - stable type" and "medium-concentration dust and volatile organic compound emission of machinery manufacturing enterprises - local environmental impact - stable type". Statistical results show that there are 50 sets of machine learning data marked with "high-concentration complex pollution emission of chemical enterprises - extensive environmental impact - stable type", which is the template quantity of this pollution emission description label; while there are only 10 sets of machine learning data marked with "low-concentration dust emission of machinery manufacturing enterprises - limited environmental impact - fluctuating type", which is its corresponding template quantity. The template quantity of each pollution emission description label represents the quantity of machine learning data containing this pollution emission description label among the multiple machine learning data.

[0114] Step A130, based on the template quantity of each pollution emission description label, determine the marginal pollution emission description labels from the multiple pollution emission description labels. The marginal pollution emission description labels are used to represent: the pollution emission description labels corresponding to template quantities less than the set quantity.

[0115] Assume that the set quantity is 20. Then, the pollution emission description tags with a template quantity less than 20 are determined as marginal pollution emission description tags. For example, the template quantity of "Low-concentration dust emission from machinery manufacturing enterprises - Limited environmental impact - Fluctuating type" is 10, which is less than the set quantity 20. So, it is a marginal pollution emission description tag. The machine learning data corresponding to these marginal pollution emission description tags is relatively less, probably because this type of pollution emission situation is relatively special or less common in industrial parks.

[0116] Step A140: Extract one or more template pollution emission data from the template pollution emission data included in the machine learning data containing the marginal pollution emission description tags.

[0117] For "Low-concentration dust emission from machinery manufacturing enterprises - Limited environmental impact - Fluctuating type", which is determined as a marginal pollution emission description tag, extract the template pollution emission data from the corresponding machine learning data. For example, extract the dust emission data of a small machinery manufacturing enterprise during a specific period from its machine learning data, including data such as a dust concentration of 15 milligrams per cubic meter and an emission rate of 2 kilograms per hour.

[0118] Step A150: Perform feature derivation on each selected template pollution emission data respectively to generate Y iterative pollution emission data, where Y is a positive integer. One iterative pollution emission data corresponds to one template pollution emission data, and there is a consistent pollution feature path between any iterative pollution emission data and the corresponding template pollution emission data.

[0119] Taking the dust emission data extracted from the machinery manufacturing enterprise as an example, perform feature derivation. Assume that Y is 3. For this dust emission data, through a specific feature derivation algorithm, keep its pollution feature path unchanged, that is, still perform feature derivation related to dust emission. It may generate 3 iterative pollution emission data according to some relevant factors in the enterprise production process, such as equipment operation time and equipment cleaning frequency. The first iterative pollution emission data may be the dust emission situation considering the extended equipment operation time, such as the dust concentration becoming 18 milligrams per cubic meter and the emission rate becoming 2.2 kilograms per hour; the second iterative pollution emission data may be the situation after the reduction of equipment cleaning frequency, with the dust concentration becoming 20 milligrams per cubic meter and the emission rate becoming 2.5 kilograms per hour; the third iterative pollution emission data may be the situation considering both the extended equipment operation time and the reduced cleaning frequency, with the dust concentration becoming 22 milligrams per cubic meter and the emission rate becoming 2.8 kilograms per hour.

[0120] Step A160: Use the training supervision data of the template pollution emission data corresponding to each iterative pollution emission data as the training supervision data of the corresponding iterative pollution emission data.

[0121] For the 3 iterative pollution emission data of the above-mentioned dust emission data of mechanical manufacturing enterprises, since the training supervision data of the original template pollution emission data contains the prior pollution emission description label "Low-concentration dust emission from mechanical manufacturing enterprises - Limited environmental impact - Fluctuating type" and the relevant label-guided knowledge data (such as the distribution of residential areas in the surrounding environment of the enterprise, the impact of wind direction on dust diffusion, etc.), these training supervision data are also assigned to the corresponding iterative pollution emission data.

[0122] Step A170, take each of the iterative pollution emission data as iterative template pollution emission data, and generate Y iterative machine learning data based on the Y iterative pollution emission data and the corresponding training supervision data.

[0123] For example, for the 3 generated iterative pollution emission data and their corresponding training supervision data, 3 iterative machine learning data are respectively constructed. Each iterative machine learning data contains an iterative template pollution emission data and the corresponding training supervision data, which can be understood as being the same as the original machine learning data, except that the data here is an iterative version obtained through feature derivation.

[0124] Step A180, generate the training data sequence based on the multiple machine learning data and the Y iterative machine learning data.

[0125] For example, integrate all the previously collected multiple machine learning data about the enterprises in the industrial park and the newly generated Y iterative machine learning data to form a complete training data sequence. This training data sequence contains the pollution emission data of various enterprises in the industrial park, prior pollution emission description labels, label-guided knowledge data, and iterative data obtained through feature derivation, providing rich and comprehensive data resources for the target machine learning network model, enabling the model to learn the characteristics of various pollution emission situations, thereby improving the accuracy and reliability of predicting pollution emissions in the industrial park, and helping the environmental management department of the industrial park better supervise the pollution emission behaviors of enterprises and protect the environmental quality of the industrial park.

[0126] In a possible implementation manner, step A150 includes:

[0127] Step A151, obtain a feature derivation request of the feature derivation network, where the feature derivation request represents: based on the input template pollution emission data and the prior pollution emission description label, generate iterative pollution emission data with a consistent pollution feature path for the template pollution emission data.

[0128] In this embodiment, in the industrial park, this feature derivation request has a clear representational meaning, that is, based on the input template pollution emission data and the prior pollution emission description label, iterative pollution emission data with a consistent pollution feature path is generated for the template pollution emission data. Taking a chemical enterprise in the industrial park as an example, its template pollution emission data includes various pollutant emission information within a specific time period, such as the sulfur dioxide (SO2) concentration of 50 milligrams per cubic meter, nitrogen oxides (NO x ) concentration of 80 milligrams per cubic meter, particulate matter (PM) concentration of 30 milligrams per cubic meter, etc., and the prior pollution emission description label is "medium-concentration composite pollution emission of chemical enterprises - local environmental impact - fluctuating type". The feature derivation request is to generate new pollution emission data based on such template pollution emission data and prior pollution emission description labels, and the new data should be consistent with the original data in terms of the pollution feature path, which means the new data has similarity with the original data in terms of pollutant types, relevant logical relationships of emissions, etc.

[0129] Step A152, obtain one or more derived example combination data. Any one of the derived example combination data includes an example pollution emission data, the prior pollution emission description label of the corresponding example pollution emission data, and pollution emission data having a consistent pollution feature path with the corresponding example pollution emission data.

[0130] For example, for another chemical enterprise (as an example enterprise), its example pollution emission data is sulfur dioxide (SO2) concentration of 45 milligrams per cubic meter, nitrogen oxides (NO x ) concentration of 75 milligrams per cubic meter, particulate matter (PM) concentration of 25 milligrams per cubic meter, and the prior pollution emission description label is "medium-concentration composite pollution emission of chemical enterprises - local environmental impact - fluctuating type". The pollution emission data having a consistent pollution feature path with the example pollution emission data may be emission data under different time periods but with similar production processes and equipment operating states, such as sulfur dioxide (SO2) concentration of 48 milligrams per cubic meter, nitrogen oxides (NO x ) concentration of 78 milligrams per cubic meter, particulate matter (PM) concentration of 28 milligrams per cubic meter. These derived example combination data cover the emission situations of different chemical enterprises under similar pollution feature paths, providing more reference bases for subsequent feature derivation.

[0131] Step A153, based on the feature derivation request and the one or more derived example combination data through the feature derivation network, generate iterative pollution emission data with a consistent pollution feature path for each selected template pollution emission data, and generate Y iterative pollution emission data.

[0132] In a possible implementation manner, step A153 may include:

[0133] Poll the pollution emission data of each template selected, and use the pollution emission data of the currently polled template as the current template pollution emission data.

[0134] Generate an input pollution emission data segment based on the current template pollution emission data and the corresponding prior pollution emission description label.

[0135] Integrate the feature derivation request, the one or more derived sample combination data, and the generated input pollution emission data segment to generate the derived learning data corresponding to the current template pollution emission data.

[0136] Through the feature derivation network, by learning the derived learning data corresponding to the current template pollution emission data, generate iterative pollution emission data with a pollution feature path consistent with the current template pollution emission data.

[0137] Continue polling until the pollution emission data of each selected template has been polled, and generate Y iterative pollution emission data.

[0138] Continuing with the example of the pollution emission data of the chemical enterprise template mentioned earlier, during the polling process, when it comes to the pollution emission data of the chemical enterprise template, it becomes the current template pollution emission data. Based on the current template pollution emission data and the corresponding prior pollution emission description label, generate an input pollution emission data segment. For this chemical enterprise, according to its template pollution emission data (SO2 concentration is 50 milligrams per cubic meter, NO x concentration is 80 milligrams per cubic meter, PM concentration is 30 milligrams per cubic meter) and the prior pollution emission description label ("medium-concentration composite pollution emission of chemical enterprises - local environmental impact - fluctuating type"), through specific algorithms and rules, generate an input pollution emission data segment. This data segment may have performed certain feature extraction and organization on the original template pollution emission data. For example, it has focused on extracting features related to pollution emission concentration, forming an input pollution emission data segment containing information such as the proportional relationship of specific pollutant concentrations and the total pollution concentration level.

[0139] Integrate the feature derivation request, one or more derived example combination data, and the generated input pollution emission data segment to generate the derived learning data corresponding to the current template pollution emission data. Integrate the previously obtained feature derivation request (generating iterative pollution emission data with consistent pollution feature paths based on the template pollution emission data of this chemical enterprise and the prior pollution emission description labels), the derived example combination data (such as the sample pollution emission data and its related information of other chemical enterprises), and the just-generated input pollution emission data segment. This integration process is a complex data integration operation that requires combining these data together according to a specific format and logic. For example, arrange and combine the target requirements in the feature derivation request, the emission data and label information of different chemical enterprises in the derived example combination data, the concentration features in the input pollution emission data segment, etc., according to the predefined data structure to generate the derived learning data corresponding to the current template pollution emission data. This derived learning data contains rich information, including both target requirements, reference examples, and its own feature data segments, providing comprehensive learning materials for the feature derivation network.

[0140] Through the feature derivation network, by learning the derived learning data corresponding to the current template pollution emission data, generate iterative pollution emission data with the same pollution feature path as the current template pollution emission data. The feature derivation network learns and analyzes based on various information in the derived learning data. For example, according to the requirement of consistent pollution feature paths in the feature derivation request, refer to the emission situations of similar enterprises in the derived example combination data, combine information such as the concentration features in the input pollution emission data segment, and use its internal algorithms and model structures to generate iterative pollution emission data with the same pollution feature path as the current template pollution emission data. Suppose the generated iterative pollution emission data is 52 milligrams per cubic meter of sulfur dioxide (SO2) concentration, 82 milligrams per cubic meter of nitrogen oxides (NO x ) concentration, and 32 milligrams per cubic meter of particulate matter (PM) concentration. This data is consistent with the original template pollution emission data in terms of pollutant types and emission logical relationships, only with changes in the specific pollutant concentrations, reflecting the derivation of pollution emission data under specific conditions. Continue polling until all the selected template pollution emission data are polled, generating Y iterative pollution emission data. For example, if it is required to generate Y = 3 iterative pollution emission data, then perform the same operations on the other selected template pollution emission data (possibly from different enterprises or the same enterprise at different time periods) in the above polling and generating manner until 3 such iterative pollution emission data are generated.

[0141] In a possible implementation manner, step A153 may further include:

[0142] Fuse the feature derivation request and the combined data of the one or more derived examples to generate training samples for the feature derivation network.

[0143] Respectively use each selected template pollution emission data and the corresponding prior pollution emission description label to generate input pollution emission data segments corresponding to the respective template pollution emission data.

[0144] Train the feature derivation network based on the training samples to generate a trained feature derivation network.

[0145] Respectively, through the trained feature derivation network, based on the input pollution emission data segments corresponding to the respective template pollution emission data, generate iterative pollution emission data with consistent pollution feature paths for the corresponding template pollution emission data, generating Y iterative pollution emission data.

[0146] In another possible implementation, fuse the feature derivation request (generating iterative pollution emission data with consistent pollution feature paths based on template pollution emission data and prior pollution emission description labels) and the combined data of derived examples (sample pollution emission data and their related information of multiple chemical enterprises). This fusion process needs to consider how to effectively combine the target information in the feature derivation request with the actual emission data and label information in the combined data of derived examples. For example, the target requirements in the feature derivation request can be transformed into specific marker information, and then combined with the pollution emission data of each enterprise, prior pollution emission description labels, etc. in the combined data of derived examples in a certain order to form training samples for the feature derivation network. This training sample is a dataset that combines target requirements and actual reference data, providing a basis for the training of the feature derivation network.

[0147] Respectively use each selected template pollution emission data and the corresponding prior pollution emission description label to generate input pollution emission data segments corresponding to the respective template pollution emission data. For each selected template pollution emission data (such as the template pollution emission data of chemical enterprises and the template pollution emission data of other enterprises mentioned before), combine its corresponding prior pollution emission description label and generate input pollution emission data segments according to a specific algorithm. This process is similar to the one mentioned before, which is to extract and organize the features of the template pollution emission data to form an input pollution emission data segment containing key pollution emission feature information.

[0148] The feature derivation network is trained based on training samples to generate a trained feature derivation network. The generated training samples are input into the feature derivation network, and the feature derivation network adjusts its own parameters according to the information in the training samples. For example, according to the sample pollution emission data of different chemical enterprises, the prior pollution emission description labels in the training samples, and the target requirements in the feature derivation request, the feature derivation network adjusts parameters such as the connection weights between neurons through an internal algorithm mechanism, so that the network can accurately generate iterative pollution emission data with consistent pollution feature paths based on the input template pollution emission data and its prior pollution emission description labels. After multiple training iterations, the feature derivation network gradually optimizes its own performance and finally generates a trained feature derivation network.

[0149] Based on the input pollution emission data segments corresponding to each template pollution emission data through the trained feature derivation network, iterative pollution emission data with consistent pollution feature paths are generated for the corresponding template pollution emission data, generating Y iterative pollution emission data. For each template pollution emission data, its corresponding input pollution emission data segment is input into the trained feature derivation network. For example, for the input pollution emission data segment corresponding to the template pollution emission data of a chemical enterprise, the trained feature derivation network generates iterative pollution emission data with a consistent pollution feature path with the template pollution emission data of this chemical enterprise according to the feature information in this segment, combined with the parameters and model structure obtained from its own training. In this way, all template pollution emission data for which iterative pollution emission data need to be generated are operated on until Y iterative pollution emission data are generated. These iterative pollution emission data are of great significance in the environmental management of industrial parks. They can enrich the types and variation situations of pollution emission data, provide more diverse data for the target machine learning network model, thereby improving the accuracy and comprehensiveness of the model's prediction of pollution emissions in the industrial park, and helping to more precisely supervise the pollution emission behaviors of enterprises and protect the environmental quality of the industrial park.

[0150] In a possible implementation manner, the method further includes:

[0151] Step S150, after receiving a prediction instruction for target pollution emission data, generate candidate pollution emission data based on the model enabling definition information and the target pollution emission data, and configure the target pollution emission data as input pollution emission data in the candidate pollution emission data.

[0152] Step S160, perform pollution emission prediction on the target pollution emission data based on the candidate pollution emission data through the pollution emission prediction model, and generate a pollution emission description label and corresponding label guidance knowledge data for the target pollution emission data.

[0153] In this embodiment, for example, there is a newly established electronic enterprise in an industrial park. Its target pollution emission data includes various information in waste gas emissions, such as the fluoride concentration of 10 milligrams per cubic meter, the ammonia concentration of 5 milligrams per cubic meter, etc., as well as relevant data such as the emission flow rate and temperature. The model activation definition information covers the rules and requirements for processing various pollution emission data, such as the weight distribution principle for different pollutants, the consideration of the importance of different emission time periods, and the calculation method of the impact of meteorological conditions on pollution diffusion. When generating candidate pollution emission data, the target pollution emission data of the electronic enterprise is integrated according to the rules in the model activation definition information. For example, according to the weight distribution principle in the model activation definition information, the weight of the fluoride concentration in the overall pollution assessment is 0.3, the weight of the ammonia concentration is 0.2, and the weight of other factors is 0.5. These weights are calculated and integrated with the actual pollution emission data of the electronic enterprise. This integration is not a simple numerical addition, but is based on pre-set complex calculation rules, fully considering the mutual relationship between various factors and the comprehensive effect on the environmental impact, and finally generating candidate pollution emission data. This candidate pollution emission data is a standardized data form that combines the model requirements and the actual pollution emission characteristics, providing input data that conforms to the model logical structure for subsequent pollution emission prediction.

[0154] Next, the pollution emission prediction model is used to predict the target pollution emission data based on the candidate pollution emission data, and the pollution emission description label of the target pollution emission data and the corresponding label-guided knowledge data are generated. The pollution emission prediction model takes the candidate pollution emission data as input, which contains the fused pollution emission characteristics of the electronic enterprise and the relevant information required by the model. The model analyzes and predicts the target pollution emission data of the electronic enterprise based on its own network structure and existing parameters. For example, when processing the pollutant concentration, flow rate, temperature and other data of the electronic enterprise, the model combines the pollution emission laws and related knowledge learned by itself to generate pollution emission prediction results. This pollution emission prediction result includes the pollution emission description label of the corresponding target pollution emission data generated, as well as the corresponding label-guided knowledge data. Assume that the generated pollution emission description label is "Electronic enterprise low concentration special gas emission-specific regional impact-stable type". Here, "specific regional impact" is the environmental impact assessment based on the analysis of pollutant concentration and surrounding environmental sensitive areas (such as nearby high-precision electronic equipment production workshops that are more sensitive to these gases); the corresponding label-guided knowledge data may contain more detailed predictions of surrounding environmental changes, such as the trend of surrounding air quality changes in the future under the current pollution emission level, the potential impact risk on the product quality of nearby high-precision electronic equipment production workshops, and suggestions on measures that enterprises can take to reduce pollution emissions. These prediction results can provide valuable information for the management department of the industrial park, helping it to timely understand the pollution emissions of newly settled enterprises, so as to formulate corresponding environmental management strategies in advance, ensure the overall environmental quality of the industrial park, and protect the health and safety of surrounding residents and the ecological environment.

[0155] Figure 2 The hardware structure of the pollution emission prediction system 100 based on machine learning for implementing the above-mentioned pollution emission prediction method based on machine learning provided by an embodiment of the present invention is shown as follows: Figure 2 As shown, the machine learning-based pollution emission prediction system 100 may include a processor 110 , a machine-readable storage medium 120 , a bus 130 , and a communication unit 140 .

[0156] The machine-readable storage medium 120 may store data and / or instructions. In some embodiments, the machine-readable storage medium 120 may store data acquired from an external terminal. In some embodiments, the machine-readable storage medium 120 may store data and / or instructions that the machine-learning-based pollution emission prediction system 100 uses to execute or use to complete the exemplary method described in the present invention.

[0157] In the specific implementation process, one or more processors 110 execute computer-executable instructions stored in the machine-readable storage medium 120, enabling the processors 110 to execute the machine learning-based pollution emission prediction method as described in the above method embodiments. The processors 110, the machine-readable storage medium 120, and the communication unit 140 are connected through the bus 130, and the processors 110 can be used to control the transceiver actions of the communication unit 140.

[0158] For the specific implementation process of the processors 110, reference can be made to the respective method embodiments executed by the above machine learning-based pollution emission prediction system 100. Their implementation principles and technical effects are similar, and will not be elaborated herein.

[0159] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above machine learning-based pollution emission prediction method is implemented.

[0160] It should be noted that, in order to simplify the description of the present invention disclosure and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes multiple features are merged into one embodiment, drawing, or description thereof.

Claims

1. A pollution emission prediction method based on machine learning, characterized in that, The method includes: Obtaining target machine learning data of a target machine learning network model, where the target machine learning data includes: template pollution emission data, a prior pollution emission description label corresponding to the template pollution emission data, and corresponding label guiding knowledge data, and the prior pollution emission description label of the template pollution emission data is selected from a plurality of pre-defined pollution emission description labels; Obtaining model enabling definition information of the target machine learning network model; the model enabling definition information characterizes that based on the plurality of pollution emission description labels, pollution emission prediction is performed on input pollution emission data, and a pollution emission description label and corresponding label guiding knowledge data of the input pollution emission data are generated; Generating example learning data for updating the target machine learning network model according to the model enabling definition information and the target machine learning data, and configuring the template pollution emission data in the target machine learning data as input pollution emission data in the example learning data; Performing model parameter learning on the target machine learning network model according to the example learning data, and using the target machine learning network model that has completed model parameter learning as a pollution emission prediction model; the pollution emission prediction model is used to perform pollution emission prediction on any input pollution emission data and generate a corresponding pollution emission description label and label guiding knowledge data.

2. The pollution emission prediction method based on machine learning according to claim 1, wherein The example learning data is a pollution emission data generated by fusing the model enabling definition information and the target machine learning data; performing model parameter learning on the target machine learning network model according to the example learning data includes: Performing feature encoding on the example learning data to generate X encoded feature node data, where X is a positive integer; Based on the time series information of the X encoded feature node data, generating a to-be-learned feature according to the first X - 1 encoded feature node data among the X encoded feature node data; and generating an expected output feature corresponding to the to-be-learned feature according to the X - 1 encoded feature node data except the first encoded feature node data among the X encoded feature node data; Deriving the to-be-learned feature through the target machine learning network model for each encoded feature node data to generate a derivation result; the derivation result includes X - 1 encoded feature node data derived, and the a-th encoded feature node data in the derivation result is derived based on the first a encoded feature node data in the to-be-learned feature, where a is a positive integer between 1 and X - 1; Based on the expected output feature and the derivation result, calculating error parameters corresponding to each encoded feature node data in the derivation result, and the error parameter corresponding to the a-th encoded feature node data in the derivation result characterizes the loss result between the a-th encoded feature node data in the derivation result and the a-th encoded feature node data in the expected output feature; Fusing the error parameters corresponding to each encoded feature node data in the derivation result to generate a training error parameter of the target machine learning network model; Update the neuron weight information of the target machine learning network model according to a training objective that minimizes the training error parameter, so as to learn the model parameters of the target machine learning network model.

3. The pollution emission prediction method based on machine learning according to claim 1, characterized in that, Generating example learning data for updating the target machine learning network model based on the model enabling definition information and the target machine learning data includes: Fusing the model enabling definition information and the template pollution emission data in the target machine learning data and outputting them as a pollution emission data to generate a first pollution emission data; Fusing the prior pollution emission description label in the target machine learning data and the label guiding knowledge data in the target machine learning data and outputting them as a pollution emission data to generate a second pollution emission data; Generate example learning data for updating the target machine learning network model based on the first pollution emission data and the second pollution emission data.

4. The method for predicting pollution emissions based on machine learning according to claim 3, wherein The learning of the model parameters of the target machine learning network model based on the example learning data includes: Based on the first pollution emission data in the example learning data through the target machine learning network model, perform pollution emission prediction on the template pollution emission data in the target machine learning data to generate a pollution emission prediction result; the pollution emission prediction result includes: the pollution emission description label of the generated corresponding template pollution emission data, and the corresponding label guiding knowledge data; Based on the loss result between the pollution emission prediction result and the second pollution emission data in the example learning data, update the neuron weight information of the target machine learning network model, so as to learn the model parameters of the target machine learning network model.

5. The pollution emission prediction method based on machine learning according to claim 1, characterized in that, The target machine learning data is a machine learning data obtained from a training data sequence. The generation steps of the training data sequence include: Obtain multiple machine learning data of the target machine learning network model. A machine learning data includes: a template pollution emission data and the training supervision data of the corresponding template pollution emission data; the training supervision data of any template pollution emission data includes the prior pollution emission description label of the corresponding template pollution emission data and the corresponding label guiding knowledge data. The prior pollution emission description label in any machine learning data is selected from the multiple pollution emission description labels; Determine the template quantity of each pollution emission description label in the multiple pollution emission description labels according to the multiple machine learning data; the template quantity of any pollution emission description label is used to represent: the number of machine learning data containing the any pollution emission description label in the multiple machine learning data; Based on the template quantity of each pollution emission description label, determine the marginal pollution emission description labels from the multiple pollution emission description labels; the marginal pollution emission description labels are used to represent: the pollution emission description labels corresponding to the template quantities less than the set quantity; Extract one or more template pollution emission data from the template pollution emission data included in the machine learning data containing the marginal pollution emission description label; Feature derivation is respectively performed on each selected template pollution emission data to generate Y iterative pollution emission data, where Y is a positive integer; one iterative pollution emission data corresponds to one template pollution emission data, and there is a consistent pollution feature path between any iterative pollution emission data and the corresponding template pollution emission data; The training supervision data of the template pollution emission data corresponding to each iterative pollution emission data is used as the training supervision data of the corresponding iterative pollution emission data; Each of the iterative pollution emission data is used as iterative template pollution emission data, and based on the Y iterative pollution emission data and the corresponding training supervision data, Y iterative machine learning data are generated; Based on the multiple machine learning data and the Y iterative machine learning data, the training data sequence is generated.

6. The method for predicting pollution emissions based on machine learning according to claim 5, wherein The step of respectively performing feature derivation on each selected template pollution emission data to generate Y iterative pollution emission data includes: Obtain a feature derivation request of a feature derivation network, where the feature derivation request represents: based on the input template pollution emission data and the prior pollution emission description label, generate iterative pollution emission data with a consistent pollution feature path for the template pollution emission data; Obtain one or more derivative example combination data, and any one of the derivative example combination data includes an example pollution emission data, the prior pollution emission description label of the corresponding example pollution emission data, and a pollution emission data having a consistent pollution feature path with the corresponding example pollution emission data; Through the feature derivation network, based on the feature derivation request and the one or more derivative example combination data, iterative pollution emission data with a consistent pollution feature path are respectively generated for each selected template pollution emission data, generating Y iterative pollution emission data.

7. The method for predicting pollution emissions based on machine learning according to claim 6, wherein The step of through the feature derivation network, based on the feature derivation request and the one or more derivative example combination data, respectively generating iterative pollution emission data with a consistent pollution feature path for each selected template pollution emission data, generating Y iterative pollution emission data, includes: Poll each selected template pollution emission data in turn, and use the currently polled template pollution emission data as the current template pollution emission data; Generate an input pollution emission data segment according to the current template pollution emission data and the corresponding prior pollution emission description label; Integrate the feature derivation request, the one or more derivative example combination data, and the generated input pollution emission data segment to generate derivative learning data corresponding to the current template pollution emission data; Through the feature derivation network, by learning the derivative learning data corresponding to the current template pollution emission data, generate iterative pollution emission data having a consistent pollution feature path with the current template pollution emission data; Continue polling until each selected template pollution emission data has been polled, generating Y iterative pollution emission data.

8. The method for predicting pollution emissions based on machine learning according to claim 6, wherein The feature derivation network generates iterative pollution emission data with consistent pollution feature paths for each selected template pollution emission data based on the feature derivation request and the one or more derived example combination data, generating Y iterative pollution emission data, including: Fusing the feature derivation request and the one or more derived example combination data to generate a training sample for the feature derivation network; Generating input pollution emission data segments corresponding to the respective template pollution emission data by using each selected template pollution emission data and the corresponding prior pollution emission description label; Training the feature derivation network based on the training sample to generate a trained feature derivation network; Generating iterative pollution emission data with consistent pollution feature paths for the corresponding template pollution emission data respectively through the trained feature derivation network based on the input pollution emission data segments corresponding to the respective template pollution emission data, generating Y iterative pollution emission data.

9. The pollution emission prediction method based on machine learning according to claim 1, wherein The method further includes: After receiving a prediction instruction for target pollution emission data, generating candidate pollution emission data according to the model enabling definition information and the target pollution emission data, and configuring the target pollution emission data as input pollution emission data in the candidate pollution emission data; Performing pollution emission prediction on the target pollution emission data based on the candidate pollution emission data through the pollution emission prediction model to generate a pollution emission description label and corresponding label guiding knowledge data for the target pollution emission data.

10. A pollution emission prediction system based on machine learning, characterized in that, The machine learning-based pollution emission prediction system includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the machine learning-based pollution emission prediction method according to any one of claims 1-9 above.

Citation Information

Patent Citations

  • Enterprise pollutant emission prediction method and prediction system

    CN116862079A

  • Method of predicting fine dust concentration and inferring source by using local public data and prediction and inference device

    US20240029457A1