Pollutant emission prediction method and system based on machine learning

By combining template pollution emission data, prior pollution emission description labels and label guidance knowledge data, sample learning data is generated, and parameter learning of the target machine learning network model is solved, and pollution emission prediction in the existing technology is solved, and more efficient and reliable pollution emission prediction is achieved.

CN119940632AActive Publication Date: 2025-05-06CHINESE ACAD OF ENVIRONMENTAL PLANNING

Patent Information

Application Number
CN202510028500.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The existing machine learning-based pollution emission prediction methods are difficult to accurately reflect the actual situation and future trends of pollution emissions when processing complex and changeable pollution emission data, and lack detailed description and explanatory information on the characteristics of pollution emissions.

Method used

By obtaining the target machine learning data of the target machine learning network model, including template pollution emission data, prior pollution emission description tags and label guidance knowledge data, combining the model enable definition information to generate sample learning data, and learning model parameters of the target machine learning network model to generate a model for pollution emission prediction.

Benefits of technology

It improves the accuracy and reliability of pollution emission prediction, optimizes the learning efficiency of the model, enhances the ability to identify and predict complex pollution emission patterns, and provides detailed pollution emission description and explanatory information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940632A_ABST
    Figure CN119940632A_ABST
Patent Text Reader

Abstract

The invention provides a pollution emission prediction method and system based on machine learning. The accuracy and reliability of pollution emission prediction are remarkably improved. The method comprises the following steps: firstly, acquiring target machine learning data containing rich priori knowledge, and generating sample learning data for model updating in combination with model starting definition information; and parameter learning is performed on the target machine learning network model by using the sample data, so that the obtained pollution emission prediction model can predict randomly input pollution emission data more accurately, and corresponding pollution emission description labels and label guide knowledge data are generated. The process not only optimizes the learning efficiency of the model, but also enhances the recognition and prediction ability of the model to the complex pollution emission mode, thereby facilitating the monitoring, treatment and decision making of environmental pollution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular to a method and system for predicting pollution emissions based on machine learning. Background Art

[0002] With the rapid development of industrialization and urbanization, environmental pollution problems are becoming increasingly serious. Accurate prediction of pollution emissions is crucial for environmental protection, policy making, and public health. Traditional pollution emission prediction methods often rely on statistical models or physical models, which have limitations in dealing with complex and changeable pollution emission data and are difficult to accurately reflect the actual situation and future trends of pollution emissions.

[0003] In recent years, machine learning technology has gradually shown great potential in the field of pollution emission prediction due to its powerful data processing and pattern recognition capabilities. However, existing pollution emission prediction methods based on machine learning still face some challenges. On the one hand, these methods usually rely on a large amount of historical pollution emission data for model training, but the acquisition of historical data is often subject to many restrictions, and the data quality is uneven, which affects the prediction accuracy of the model. On the other hand, traditional machine learning methods often find it difficult to effectively extract key information when dealing with pollution emission data with complex characteristics and high nonlinearity, resulting in instability in the prediction results.

[0004] In addition, when predicting pollution emissions, existing machine learning models usually only output a single pollution emission value or a simple classification label, lacking detailed descriptions and explanatory information about pollution emission characteristics, which limits the value of the prediction results in practical applications. Therefore, how to build a machine learning prediction method that can make full use of limited data resources, effectively extract pollution emission characteristics, and provide rich prediction information has become an urgent problem to be solved in the current field of pollution emission prediction. Summary of the invention

[0005] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a method for predicting pollution emissions based on machine learning, the method comprising:

[0006] Obtaining target machine learning data of a target machine learning network model, wherein the target machine learning data includes: template pollution emission data, a priori pollution emission description label of the corresponding template pollution emission data, and corresponding label-guided knowledge data, wherein the priori pollution emission description label of the template pollution emission data is selected from a plurality of pollution emission description labels previously defined;

[0007] Obtaining model activation definition information of the target machine learning network model; the model activation definition information represents: performing pollution emission prediction on input pollution emission data based on the multiple pollution emission description labels, and generating pollution emission description labels of the input pollution emission data and corresponding label guidance knowledge data;

[0008] Generate sample learning data for updating the target machine learning network model according to the model activation definition information and the target machine learning data, wherein the template pollution emission data in the target machine learning data is configured as input pollution emission data in the sample learning data;

[0009] The target machine learning network model is subjected to model parameter learning based on the sample learning data, and the target machine learning network model that has completed the model parameter learning is used as a pollution emission prediction model; the pollution emission prediction model is used to perform pollution emission prediction for any input pollution emission data, and generate corresponding pollution emission description labels and label-guided knowledge data.

[0010] In a possible implementation of the first aspect, the sample learning data is pollution emission data generated by fusing the model activation definition information and the target machine learning data; and the model parameter learning of the target machine learning network model based on the sample learning data includes:

[0011] Performing feature encoding on the sample learning data to generate X encoded feature node data, where X is a positive integer;

[0012] Based on the time series information of the X encoding feature node data, generating a feature to be learned according to the first X-1 encoding feature node data among the X encoding feature node data; and generating an expected output feature corresponding to the feature to be learned according to the X-1 encoding feature node data except the first encoding feature node data among the X encoding feature node data;

[0013] The target machine learning network model is used to derive the coded feature node data of the feature to be learned one by one to generate a derivation result; the derivation result includes X-1 derived coded feature node data, and the a-th coded feature node data in the derivation result is derived based on the first a coded feature node data in the feature to be learned, and a is a positive integer between 1 and X-1;

[0014] Based on the expected output feature and the derivation result, the error parameters corresponding to each coding feature node data in the derivation result are calculated, and the error parameter corresponding to the a-th coding feature node data in the derivation result represents: the loss result between the a-th coding feature node data in the derivation result and the a-th coding feature node data in the expected output feature;

[0015] Fusion of the error parameters corresponding to the respective encoding feature node data in the derivation results to generate training error parameters of the target machine learning network model;

[0016] According to the training goal of minimizing the training error parameters, the neuron weight information of the target machine learning network model is updated, thereby learning the model parameters of the target machine learning network model.

[0017] In a possible implementation manner of the first aspect, generating sample learning data for updating the target machine learning network model according to the model activation definition information and the target machine learning data includes:

[0018] fusing the model enabling definition information and the template pollution emission data in the target machine learning data into one pollution emission data to generate first pollution emission data;

[0019] The prior pollution emission description label in the target machine learning data and the label-guided knowledge data in the target machine learning data are fused and outputted into one pollution emission data to generate second pollution emission data;

[0020] Based on the first pollution emission data and the second pollution emission data, sample learning data for updating the target machine learning network model is generated.

[0021] In a possible implementation of the first aspect, the performing model parameter learning on the target machine learning network model according to the sample learning data includes:

[0022] Based on the first pollution emission data in the sample learning data, the target machine learning network model performs pollution emission prediction on the template pollution emission data in the target machine learning data to generate a pollution emission prediction result; the pollution emission prediction result includes: a pollution emission description label of the generated corresponding template pollution emission data, and corresponding label-guided knowledge data;

[0023] Based on the loss result between the pollution emission prediction result and the second pollution emission data in the sample learning data, the neuron weight information of the target machine learning network model is updated, thereby learning the model parameters of the target machine learning network model.

[0024] In a possible implementation of the first aspect, the target machine learning data is machine learning data obtained from a training data sequence, and the step of generating the training data sequence includes:

[0025] Acquire multiple machine learning data of the target machine learning network model, wherein one machine learning data includes: a template pollution emission data and training supervision data of the corresponding template pollution emission data; the training supervision data of any template pollution emission data includes a priori pollution emission description label of the corresponding template pollution emission data and corresponding label-guided knowledge data, and the priori pollution emission description label in any machine learning data is selected from the multiple pollution emission description labels;

[0026] Determine the number of templates of each pollution emission description label in the multiple pollution emission description labels based on the multiple machine learning data; the number of templates of any one pollution emission description label is used to represent: the number of machine learning data containing any one pollution emission description label in the multiple machine learning data;

[0027] Based on the number of templates of each pollution emission description label, determining an edge pollution emission description label from the multiple pollution emission description labels; the edge pollution emission description label is used to represent: the pollution emission description label corresponding to the number of templates less than the set number;

[0028] Extract one or more template pollution emission data from the template pollution emission data included in the machine learning data containing the edge pollution emission description label;

[0029] Derivation of features is performed on each selected template pollution emission data to generate Y iterative pollution emission data, where Y is a positive integer; one iterative pollution emission data corresponds to one template pollution emission data, and any iterative pollution emission data has a consistent pollution feature path with the corresponding template pollution emission data;

[0030] Using the training supervision data of the template pollution emission data corresponding to each iterative pollution emission data as the training supervision data of the corresponding iterative pollution emission data;

[0031] Each of the iterative pollution emission data is used as iterative template pollution emission data, and Y iterative machine learning data is generated according to the Y iterative pollution emission data and the corresponding training supervision data;

[0032] The training data sequence is generated according to the multiple machine learning data and the Y iterative machine learning data.

[0033] In a possible implementation manner of the first aspect, the step of performing feature derivation on each selected template pollution emission data to generate Y iterative pollution emission data includes:

[0034] Obtaining a feature derivation request of a feature derivation network, wherein the feature derivation request indicates: generating iterative pollution emission data having a consistent pollution feature path for the template pollution emission data based on input template pollution emission data and a priori pollution emission description label;

[0035] Obtain one or more derived sample combination data, any one of the derived sample combination data includes a sample pollution emission data and a priori pollution emission description label of the corresponding sample pollution emission data and pollution emission data having a consistent pollution characteristic path with the corresponding sample pollution emission data;

[0036] Based on the feature derivation request and the one or more derived sample combination data, the feature derivation network generates iterative pollution emission data with consistent pollution feature paths for each selected template pollution emission data, thereby generating Y iterative pollution emission data.

[0037] In a possible implementation manner of the first aspect, generating iterative pollution emission data having a consistent pollution feature path for each selected template pollution emission data based on the feature derivation request and the one or more derived sample combination data through the feature derivation network to generate Y iterative pollution emission data includes:

[0038] Polling each selected template pollution emission data, and taking the currently polled template pollution emission data as the current template pollution emission data;

[0039] Generating an input pollution emission data segment according to the current template pollution emission data and the corresponding prior pollution emission description label;

[0040] Integrate the feature derivation request, the one or more derived sample combination data, and the generated input pollution emission data fragment to generate derived learning data corresponding to the current template pollution emission data;

[0041] Generate iterative pollution emission data having a consistent pollution feature path with the current template pollution emission data by learning the derived learning data corresponding to the current template pollution emission data through the feature derivation network;

[0042] Continue polling until all selected template pollution emission data are polled, generating Y iterative pollution emission data.

[0043] In a possible implementation manner of the first aspect, generating iterative pollution emission data having a consistent pollution feature path for each selected template pollution emission data based on the feature derivation request and the one or more derived sample combination data through the feature derivation network to generate Y iterative pollution emission data includes:

[0044] fusing the feature derivation request and the one or more derived sample combination data to generate a training sample for the feature derivation network;

[0045] Using each selected template pollution emission data and the corresponding prior pollution emission description label respectively, generating an input pollution emission data segment corresponding to each template pollution emission data;

[0046] Training the feature derivation network based on the training samples to generate a trained feature derivation network;

[0047] The trained feature derivation network is used to generate iterative pollution emission data with consistent pollution feature paths for the corresponding template pollution emission data based on the input pollution emission data segments corresponding to the respective template pollution emission data, thereby generating Y iterative pollution emission data.

[0048] In a possible implementation of the first aspect, the method further includes:

[0049] After receiving a prediction instruction for target pollution emission data, generating candidate pollution emission data according to the model activation definition information and the target pollution emission data, wherein the target pollution emission data is configured as input pollution emission data in the candidate pollution emission data;

[0050] The pollution emission prediction model is used to perform pollution emission prediction on the target pollution emission data based on the candidate pollution emission data, and a pollution emission description label of the target pollution emission data and corresponding label-guided knowledge data are generated.

[0051] On the other hand, an embodiment of the present invention also provides a pollution emission prediction system based on machine learning, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0052] Based on the above aspects, the embodiments of the present application significantly improve the accuracy and reliability of pollution emission prediction. The method first obtains target machine learning data containing rich prior knowledge, and combines the model activation definition information to generate sample learning data for model updating. By using these sample data to learn the parameters of the target machine learning network model, the resulting pollution emission prediction model can more accurately predict any input pollution emission data, and generate corresponding pollution emission description labels and label-guided knowledge data. This process not only optimizes the learning efficiency of the model, but also enhances the model's ability to recognize and predict complex pollution emission patterns, thereby facilitating environmental pollution monitoring, governance and decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a schematic diagram of the execution flow of the pollution emission prediction method based on machine learning provided in an embodiment of the present invention.

[0054] Figure 2 It is a schematic diagram of the hardware architecture of the pollution emission prediction system based on machine learning provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0055] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 It is a flow chart of a pollution emission prediction method based on machine learning provided by an embodiment of the present invention. The pollution emission prediction method based on machine learning is introduced in detail below.

[0056] Step S110, obtaining target machine learning data of the target machine learning network model, wherein the target machine learning data includes: template pollution emission data, prior pollution emission description labels of corresponding template pollution emission data, and corresponding label-guided knowledge data, wherein the prior pollution emission description labels of the template pollution emission data are selected from multiple pollution emission description labels previously defined.

[0057] In this embodiment, in the environmental monitoring scenario of an industrial park, the target machine learning network model aims to accurately analyze and predict the air pollution emissions of enterprises in the park. The target machine learning data is the basic sample for model learning.

[0058] The template pollution emission data can be data collected from the exhaust emission monitoring equipment of various enterprises in the industrial park. For example, the exhaust gas emitted by the chimney of a chemical enterprise in a specific period of time has a sulfur dioxide (SO2) concentration of 50 mg per cubic meter and nitrogen oxides (NO x ) concentration is 80 mg per cubic meter, particulate matter (PM) concentration is 30 mg per cubic meter, etc. These data constitute a template pollution emission data sample.

[0059] The prior pollution emission description label is selected from a complex pre-defined multiple pollution emission description labels. These pollution emission description labels are not just simple classifications of pollutant concentrations, but also involve many factors such as the nature of the emission source and the potential impact of the emission pattern on the environment. For example, for the waste gas emissions of the above-mentioned chemical enterprises, the prior pollution emission description label may be "high-concentration complex waste gas emissions from chemical enterprises-potentially large environmental hazards-short-term fluctuations". This label indicates that the emission source is a chemical enterprise, and the concentrations of various pollutants in the waste gas are high. Due to the types and concentrations of pollutants emitted, there is a great potential harm to the environment, and the emissions fluctuate in the short term, which may be related to the intermittent nature of the company's production process.

[0060] The corresponding label-guided knowledge data is auxiliary information related to the prior pollution emission description label. For example, it may include the relationship between the production process of chemical companies and the generation of pollutants, historical change data of the surrounding environment under similar emission conditions (such as damage to nearby vegetation, long-term trend of air quality, etc.), the existing measures taken by the company to control pollution and their effectiveness evaluation, etc. These label-guided knowledge data help the model to have a deeper understanding of the meaning behind the prior pollution emission description label, so as to better learn and predict.

[0061] By collecting a large number of such template pollution emission data, the corresponding prior pollution emission description labels and label-guided knowledge data, the target machine learning data of the target machine learning network model is formed. Each set of data in it is a detailed record and interpretation of different pollution emission conditions in the industrial park, providing rich material for subsequent model training and optimization.

[0062] Step S120, obtaining model activation definition information of the target machine learning network model. The model activation definition information represents: performing pollution emission prediction on the input pollution emission data based on the multiple pollution emission description labels, and generating pollution emission description labels of the input pollution emission data and corresponding label guidance knowledge data.

[0063] For example, in the environmental management system of an industrial park, the model activation definition information of the target machine learning network model specifies how the target machine learning network model processes the input pollution emission data based on complex multiple pollution emission description labels.

[0064] For example, in an industrial park, there are many types of enterprises, including chemical, machinery manufacturing, electronics, etc. The pollution emission characteristics of each enterprise are different, and the form and meaning of its pollution emission data are also very complex. Model activation definition information is used to guide the target machine learning network model on how to predict pollution emissions when faced with pollution emission data of different types of enterprises.

[0065] For the exhaust gas emission data collected from various enterprises (which is a type of input pollution emission data), the model activation definition information requires the model to conduct a comprehensive analysis based on multiple pre-defined pollution emission description tags. These tags may cover information on multiple dimensions, from the type, concentration, and emission rate of pollutants to the time pattern of emissions (such as whether it is intermittent emission, whether it is related to production shifts, etc.), to the potential impact of emissions on different areas of the surrounding environment (such as residential areas, farmland, near water bodies, etc.).

[0066] Taking a chemical enterprise as an example, after receiving the waste gas emission data of the enterprise, the model starts to use the definition information to guide the model to consider the various complex pollutants (such as volatile organic compounds, heavy metals, etc.) generated in the chemical production process, as well as the diffusion laws and environmental impacts of these pollutants under different meteorological conditions (such as wind direction, wind speed, temperature, humidity, etc.). At the same time, for the waste gas emissions of machinery manufacturing enterprises, the model should pay attention to the emission of pollutants such as metal dust and lubricating oil volatiles, and combine the production scale of the enterprise, equipment operation status and other factors to make accurate pollution emission predictions based on pollution emission description labels.

[0067] While predicting pollution emissions, the model also generates pollution emission description labels and corresponding label-guided knowledge data for the input pollution emission data. For example, for the waste gas emission data of a machinery manufacturing enterprise, after analysis, the model may generate a pollution emission description label of "Medium-concentration dust and volatile organic compound emissions from machinery manufacturing enterprises-local environmental impact-stable type", and the corresponding label-guided knowledge data may include wind direction frequency maps of the area where the enterprise is located, historical data comparisons of surrounding air quality monitoring stations, maintenance records and effect analysis of environmental protection equipment within the enterprise, etc. These generated labels and knowledge data can provide more detailed and targeted decision-making basis for the environmental management department of the park, helping them to better supervise the pollution emission behavior of enterprises and formulate more reasonable environmental protection policies.

[0068] Step S130, based on the model activation definition information and the target machine learning data, generate sample learning data for updating the target machine learning network model, and the template pollution emission data in the target machine learning data is configured as input pollution emission data in the sample learning data.

[0069] Assume that the target machine learning data of the target machine learning network model has been obtained, where the template pollution emission data comes from actual monitoring of different enterprises in the park. For example, there is waste gas emission data of an electronics company, which contains information such as fluoride concentration of 10 mg per cubic meter and ammonia concentration of 5 mg per cubic meter. This is the template pollution emission data part of the target machine learning data. The corresponding prior pollution emission description label is "Electronic enterprise low-concentration special gas emission-specific regional impact-stable type", and the label-guided knowledge data includes the types of chemical agents used in the production process of the enterprise, the distribution of surrounding environmental sensitive areas (such as nearby high-precision electronic equipment production workshops that are more sensitive to these gases), etc.

[0070] The model activation definition information requires that the input pollution emission data be processed in a specific way to generate the information required for pollution emission prediction. According to these requirements, the model activation definition information and the template pollution emission data in the target machine learning data are fused. For example, the model activation definition information specifies the weight distribution method for different types of pollutants (such as the weight of fluoride in a specific environmental impact assessment is 0.3, ammonia is 0.2, etc.). In this way, the waste gas emission data of the electronic enterprise is fused with the model activation definition information to generate the first pollution emission data. This data combines the model's requirements for data processing and the actual pollution emission data, and is a standardized data form.

[0071] Then, the prior pollution emission description label and label-guided knowledge data in the target machine learning data are fused and output as a pollution emission data to generate the second pollution emission data. For the above-mentioned electronics enterprise example, the "electronic enterprise low-concentration special gas emission-specific regional impact-stable type" in the prior pollution emission description label is integrated with the information such as the types of pharmaceuticals produced by the enterprise and the distribution of surrounding sensitive areas in the label-guided knowledge data to form a second pollution emission data containing more semantic information.

[0072] Finally, based on the first pollution emission data and the second pollution emission data, sample learning data for updating the target machine learning network model is generated. Sample learning data is a special data structure that integrates model requirements, actual pollution emission data, prior labels and related knowledge data. It can be understood as a learning sample customized for the target machine learning network model. It contains both the original pollution emission data characteristics and the processing rules required by the model and the prior understanding of the data, so that the model can better adapt to the complex pollution emission situation in the industrial park in the subsequent learning process, and improve the accuracy of pollution emission prediction for different enterprises.

[0073] Step S140, the target machine learning network model is subjected to model parameter learning based on the sample learning data, and the target machine learning network model that has completed the model parameter learning is used as a pollution emission prediction model. The pollution emission prediction model is used to perform pollution emission prediction on any input pollution emission data, and generate corresponding pollution emission description labels and label guidance knowledge data.

[0074] Take the previously generated sample learning data as an example, which contains various information after fusion. First, feature encode the sample learning data. Assume that the sample learning data is about the pollution emissions of a chemical enterprise. After feature encoding, X coded feature node data are generated. For example, the concentration, emission rate, emission time and other information of different pollutants in the waste gas of a chemical enterprise are encoded into a series of feature node data. If X=5, then these 5 coded feature node data may represent the sulfur dioxide concentration characteristics, nitrogen oxide emission rate characteristics, emission time period characteristics, meteorological conditions (such as wind direction) on emission characteristics, and enterprise production load and emission relationship characteristics.

[0075] Based on the time series information of these X coded feature node data, the features to be learned are generated based on the first X-1 coded feature node data. For example, for the above 5 coded feature node data, the features to be learned are constructed based on the first 4 (i.e., sulfur dioxide concentration features, nitrogen oxide emission rate features, emission time period features, and meteorological conditions on emission impact features). At the same time, based on the X-1 coded feature node data except the first coded feature node data, the expected output features corresponding to the features to be learned are generated. In other words, the expected output features are constructed based on the following 4 feature node data (nitrogen oxide emission rate features, emission time period features, meteorological conditions on emission impact features, and enterprise production load and emission relationship features).

[0076] Then, the target machine learning network model is used to derive the coded feature node data one by one for the feature to be learned. For the first coded feature node data (sulfur dioxide concentration feature), the model derives according to its own network structure and initial parameters to obtain a preliminary result. Then, based on this preliminary result and the next coded feature node data (nitrogen oxide emission rate feature), derivation is performed again to obtain the derivation result of the second coded feature node data. And so on, a derivation result is finally generated, which includes X-1 derived coded feature node data, and the ath coded feature node data is derived based on the first a coded feature node data in the feature to be learned (here a is a positive integer between 1 and X-1).

[0077] Based on the expected output characteristics and the derivation results, the error parameters corresponding to each coded feature node data in the derivation results are calculated. For example, for the second coded feature node data (nitrogen oxide emission rate feature) in the derivation result, it is compared with the corresponding data in the expected output characteristics, and the loss result between the two is calculated. This loss result is the error parameter corresponding to the coded feature node data.

[0078] The error parameters corresponding to each encoded feature node data in the derivation result are fused to generate the training error parameters of the target machine learning network model. This training error parameter comprehensively reflects the overall difference between the model derivation result and the expected output feature.

[0079] According to the training goal of minimizing the training error parameter, the neuron weight information of the target machine learning network model is updated. For example, if a neuron is found to have a large error in the derivation of the nitrogen oxide emission rate characteristics, the weight associated with this neuron is adjusted so that the model can deduce more accurately the next time it processes similar data. By continuously processing the sample learning data in this way, the parameters of the model are gradually adjusted, and finally the target machine learning network model that has completed the model parameter learning is used as the pollution emission prediction model. This pollution emission prediction model can predict pollution emissions for any pollution emission data input in the industrial park, and generate corresponding pollution emission description labels and label-guided knowledge data. For example, for the newly input waste gas emission data of a machinery manufacturing enterprise, the model can accurately predict that its pollution emission description label is "Medium-concentration dust and volatile organic compound emissions of machinery manufacturing enterprises-local environmental impact-stable type", and generate label-guided knowledge data containing the trend of air quality changes around the enterprise and the potential impact assessment on the health of nearby residents, providing strong decision-making support for the environmental management of the industrial park.

[0080] Among them, the pollution emission prediction model plays an important role in the daily environmental management of industrial parks. For example, when the environmental monitoring platform receives a prediction instruction for the target pollution emission data of a newly settled enterprise (assuming it is a small metal processing enterprise), it will generate candidate pollution emission data based on the previous model activation definition information and the target pollution emission data. The target pollution emission data contains information such as the metal dust concentration of 20 mg per cubic meter and the volatile organic compound concentration of 15 mg per cubic meter in the exhaust gas emitted from the chimney of the enterprise. These data are configured in the candidate pollution emission data as input pollution emission data.

[0081] Then, the trained pollution emission prediction model is used to predict the target pollution emission data based on the candidate pollution emission data. After complex calculations and analysis, the model generates a pollution emission description label for the target pollution emission data, which may be "medium-concentration composite pollutant emissions from small metal processing enterprises-local environmental risks-fluctuation type", as well as corresponding label-guided knowledge data, such as analysis of the impact of wind direction and wind speed changes in the area where the enterprise is located on the diffusion of pollutants, cross-contamination risk assessments that other surrounding enterprises may be subject to, and recommendations on environmental protection measures that the enterprise should take.

[0082] These prediction results can help the management departments of industrial parks to timely understand the pollution emissions of newly settled enterprises, formulate corresponding environmental management strategies in advance, ensure the overall environmental quality of industrial parks, and protect the health and safety of surrounding residents and the ecological environment. At the same time, as more enterprises' pollution emission data are continuously input into the model for prediction, the model can also continuously optimize and update itself, improving the accuracy and reliability of pollution emission predictions for various enterprises in the industrial park.

[0083] Based on the above steps, the embodiment of the present application significantly improves the accuracy and reliability of pollution emission prediction. The method first obtains target machine learning data containing rich prior knowledge, and combines the model activation definition information to generate sample learning data for model updating. By using these sample data to learn the parameters of the target machine learning network model, the resulting pollution emission prediction model can more accurately predict any input pollution emission data, and generate corresponding pollution emission description labels and label-guided knowledge data. This process not only optimizes the learning efficiency of the model, but also enhances the model's ability to recognize and predict complex pollution emission patterns, thereby facilitating environmental pollution monitoring, governance and decision-making.

[0084] In a possible implementation manner, the sample learning data is a pollution emission data generated by fusing the model enabling definition information and the target machine learning data. Step S140 includes:

[0085] Step S141, feature encoding is performed on the sample learning data to generate X encoded feature node data, where X is a positive integer.

[0086] In this embodiment, for example, there is a large chemical company in the industrial park, and its pollution emission data contains a variety of complex information. The sample learning data may integrate various information such as the concentration of various pollutants in the company's waste gas emissions, emission rate, emission time, and meteorological conditions. When encoding features, if X is 5, these 5 coded feature node data may correspond to different feature aspects. The first coded feature node data may be about the concentration characteristics of the main pollutant sulfur dioxide (SO2), which accurately records the SO2 concentration value emitted by the company in a specific time period; the second coded feature node data may be nitrogen oxides (NO x ) emission rate characteristics, which reflects in detail the NO x emissions; the third coded feature node data is the emission time period feature, which clarifies whether it is the emission during the daytime production peak period or the emission during the nighttime low-load period; the fourth coded feature node data is the influence of wind direction on emission diffusion in meteorological conditions, taking into account that different wind directions will lead to great differences in the diffusion path and range of pollutants in the park; the fifth coded feature node data is the relationship between enterprise production load and emissions, which reflects the relationship between the size of the enterprise's production task and the amount of pollutant emissions.

[0087] Step S142: Based on the time series information of the X coded feature node data, the feature to be learned is generated according to the first X-1 coded feature node data among the X coded feature node data. And based on the X-1 coded feature node data except the first coded feature node data among the X coded feature node data, the expected output feature corresponding to the feature to be learned is generated.

[0088] For the chemical enterprise example above, the feature to be learned is constructed based on the first four coded feature node data (SO2 concentration feature, NOx emission rate feature, emission time period feature, and wind direction on emission diffusion feature). This feature to be learned comprehensively covers the main pollution emission-related factors except for the relationship between production load and emission, forming a feature combination with specific time series logic. At the same time, based on the X-1 coded feature node data except for the first coded feature node data, the expected output feature corresponding to the feature to be learned is generated. That is, according to NO x The expected output feature is constructed by the emission rate feature, emission time period feature, wind direction effect on emission diffusion feature, and enterprise production load and emission relationship feature. This expected output feature is the result that the model should output under ideal conditions. It has a close logical connection with the feature to be learned and is determined based on a deep understanding of the pollution emission laws and influencing factors of chemical enterprises.

[0089] Step S143, deriving the coded feature node data of the feature to be learned one by one through the target machine learning network model to generate a derivation result. The derivation result includes derived X-1 coded feature node data, and the a-th coded feature node data in the derivation result is derived based on the first a coded feature node data in the feature to be learned, where a is a positive integer between 1 and X-1.

[0090] In this process, the target machine learning network model starts with the first coded feature node data of the feature to be learned. For the first coded feature node data (SO2 concentration feature), the target machine learning network model processes it according to its own existing network structure and initial parameters to obtain a preliminary derivation result. Then, combined with the second coded feature node data (NO x Emission rate characteristics), the model is deduced again to obtain the derivation result of the second coded feature node data based on the first two coded feature node data. In this way, the subsequent coded feature node data are combined in sequence to generate a derivation result. This derivation result contains X-1 derived coded feature node data, and the a-th coded feature node data in the derivation result is derived based on the first a coded feature node data in the feature to be learned. For example, when a is 3, the third coded feature node data in the derivation result is a combination of the first three coded feature node data in the feature to be learned (SO2 concentration characteristics, NO x It is obtained after considering the emission rate characteristics and emission time period characteristics, which reflects the model derivation results under the joint action of these three factors.

[0091] Step S144, based on the expected output feature and the derivation result, calculate the error parameters corresponding to each coding feature node data in the derivation result, the error parameter corresponding to the a-th coding feature node data in the derivation result represents: the loss result between the a-th coding feature node data in the derivation result and the a-th coding feature node data in the expected output feature.

[0092] Continuing with the chemical industry as an example, for the second coded feature node data (NO x Emission rate characteristics) and compare them with the corresponding data in the expected output characteristics. x The emission rate is 100 kg per hour, and the desired output characteristic is NO xThe emission rate is 120 kg per hour, so the difference between the two (20 kg per hour) is the error parameter corresponding to the coded feature node data. This error parameter represents the loss result between the coded feature node data in the derivation result and the corresponding data in the expected output feature, which intuitively reflects the derivation accuracy of the model on this feature node.

[0093] Step S145, fusing the error parameters corresponding to each encoded feature node data in the derivation result to generate the training error parameters of the target machine learning network model.

[0094] For the error parameters corresponding to the five coded feature node data of the chemical enterprise mentioned above, they are combined through a specific fusion method (such as weighted summation or root mean square calculation, etc.). This fusion method comprehensively considers the error conditions of each coded feature node data and forms a training error parameter that fully reflects the overall difference between the model derivation result and the expected output feature.

[0095] Step S146, based on the training goal of minimizing the training error parameters, the neuron weight information of the target machine learning network model is updated, thereby learning the model parameters of the target machine learning network model.

[0096] For example, in the chemical industry example, if a neuron is found to be responding to NO x A large error in the derivation of the emission rate signature indicates that the weight of this neuron may need to be adjusted. By analyzing the training error parameters, the model can determine which neurons contribute more to the error and adjust the weights of those neurons. For example, if a neuron is associated with NO x If the associated weight of the emission rate feature is too large, causing the derivation result of the model on this feature to deviate greatly from the expected output feature, then the weight of the neuron should be appropriately reduced. By continuously adjusting the neuron weight information in this way, the target machine learning network model can gradually improve its processing capabilities for sample learning data, thereby improving the accuracy of pollution emission predictions for enterprises in the industrial park. In the environmental management system of the entire industrial park, this model parameter learning process based on sample learning data helps to build an accurate and reliable pollution emission prediction model, providing strong technical support for the park's environmental supervision, enterprise pollution control, and overall environmental protection.

[0097] In a possible implementation, step S130 includes:

[0098] Step S131, fusing the model activation definition information and the template pollution emission data in the target machine learning data into one pollution emission data to generate first pollution emission data.

[0099] Step S132: The prior pollution emission description labels in the target machine learning data and the label-guided knowledge data in the target machine learning data are fused and output as one pollution emission data to generate second pollution emission data.

[0100] Step S133: Generate sample learning data for updating the target machine learning network model based on the first pollution emission data and the second pollution emission data.

[0101] In a possible implementation, step S140 may further include:

[0102] The target machine learning network model performs pollution emission prediction on the template pollution emission data in the target machine learning data based on the first pollution emission data in the sample learning data to generate a pollution emission prediction result. The pollution emission prediction result includes: a pollution emission description label of the corresponding template pollution emission data generated, and corresponding label-guided knowledge data.

[0103] Based on the loss result between the pollution emission prediction result and the second pollution emission data in the sample learning data, the neuron weight information of the target machine learning network model is updated, thereby learning the model parameters of the target machine learning network model.

[0104] In this example, a large steel company in an industrial park is used as an example. The model activation definition information includes the rules and requirements for processing various types of pollution emission data, such as the weight distribution principle for different pollutants, the importance of different emission time periods, and the calculation method of the impact of meteorological conditions on pollution diffusion. The template pollution emission data of the steel company contains rich information, such as the concentration of particulate matter (PM) emitted from its chimney is 80 mg per cubic meter, the concentration of sulfur dioxide (SO2) is 150 mg per cubic meter, and the concentration of nitrogen oxides (NO x ) concentration is 200 mg per cubic meter, and there are also relevant data such as emission flow, temperature, etc. In the fusion process, according to the weight allocation principle in the model activation definition information, for example, the weight of SO2 concentration in the overall pollution assessment is 0.3, and NO x The concentration weight is 0.4, the PM concentration weight is 0.2, and the other factors weight is 0.1. These weights are calculated and integrated with the actual pollution emission data of steel enterprises. This integration is not a simple addition of values, but is based on pre-set complex calculation rules, taking into full account the relationship between various factors and the comprehensive effects on the environment, and finally generating the first pollution emission data. This first pollution emission data is a standardized data form that integrates the model requirements and the actual pollution emission characteristics, providing a kind of input data that conforms to the logical structure of the model for subsequent model learning.

[0105] Next, the prior pollution emission description labels in the target machine learning data and the label-guided knowledge data in the target machine learning data are fused and output as one pollution emission data to generate the second pollution emission data. For the above-mentioned steel enterprise, its prior pollution emission description label may be "large steel enterprise high-concentration complex pollution emissions-wide environmental impact-stable type". This label not only indicates the scale of the enterprise and the concentration type of pollutants emitted, but also implies the scope of environmental impact and the stability of emissions. The corresponding label-guided knowledge data contains a lot of information, such as the relationship between the production process of steel enterprises and the generation of pollution, that is, which links in the iron and steelmaking process will produce large amounts of SO2 and NO x and PM; the state of the environment around the enterprise, such as what sensitive areas are there in the surrounding area (such as residential areas, farmland, water sources, etc.) and the relative position of these areas to the enterprise; and the pollution control measures that the enterprise has taken and their effect evaluation, such as the operating efficiency of the installed desulfurization, denitrification, and dust removal equipment and the actual contribution to pollutant emission reduction. When integrating the prior pollution emission description labels and label-guided knowledge data, it is necessary to organically integrate the semantic information in the labels with the specific content in the knowledge data. For example, combining "widespread environmental impact" with the distribution of surrounding sensitive areas, and linking "stable" emissions with the stable operation of the company's existing pollution control equipment, in this way, a second pollution emission data containing rich semantics and actual conditions is constructed.

[0106] Then, based on the first pollution emission data and the second pollution emission data, sample learning data for updating the target machine learning network model is generated. The sample learning data is a special data structure that combines the first pollution emission data and the second pollution emission data. It contains both the actual data characteristics of pollution emissions after being processed by the model requirements (first pollution emission data) and the prior pollution emission description labels and related knowledge data (second pollution emission data). This structure enables the sample learning data to provide comprehensive learning information for the target machine learning network model, allowing the model to understand the actual numerical situation of pollution emissions, and the meaning behind these numerical values ​​and related environmental influencing factors, thereby laying the foundation for accurate learning and parameter updating of the model.

[0107] In terms of learning model parameters of the target machine learning network model based on the sample learning data, the target machine learning network model predicts the pollution emission of the template pollution emission data in the target machine learning data based on the first pollution emission data in the sample learning data, and generates a pollution emission prediction result. Still taking the steel enterprise as an example, the target machine learning network model takes the first pollution emission data as input, and this first pollution emission data contains the pollution emission characteristics of the steel enterprise after fusion and the relevant information required by the model. The target machine learning network model analyzes and predicts the template pollution emission data of the steel enterprise based on its network structure and existing parameters. For example, when the model processes the pollutant concentration, flow rate, temperature and other data of the steel enterprise, it combines the pollution emission laws and related knowledge learned by itself to generate a pollution emission prediction result. This pollution emission prediction result includes the pollution emission description label of the corresponding template pollution emission data generated, and the corresponding label-guided knowledge data. Suppose the generated pollution emission description label is "high-concentration complex pollution emissions from large steel enterprises - serious environmental impact - stable type". The "serious environmental impact" here is a more accurate environmental impact assessment based on the model's analysis of pollutant concentrations and surrounding environmental sensitive areas; the corresponding label-guided knowledge data may contain more detailed predictions of surrounding environmental changes, such as the deterioration trend of surrounding air quality in the future under the current pollution emission level, the potential increase in pollution risks to water quality in nearby water sources, and the increase in health risks to surrounding residents.

[0108] Finally, based on the loss result between the pollution emission prediction result and the second pollution emission data in the sample learning data, the neuron weight information of the target machine learning network model is updated, so as to learn the model parameters of the target machine learning network model. The pollution emission prediction results of the above steel enterprises are compared with the second pollution emission data, and the loss results between the two are calculated. For example, in terms of pollution emission description labels, there is a semantic difference between "serious environmental impact" and "extensive environmental impact" in the second pollution emission data, which reflects the deviation between the model prediction and prior knowledge; in terms of label-guided knowledge data, the predicted trend of deterioration of surrounding air quality, the increase in potential pollution risk of water quality in water sources, etc., may also have different degrees of deviation from the actual situation in the second pollution emission data (such as the results obtained based on historical data and actual monitoring). These deviations are quantified and calculated through a specific loss function to obtain an overall loss result. According to this loss result, analyze which neurons in the model contribute more to the deviation in the prediction process. For example, if it is found that the neurons related to pollutant concentration analysis play a key role in predicting "serious environmental impact" and their prediction results deviate greatly from the actual situation, then the weight information of these neurons is adjusted. By continuously updating the neuron weight information according to this loss result, the target machine learning network model can gradually optimize its own parameters and improve the accuracy of pollution emission predictions for enterprises in the industrial park, thereby better adapting to the complex environmental management needs of the industrial park and providing strong technical support for the environmental protection and sustainable development of the park.

[0109] In a possible implementation, the target machine learning data is machine learning data obtained from a training data sequence, and the step of generating the training data sequence includes:

[0110] Step A110, obtaining multiple machine learning data of the target machine learning network model, wherein one machine learning data includes: one template pollution emission data and training supervision data of the corresponding template pollution emission data. The training supervision data of any template pollution emission data includes a priori pollution emission description label of the corresponding template pollution emission data and corresponding label-guided knowledge data, and the priori pollution emission description label in any machine learning data is selected from the multiple pollution emission description labels.

[0111] In this embodiment, in the industrial park, different enterprises have different pollution emission conditions, which are collected and organized into machine learning data. A machine learning data contains a template pollution emission data and corresponding training supervision data. For example, for a chemical enterprise, its template pollution emission data may be information on various pollutants emitted from chimneys during a specific time period, such as sulfur dioxide (SO2) concentration of 50 mg per cubic meter, nitrogen oxides (NOx ) concentration is 80 mg per cubic meter, particulate matter (PM) concentration is 30 mg per cubic meter, etc. The training supervision data of the corresponding template pollution emission data contains prior pollution emission description labels and label-guided knowledge data. The prior pollution emission description label may be "Medium-concentration complex pollution emissions from chemical enterprises-local environmental impact-fluctuation type". This label describes in detail the type of enterprise, pollution emission concentration level, scope of environmental impact, and stability of emissions. Label-guided knowledge data may include links related to pollution generation in the production process of chemical enterprises, such as specific chemical reactions that produce a large amount of SO2, as well as relevant information about the surrounding environment, such as nearby farmland and the impact of wind direction on the diffusion direction of pollutants. The prior pollution emission description label in each machine learning data is selected from multiple pollution emission description labels, which cover the classification description of various possible pollution emission situations in the industrial park.

[0112] Step A120, determining the number of templates of each of the multiple pollution emission description labels based on the multiple machine learning data. The number of templates of any one pollution emission description label is used to represent: the number of machine learning data containing any one of the pollution emission description labels in the multiple machine learning data.

[0113] After collecting machine learning data from many enterprises in the industrial park, these machine learning data were analyzed and counted. For example, there are many pollution emission description labels such as "high-concentration complex pollution emissions from chemical enterprises-wide environmental impact-stable type" and "medium-concentration dust and volatile organic compound emissions from machinery manufacturing enterprises-local environmental impact-stable type". Statistics show that there are 50 groups of machine learning data marked as "high-concentration complex pollution emissions from chemical enterprises-wide environmental impact-stable type", which is the number of templates for this pollution emission description label; and there are only 10 groups of machine learning data marked as "low-concentration dust emissions from machinery manufacturing enterprises-limited environmental impact-fluctuation type", which is the number of corresponding templates. The number of templates for each pollution emission description label indicates the number of machine learning data that contain this pollution emission description label in multiple machine learning data.

[0114] Step A130: Based on the number of templates of each pollution emission description label, determine an edge pollution emission description label from the plurality of pollution emission description labels. The edge pollution emission description label is used to represent: the pollution emission description label corresponding to the number of templates less than the set number.

[0115] Assuming the set number is 20, the pollution emission description labels with a template number less than 20 are determined to be marginal pollution emission description labels. For example, the template number of "Low-concentration dust emissions from machinery manufacturing enterprises-limited environmental impact-fluctuation type" is 10, which is less than the set number of 20, so it is a marginal pollution emission description label. There are relatively few machine learning data corresponding to these marginal pollution emission description labels, probably because this type of pollution emission situation is relatively special or rare in industrial parks.

[0116] Step A140, extracting one or more template pollution emission data from the template pollution emission data included in the machine learning data containing the edge pollution emission description label.

[0117] For the marginal pollution emission description label "Low-concentration dust emission from machinery manufacturing enterprises - limited environmental impact - volatile type", the template pollution emission data is extracted from the corresponding machine learning data. For example, the dust emission data of a small machinery manufacturing enterprise in a specific time period is extracted from the machine learning data, including data such as dust concentration of 15 mg per cubic meter and emission rate of 2 kg per hour.

[0118] Step A150, feature derivation is performed on each selected template pollution emission data to generate Y iterative pollution emission data, where Y is a positive integer. One iterative pollution emission data corresponds to one template pollution emission data, and any iterative pollution emission data has a consistent pollution feature path with the corresponding template pollution emission data.

[0119] Take the dust emission data extracted from machinery manufacturing enterprises as an example to perform feature derivation. Assuming Y is 3, for the dust emission data, through a specific feature derivation algorithm, keep its pollution feature path unchanged, that is, it is still a feature derivation related to dust emission. Three iterative pollution emission data may be generated based on some relevant factors in the production process of the enterprise, such as equipment operation time, equipment cleaning frequency, etc. The first iterative pollution emission data may be the dust emission situation after the equipment operation time is extended, such as the dust concentration becomes 18 mg per cubic meter and the emission rate becomes 2.2 kg per hour; the second iterative pollution emission data may be the situation after the equipment cleaning frequency is reduced, the dust concentration becomes 20 mg per cubic meter, and the emission rate becomes 2.5 kg per hour; the third iterative pollution emission data may be the situation after the comprehensive equipment operation time is extended and the cleaning frequency is reduced, the dust concentration becomes 22 mg per cubic meter, and the emission rate becomes 2.8 kg per hour.

[0120] Step A160, using the training supervision data of the template pollution emission data corresponding to each iterative pollution emission data as the training supervision data of the corresponding iterative pollution emission data.

[0121] For the three iterative pollution emission data of the above-mentioned mechanical manufacturing enterprise dust emission data, since the training supervision data of the original template pollution emission data contains the prior pollution emission description label "low-concentration dust emissions from mechanical manufacturing enterprises-limited environmental impact-fluctuating type" and related label-guided knowledge data (such as the distribution of residential areas in the surrounding environment of the enterprise, the impact of wind direction on dust diffusion, etc.), these training supervision data are also assigned to the corresponding iterative pollution emission data.

[0122] Step A170, taking each of the iterative pollution emission data as iterative template pollution emission data, and generating Y iterative machine learning data based on the Y iterative pollution emission data and the corresponding training supervision data.

[0123] For example, for the three generated iterative pollution emission data and their corresponding training supervision data, three iterative machine learning data are constructed respectively. Each iterative machine learning data contains an iterative template pollution emission data and corresponding training supervision data, which can be understood as the same as the original machine learning data, except that the data here is an iterative version derived from features.

[0124] Step A180, generating the training data sequence based on the multiple machine learning data and the Y iterative machine learning data.

[0125] For example, all the previously collected machine learning data about enterprises in the industrial park and the newly generated Y iterative machine learning data are integrated together to form a complete training data sequence. This training data sequence contains the pollution emission data of various enterprises in the industrial park, prior pollution emission description labels, label-guided knowledge data, and iterative data after feature derivation, providing rich and comprehensive data resources for the target machine learning network model, enabling the model to learn the characteristics of various pollution emission conditions, thereby improving the accuracy and reliability of pollution emission predictions in the industrial park, and helping the environmental management department of the industrial park to better supervise the pollution emission behavior of enterprises and protect the environmental quality of the industrial park.

[0126] In a possible implementation, step A150 includes:

[0127] Step A151, obtaining a feature derivation request of a feature derivation network, wherein the feature derivation request indicates: based on input template pollution emission data and a priori pollution emission description label, generating iterative pollution emission data with a consistent pollution feature path for the template pollution emission data.

[0128] In this embodiment, in the industrial park, this feature derivation request has a clear representational meaning, that is, based on the input template pollution emission data and the prior pollution emission description label, it generates iterative pollution emission data with consistent pollution feature path for the template pollution emission data. Taking a chemical enterprise in the industrial park as an example, its template pollution emission data contains various pollutant emission information within a specific time period, such as sulfur dioxide (SO2) concentration of 50 mg per cubic meter, nitrogen oxides (NO x ) concentration is 80 mg per cubic meter, particulate matter (PM) concentration is 30 mg per cubic meter, etc., and the prior pollution emission description label is "medium-concentration complex pollution emission of chemical enterprises-local environmental impact-fluctuation type". The feature derivation request is to generate new pollution emission data based on such template pollution emission data and prior pollution emission description labels, and the new data must maintain the same pollution feature path as the original data, which means that the new data is similar to the original data in terms of pollutant types, related logical relationships of emissions, etc.

[0129] Step A152, obtaining one or more derived sample combination data, any one of which includes a sample pollution emission data and a priori pollution emission description label of the corresponding sample pollution emission data, and pollution emission data having a consistent pollution characteristic path with the corresponding sample pollution emission data.

[0130] For example, for another chemical enterprise (as a sample enterprise), its sample pollution emission data is sulfur dioxide (SO2) concentration of 45 mg per cubic meter, nitrogen oxides (NO x ) concentration is 75 mg per cubic meter, and the particulate matter (PM) concentration is 25 mg per cubic meter. The prior pollution emission description label is "medium-concentration complex pollution emission of chemical enterprises-local environmental impact-fluctuation type". The pollution emission data with the same pollution characteristic path as the sample pollution emission data may be emission data under different time periods but similar production processes and equipment operating conditions, such as sulfur dioxide (SO2) concentration of 48 mg per cubic meter, nitrogen oxides (NO x ) concentration is 78 mg per cubic meter, and the particulate matter (PM) concentration is 28 mg per cubic meter. These derived sample combination data cover the emissions of different chemical companies under similar pollution characteristic pathways, providing more reference for subsequent characteristic derivation.

[0131] Step A153, based on the feature derivation request and the one or more derived sample combination data, the feature derivation network generates iterative pollution emission data with consistent pollution feature paths for each selected template pollution emission data, thereby generating Y iterative pollution emission data.

[0132] In a possible implementation, step A153 may include:

[0133] Poll each selected template pollution emission data, and use the currently polled template pollution emission data as the current template pollution emission data.

[0134] An input pollution emission data segment is generated according to the current template pollution emission data and the corresponding prior pollution emission description label.

[0135] The feature derivation request, the one or more derived sample combination data, and the generated input pollution emission data fragment are integrated to generate derived learning data corresponding to the current template pollution emission data.

[0136] Through the feature derivation network, by learning the derived learning data corresponding to the current template pollution emission data, iterative pollution emission data having a consistent pollution feature path with the current template pollution emission data is generated.

[0137] Continue polling until all selected template pollution emission data are polled, generating Y iterative pollution emission data.

[0138] Continuing with the previously mentioned chemical enterprise template pollution emission data as an example, during the polling process, when the chemical enterprise's template pollution emission data comes, it becomes the current template pollution emission data. Based on the current template pollution emission data and the corresponding prior pollution emission description label, an input pollution emission data segment is generated. For this chemical enterprise, based on its template pollution emission data (SO2 concentration is 50 mg per cubic meter, NO x The concentration is 80 mg per cubic meter, the PM concentration is 30 mg per cubic meter) and the prior pollution emission description label ("medium-concentration composite pollution emission of chemical enterprises-local environmental impact-fluctuation type"), and through specific algorithms and rules, an input pollution emission data segment is generated. This data segment may be a certain feature extraction and arrangement of the original template pollution emission data, for example, focusing on extracting features related to pollution emission concentration, forming an input pollution emission data segment containing information such as the ratio relationship of specific pollutant concentrations and the total pollution concentration level.

[0139] The feature derivation request, one or more derived sample combination data, and the generated input pollution emission data fragment are integrated to generate the derived learning data corresponding to the current template pollution emission data. The previously obtained feature derivation request (generating iterative pollution emission data with consistent pollution feature path based on the template pollution emission data of the chemical enterprise and the prior pollution emission description label), the derived sample combination data (such as the sample pollution emission data of other chemical enterprises and related information) and the newly generated input pollution emission data fragment are integrated. This integration process is a complex data integration operation, which requires combining these data together according to a specific format and logic. For example, the target requirements in the feature derivation request, the emission data and label information of different chemical enterprises in the derived sample combination data, and the concentration characteristics in the input pollution emission data fragment are arranged and combined according to the predefined data structure to generate the derived learning data corresponding to the current template pollution emission data. This derived learning data contains rich information, including both target requirements and reference samples and its own feature data fragments, providing comprehensive learning materials for the feature derivation network.

[0140] Through the feature derivation network, by learning the derived learning data corresponding to the current template pollution emission data, iterative pollution emission data with a pollution characteristic path consistent with the current template pollution emission data is generated. The feature derivation network learns and analyzes according to various information in the derived learning data. For example, according to the pollution characteristic path consistency requirements in the feature derivation request, it can refer to the emissions of similar enterprises in the derived sample combination data, and combine the concentration characteristics and other information in the input pollution emission data fragment, and use its internal algorithm and model structure to generate iterative pollution emission data with a pollution characteristic path consistent with the current template pollution emission data. Assume that the generated iterative pollution emission data is a sulfur dioxide (SO2) concentration of 52 mg per cubic meter and nitrogen oxides (NO x ) concentration is 82 mg per cubic meter, and the particulate matter (PM) concentration is 32 mg per cubic meter. This data is consistent with the original template pollution emission data in terms of pollutant types and emission logic relationships, but has changed in specific pollutant concentrations, reflecting the derivation of pollution emission data under specific conditions. Continue polling until all selected template pollution emission data are polled and Y iterative pollution emission data are generated. For example, if Y = 3 iterative pollution emission data need to be generated, then perform the same operation on other selected template pollution emission data (which may be from different enterprises or the same enterprise in different time periods) in the same manner as above, until 3 such iterative pollution emission data are generated.

[0141] In a possible implementation, step A153 may further include:

[0142] The feature derivation request and the one or more derived sample combination data are fused to generate a training sample for the feature derivation network.

[0143] The selected template pollution emission data and the corresponding prior pollution emission description labels are respectively used to generate input pollution emission data segments corresponding to the template pollution emission data.

[0144] The feature derivation network is trained based on the training samples to generate a trained feature derivation network.

[0145] The trained feature derivation network is used to generate iterative pollution emission data with consistent pollution feature paths for the corresponding template pollution emission data based on the input pollution emission data segments corresponding to the respective template pollution emission data, thereby generating Y iterative pollution emission data.

[0146] In another possible implementation, the feature derivation request (iterative pollution emission data with consistent pollution feature paths generated based on template pollution emission data and prior pollution emission description labels) and the derived sample combination data (sample pollution emission data of multiple chemical companies and related information) are fused. This fusion process needs to consider how to effectively combine the target information in the feature derivation request with the actual emission data and label information in the derived sample combination data. For example, the target requirements in the feature derivation request can be converted into specific labeling information, and then arranged and combined in a certain order with the pollution emission data of each company in the derived sample combination data, prior pollution emission description labels, etc., to form a training sample for the feature derivation network. This training sample is a data set that combines target requirements and actual reference data, which provides a basis for the training of the feature derivation network.

[0147] The selected template pollution emission data and the corresponding prior pollution emission description labels are respectively used to generate the input pollution emission data segments corresponding to the template pollution emission data. For each selected template pollution emission data (such as the template pollution emission data of the chemical enterprise mentioned above and the template pollution emission data of other enterprises), combined with its corresponding prior pollution emission description label, an input pollution emission data segment is generated according to a specific algorithm. This process is similar to the previous one, which is to extract and organize the features of the template pollution emission data to form an input pollution emission data segment containing key pollution emission feature information.

[0148] The feature derivation network is trained based on the training samples to generate a trained feature derivation network. The generated training samples are input into the feature derivation network, and the feature derivation network adjusts its own parameters according to the information in the training samples. For example, according to the sample pollution emission data of different chemical companies in the training samples, the prior pollution emission description labels, and the target requirements in the feature derivation request, the feature derivation network continuously adjusts the connection weights and other parameters between neurons through the internal algorithm mechanism, so that the network can accurately generate iterative pollution emission data with consistent pollution feature paths based on the input template pollution emission data and its prior pollution emission description labels. After multiple training iterations, the feature derivation network gradually optimizes its own performance and finally generates a trained feature derivation network.

[0149] The trained feature derivation network is used to generate iterative pollution emission data with consistent pollution feature paths for the corresponding template pollution emission data based on the input pollution emission data fragments corresponding to each template pollution emission data, and Y iterative pollution emission data are generated. For each template pollution emission data, the corresponding input pollution emission data fragment is input into the trained feature derivation network. For example, for the input pollution emission data fragment corresponding to the template pollution emission data of a chemical enterprise, the trained feature derivation network generates iterative pollution emission data with consistent pollution feature paths with the template pollution emission data of the chemical enterprise according to the feature information in this fragment, combined with the parameters and model structure obtained by its own training. In this way, all template pollution emission data that need to generate iterative pollution emission data are operated until Y iterative pollution emission data are generated. These iterative pollution emission data are of great significance in the environmental management of industrial parks. They can enrich the types and changes of pollution emission data, provide more diverse data for the target machine learning network model, thereby improving the accuracy and comprehensiveness of the model's prediction of pollution emissions in industrial parks, and help to more accurately supervise the pollution emission behavior of enterprises and protect the environmental quality of industrial parks.

[0150] In a possible implementation, the method further includes:

[0151] Step S150, after receiving the prediction instruction for the target pollution emission data, generating candidate pollution emission data according to the model activation definition information and the target pollution emission data, and the target pollution emission data is configured in the candidate pollution emission data as input pollution emission data.

[0152] Step S160 , performing pollution emission prediction on the target pollution emission data based on the candidate pollution emission data through the pollution emission prediction model, and generating pollution emission description labels of the target pollution emission data and corresponding label guidance knowledge data.

[0153] In this embodiment, for example, there is a newly settled electronic enterprise in the industrial park, and its target pollution emission data includes various information in the waste gas emission, such as fluoride concentration of 10 mg per cubic meter, ammonia concentration of 5 mg per cubic meter, etc., as well as relevant data such as emission flow and temperature. The model activation definition information covers the rules and requirements for processing various types of pollution emission data, such as the weight allocation principle of different pollutants, the importance of different emission time periods, and the calculation method of the impact of meteorological conditions on pollution diffusion. When generating candidate pollution emission data, the target pollution emission data of the electronic enterprise are integrated according to the rules in the model activation definition information. For example, according to the weight allocation principle in the model activation definition information, the weight of fluoride concentration in the overall pollution assessment is 0.3, the weight of ammonia concentration is 0.2, and the weight of other factors is 0.5. These weights are calculated and fused with the actual pollution emission data of the electronic enterprise. This fusion is not a simple numerical addition, but is based on pre-set complex calculation rules, taking into full account the relationship between various factors and the comprehensive effect of environmental impact, and finally generating candidate pollution emission data. This candidate pollution emission data is a normalized data form that integrates the model requirements and actual pollution emission characteristics, providing input data that conforms to the logical structure of the model for subsequent pollution emission predictions.

[0154] Next, the pollution emission prediction model is used to predict the target pollution emission data based on the candidate pollution emission data, and the pollution emission description label of the target pollution emission data and the corresponding label-guided knowledge data are generated. The pollution emission prediction model takes the candidate pollution emission data as input, which contains the fused pollution emission characteristics of the electronic enterprise and the relevant information required by the model. The model analyzes and predicts the target pollution emission data of the electronic enterprise based on its own network structure and existing parameters. For example, when processing the pollutant concentration, flow rate, temperature and other data of the electronic enterprise, the model combines the pollution emission laws and related knowledge learned by itself to generate pollution emission prediction results. This pollution emission prediction result includes the pollution emission description label of the corresponding target pollution emission data generated, as well as the corresponding label-guided knowledge data. Assume that the generated pollution emission description label is "Electronic enterprise low concentration special gas emission-specific regional impact-stable type". Here, "specific regional impact" is the environmental impact assessment based on the analysis of pollutant concentration and surrounding environmental sensitive areas (such as nearby high-precision electronic equipment production workshops that are more sensitive to these gases); the corresponding label-guided knowledge data may contain more detailed predictions of surrounding environmental changes, such as the trend of surrounding air quality changes in the future under the current pollution emission level, the potential impact risk on the product quality of nearby high-precision electronic equipment production workshops, and suggestions on measures that enterprises can take to reduce pollution emissions. These prediction results can provide valuable information for the management department of the industrial park, helping it to timely understand the pollution emissions of newly settled enterprises, so as to formulate corresponding environmental management strategies in advance, ensure the overall environmental quality of the industrial park, and protect the health and safety of surrounding residents and the ecological environment.

[0155] Figure 2 The hardware structure of the pollution emission prediction system 100 based on machine learning for implementing the above-mentioned pollution emission prediction method based on machine learning provided by an embodiment of the present invention is shown as follows: Figure 2 As shown, the machine learning-based pollution emission prediction system 100 may include a processor 110 , a machine-readable storage medium 120 , a bus 130 , and a communication unit 140 .

[0156] The machine-readable storage medium 120 may store data and / or instructions. In some embodiments, the machine-readable storage medium 120 may store data acquired from an external terminal. In some embodiments, the machine-readable storage medium 120 may store data and / or instructions that the machine-learning-based pollution emission prediction system 100 uses to execute or use to complete the exemplary method described in the present invention.

[0157] During the specific implementation process, one or more processors 110 execute computer executable instructions stored in the machine-readable storage medium 120, so that the processor 110 can execute the pollution emission prediction method based on machine learning in the above method embodiment. The processor 110, the machine-readable storage medium 120 and the communication unit 140 are connected through the bus 130, and the processor 110 can be used to control the sending and receiving actions of the communication unit 140.

[0158] The specific implementation process of the processor 110 can refer to the various method embodiments executed by the above-mentioned machine learning-based pollution emission prediction system 100. The implementation principles and technical effects are similar, and this embodiment will not be repeated here.

[0159] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer executable instructions are preset. When a processor executes the computer executable instructions, the above-mentioned pollution emission prediction method based on machine learning is implemented.

[0160] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, various features are sometimes combined into one embodiment, drawing or description thereof.

Claims

1. A method for predicting pollution emissions based on machine learning, characterized in that: The method comprises: Obtaining target machine learning data of a target machine learning network model, wherein the target machine learning data includes: template pollution emission data, a priori pollution emission description label of the corresponding template pollution emission data, and corresponding label-guided knowledge data, wherein the priori pollution emission description label of the template pollution emission data is selected from a plurality of pollution emission description labels previously defined; Obtaining model activation definition information of the target machine learning network model; the model activation definition information represents: performing pollution emission prediction on input pollution emission data based on the multiple pollution emission description labels, and generating pollution emission description labels of the input pollution emission data and corresponding label guidance knowledge data; Generate sample learning data for updating the target machine learning network model according to the model activation definition information and the target machine learning data, wherein the template pollution emission data in the target machine learning data is configured as input pollution emission data in the sample learning data; The target machine learning network model is subjected to model parameter learning based on the sample learning data, and the target machine learning network model that has completed the model parameter learning is used as a pollution emission prediction model; the pollution emission prediction model is used to perform pollution emission prediction for any input pollution emission data, and generate corresponding pollution emission description labels and label-guided knowledge data.

2. The method for predicting pollution emissions based on machine learning according to claim 1, characterized in that: The sample learning data is a pollution emission data generated by fusing the model activation definition information and the target machine learning data; and the model parameter learning of the target machine learning network model based on the sample learning data includes: Performing feature encoding on the sample learning data to generate X encoded feature node data, where X is a positive integer; Based on the time series information of the X encoding feature node data, generating a feature to be learned according to the first X-1 encoding feature node data among the X encoding feature node data; and generating an expected output feature corresponding to the feature to be learned according to the X-1 encoding feature node data except the first encoding feature node data among the X encoding feature node data; The target machine learning network model is used to derive the coded feature node data of the feature to be learned one by one to generate a derivation result; the derivation result includes X-1 derived coded feature node data, and the a-th coded feature node data in the derivation result is derived based on the first a coded feature node data in the feature to be learned, and a is a positive integer between 1 and X-1; Based on the expected output feature and the derivation result, the error parameters corresponding to each coding feature node data in the derivation result are calculated, and the error parameter corresponding to the a-th coding feature node data in the derivation result represents: the loss result between the a-th coding feature node data in the derivation result and the a-th coding feature node data in the expected output feature; Fusion of the error parameters corresponding to the respective encoding feature node data in the derivation results to generate training error parameters of the target machine learning network model; According to the training goal of minimizing the training error parameters, the neuron weight information of the target machine learning network model is updated, thereby learning the model parameters of the target machine learning network model.

3. The method for predicting pollution emissions based on machine learning according to claim 1, characterized in that: The step of generating sample learning data for updating the target machine learning network model according to the model activation definition information and the target machine learning data includes: fusing the model enabling definition information and the template pollution emission data in the target machine learning data into one pollution emission data to generate first pollution emission data; The prior pollution emission description label in the target machine learning data and the label-guided knowledge data in the target machine learning data are fused and outputted into one pollution emission data to generate second pollution emission data; Based on the first pollution emission data and the second pollution emission data, sample learning data for updating the target machine learning network model is generated.

4. The method for predicting pollution emissions based on machine learning according to claim 3, characterized in that: The performing model parameter learning on the target machine learning network model according to the sample learning data includes: Based on the first pollution emission data in the sample learning data, the target machine learning network model performs pollution emission prediction on the template pollution emission data in the target machine learning data to generate a pollution emission prediction result; the pollution emission prediction result includes: a pollution emission description label of the generated corresponding template pollution emission data, and corresponding label-guided knowledge data; Based on the loss result between the pollution emission prediction result and the second pollution emission data in the sample learning data, the neuron weight information of the target machine learning network model is updated, thereby learning the model parameters of the target machine learning network model.

5. The method for predicting pollution emissions based on machine learning according to claim 1, characterized in that: The target machine learning data is machine learning data obtained from a training data sequence, and the steps of generating the training data sequence include: Acquire multiple machine learning data of the target machine learning network model, wherein one machine learning data includes: a template pollution emission data and training supervision data of the corresponding template pollution emission data; the training supervision data of any template pollution emission data includes a priori pollution emission description label of the corresponding template pollution emission data and corresponding label-guided knowledge data, and the priori pollution emission description label in any machine learning data is selected from the multiple pollution emission description labels; Determine the number of templates of each pollution emission description label in the multiple pollution emission description labels based on the multiple machine learning data; the number of templates of any one pollution emission description label is used to represent: the number of machine learning data containing any one pollution emission description label in the multiple machine learning data; Based on the number of templates of each pollution emission description label, determining an edge pollution emission description label from the multiple pollution emission description labels; the edge pollution emission description label is used to represent: the pollution emission description label corresponding to the number of templates less than the set number; Extract one or more template pollution emission data from the template pollution emission data included in the machine learning data containing the edge pollution emission description label; Derivation of features is performed on each selected template pollution emission data to generate Y iterative pollution emission data, where Y is a positive integer; one iterative pollution emission data corresponds to one template pollution emission data, and any iterative pollution emission data has a consistent pollution feature path with the corresponding template pollution emission data; Using the training supervision data of the template pollution emission data corresponding to each iterative pollution emission data as the training supervision data of the corresponding iterative pollution emission data; Each of the iterative pollution emission data is used as iterative template pollution emission data, and Y iterative machine learning data is generated according to the Y iterative pollution emission data and the corresponding training supervision data; The training data sequence is generated according to the multiple machine learning data and the Y iterative machine learning data.

6. The method for predicting pollution emissions based on machine learning according to claim 5, characterized in that: The step of performing feature derivation on each selected template pollution emission data to generate Y iterative pollution emission data includes: Obtaining a feature derivation request of a feature derivation network, wherein the feature derivation request indicates: generating iterative pollution emission data having a consistent pollution feature path for the template pollution emission data based on input template pollution emission data and a priori pollution emission description label; Obtain one or more derived sample combination data, any one of the derived sample combination data includes a sample pollution emission data and a priori pollution emission description label of the corresponding sample pollution emission data and pollution emission data having a consistent pollution characteristic path with the corresponding sample pollution emission data; Based on the feature derivation request and the one or more derived sample combination data, the feature derivation network generates iterative pollution emission data with consistent pollution feature paths for each selected template pollution emission data, thereby generating Y iterative pollution emission data.

7. The method for predicting pollution emissions based on machine learning according to claim 6, characterized in that: The generating of iterative pollution emission data with consistent pollution feature paths for each selected template pollution emission data by the feature derivation network based on the feature derivation request and the one or more derived sample combination data, and generating Y iterative pollution emission data, comprises: Polling each selected template pollution emission data, and taking the currently polled template pollution emission data as the current template pollution emission data; generating an input pollution emission data segment according to the current template pollution emission data and the corresponding prior pollution emission description label; Integrate the feature derivation request, the one or more derived sample combination data, and the generated input pollution emission data fragment to generate derived learning data corresponding to the current template pollution emission data; Generate iterative pollution emission data having a consistent pollution feature path with the current template pollution emission data by learning the derived learning data corresponding to the current template pollution emission data through the feature derivation network; Continue polling until all selected template pollution emission data are polled, generating Y iterative pollution emission data.

8. The method for predicting pollution emissions based on machine learning according to claim 6, characterized in that: The generating of iterative pollution emission data with consistent pollution feature paths for each selected template pollution emission data by the feature derivation network based on the feature derivation request and the one or more derived sample combination data, and generating Y iterative pollution emission data, comprises: fusing the feature derivation request and the one or more derived sample combination data to generate a training sample for the feature derivation network; Using each selected template pollution emission data and the corresponding prior pollution emission description label respectively, generating an input pollution emission data segment corresponding to each template pollution emission data; Training the feature derivation network based on the training samples to generate a trained feature derivation network; The trained feature derivation network is used to generate iterative pollution emission data with consistent pollution feature paths for the corresponding template pollution emission data based on the input pollution emission data segments corresponding to the respective template pollution emission data, thereby generating Y iterative pollution emission data.

9. The method for predicting pollution emissions based on machine learning according to claim 1, characterized in that: The method further comprises: After receiving a prediction instruction for target pollution emission data, generating candidate pollution emission data according to the model activation definition information and the target pollution emission data, wherein the target pollution emission data is configured as input pollution emission data in the candidate pollution emission data; The pollution emission prediction model is used to perform pollution emission prediction on the target pollution emission data based on the candidate pollution emission data, and a pollution emission description label of the target pollution emission data and corresponding label-guided knowledge data are generated.

10. A pollution emission prediction system based on machine learning, characterized in that: The machine learning-based pollution emission prediction system includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the machine learning-based pollution emission prediction method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Pollution emission prediction method and device

    CN110807577A

  • Vehicle exhaust emission prediction method and system based on machine learning algorithm

    CN114282680A

  • Enterprise pollutant emission prediction method and prediction system

    CN116862079A

  • Soil arsenic pollution risk assessment method, device and equipment based on machine learning

    CN118506922A

  • Method of predicting fine dust concentration and inferring source by using local public data and prediction and inference device

    US20240029457A1

Cited By

  • Sewage detection and analysis method based on Internet of Things

    CN120217125A

  • Sewage detection and analysis method based on the Internet of Things

    CN120217125B