Dynamic construction method for confidence interval of automatic monitoring data of pollution source
By introducing highly adaptive energy management algorithms and machine learning technology into the pollution source automatic monitoring system, the data confidence interval is dynamically constructed, and the problem of the pollution source automatic monitoring system processing abnormal data is solved, and more efficient, safe and economical pollution source management is achieved.
Patent Information
- Application Number
- CN202510208363.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The automatic pollution source monitoring system is difficult to effectively process abnormal data during operation, resulting in risks and hidden dangers in environmental protection. The existing technology is difficult to meet the needs of non-site law enforcement, and it is necessary to dynamically build confidence intervals to improve data credibility.
A highly adaptive energy management algorithm is adopted, combined with machine learning (especially reinforcement learning) to predict energy demand and dynamically adjust energy allocation, identify behavior patterns through time-series correlation analysis and behavioral correlation diagram, integrate multiple data sources for time series analysis, build demand prediction models and environmental impact models, and optimize energy allocation and resource scheduling.
It improves the accuracy and real-time nature of automatic monitoring of pollution sources, effectively prevents the occurrence of safety accidents, reduces energy waste, reduces costs, improves equipment operation efficiency, and provides more efficient, safe and economical regional management solutions for pollution sources.
Smart Images

Figure CN120146469A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of automatic monitoring of pollution sources, and particularly relates to a method for dynamically constructing a confidence interval for automatic monitoring data of pollution sources. Background Art
[0002] The problem of environmental pollution has always been one of the serious challenges faced by the world, especially in the context of the accelerating process of industrialization and urbanization. Emissions of waste gas, waste water and solid waste generated by activities such as industrial production, transportation and energy utilization not only cause serious pollution to the atmosphere, water bodies and soil, but also have irreversible impacts on human health and the ecosystem. In order to address this challenge, countries have strengthened environmental supervision, and online monitoring of pollution sources has become a key measure.
[0003] Online monitoring of pollution sources is an important task in the field of environmental protection. Its key significance lies not only in the comprehensive understanding and monitoring of environmental pollution, but also in timely warning and taking effective measures to reduce pollutant emissions. The online monitoring system of pollution sources can collect and transmit data in real time, and can immediately detect abnormal emissions of pollution sources, so as to quickly take measures for adjustment and emission restriction. This helps to prevent the further expansion of environmental pollution, reduce the risks to the environment and people's health, and protect the ecological environment and people's health. With the continuous progress of information technology and sensor technology, the technology of online monitoring of pollution sources has developed rapidly. The accuracy and stability of sensors have been continuously improved, and data transmission and storage technologies have also been improved, making the online monitoring system more feasible and cost-effective.
[0004] Automatic monitoring of pollution sources refers to the use of advanced sensors, controllers and data acquisition devices to monitor and control pollution sources in the industrial production process in real time, continuously and automatically. The application of this technology can help enterprises make rational use of resources, protect the environment and reduce production costs, which is of great significance for environmental protection work. However, abnormal data will inevitably appear during the operation of the automatic monitoring system of pollution sources. These abnormal data may bring risks and potential hazards to environmental protection. Therefore, it is very important to analyze and dispose of abnormal data. In order to avoid sampling errors and data manipulation problems that may occur in traditional regular monitoring and improve data credibility, it is very urgent to establish a confidence interval so that the online monitoring data of pollution sources can become an important basis for environmental protection decision-making and law enforcement. The rules such as constant values, outliers, and steep drops provided by online monitoring of pollution sources are difficult to meet the needs of non-site law enforcement in the new situation. It is necessary to dynamically construct a confidence interval, deeply excavate and comprehensively judge. Only by timely discovering the causes of abnormal data and taking effective measures to deal with them can the normal operation of the automatic monitoring system of pollution sources be ensured, the environment be effectively protected, and the occurrence of environmental pollution be prevented. Summary of the Invention
[0005] The object of the present invention is to design a dynamic construction method for the confidence interval of automatic monitoring data of pollution sources, introducing a highly adaptive energy management algorithm, using machine learning algorithms (especially reinforcement learning) to predict energy demand and dynamically adjust energy allocation. Compared with the prior art, this method can improve the accuracy and real-time performance of safety monitoring and effectively prevent the occurrence of safety accidents.
[0006] To achieve the above object, in the first aspect of the present invention, a dynamic construction method for the confidence interval of automatic monitoring data of pollution sources is provided, and the method includes the following steps:
[0007] S1. Collect pollution source data, preprocess the data and extract behavioral characteristics, vectorize the behavioral characteristics for time series correlation analysis, construct a behavioral association graph, and train the behavioral association graph to identify behavioral patterns;
[0008] S2. Integrate the data of the data sources and perform time series analysis on the data to generate a statistical report; wherein, the data sources include sensor data, equipment status data, and safety monitoring output data;
[0009] S3. Use machine learning algorithms to monitor and analyze the total equipment data, predict pollution sources and prevent the occurrence of environmental pollution, and the total equipment is all pollution treatment equipment;
[0010] S4. Construct a demand prediction model and an environmental impact model based on real-time data and analyze the data to optimize energy allocation; the real-time data includes the temperature T, humidity H, and flow rate L at the industrial production discharge port, the current period energy consumption E of the industrial production equipment, and the past same period energy consumption O, and then smooth the short-term fluctuations through the moving average method of real-time data, and calculate the sliding average value of each data as It is expressed as
[0011]
[0012] wherein, n represents the moving window size, x i represents the real-time data to be calculated, t represents the original timestamp, and i represents the lower time bound of the window;
[0013] The demand prediction model is
[0014]
[0015] wherein, a 1 , a 2 , a 3 , a 4 represent model parameters, learned from historical data through machine learning algorithms; b represents the bias term; represents the average temperature, represents the average humidity, represents the average flow rate, represents the average energy consumption in the same period in the past; E pred represents the predicted demand;
[0016] S5. Construct a multi-resource demand model optimization scheduling strategy based on the predicted demand and real-time data, define reinforcement learning to minimize the total cost, and perform resource management and scheduling; the resources include power, human, and mechanical equipment resources; the power resource preferably includes solar energy P s , wind energy P w and mains power P g . The power resources are not limited to the above-mentioned solar energy, wind energy, and ordinary mains power, and can also include new energy such as hydrogen energy; the human resources include the number of staff N w ; the mechanical equipment resources include the number of in-use machines N m and the machine status; among them, the definition of reinforcement learning and the design of an Actor-Critic network to train the reinforcement learning specifically include:
[0017] The state includes the energy supply P s , P w , P g at the current time step and the human and mechanical usage status N w , N m ;
[0018] The action includes the decision to adjust various energy inputs and resource allocations;
[0019] The optimization function is the negative value of the cost function, and the goal is to minimize the total cost, expressed as:
[0020] R = -(c s P s + c w P w + c g P g + c w N w + c m N m )
[0021] where R represents the optimization function, and c s , c w , c g , c w , c m represent the unit costs of solar energy, wind energy, mains power, human, and machinery respectively;
[0022] S6. Based on the data from the data source in S2, through machine learning, timely capture the constant values, exceeded standard values, and clues of abnormal data fluctuations in the online monitoring data of pollution sources to prevent environmental pollution.
[0023] Further, each node of the constructed behavior association graph represents the pollution emission behavior state at a time point, the edges between the nodes represent the possibility and similarity of behavior conversion, and the weights of the edges are calculated using any of the following methods:
[0024] 1. Variance:
[0025] 2. Average:
[0026] 3. The median is the number in the middle position among a set of data arranged in order. If the number of data is odd, the median is the (n + 1) / 2 - th number; if the number of data is even, the median is the average of the n / 2 - th number and the (n / 2 + 1) - th number, where x1 is the i - th data point in the data set, is the average value of the data set, n is the total number of data points. Σ represents the sum of all data points from i = 1 to n.
[0027]
[0028] In the formula, u k is the clustering center value;
[0029] 5. Constant value algorithm:
[0030] The algorithm for identifying consecutive identical values in a set of data can be implemented by traversing the data and comparing adjacent elements. Each tuple contains a value and the number of consecutive occurrences of that value.
[0031] Further, training the behavior association graph to identify behavior patterns, where the detection algorithm is as follows:
[0032]
[0033] where ω j represents the weight adjusted according to time and space proximity, N(x) represents the set of adjacent nodes of node x; S(x) represents the currently calculated risk value. When S(x) exceeds the dynamically calculated threshold, the system automatically triggers an alarm.
[0034] Further, performing trend analysis and periodic analysis on the data using time - series analysis, where the trend analysis is expressed as follows
[0035]
[0036]
[0037] Among them, represents the predicted value, and X t represents the actual observed value, and b t represents the trend component, and α and β represent the smoothing parameters.
[0038] Furthermore, the said S3 specifically includes:
[0039] S301. Collect the total equipment data and perform standardization processing on the data;
[0040] S302. Construct a physical model and estimate the model parameters; among them, the least squares method is used to estimate the model parameters; among them, the physical model is expressed as follows:
[0041]
[0042] Among them, T represents the operating temperature of the production equipment, P represents the heat generated during the operation of the production equipment, h represents the convective heat transfer coefficient, A represents the surface area of the production equipment, and T env represents the ambient temperature, and m and c respectively represent environmental changes and sensor failures;
[0043] S303. Define the total equipment health index H according to the physical model and real-time sensor data, which is expressed as follows:
[0044]
[0045] Among them, T nom represents the nominal operating temperature of the equipment, and T max represents the maximum safe operating temperature;
[0046] S304. Monitor the total equipment health index H and trigger an early warning according to the monitoring situation.
[0047] Furthermore, the said S4 also includes periodically identifying the data, predicting the demand peaks and troughs, which is expressed as follows:
[0048]
[0049] Among them, x n represents the nth data point in the time series, and X k represents the complex amplitude of the kth frequency component.
[0050] Furthermore, the said multi-resource demand model uses a vector autoregressive model to jointly predict environmental changes, historical comparison of data, and feature analysis of data, which is expressed as follows:
[0051]
[0052] Among them, Φ1 ,…, Φ p is the model coefficient matrix, c is the constant vector, and ∈ is the error vector; respectively represent the environmental change, the historical comparison of data, and the feature analysis of data; E t-i , N w,t-i , N m,t-i respectively represent the environmental change, the historical comparison of data, and the feature analysis of data at the i-th past time point, where i ∈ [1, 2, 3... p].
[0053] Furthermore, an optimization model is constructed according to minimizing the total cost, including energy cost, labor cost, and equipment operation cost, which is expressed as follows:
[0054] minimize C = c s P s + c w P w + c g P g + c w N w + c m N m
[0055] subject to P s + P w + P g ≥ E required , N w ≥ N w,required , N m ≥ N m,required
[0056] where c s , c w , c g , c w , c m respectively represent the unit costs of solar energy, wind energy, commercial power, labor, and machinery.
[0057] Furthermore, the specific steps of S6 include:
[0058] S601. Collect the data of the data source and perform standardization processing;
[0059] S602. Construct a load balancing network flow model, including:
[0060] Convert the concentration value of the pollution source monitoring data into a weighted graph, where the nodes represent the key monitoring values, the edges represent the feasible paths, design the dynamic adjustment of the edge weights, and introduce a time variable to adjust the weights to reflect the flow changes in different time periods. Apply the improved Dijkstra algorithm, where the path cost function considers time, distance, and congestion degree, and is expressed as follows:
[0061] Cost i,j = time i,j + λ × congestion i,j
[0062] where λ represents the congestion sensitivity parameter, time i,j and congestion i,j respectively represent the travel time from node i to j and the current congestion level; Cost i,j represents the path cost;
[0063] S603. Perform real-time path optimization and resource scheduling according to the path cost, including:
[0064] Define the path cost function based on time step t as:
[0065] Cost i,j (t) = t i,j + λ·c i,j (t)
[0066] where t i,j represents the basic travel time from node i to j, c i,j (t) represents the congestion factor based on time t, and λ represents the adjustment coefficient for balancing the weights of time and congestion;
[0067] Design a dynamic path optimization algorithm:
[0068] f(n) = g(n) + h(n)
[0069] where g(n) represents the known cost from the starting point to node n, and h(n) represents the estimated lowest cost from node n to the target;
[0070] Real-time resource scheduling:
[0071] R k (t) = α·D k (t) + β·S k (t)
[0072] where R k (t) represents the adjustment level of resource k at time t, D k (t) represents the demand forecast, S k (t) represents the current state of resource k, and α and β represent the adjustment parameters;
[0073] where the S6 further includes at least 0 of the following steps:
[0074] Iterate based on system performance metric analysis and system feedback data.
[0075] In a second aspect of the present invention, a dynamic construction system for the confidence interval of automatic monitoring data of pollution sources is provided. The system includes:
[0076] A monitoring data and video data acquisition unit, which is used to collect pollution source data, preprocess the data and extract behavioral characteristics, vectorize the behavioral characteristics for time series correlation analysis, construct a behavioral association graph, and train the behavioral association graph to identify behavioral patterns;
[0077] A time series analysis unit, which is used to integrate the data of data sources and perform time series analysis on the data, and generate a statistical report; wherein, the data sources include sensor data, equipment status, and safety monitoring outputs;
[0078] A machine learning prediction unit, which is used to monitor and analyze the total equipment data using machine learning algorithms, and predict pollution sources and prevent environmental pollution;
[0079] An optimized energy distribution unit, which is used to construct a demand prediction model and an environmental impact model based on real-time data and analyze the data to optimize energy distribution; the real-time data includes the temperature T, humidity H, and flow rate L at the industrial production emission outlet, the current period energy consumption E of the industrial production equipment, and the energy consumption O in the same period in the past. Then, the short-term fluctuations are smoothed by the moving average method of the real-time data, and the sliding average value of each data is calculated as shown as follows
[0080]
[0081] wherein, n represents the moving window size, x i represents the real-time data to be calculated, t represents the original timestamp, and i represents the lower time bound of the window;
[0082] The demand prediction model is
[0083]
[0084] wherein, a 1 , a 2 , a 3 , a 4 represent model parameters, which are learned from historical data through machine learning algorithms; b represents the bias term; represents the average temperature, represents the average humidity, represents the average flow rate, represents the average energy consumption in the same period in the past; E pred represents the predicted demand;
[0085] Optimized scheduling strategy unit, which is used to construct an optimized scheduling strategy for a multi-resource demand model according to predicted demands and real-time data, define reinforcement learning to minimize the total cost, and then manage and schedule resources; the resources include power, human resources, and mechanical equipment resources, and the power resources include solar energy P s , wind energy P w , and commercial power P g , the human resources include the number of staff N w , and the mechanical equipment resources include the number of in-use machines N m and the machine status; among them, the definition of reinforcement learning and the design of an Actor-Critic network to train the reinforcement learning specifically include:
[0086] The state includes the energy supply P s , P w , P g at the current time step and the human and mechanical usage status N w , N m ;
[0087] The action includes the decision to adjust various energy inputs and resource allocations;
[0088] The optimization function is the negative value of the cost function, and the goal is to minimize the total cost, which is expressed as:
[0089] R = -(c s P s + c w P w + c g P g + c w N w + c m N m )
[0090] Among them, R represents the optimization function, and c s , c w , c g , c w , c m respectively represent the unit costs of solar energy, wind energy, commercial power, human resources, and machinery;
[0091] Optimized flow route unit, which is used to optimize the prevention of environmental pollution through machine learning according to the data from the data source in the time series analysis unit.
[0092] The beneficial technical effects of the present invention are at least as follows:
[0093] (1) In the present invention, a highly adaptive energy management algorithm is introduced. This algorithm uses machine learning algorithms (especially reinforcement learning) to predict energy demand and dynamically adjust energy allocation. The system can integrate energy inputs from different sources, such as solar energy, wind energy, and the power grid, and optimize the energy usage strategy in real time according to energy prices and usage demands, significantly reducing energy waste and costs.
[0094] (2) The present invention adopts advanced deep learning technologies (such as convolutional neural networks and long short-term memory networks) to monitor and analyze the emission behavior patterns of pollution source monitoring data in real time, automatically identify abnormal behaviors, and respond quickly. Compared with the prior art, this method can improve the accuracy and real-time performance of safety monitoring and effectively prevent the occurrence of safety accidents.
[0095] (3) The present invention integrates advanced data analysis tools, uses a large amount of device operation data collected by the Internet of Things, and predicts potential device failures through data mining and machine learning technologies. This enables the managers in the pollution source area to implement preventive maintenance, reduce the device failure rate, lower the maintenance cost, and improve the overall operation efficiency of the device.
[0096] (4) The present invention not only solves the limitations of the prior art in energy management, safety monitoring, and device maintenance, but also provides a more efficient, safer, and more economical solution for the management of pollution source areas. These innovative points precisely solve the core problems existing in the prior art and improve the overall performance and sustainability of the management of pollution source areas. Description of the Drawings
[0097] Figure 1 It is a flowchart of a method for dynamically constructing the confidence interval of automatic monitoring data of pollution sources in an embodiment of the present invention.
[0098] Figure 2 It is a framework diagram of a method for dynamically constructing the confidence interval of automatic monitoring data of pollution sources in an embodiment of the present invention. Detailed Embodiment
[0099] The technical solution of the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0100] In one or more embodiments, as Figure 1 shown, a method for dynamically constructing the confidence interval of automatic monitoring data of pollution sources according to the present invention is disclosed. The method includes steps 1 to 6, including:
[0101] S1. Collect pollution source data, preprocess the data and extract behavioral characteristics, vectorize the behavioral characteristics for time series correlation analysis, construct a behavioral correlation graph, and train the behavioral correlation graph to identify behavioral patterns.
[0102] Specifically, configure a high-resolution camera system to cover all key areas of the pollution source area, collect video streams in real time, and apply real-time image segmentation algorithms (such as U-Net or MaskR-CNN) to distinguish the foreground (individuals) and the background, so as to reduce noise and errors in subsequent analysis. Use a pre-trained deep convolutional neural network (such as VGG-16 or ResNet) to extract behavioral features from the segmented images. The feature vector F i = [f i1 , f i2 , …, f in generated by each image frame, where f ij represents the output of the j-th feature channel after ReLU activation, ensuring that all feature values are non-negative and enhancing the non-linear expression ability of the model.
[0103] Conduct temporal correlation analysis on the feature vectors to construct a behavior graph. Each node represents the behavior state at a time point, and the edges between nodes represent the possibility and similarity of behavior transitions. The weight G ij of the edge is calculated as follows:
[0104]
[0105] where α represents an adaptive adjustment parameter, which is dynamically adjusted according to the scenario to adapt to the sensitivity of behavior changes at different times and in different environments, f ik represents the output of the k-th feature channel after ReLU activation, and f ik represents the feature vector.
[0106] Specifically, apply a graph convolutional network (GCN) to perform deep learning on the behavior graph to identify behavior patterns. The anomaly detection algorithm is as follows,
[0107]
[0108] where ω j is the weight adjusted according to temporal and spatial proximity, and N(x) is the set of adjacent nodes of node x. When S(x) exceeds the dynamically calculated threshold, the system automatically triggers an alarm.
[0109] As an embodiment of the present invention, the system monitors through real-time data streams. Once an abnormal behavior is detected, a visual alarm is immediately issued at the control center, pointing to the specific camera location and timestamp, facilitating rapid response and handling. At the same time, for specific types of anomalies that occur frequently, the system adjusts the model parameters through a self-learning mechanism to improve the accuracy of prediction and the timeliness of response.
[0110] S2. Integrate the data from the data sources and perform time series analysis on the data, and generate a statistical report; wherein, the data sources include sensor data, device status, and security monitoring outputs.
[0111] Specifically, the data sources include, but are not limited to, environmental sensor data (temperature T, humidity H, flow intensity L), device status (on / off status S, energy consumption E), and safety monitoring outputs (number of people in the area N, alarm status A). Then, timestamp normalization is performed on all sensor outputs to ensure data synchronization, as shown below:
[0112] t′ = t - t 0
[0113] where t is the original timestamp, t 0 is the starting time of the day when the data is collected, and t' is the aligned timestamp.
[0114] As an embodiment of the present invention, the relationship between physical quantities is established through basic physical laws and actual measurements. For example, the relationship between the device energy consumption E, environmental temperature T, and device status S can be expressed as:
[0115] E = β 1 TS + β 2 S
[0116] where β 1 and β 2 are coefficients obtained through historical data regression analysis, representing the impact of temperature and device on / off status on energy consumption.
[0117] Specifically, the integrated model provides in-depth data insights and trend predictions to support the monitoring of pollution source monitoring data.
[0118] Furthermore, for data with obvious seasonal variations in environmental parameters such as temperature and humidity, a seasonal adjustment method is adopted to exclude the influence of seasonal effects. This can be achieved through the seasonal decomposition of time series (STL), and the formula is:
[0119] X tadjusted = X t - S t
[0120] where X t is the original time series data, and S t is the seasonal component obtained from the STL decomposition.
[0121] Among them, time series analysis includes trend analysis and periodic analysis. Time series analysis uses time series analysis methods (such as the Holt-Winters method) to predict the long-term trend of data, and the formula is:
[0122]
[0123]
[0124] Among them, is the predicted value, X t is the actual observed value, b t is the trend component, and α and β are smoothing parameters.
[0125] Periodic analysis is to identify repeated patterns or periodic changes in the data to help predict demand fluctuations or peaks in resource usage. Periodic analysis can be performed using the autocorrelation function (ACF) and the partial autocorrelation function (PACF) to identify.
[0126] Furthermore, generating statistical reports specifically provides descriptive statistical information of key indicators for management, including mean, median, standard deviation, skewness, and kurtosis, etc. This helps to depict the distribution characteristics and variability of the data. At the same time, based on the results of trend and periodic analysis, detailed reports are generated, pointing out possible risk areas and optimization opportunities. These reports will particularly emphasize the optimization of resource usage and cost control.
[0127] S3. Use machine learning algorithms to monitor and analyze the total equipment data to predict the generation of pollution sources and prevent environmental pollution incidents, including:
[0128] S301. Collect the total equipment data and perform standardization processing on the data;
[0129] Specifically, the total equipment data includes sensor data such as temperature T (in degrees Celsius), pressure P (in Pascals), rotational speed R (in revolutions per minute), current I (in amperes), etc. The data is collected from various types of equipment such as pumps, motors, compressors, etc.
[0130] The data is first received from each sensor through a real-time data acquisition system, and then standardization processing is performed on all the data. The formula is:
[0131]
[0132] Among them, x' is the value after standardization, making the sensor data of different types and ranges comparable.
[0133] S302. Build a physical model and estimate model parameters (the model describes the thermal behavior of the motor, considering energy input, heat dissipation, and environmental interaction); among them, the least squares method is used to estimate the model parameters, such as h, A, m, and c; among them, the physical model is expressed as follows:
[0134]
[0135] Among them, T represents the operating temperature of the production equipment, P represents the heat generated by the operation of the production equipment, h represents the convective heat transfer coefficient, A represents the surface area of the production equipment, T env represents the environmental temperature, and m and c represent environmental changes and sensor failures respectively;
[0136] S303. Define the overall equipment health indicator H based on the physical model and real-time sensor data, which is expressed as follows:
[0137]
[0138] where T nom represents the nominal operating temperature of the equipment, and T max represents the maximum safe operating temperature;
[0139] Among them, the operating data is compared with the model prediction data in real time to calculate the health indicator. When the monitoring indicator crosses the preset threshold, such as when H is lower than 0.8, a warning is triggered.
[0140] S304. Monitor the overall equipment health indicator H and trigger a warning according to the monitoring situation;
[0141] Specifically, when the health indicator drops below the warning threshold, the system automatically triggers a warning. The warning information includes the equipment identifier, the current health indicator, and the recommended maintenance operations. According to the specific value of the health indicator, the system automatically generates maintenance suggestions, such as replacing worn components, adjusting operating parameters, or performing inspections. The maintenance suggestions are intended to prevent possible equipment failures and ensure that the equipment operates in the best condition.
[0142] S4. Analyze the data by constructing a demand prediction model and an environmental impact model based on real-time data to optimize energy distribution; the real-time data includes the temperature T, humidity H, and flow rate L at the industrial production emission outlet, the current-period energy consumption E of the industrial production equipment, and the energy consumption O in the same period in the past. Then, smooth the short-term fluctuations of the real-time data by the moving average method, and calculate the moving average value of each type of data as follows which is expressed as
[0143]
[0144] where n represents the moving window size, x i represents the real-time data to be calculated, t represents the original timestamp, and i represents the lower bound of the window time;
[0145] Use Fourier transform to analyze the periodicity of the energy usage data to predict possible demand peaks and valleys:
[0146]
[0147] where x n is the nth data point in the time series, and X k is the complex amplitude of the kth frequency component;
[0148] Reconstruct the environmental impact model and establish an impact model of environmental factors on energy demand, such as the impact of temperature and flow rate on the HVAC system:
[0149] E HVAC = αT + βL + γ
[0150] Where α and β are coefficients obtained by historical data regression, reflecting the sensitivity of energy demand to temperature and flow rate changes, and γ is a constant term
[0151] Use a multivariable linear regression model to combine the above analysis results for energy demand forecasting. The demand forecasting model is
[0152]
[0153] Where a 1 , a 2 , a 3 , a 4 represent model parameters learned from historical data through machine learning algorithms; b represents the bias term; represents the average temperature, represents the average humidity, represents the average flow rate, represents the average energy consumption in the same past period; E pred represents the predicted demand.
[0154] Specifically, based on the predicted demand E pred , the system automatically generates energy scheduling suggestions, such as adjusting air conditioner settings, lighting control, etc., to optimize energy use efficiency.
[0155] Among them, the collected data is uploaded in real time through IoT devices and organized in a predetermined format to ensure the immediacy and integrity of the data.
[0156] S5. Construct a multi-resource demand model based on the predicted demand and real-time data to optimize the scheduling strategy, define reinforcement learning to minimize the total cost, and then manage and schedule resources; the resources include electric power, human resources, and mechanical equipment resources. The electric power resources include solar power P s , wind power P w and mains power P g , the human resources include the number of staff N w , and the mechanical equipment resources include the number of in-use machines N m and the machine status;
[0157] The multi-resource demand model uses a vector autoregressive model to jointly predict environmental changes, historical comparison of data, and feature analysis of data, as shown below:
[0158]
[0159] Among them, Φ 1 , …, Φ p is the model coefficient matrix, c is the constant vector, and ∈ is the error vector; respectively represent environmental changes, historical comparison of data, and feature analysis of data; E t-i , N w,t-i , N m,t-i respectively represent environmental changes, historical comparison of data, and feature analysis of data at the i-th past time point, where i ∈ [1, 2, 3, …, p].
[0160] An optimization model is constructed according to minimizing the total cost, including energy cost, labor cost, and equipment operation cost, which is expressed as follows:
[0161] minimize C = c s P s + c w P w + c g P g + c w N w + c m N m
[0162] subject to P s + P w + P g ≥ E required , N w ≥ N w,required , N m ≥ N m,required
[0163] Among them, c s , c w , c g , c w , c m respectively represent the unit costs of solar energy, wind energy, mains electricity, labor, and machinery;
[0164] Among them, the defined reinforcement learning is designed, and an Actor - Critic network is designed to train the reinforcement learning, specifically including:
[0165] The state includes the energy supply P s , P w , P g and the usage status of labor and machinery N w , N m ;
[0166] Actions include decisions to adjust various energy inputs and resource allocations; actions that the agent can take include decisions to adjust various energy inputs and resource allocations, such as increasing or decreasing solar output ΔP s , wind output ΔP w , mains power usage ΔP g , adjusting the workforce ΔN w and machinery usage ΔN m ;
[0167] The optimization function is the negative of the cost function, and the goal is to minimize the total cost, expressed as:
[0168] R = -(c s P s + c w P w + c g P g + c w N w + c m N m )
[0169] where R represents the optimization function, and c s , c w , c g , c w , c m represent the unit costs of solar energy, wind energy, mains power, workforce, and machinery respectively.
[0170] Specifically, data integration includes real-time collection of the usage status and availability of all resources and synchronous updating of the resource database to reflect the current resource configuration.
[0171] Specifically, the Actor network is responsible for generating actions in a given state, with the network input being the current state and the output being the corresponding action. The Critic network evaluates the value of the current state and the action selected by the Actor, with the input being the state and the action and the output being the value function estimate of the state-action pair. Then, the experience replay mechanism is used for training, storing the transitions (state, action, reward, new state) at each time step, and using random sample batches for network training to break temporal correlations and reduce variance. The Actor network is responsible for generating the optimal action, and the Critic network is responsible for optimizing the value function estimate through the gradient descent method, using a soft update strategy to gradually adjust the weights of the target network to those of the main network to maintain the stability of the learning process.
[0172] In each decision-making cycle, the optimal action is generated through the actor network based on the current state, and the resource allocation is adjusted in real time. As the external conditions (such as weather changes, market price fluctuations, etc.) and internal conditions (such as equipment aging, prevention of environmental pollution, etc.) change, the system continuously receives new state information and dynamically adjusts the strategy to cope with real-time changes.
[0173] S6. Optimize the prevention of environmental pollution through machine learning according to the data from the data source in S2, specifically including:
[0174] S601. Collect the data from the data source and perform standardization processing;
[0175] Specifically, real-time pedestrian flow data: collected through Wi-Fi tracking, video surveillance, access control systems, etc.
[0176] Vehicle flow data: collected through geomagnetic sensors, ANPR (Automatic Number Plate Recognition System), etc.
[0177] Resource status data: including meeting room occupancy, restaurant seat usage, parking space availability, etc.
[0178] The preprocessing includes using the Central Data Processing Center (CDC) to synchronize and integrate the real-time information from each data point, including cleaning, normalization, and time series analysis, to ensure data quality and consistency.
[0179] S602. Construct a load balancing network flow model, including:
[0180] Convert the concentration value of the pollution source monitoring data into a weighted graph, where the nodes represent the key monitoring values and the edges represent the feasible paths. Design the dynamic adjustment of the edge weights, which reflect the congestion degree of the path or the resource utilization rate based on the real-time data. At the same time, introduce a time variable to adjust the weights to reflect the traffic changes in different time periods. Apply the improved Dijkstra algorithm, where the path cost function considers time, distance, and congestion degree, and is expressed as follows:
[0181] Cost i,j =time i,j +λ×congestion i,j
[0182] Among them, λ represents the congestion sensitivity parameter, time i,j and congestion i,j respectively represent the travel time from node i to j and the current congestion level; Cost i,j represents the path cost;
[0183] S603. Perform real-time path optimization and resource scheduling according to the path cost, including:
[0184] Define the path cost function based on time step t as follows:
[0185] Cost i,j (t) = t i,j + λ·c i,j (t)
[0186] where t i,j represents the basic travel time from node i to j, c i,j (t) represents the congestion factor based on time t, and λ represents the adjustment coefficient used to balance the weights of time and congestion;
[0187] Design a dynamic path optimization algorithm that uses the A* search algorithm and considers the actual travel time and expected congestion between nodes:
[0188] f(n) = g(n) + h(n)
[0189] where g(n) represents the known cost from the starting point to node n, and h(n) represents the estimated lowest cost from node n to the target;
[0190] Adjust the availability and configuration of key resources based on real-time and predicted data:
[0191] R k (t) = α·D k (t) + β·S k (t)
[0192] where R k (t) represents the adjustment level of resource k at time t, D k (t) represents the demand forecast, S k (t) represents the current state of resource k, and α and β represent the adjustment parameters;
[0193] Iterate based on system performance metric analysis and system feedback data:
[0194] Calculate the system performance metric P, which measures the proportion of optimal path selections to evaluate the effectiveness of the system:
[0195]
[0196] Adjust the model parameters based on the feedback data:
[0197]
[0198] where θ is the model parameter, η is the learning rate, is the gradient of the cost function J with respect to the parameter θ, and D is the collected feedback data.
[0199] As an embodiment of the present invention, periodic model retraining and parameter adjustment are implemented to adapt to new data and operating environments, monitor key performance indicators, and adjust system configurations based on historical trends and prediction models.
[0200] Through these refined steps and formulas, not only can real-time and accurate navigation and resource scheduling solutions be provided, but also self-optimization can be carried out through a powerful feedback mechanism to ensure long-term and continuous system performance improvement. The implementation of this method will greatly improve the efficiency of pollution source monitoring.
[0201] In one or more embodiments, as Figure 2 shown, a dynamic construction method for the confidence interval of automatic pollution source monitoring data of the present invention is disclosed. The system includes:
[0202] A video data acquisition unit 101, configured to acquire pollution source data, preprocess the data and extract behavioral characteristics, vectorize the behavioral characteristics for time series correlation analysis, construct a behavioral association graph, and train the behavioral association graph to identify behavioral patterns;
[0203] A time series analysis unit 102, configured to integrate the data of the data sources and perform time series analysis on the data, and generate a statistical report; wherein, the data sources include sensor data, device status, and safety monitoring outputs;
[0204] A machine learning prediction unit 103, configured to monitor and analyze the total device data using machine learning algorithms, and predict pollution sources and prevent the occurrence of environmental pollution;
[0205] An optimized energy distribution unit 104, configured to analyze the data by constructing a demand prediction model and an environmental impact model based on real-time data, and optimize energy distribution; the real-time data includes the temperature T, humidity H, and flow rate L at the industrial production emission port, the current period energy consumption E of the industrial production equipment, and the energy consumption O in the same period in the past. Then, the short-term fluctuations are smoothed by the moving average method for the real-time data, and the sliding average value of each data is calculated as shown as
[0206]
[0207] wherein, n represents the moving window size, x i represents the real-time data to be calculated, t represents the original timestamp, and i represents the lower time bound of the window;
[0208] The demand prediction model is
[0209]
[0210] wherein, a 1 , a 2 , a3 , a 4 represents the model parameters, which are learned from historical data through machine learning algorithms; b represents the bias term; represents the average temperature, represents the average humidity, represents the average flow rate, represents the average energy consumption in the same period in the past; E pred represents the predicted demand;
[0211] The optimal scheduling strategy unit 105 is used to construct an optimal scheduling strategy for the multi-resource demand model based on the predicted demand and real-time data, define reinforcement learning to minimize the total cost, and then manage and schedule resources; the resources include power, human, and mechanical equipment resources, and the power resources include solar energy P s , wind energy P w and mains power P g , the human resources include the number of staff N w , and the mechanical equipment resources include the number of in-use machines N m and the machine status; among them, the definition of reinforcement learning and the design of the Actor-Critic network to train the reinforcement learning specifically include:
[0212] The state includes the energy supply P s , P w , P g at the current time step and the human and mechanical usage status N w , N m ;
[0213] The action includes the decision to adjust various energy inputs and resource allocations;
[0214] The optimization function is the negative value of the cost function, and the goal is to minimize the total cost, expressed as:
[0215] R = -(c s P s + c w P w + c g P g + c w N w + c m N m )
[0216] Among them, R represents the optimization function, and c s , c w , c g , c w , c m respectively represent the unit costs of solar energy, wind energy, mains power, human, and machinery;
[0217] The optimized flow route unit 106 is used to optimize the prevention of environmental pollution according to the data from the data source in the time series analysis unit through machine learning.
[0218] In summary, the present invention uses advanced deep learning technologies such as convolutional neural networks and recurrent neural networks to monitor the pollution degree of pollution sources in real time, ensuring that abnormal situations can be identified and responded to in a timely manner. This step not only improves the acquisition effect of pollution sources but also provides rich behavioral data and safety status information for subsequent steps. By using advanced deep learning technologies such as convolutional neural networks and recurrent neural networks to monitor the pollution degree of pollution sources in real time, ensuring that abnormal situations can be identified and responded to in a timely manner. This step not only improves the acquisition effect of pollution sources but also provides rich behavioral data and safety status information for subsequent steps. By using advanced deep learning technologies such as convolutional neural networks and recurrent neural networks to monitor the pollution degree of pollution sources in real time, ensuring that abnormal situations can be identified and responded to in a timely manner. This step not only improves the acquisition effect of pollution sources but also provides rich behavioral data and safety status information for subsequent steps. Based on the data generated in the previous steps, especially the comprehensive information from the intelligent analysis platform, this step uses time series analysis and prediction models to predict energy demand. By accurately predicting the energy demand in different time periods, it provides decision-making support for energy allocation to ensure the optimization of energy use. After accurately predicting the energy demand, an energy management system based on optimization algorithms such as genetic algorithms dynamically adjusts the energy supply, integrating multiple energy resources such as solar energy and wind energy to achieve the lowest cost and highest efficiency of energy use. This directly solves the problems of low efficiency and waste in energy management in the prior art. Finally, the system optimizes the acquisition effect of pollution sources according to real-time data and prediction results, provides real-time route suggestions and resource scheduling plans, and improves the monitoring efficiency of pollution sources.
[0219] The above-disclosed are only some preferred embodiments of the present invention, and the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. A dynamic construction method for confidence interval of pollution source automatic monitoring data, characterized in that: The method comprises the following steps: S1. Collect pollution source data, pre-process the data and extract behavior features, quantify the behavior features for time series correlation analysis, build a behavior correlation graph, train the behavior correlation graph, and identify behavior patterns; S2. Integrate data from data sources and perform time series analysis on the data, and generate statistical reports; wherein the data sources include sensor data, equipment status data, and safety monitoring output data; S3. Use machine learning algorithms to monitor and analyze total equipment data, predict pollution sources in industrial production processes, and prevent environmental pollution; S4. Build demand forecasting models and environmental impact models based on real-time data, analyze the data, and optimize energy allocation and process adjustment; the real-time data includes the temperature T, humidity H, flow L at the industrial production outlet, the energy consumption E of the industrial production equipment in the current period, and the energy consumption O in the same period in the past, and then smooth short-term fluctuations by moving the real-time data average method to calculate the sliding average of each data It is expressed as follows Among them, n represents the moving window size, x i represents the real-time data to be calculated, t represents the original timestamp, and i represents the lower bound of the window time; The demand forecasting model is Among them, a1, a2, a3, and a4 represent model parameters, which are learned from historical data through machine learning algorithms; b represents the bias term; represents the average temperature, Represents the average humidity value, represents the average flow rate. Indicates the average energy consumption in the same period in the past; E pred represents the predicted demand; S5. Build a multi-resource demand model to optimize the scheduling strategy based on the predicted demand and real-time data, define reinforcement learning to minimize the total cost, and then manage and schedule resources; the resources include power resources, human resources, and mechanical equipment resources, and the power resources include solar P s 、Wind Energy w and mains electricity P g , the human resources include the number of staff N w The mechanical equipment resources include the number of machines in use N m and machine state; wherein, the definition of reinforcement learning and the design of an Actor-Critic network to train reinforcement learning specifically include: The state includes the energy supply P at the current time step s ,P w ,P g and the human and mechanical usage status N w ,N m ; Actions include decisions to adjust various energy inputs and resource allocations; The optimization function is the negative value of the cost function, and the goal is to minimize the total cost, expressed as: R=-(c s P s +c w P w +c g P g +c w N w +c m N m ) Among them, R represents the optimization function, c s ,c w ,c g ,c w ,c m Represent the unit costs of solar energy, wind energy, city electricity, manpower and machinery respectively; S6. Based on the data from the data source in S2, through machine learning, timely capture the constant value, exceeded value and abnormal data fluctuation clues of the online monitoring data of the pollution source, and optimize the prevention of environmental pollution.
2. According to claim 1, a method for dynamically constructing confidence intervals of pollution source automatic monitoring data is characterized in that: Each node x in the behavior association graph represents a behavior state at a time point, and the edges between nodes represent the possibility and similarity of behavior conversion. The weight of the edge is calculated using any of the following methods: Method 1: Variance method: in Method 2: Average value method: Method 3: Median method: The median is the number in the middle of a set of data arranged in order. When the number of data is an odd number, the median is: (n+1) / 2 numbers; when the number of data is an even number, the median is the average of the n / 2th number and the (n / 2)+1th number; Where xi is the i-th data point in the dataset, is the average value of the data set, n is the total number of data points; Σ means summing all data points with i ranging from 1 to n; Method 4: K-means algorithm In the formula, u k is the cluster center value; Method 5: Constant value algorithm: Identify consecutive identical values in a set of data by traversing the data and comparing adjacent elements. Each tuple contains a value and the number of times the value appears consecutively.
3. The method for dynamically constructing confidence intervals of pollution source automatic monitoring data according to claim 2 is characterized in that: The behavior association graph is trained to identify the behavior pattern, wherein the detection algorithm is as follows: Among them, ω j It represents the weight adjusted according to the proximity in time and space, N(x) represents the set of adjacent nodes of node x, S(x) represents the current calculated danger value, when S(x) exceeds the dynamically calculated threshold, the system automatically triggers an alarm.
4. The method for dynamically constructing confidence intervals of pollution source automatic monitoring data according to claim 1 is characterized in that: The data is subjected to time series analysis, trend analysis and periodicity analysis, wherein the trend analysis is expressed as follows: in, represents the predicted value, X t represents the actual observed value, b t represents the trend component, and α and β represent smoothing parameters.
5. The method for dynamically constructing confidence intervals of pollution source automatic monitoring data according to claim 1 is characterized in that: The S3 specifically includes: S301, collect total equipment data and standardize the data; S302, constructing a physical model and estimating model parameters; wherein the model parameters are estimated using the least squares method; wherein the physical model is expressed as follows: Where T represents the operating temperature of the production equipment, P represents the heat generated by the operation of the production equipment, h represents the convection heat transfer coefficient, A represents the surface area of the production equipment, and T env represents the ambient temperature, m and c represent the environmental change and sensor failure respectively; S303. Based on the physical model and real-time sensor data, define the total equipment health index H, which is expressed as follows: Among them, T nom Indicates the nominal operating temperature of the equipment, T max Indicates the maximum safe operating temperature; S304: Monitor the total equipment health index H and trigger an early warning based on the monitoring situation.
6. The method for dynamically constructing confidence intervals of pollution source automatic monitoring data according to claim 1 is characterized in that: The S4 also includes periodic identification of data and prediction of demand peaks and valleys, as shown below: Among them, x n represents the nth data point in the time series, X k Represents the complex amplitude of the kth frequency component.
7. The method for dynamically constructing confidence intervals of pollution source automatic monitoring data according to claim 1 is characterized in that: The multi-resource demand model uses a vector autoregression model to jointly predict environmental changes, historical comparison of data, and feature analysis of data, as shown below: Among them, Φ1,…,Φ p is the model coefficient matrix, c is the constant vector, ∈ is the error vector; They represent environmental changes, historical comparison of data, and feature analysis of data respectively; E t-i ,N w,t-i ,N m,t-i They represent the environmental changes at the i-th time point in the past, historical comparison of data, and feature analysis of data, respectively, i∈[1.2.3...p].
8. The method for dynamically constructing confidence intervals of pollution source automatic monitoring data according to claim 7 is characterized in that: Based on minimizing the total cost, including energy cost, labor cost and equipment operation cost, an optimization model is constructed, which is expressed as follows: minimizeC=c s P s +c w P w +c g P g +c w N w +c m N m subjectto P s +P w +P g ≥E required ,N w ≥N w,required ,N m ≥N m,required Among them, c s ,c w ,c g ,c w ,c m Represent the unit costs of solar energy, wind energy, city electricity, manpower and machinery respectively.
9. The method for dynamically constructing confidence intervals of pollution source automatic monitoring data according to claim 1, characterized in that: The S6 specifically includes: S601, collect data from the data source and perform standardization processing; S602: Building a load balancing network flow model, including: The numerical values of pollution source monitoring data are converted into a weighted graph. Nodes represent key monitoring values, edges represent feasible paths, and the edge weights are designed to be dynamically adjusted. At the same time, time variables are introduced to adjust the weights to reflect the flow changes in different time periods. The improved Dijkstra algorithm is applied, in which the path cost function considers time, distance, and congestion level, as shown below: Cost i,j =time i,j +λ×congestion i,j Among them, λ represents the congestion sensitivity parameter, time i,j and congestion i,j Respectively represent the travel time from node i to j and the current congestion level; Cost i,j represents the path cost; S603: Perform real-time path optimization and resource scheduling according to the path cost, including: The path cost function based on time step t is defined as: Cost i,j (t)=t i,j +λ·c i,j (t) Among them, t i,j represents the basic travel time from node i to j, c i,j (t) represents the congestion factor based on time t, λ represents the adjustment coefficient, which is used to balance the weight of time and congestion; Design a dynamic path optimization algorithm: f(n)=g(n)+h(n) Where g(n) represents the known cost from the starting point to node n, and h(n) represents the estimated minimum cost from node n to the target; Real-time resource scheduling: R k (t)=α·D k (t)+β·S k (t) Among them, R k (t) represents the adjustment level of resource k at time t, D k (t) represents demand forecast, S k (t) represents the current state of resource k, α and β represent adjustment parameters; Wherein, the S6 at least further comprises the following steps: Iterate based on system performance indicator analysis and system feedback data.
10. A dynamic construction system for confidence intervals of pollution source automatic monitoring data, characterized in that: The system comprises: The monitoring data and video data acquisition unit is used to collect pollution source data, pre-process the data and extract behavioral features, quantify the behavioral features for time series correlation analysis, build a behavioral correlation graph, train the behavioral correlation graph, and identify behavioral patterns; A time series analysis unit, used to integrate data from data sources, perform time series analysis on the data, and generate statistical reports; wherein the data sources include sensor data, device status, and safety monitoring output; A machine learning prediction unit, which is used to monitor and analyze total equipment data using machine learning algorithms, predict pollution sources, and prevent environmental pollution; The energy allocation optimization unit is used to build a demand forecasting model and an environmental impact model based on real-time data and analyze the data to optimize energy allocation; the real-time data includes the temperature T, humidity H, flow L at the industrial production outlet, the energy consumption E of the industrial production equipment in the current period, and the energy consumption O in the same period in the past. Then, the real-time data is moved averaged to smooth short-term fluctuations, and the sliding average of each data is calculated. It is expressed as follows Among them, n represents the moving window size, x i represents the real-time data to be calculated, t represents the original timestamp, and i represents the lower bound of the window time; The demand forecasting model is Among them, a1, a2, a3, and a4 represent model parameters, which are learned from historical data through machine learning algorithms; b represents the bias term; represents the average temperature, Represents the average humidity value, represents the average flow rate. Indicates the average energy consumption in the same period in the past; E pred represents the predicted demand; The optimization scheduling strategy unit is used to build a multi-resource demand model to optimize the scheduling strategy based on the predicted demand and real-time data, define reinforcement learning to minimize the total cost, and then manage and schedule resources; the resources include electricity, manpower and mechanical equipment resources, and the power resources include solar P s 、Wind Energy w and mains electricity P g , the human resources include the number of staff N w The mechanical equipment resources include the number of machines in use N m and machine state; wherein, the definition of reinforcement learning and the design of an Actor-Critic network to train reinforcement learning specifically include: The state includes the energy supply P at the current time step s ,P w ,P g and the human and mechanical usage status N w ,N m ; Actions include decisions to adjust various energy inputs and resource allocations; The optimization function is the negative value of the cost function, and the goal is to minimize the total cost, expressed as: R=-(c s P s +c w P w +c g P g +c w N w +c m N m ) Among them, R represents the optimization function, c s ,c w ,c g ,c w ,c m Represent the unit costs of solar energy, wind energy, city electricity, manpower and machinery respectively; The flow route optimization unit is used to optimize and prevent the occurrence of environmental pollution through machine learning based on the data from the data source in the time series analysis unit.