A pollution visualization data processing system and method applying cloud computing technology

By analyzing the data correlation between monitoring stations and determining the data processing sequence, the problem that cloud platforms cannot efficiently process air pollution data under large-scale monitoring is solved, and more efficient data processing and more accurate monitoring results are achieved.

CN119202063BActive Publication Date: 2025-05-27呼和浩特市生态环境监控中心
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411403020.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-09
Publication Date
2025-05-27
Estimated Expiration
2044-10-09

AI Technical Summary

Technical Problem

The existing cloud platforms cannot take into account the processing of air pollution data at all monitoring stations in the first time under large-scale monitoring, resulting in insufficient processing efficiency and real-time performance.

Method used

By analyzing the correlation between data between monitoring stations, the relevant monitoring stations of each monitoring station are determined, and the data processing sequence is determined based on the parallel data processing capabilities of the cloud platform, and the method of processing part of the monitoring station data in real time and processing all monitoring station data regularly is adopted.

Benefits of technology

It has achieved the improvement of the efficiency of data processing of monitoring stations under large-scale monitoring in terms of taking into account real-time and accuracy, reducing the risk of abnormal situations being ignored at monitoring locations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119202063B_ABST
    Figure CN119202063B_ABST
Patent Text Reader

Abstract

The present invention discloses a pollution visualization data processing system and method applying cloud computing technology, which relates to the technical field of pollution data processing, obtains air pollution data from monitoring stations; analyzes the correlation between data among monitoring stations to determine the relevant monitoring stations of each monitoring station; determines the contribution rate of a single monitoring station to all monitoring stations and the contribution rate of all monitoring stations to a single monitoring station based on the correlation between data among monitoring stations; determines the processing sequence of the air pollution data of the monitoring stations by the cloud platform, generates a first processing result and a second processing result and performs visual display; analyzes the correlation between air pollution data among different monitoring stations, performs real-time processing on the data of some monitoring stations to reduce the data processing burden of the cloud platform; takes into account both real-time performance and security, regularly processes the data of all monitoring stations to reduce the possibility of ignoring abnormal situations at monitoring locations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pollution data processing, and specifically to a pollution visualization data processing system and method applying cloud computing technology. Background Art

[0002] The atmospheric environment is closely related to human life. By monitoring air pollution, the pollution status in the environment can be understood and grasped, the environmental quality and ecological risks can be evaluated, providing a scientific basis for environmental protection and sustainable development; under large-scale monitoring, a large number of standard monitoring stations and micro-monitoring stations are designed. The speed at which the monitoring stations obtain air pollution data is often very high, which requires the system to have extremely high processing efficiency and real-time performance. However, the parallel data processing ability of the cloud platform is limited and cannot take into account all the monitoring stations in the first place. Therefore, how to improve the data processing effect of the monitoring stations under large-scale monitoring has become an urgent problem to be solved. Summary of the Invention

[0003] The purpose of the present invention is to provide a pollution visualization data processing system and method applying cloud computing technology to solve the problems raised in the above background art.

[0004] In one aspect of the present invention, a pollution visualization data processing method applying cloud computing technology is provided, including:

[0005] S11, obtaining air pollution data from the monitoring stations;

[0006] S12, based on the air pollution data of the monitoring stations, analyzing the correlation between the data of the monitoring stations, and determining the relevant monitoring stations of each monitoring station;

[0007] S13, based on the correlation between the data of the monitoring stations, determining the contribution rate of a single monitoring station to all monitoring stations and the contribution rate of all monitoring stations to a single monitoring station;

[0008] S14, based on the parallel data processing ability of the cloud platform and the correlation between the data of the monitoring stations, determining the processing order of the air pollution data of the monitoring stations by the cloud platform, generating a first processing result and a second processing result and performing visual display; the first processing result is a real-time processing result generated by the cloud platform using the air pollution data of some monitoring stations, and the second processing result is a processing result generated by the cloud platform using the air pollution data of all monitoring stations.

[0009] Air pollution data has a huge amount of data and a high degree of complexity, including various types of pollution factors, time series data, etc. This brings great pressure to the analysis and transmission of data. The air pollution data monitoring activities carried out in a large geographical area involve a large number of ground monitoring stations, making the characteristics of air pollution data more prominent. In the case of limited cloud computing resources, in order to reflect the air pollution situation in real time, the data of some monitoring stations are analyzed to obtain the real-time air pollution situation, and the data of all monitoring stations are analyzed in the idle time, taking into account the requirements of real-time and accuracy.

[0010] In step S12, the steps of analyzing the correlation between the data of the monitoring stations and determining the relevant monitoring stations of each monitoring station further include:

[0011] S20. For i = 1, 2,..., n, where n is the number of monitoring stations, execute steps S21 and S22 to obtain each monitoring station and its corresponding association set. The monitoring stations in the association set are the relevant monitoring stations of each monitoring station;

[0012] S21. For the air pollution data A of the i-th monitoring station i , regress the air pollution data of the i-th monitoring station through the air pollution data of other monitoring stations to obtain the first regression error e1, A i = w 1 A 1 + w 2 A 2 +... + w n A n , where w 1 , w 2 ,..., w n are regression coefficients, and A 1 , A 2 ,..., A n are the air pollution data of the monitoring stations;

[0013] S22. If the regression coefficient is greater than the significance threshold, add the monitoring station corresponding to the regression coefficient to the association set C i of the i-th monitoring station; if the regression coefficient w j is not greater than the significance threshold, perform a significance monitoring on the air pollution data A j of the monitoring station corresponding to the regression coefficient: Remove the air pollution data A j from the input, and regress the air pollution data of the i-th monitoring station through the air pollution data of other monitoring stations except i and j to obtain the second regression error e2; j is different from i;

[0014] If the increase amplitude of e2 relative to e1 is less than the set value, then the air pollution data Aj Fail to pass the significance test, and the j-th monitoring station does not belong to the association set C of the i-th monitoring station i ; If the increase amplitude of e2 is not less than the set value, then the air pollution data A j passes the significance test, and the j-th monitoring station is added to the association set C of the i-th monitoring station i .

[0015] When the regression coefficient w j is too small, it indicates that the air pollution data A j may not contribute to A i , or it may be due to the large value of the air pollution data A j itself, which makes the regression coefficient w j too small; Therefore, a significance test is performed on the air pollution data A j to determine whether the air pollution data A j and the regression coefficient w j are really needed.

[0016] In step S13, the determination of the contribution rate of a single monitoring station to all monitoring stations and the contribution rate of all monitoring stations to a single monitoring station further includes the following steps:

[0017] S31, for the i-th monitoring station B i and the monitoring stations B i in the association set C k , where k is different from i, use the air pollution data of all monitoring stations in the association set C i to perform regression on the air pollution data of the monitoring station B i to obtain the first regression value P1A i and the first regression accuracy Acc1, The first regression accuracy Acc1 is the contribution rate B i of all monitoring stations to a single monitoring station B sum→i ; Change the value of i to obtain the contribution rate of all monitoring stations to any monitoring station; Use the air pollution data of all monitoring stations outside the monitoring station B i in the association set C k to perform regression on the air pollution data of the monitoring station B i to obtain the second regression value P2A i and the second regression accuracy Acc2, Calculate the contribution rate B k of the monitoring station B i to the monitoring station B k→i according to Acc1 and Acc2, B k→i =Acc1 - Acc2; When the monitoring station B k is not in the monitoring station Bi associated set C i When B k→i is 0, there is no need to perform calculations;

[0018] S32. While keeping k unchanged, perform the above steps for different cases of i to obtain the contribution rate B of monitoring station B k to all monitoring stations k→sum ,

[0019] S33. Change the value of k and execute steps S21 and S22 to obtain the contribution rate of any monitoring station to all monitoring stations.

[0020] When there is no air pollution data for some monitoring stations, partial information of the monitoring stations without air pollution data can be obtained from the air pollution data of the existing monitoring stations. The information volume of all other monitoring stations is obtained from the existing monitoring stations as an index of the monitoring station, that is, the contribution rate of a single monitoring station to all monitoring stations. There is no need to obtain the specific information of all monitoring stations, and the information of all monitoring stations can be obtained through some monitoring stations; the contribution rate of all monitoring stations to a single monitoring station is used as another index to reflect the particularity of the information of a single monitoring station. If the contribution rate of all monitoring stations to a single monitoring station is low, it means that the information of a single monitoring station is difficult to obtain from other monitoring stations, and it is more important to maintain real-time monitoring of a single monitoring station.

[0021] In step S14, the determination of the processing order of the air pollution data of the monitoring stations by the cloud platform further includes the following steps:

[0022] S41. Set the initial temperature as T0, randomly select m monitoring stations from n monitoring stations as the alternative monitoring station set, where m is less than n. Take the alternative monitoring station set as the initial solution, and determine the loss value L0 and contribution value F0 corresponding to the initial solution; and take the initial solution as the current solution and the initial temperature as the current temperature;

[0023] S42. For u = 1, 2,..., num, repeat steps S43 to S44; L is the set number of loops;

[0024] S43. Generate a perturbation based on the current solution to change the alternative monitoring station set, and take the generated alternative monitoring station set after perturbation as the new solution, and determine the loss value L and contribution value F corresponding to the new solution;

[0025] S44. Calculate the contribution value increment ΔF brought by the new solution. If the increment ΔF is less than 0, directly accept the new solution

[0026] as the new current solution. If the increment ΔF is not less than 0, then accept the new solution as the new current solution with probability where T represents the current temperature;

[0027] S45. Lower the current temperature. If the current temperature is not less than the set temperature threshold, go to step S42; if the current temperature is less than the set temperature threshold, determine the set of alternative monitoring stations based on the current solution, obtain the optimization result, and complete the optimization of the processing sequence of air pollution data for the monitoring stations.

[0028] The cloud platform processes the air pollution data of the set of alternative monitoring stations in real time to generate a first processing result. For other monitoring stations outside the set of alternative monitoring stations, the air pollution data may not be processed for a long time, and abnormal situations cannot be detected in a timely manner. Therefore, it is necessary to compensate for the optimization result to reduce the possibility of ignoring abnormal situations.

[0029] S46. Compensate for the optimization result. The cloud platform processes the air pollution data in real time according to the processing sequence of the air pollution data of the compensated monitoring stations to generate a first processing result.

[0030] It is characterized in that in steps S41 and S43, the loss value and the contribution value are determined in the following manner:

[0031] In step S41, let the set of alternative monitoring stations corresponding to the initial solution be D, and the loss value L of the initial solution is calculated by the following formula: LOS i is the loss rate of the air pollution data of monitoring station B i . In the formula, B represents the contribution rate of set D to monitoring station B D→i . Determine the intersection C i ∩D of the associated set C i of set D and monitoring station B i . Use the air pollution data of the monitoring stations in the intersection C i ∩D to perform regression on monitoring station B i . The obtained accuracy is B i . After flipping the loss value L, the contribution value F0 is obtained. D→i ;

[0032] In step S43, the generation methods of the loss value L and the contribution value F corresponding to the new solution are the same as those of the loss value L0 and the contribution value F0 of the initial solution.

[0033] When the air pollution data of monitoring station B i is processed in real time by the cloud platform, no information loss occurs, so the loss rate is 0; when the air pollution data of monitoring station B i is not processed in real time by the cloud platform, only part of the information of monitoring station B i can be reflected by the air pollution data of the monitoring stations processed in real time by the cloud platform, and a loss rate will occur.

[0034] In step S43, the generation of perturbations based on the current solution further includes the following steps:

[0035] Add the contribution rate of a single monitoring station to all monitoring stations and the contribution rate of all monitoring stations to a single monitoring station as the score of the single monitoring station; among the monitoring stations included in set D, determine the monitoring stations to exit set D based on the score and a random number. Then, among the monitoring stations outside set D, determine the monitoring stations to join set D based on the score and a random number.

[0036] Compensating the optimization result further includes the following steps:

[0037] Obtain the information of the set of alternative monitoring stations and the contribution value information generated during the optimization process other than the optimization result. Let the set of alternative monitoring stations corresponding to the optimization result be G, and the sets of alternative monitoring stations generated during the optimization process be G1, G2, … GX, where X is the number of sets of alternative monitoring stations generated during the optimization process. Let the set formed by all monitoring stations be the universal set U, and the complement set C of G in U U G includes all monitoring stations outside set G. Respectively determine the complement sets C of G1, G2, …, GX in U U G1, C U G2, …, C U GX. Select sets GL1, GL2, … GLX from sets G1, G2, … GX such that the union of the selected partial sets in the complement set C of U U GL1, C U GL2, …, C U formed by GLX contains C U GL1 ∪ C U GL2 ∪, …, ∪ C U GLX contains C U G, Use the monitoring station data included in G, GL1, GL2, … GLX as the real-time processing objects of the cloud platform.

[0038] The cloud platform needs to monitor the air pollution data of all monitoring stations. For the data generated by other monitoring stations outside set G, regular monitoring is also required, which requires the selected set to include all the remaining other monitoring stations outside set G. After determining the set to be processed, only the set needs to be updated regularly without repeating the determination.

[0039] In another aspect of the present invention, there is provided a pollution visualization data processing system applying cloud computing technology, including a ground monitoring station module, a cloud platform, and a visualization module; the output end of the ground monitoring station module is connected to the input end of the cloud platform, and is used to obtain air pollution data at the monitoring location; the output end of the cloud platform is connected to the visualization module, and is used to determine the processing sequence of the air pollution data of the monitoring stations, perform real-time processing on the air pollution data of the monitoring stations, and send the processing results to the visualization module; the visualization module is used to render and draw the air pollution data, and display the analysis results of the air pollution data.

[0040] The ground monitoring station module further includes an air pollution data collection unit, a data transmission unit, and a time unit; the air pollution data collection unit is used to obtain air pollution data at the monitoring location, the data transmission unit is used to send the air pollution data and the collection time to the cloud platform; the time unit is used to determine the collection time of the air pollution data. The cloud platform further includes an air pollution data real-time analysis unit, an air pollution data precise analysis unit, a data storage unit, a regression unit, an optimization unit, a compensation unit, and an output unit; the air pollution data real-time analysis unit is used to perform real-time analysis on the air pollution data transmitted by the monitoring stations; the air pollution data precise analysis unit is used to analyze the air pollution data transmitted by all monitoring stations during idle time; the data storage unit is used to store the air pollution data transmitted by the monitoring stations; the regression unit is used to determine the air pollution data of the monitoring stations, determine the correlation between the monitoring stations, the contribution rate of a single monitoring station to all monitoring stations, and the contribution rate of all monitoring stations to a single monitoring station; the optimization unit is used to determine the best combination method of the monitoring station data to be processed; the compensation unit is used to compensate the optimization result to ensure that all monitoring station data can be processed by the cloud platform; the output unit is used to send the air pollution data analysis results to the visualization module. The visualization module further includes a first visualization unit and a second visualization unit; the first visualization unit is used to render and draw the air pollution data analysis results in real time; the second visualization unit is used to render and draw the air pollution data analysis results of all monitoring stations.

[0041] Compared with the prior art, the beneficial effects achieved by the present invention are: analyzing the correlation of air pollution data between different monitoring stations, performing real-time processing on the data of some monitoring stations, reducing the data processing burden of the cloud platform; taking into account real-time performance and security, processing the data of all monitoring stations regularly, reducing the possibility of ignoring abnormal situations at the monitoring locations. Description of the Drawings

[0042] The accompanying drawings are used to provide a further understanding of the present invention and form a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0043] Figure 1 It is a schematic structural diagram of a pollution visualization data processing system applying cloud computing technology according to an embodiment of the present invention. Detailed implementation manners

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0045] In an embodiment of the present invention, please refer to Figure 1 , a pollution visualization data processing system applying cloud computing technology is provided, including: a ground monitoring station module, a cloud platform, and a visualization module; the output end of the ground monitoring station module is connected to the input end of the cloud platform for obtaining air pollution data at the monitoring location; the output end of the cloud platform is connected to the visualization module for determining the processing sequence of the air pollution data of the monitoring station, performing real-time processing on the air pollution data of the monitoring station, and sending the processing result to the visualization module; the visualization module is used for rendering and drawing the air pollution data and displaying the analysis result of the air pollution data.

[0046] The ground monitoring station module further includes an air pollution data acquisition unit, a data transmission unit, and a time unit; the air pollution data acquisition unit is used for obtaining air pollution data at the monitoring location, the data transmission unit is used for sending the air pollution data and the acquisition time to the cloud platform; the time unit is used for determining the acquisition time of the air pollution data.

[0047] The cloud platform further includes a real-time air pollution data analysis unit, a precise air pollution data analysis unit, a data storage unit, a regression unit, an optimization unit, a compensation unit, and an output unit; the real-time air pollution data analysis unit is used to perform real-time analysis on the air pollution data transmitted by the monitoring stations; the precise air pollution data analysis unit is used to analyze the air pollution data transmitted by all monitoring stations during idle time; the data storage unit is used to store the air pollution data transmitted by the monitoring stations; the regression unit is used to determine the air pollution data of the monitoring stations, determine the correlation between the monitoring stations, the contribution rate of a single monitoring station to all monitoring stations, and the contribution rate of all monitoring stations to a single monitoring station; the optimization unit is used to determine the best combination method of the monitoring station data to be processed; the compensation unit is used to compensate the optimization result to ensure that all monitoring station data can be processed by the cloud platform; the output unit is used to send the air pollution data analysis result to the visualization module.

[0048] The visualization module further includes a first visualization unit and a second visualization unit; the first visualization unit is used to render and draw the air pollution data analysis result in real time; the second visualization unit is used to render and draw the air pollution data analysis result of all monitoring stations.

[0049] In an embodiment of the present invention, a pollution visualization data processing method applying cloud computing technology is provided, including:

[0050] S11, obtaining air pollution data from the monitoring stations;

[0051] S12, based on the air pollution data of the monitoring stations, analyzing the correlation between the data of the monitoring stations, and determining the relevant monitoring stations of each monitoring station;

[0052] The method includes the following steps:

[0053] S20, for i = 1, 2,..., n, where n is the number of monitoring stations, execute steps S21 and S22 to obtain each monitoring station and its corresponding association set, and the monitoring stations in the association set are the relevant monitoring stations of each monitoring station;

[0054] S21, for the air pollution data A i of the i-th monitoring station, regress the air pollution data of the i-th monitoring station through the air pollution data of other monitoring stations to obtain the first regression error e1, A i = w 1 A 1 + w 2 A 2 +... + w n A n , where w 1 , w 2 ,..., wn is the regression coefficient, A 1 , A 2 , …, A n are the air pollution data of the monitoring stations;

[0055] S22, if the regression coefficient is greater than the significance threshold, add the monitoring station corresponding to the regression coefficient to the association set C of the ith monitoring station i ; if the regression coefficient w j is not greater than the significance threshold, perform significance monitoring on the air pollution data A of the monitoring station corresponding to the regression coefficient: remove the air pollution data A j from the input, and perform regression on the air pollution data of the ith monitoring station through the air pollution data of other monitoring stations except i and j to obtain the second regression error e2; j is different from i; j If the increase amplitude of e2 relative to e1

[0056] is less than the set value, the air pollution data A fails the significance test, and the jth monitoring station does not belong to the association set C of the ith monitoring station j ; if the increase amplitude of e2 i is not less than the set value, the air pollution data A passes the significance test, and add the jth monitoring station to the association set C of the ith monitoring station j ; i

[0057] The air pollution data collected by the monitoring stations is in the form of a time series. The air pollution data generated by each monitoring station at the same time point is taken as a set of data. Taking A i as the output and the data of other monitoring stations as the input, input them into the linear regression model, and solve and iterate the regression coefficients of the linear regression model with multiple sets of data at different time points; after the regression coefficient solving and iteration are completed, obtain the regression value of A i , calculate the absolute error between the regression value of A i and A i to obtain the regression error;

[0058] The significance threshold can be set to 0.05, and the set value of the increase amplitude of e2 can be set to 5%. When the regression coefficient is less than 0.05, perform significance detection. When the increase amplitude of e2 is less than 5%, it fails the significance detection, and set the regression coefficient to 0.

[0059] S13, based on the correlation between the data of the monitoring stations, determine the contribution rate of a single monitoring station to all monitoring stations and the contribution rate of all monitoring stations to a single monitoring station, including the following steps:

[0060] ​S31. For the \(i\)-th monitoring station \(B\) i and its associated set \(C\) i in the monitoring stations \(B\) k , where \(k\neq i\), use the air pollution data of all the monitoring stations in the associated set \(C\) i to regress the air pollution data of the monitoring station \(B\) i to obtain the first regression value \(P1A\) i and the first regression accuracy \(Acc1\). The first regression accuracy \(Acc1\) is the contribution rate \(B\) i of all monitoring stations to a single monitoring station \(B\) sum→i ; change the value of \(i\) to obtain the contribution rates of all monitoring stations to any monitoring station; use the air pollution data of all the monitoring stations outside the monitoring station \(B\) i in the associated set \(C\) k to regress the air pollution data of the monitoring station \(B\) i to obtain the second regression value \(P2A\) i and the second regression accuracy \(Acc2\). Calculate the contribution rate \(B\) k of the monitoring station \(B\) i to the monitoring station \(B\) k→i according to \(Acc1\) and \(Acc2\), where \(B\) k→i = \(Acc1 - Acc2\); when the monitoring station \(B\) k is not in the associated set \(C\) i of the monitoring station \(B\) i , \(B\) k→i is 0 and no calculation is required.

[0061] S32. Without changing \(k\), perform the above steps for different values of \(i\) to obtain the contribution rate \(B\) k of the monitoring station \(B\) k→sum to all monitoring stations.

[0062] S33. Change the value of \(k\) and execute steps S21 and S22 to obtain the contribution rates of any monitoring station to all monitoring stations.

[0063] For example, there are 4 monitoring stations \(B\) 1 , \(B\) 2 , \(B\) 3 and \(B\) 4 , where \(B\) 2 , \(B\) 3 belong to the associated set \(C\) 1 of \(B\) 1 , and at the same time \(B\) 2 belongs to the associated set \(C\) 3 of \(B\) 3 , and \(B\) 2 belongs to \(B\) 4The associated set C 4 , then for B 1 , by using B 2 , B 3 to perform regression on B 1 , the contribution rate of all monitoring stations to a single monitoring station B 1 can be obtained. Similarly for B 2 , B 3 and B 4 , use the monitoring station data in the associated set to perform regression and the contribution rates of all monitoring stations to a single monitoring station B 2 , B 3 and B 4 ;

[0064] The contribution rate of a single monitoring station B 2 to all monitoring stations B k→sum , first only use B 3 to perform regression on B 1 . Compared with using B 2 , B 3 to perform regression on B 1 , the regression accuracy will decrease. The decreased part is the contribution rate of B 2 to B 1 . The contribution rate of B 2 to B 3 is determined in the same way. The contribution rate of B 2 to B 4 is 0. Add the contribution rates of B 2 to B 1 , B 3 and B 4 to obtain the contribution rate of a single monitoring station B 2 to all monitoring stations B k→sum ;

[0065] The contribution rates of a single monitoring station B 3 and B 4 to all monitoring stations can be determined by repeating the process of B 2 .

[0066] S14. Based on the parallel data processing ability of the cloud platform and the correlation of data between monitoring stations, determine the processing order of air pollution data of the cloud platform for monitoring stations, generate a first processing result and a second processing result and perform visual display; the first processing result is the real-time processing result generated by the cloud platform using the air pollution data of some monitoring stations, and the second processing result is the processing result generated by the cloud platform using the air pollution data of all monitoring stations.

[0067] Determining the processing order of air pollution data of the cloud platform for monitoring stations further includes the following steps:

[0068] S41. Set the initial temperature as T0. Randomly select m monitoring stations from n monitoring stations as the set of candidate monitoring stations, where m < n. Take the set of candidate monitoring stations as the initial solution, and determine the loss value L0 and contribution value F0 corresponding to the initial solution. Also, take the initial solution as the current solution and the initial temperature as the current temperature.

[0069] S42. For u = 1, 2,..., num, repeat steps S43 to S44; L is the set number of loops.

[0070] S43. By generating perturbations based on the current solution, change the set of candidate monitoring stations. Take the set of candidate monitoring stations generated after perturbation as the new solution, and determine the loss value L and contribution value F corresponding to the new solution.

[0071] S44. Calculate the contribution value increment ΔF brought by the new solution. If the increment ΔF < 0, directly accept the new solution

[0072] as the new current solution. If the increment ΔF ≥ 0, then accept the new solution as the new current solution with probability where T represents the current temperature.

[0073] S45. Reduce the current temperature. If the current temperature is not less than the set temperature threshold, go to step S42; if the current temperature is less than the set temperature threshold, determine the set of candidate monitoring stations according to the current solution to obtain the optimization result, and complete the optimization of the air pollution data processing sequence of the monitoring stations.

[0074] S46. Compensate for the optimization result. The cloud platform processes the air pollution data in real time according to the air pollution data processing sequence of the compensated monitoring stations to generate the first processing result.

[0075] Optionally, the initial temperature is set to 100, and the temperature is reduced by 5 after each perturbation, that is, the current temperature changes from 100 to 95 after the first perturbation. The temperature threshold is set to 10. When the current temperature is less than 10, the optimization of the air pollution data processing sequence is completed.

[0076] In step S41, let the set of candidate monitoring stations corresponding to the initial solution be D, and the loss value L of the initial solution is calculated by the following formula: LOS i is the air pollution data loss rate of monitoring station B i , where B D→i represents the contribution rate of the set D to monitoring station B i . Determine the intersection C i of the associated set C i of the set D and monitoring station B i with D, and use the intersection C iThe air pollution data of the monitoring stations in ∩D for monitoring station B i is regressed, and the obtained accuracy is B D→i ; after flipping the loss value L, the contribution value F0 is obtained;

[0077] In step S43, the generation methods of the loss value L and the contribution value F corresponding to the new solution are the same as those of the loss value L0 and the contribution value F0 of the initial solution;

[0078] When the cloud computing resources are limited, that is, the cloud platform can only process the air pollution information of 3 monitoring stations simultaneously, the set D of alternative monitoring stations corresponding to the initial solution includes B 1 、B 2 and B 3 , then the loss rates of B 1 、B 2 and B 3 are 0, while the loss rate LOS 4 of B 4 is determined according to B 1 、B 2 and B 3 's regression results for B 4 , and the final loss value L0 is LOS 4 ; the method of flipping the loss value can be selected according to requirements. Optionally, the contribution value is set to the reciprocal of the loss value, that is a is a coefficient; or, the contribution value is set to a linear function of the loss value, that is, F0 = CON - bL0, where CON is a bias and b is a coefficient;

[0079] In step S43, the generation of perturbations based on the current solution further includes the following steps:

[0080] Add the contribution rate of a single monitoring station to all monitoring stations and the contribution rate of all monitoring stations to a single monitoring station as the score of the single monitoring station; among the monitoring stations included in the set D, determine the monitoring station to exit the set D based on the score and a random number, and then, among the monitoring stations outside the set D, determine the monitoring station to join the set D based on the score and a random number.

[0081] For convenience, continue with the previous example and add 1 monitoring station B 5 , at this time there are 4 monitoring stations B 1 、B 2 、B 3 、B 4 and B 5 , and the set D corresponding to the initial solution includes monitoring stations B 1 、B 2 、B 3 ; let f 1 、f 2 、f 3, f 4 and f 5 represents B 1 , B 2 , B 3 , B 4 and B 5 's score value, f 1 = B 1→sum + B sum→1 , f 2 , f 3 , f 4 and f 5 is similar to f 1 Let f max represent the maximum value among f 1 , f 2 , f 3 , f 4 and f 5 Then generate random numbers 3 times in [0, f], corresponding to B 1 , B 2 , B 3 ; f is slightly greater than f max to prevent the situation that the monitoring station is always retained in set D; if the random number is within [0, f 1 , then B 1 continues to be retained in set D, otherwise B 1 exits set D. After judging the monitoring stations that need to exit set D among B 1 , B 2 , B 3 , when the number of monitoring stations exiting set D is 1, then generate a random number in the interval [0, f 4 + f 5 . When the random number is within [0, f 4 , then add B 4 to set D to complete one perturbation; when the number of monitoring stations exiting set D is 2, B 4 and B 5 are added to the set together; when the number of monitoring stations exiting set D is 3, regenerate random numbers to judge the monitoring stations that need to exit set D among B 1 , B 2 , B 3 .

[0082] Compensating the optimization result also includes the following steps:

[0083] Obtain the information of the set of alternative monitoring stations and the contribution value information generated during the optimization process other than the optimization result. Let the set of alternative monitoring stations corresponding to the optimization result be G, and the sets of alternative monitoring stations generated during the optimization process be G1, G2, … GX, where X is the number of sets of alternative monitoring stations generated during the optimization process. Let the set formed by all monitoring stations be the universal set U, and the complement C of G in U U G includes all monitoring stations outside the set G. Respectively determine the complements C of G1, G2, …, GX in U U G1, C U G2, …, C U GX. Select sets GL1, GL2, … GLX from the sets G1, G2, … GX such that the union of the selected partial sets in the complement C of U U GL1, C U GL2, …, C U GLX formed contains C U GL1 ∪ C U GL2 ∪, …, ∪ C U GLX contains C U G, Take the monitoring station data included in G, GL1, GL2, … GLX as the real-time processing objects of the cloud platform.

[0084] The method of selecting the sets GL1, GL2, … GLX can adopt the sampling method or the greedy algorithm to sequentially select the sets including as many monitoring stations in the complement C of G in U U The set of monitoring stations in G. The cloud platform processes the monitoring station data included in G, GL1, GL2, … GLX in real time. The processing order can be determined according to the average contribution value of the sets GL1, GL2, … GLX and the contribution value of G. When the contribution value of G is F and the average contribution value of the sets GL1, GL2, … GLX is AF, the processing rounds can be generated according to the frequency of F to AF, and the processing order of the cloud platform to process the monitoring station data can be determined according to the processing rounds; when it comes to the set G, the cloud platform reads the monitoring station data in G for analysis; when it comes to the sets GL1, GL2, … GLX, select a set from the sets GL1, GL2, … GLX in order and read the corresponding monitoring station data for analysis.

[0085] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0086] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A pollution visualization data processing method using cloud computing technology, characterized in that: The following steps are involved: S11, obtain air pollution data from monitoring stations; S12, based on the air pollution data of the monitoring stations, analyzing the correlation of the data between the monitoring stations, and determining the relevant monitoring stations of each monitoring station, including steps S20, S21 and S22; S20, executing steps S21 and S22 for i=1, 2, ..., n, where n is the number of monitoring stations, to obtain each monitoring station and a corresponding associated set, wherein the monitoring stations in the associated set are associated monitoring stations of each monitoring station; S21, for the air pollution data A of the i-th monitoring station i , the atmospheric pollution data of the ith monitoring station is regressed through the atmospheric pollution data of other monitoring stations to obtain the first regression error e1, A i =w1A1+w2A2+…+w n A n , where w1, w2, …, w n are regression coefficients, A1, A2, …, A n Air pollution data from monitoring stations; S22, if the regression coefficient is greater than the significance threshold, the monitoring station corresponding to the regression coefficient is added to the associated set C of the i-th monitoring station i If the regression coefficient w j If the regression coefficient is not greater than the significance threshold, then the atmospheric pollution data A of the monitoring station corresponding to the regression coefficient j Conduct significant monitoring: Air pollution data A j Remove from the input, regress the air pollution data of the i-th monitoring station through the air pollution data of other monitoring stations except i and j, and obtain the second regression error e2; j is different from i; If the increase of e2 relative to e1 is If the air pollution data A is less than the set value, j Failed to pass the significance test, the jth monitoring station does not belong to the associated set C of the i-th monitoring station i ; If the increase in e2 If the air pollution data A is not less than the set value, j Through saliency detection, the jth monitoring station is added to the associated set C of the i-th monitoring station. i middle; S13, based on the correlation of data between the monitoring stations, determining the contribution rate of a single monitoring station to all monitoring stations and the contribution rate of all monitoring stations to a single monitoring station; S14, based on the cloud platform's parallel data processing capabilities and the correlation between data between monitoring stations, determine the cloud platform's processing order for the air pollution data of the monitoring stations, generate a first processing result and a second processing result and display them visually; the first processing result is a real-time processing result generated by the cloud platform using the air pollution data of some monitoring stations, and the second processing result is a processing result generated by the cloud platform using the air pollution data of all monitoring stations.

2. According to claim 1, a pollution visualization data processing method using cloud computing technology is characterized in that: In step S13, the step of determining the contribution rate of a single monitoring station to all monitoring stations and the contribution rate of all monitoring stations to a single monitoring station further includes the following steps: S31, for the i-th monitoring station B i and associated set C i Monitoring station B k , k is different from i, use the associated set C i The air pollution data of all monitoring stations in the monitoring station B i The atmospheric pollution data is regressed to obtain the first regression value P1A i and the first regression accuracy Acc1, The first regression accuracy Acc1 is the value of all monitoring stations for a single monitoring station B. i Contribution rate B sum→i ; Change the value of i to obtain the contribution rate of all monitoring stations to any monitoring station; use the association set C i Central Monitoring Station B k The air pollution data of all monitoring stations except i The atmospheric pollution data is regressed to obtain the second regression value P2A i and the second regression accuracy Acc2, Calculate the monitoring station B based on Acc1 and Acc2 k For monitoring station B i Contribution rate B k→i , B k→i =Acc1-Acc2; when monitoring station B k Not at monitoring station B i The association set C i In the middle, B k→i If it is 0, no calculation is required; S32, while keeping k unchanged, perform the above steps for different situations of i to obtain monitoring station B k Contribution rate to all monitoring stations B k→sum , S33, change the value of k, execute steps S21 and S22, and obtain the contribution rate of any monitoring station to all monitoring stations.

3. The pollution visualization data processing method using cloud computing technology according to claim 2 is characterized in that: In step S14, determining the order in which the cloud platform processes the air pollution data of the monitoring station further includes the following steps: S41, setting the initial temperature to T0, randomly selecting m monitoring stations from n monitoring stations as a set of candidate monitoring stations, where m is less than n, taking the set of candidate monitoring stations as an initial solution, determining the loss value L0 and contribution value F0 corresponding to the initial solution; and taking the initial solution as the current solution, and taking the initial temperature as the current temperature; S42, for u=1, 2, ..., num, repeat steps S43 to S44; L is the number of cycles set; S43, by generating disturbance on the basis of the current solution, changing the set of candidate monitoring stations, taking the set of candidate monitoring stations generated after the disturbance as a new solution, and determining the loss value L and contribution value F corresponding to the new solution; S44, calculate the contribution value increment ΔF brought by the new solution. If the increment ΔF is less than 0, directly accept the new solution as the new current solution. If the increment ΔF is not less than 0, then Accept the new solution as the new current solution, where T represents the current temperature; S45, lowering the current temperature. If the current temperature is not less than the set temperature threshold, proceed to step S42; if the current temperature is less than the set temperature threshold, determine a set of candidate monitoring stations based on the current solution, obtain an optimization result, and complete the optimization of the air pollution data processing order of the monitoring station; S46, compensating the optimization result, the cloud platform processes the air pollution data in real time according to the air pollution data processing order of the compensated monitoring station, and generates a first processing result.

4. The method for processing pollution visualization data using cloud computing technology according to claim 3 is characterized in that: In steps S41 and S43, the loss value and contribution value are determined in the following manner: In step S41, let the set of candidate monitoring stations corresponding to the initial solution be D, and the loss value L of the initial solution is calculated by the following formula: LOS i For monitoring station B i The loss rate of air pollution data, Where B D→i Denotes the set D for monitoring station B i The contribution rate of set D and monitoring station B is determined i The association set C i The intersection of i ∩D, use intersection C i The air pollution data of the monitoring station in ∩D is i Perform regression and the accuracy obtained is B D→i ; After flipping the loss value L, we get the contribution value F0; In step S43, the loss value L and contribution value F corresponding to the new solution are generated in the same way as the loss value L0 and contribution value F0 of the initial solution; In step S43, the step of generating disturbance based on the current solution further includes the following steps: The contribution rate of a single monitoring station to all monitoring stations and the contribution rate of all monitoring stations to a single monitoring station are added together as the score of the single monitoring station; among the monitoring stations included in set D, the monitoring stations to exit set D are determined based on the scores and random numbers, and then, among the monitoring stations outside set D, the monitoring stations to join set D are determined based on the scores and random numbers.

5. The method for processing pollution visualization data using cloud computing technology according to claim 4 is characterized in that: Compensation for the optimization results also includes the following steps: Obtain the set information and contribution value information of the candidate monitoring stations generated during the optimization process in addition to the optimization results. Let the set of candidate monitoring stations corresponding to the optimization results be G, and the set of candidate monitoring stations generated during the optimization process be G1, G2, ... GX, where X is the number of sets of candidate monitoring stations generated during the optimization process. Let the set formed by all monitoring stations be the full set U, and the complement of G in U be C. U G includes all monitoring stations outside the set G, and determines the complement set C of G1, G2, ..., GX in U respectively. U G1, C U G2, ..., C U GX, select sets GL1, GL2, ... GLX from sets G1, G2, ... GX so that the selected partial sets are in the complement of U C U GL1, C U GL2, ..., C U GLX forms a union C U GL1∪C U GL2∪,…,∪C U GLX includes C U G, The monitoring station data contained in G, GL1, GL2, ... GLX are used as real-time processing objects of the cloud platform.

6. A pollution visualization data processing system using cloud computing technology, using a pollution visualization data processing method using cloud computing technology as described in any one of claims 1 to 5, characterized in that: include: A ground monitoring station module, a cloud platform and a visualization module; the output end of the ground monitoring station module is connected to the input end of the cloud platform to obtain the air pollution data of the monitoring location; the output end of the cloud platform is connected to the visualization module to determine the processing order of the air pollution data of the monitoring station, perform real-time processing on the air pollution data of the monitoring station, and send the processing results to the visualization module; the visualization module is used to render and draw the air pollution data and display the analysis results of the air pollution data.

7. The pollution visualization data processing system using cloud computing technology according to claim 6 is characterized in that: The ground monitoring station module also includes an air pollution data collection unit, a data transmission unit and a time unit; the air pollution data collection unit is used to obtain the air pollution data of the monitoring location, and the data transmission unit is used to send the air pollution data and the collection time to the cloud platform; the time unit is used to determine the collection time of the air pollution data.

8. The pollution visualization data processing system using cloud computing technology according to claim 6 is characterized in that: The cloud platform also includes an air pollution data real-time analysis unit, an air pollution data precision analysis unit, a data storage unit, a regression unit, an optimization unit, a compensation unit and an output unit; the air pollution data real-time analysis unit is used to perform real-time analysis on the air pollution data transmitted by the monitoring station; the air pollution data precision analysis unit is used to analyze the air pollution data transmitted by all monitoring stations in spare time; the data storage unit is used to store the air pollution data transmitted by the monitoring station; the regression unit is used to determine the air pollution data of the monitoring station, determine the correlation between the monitoring stations, the contribution rate of a single monitoring station to all monitoring stations and the contribution rate of all monitoring stations to a single monitoring station; the optimization unit is used to determine the best combination of monitoring station data that needs to be processed; the compensation unit is used to compensate for the optimization results to ensure that all monitoring station data can be processed by the cloud platform; the output unit is used to send the air pollution data analysis results to the visualization module.

9. The pollution visualization data processing system using cloud computing technology according to claim 8 is characterized in that: The visualization module also includes a first visualization unit and a second visualization unit; the first visualization unit is used to render and draw the air pollution data analysis results in real time; The second visualization unit is used to render and draw the air pollution data analysis results of all monitoring stations.

Citation Information

Patent Citations

  • Air quality forecasting system

    CN106651036A