Environment monitoring method and system based on big data analysis
By identifying and allocating storage nodes for industrial environmental monitoring data, a data cube is generated for multi-dimensional analysis, which solves the problems of low storage efficiency and weak correlation analysis capabilities in existing technologies, and realizes real-time and accurate monitoring of environmental conditions and accurate location of pollution sources.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies suffer from low storage efficiency and weak multi-dimensional correlation analysis capabilities when processing industrial environmental data, resulting in insufficient timeliness of environmental status monitoring and low accuracy in identifying abnormal events, thus failing to meet the needs of real-time environmental status monitoring.
By identifying the data type of monitoring data, allocating corresponding target storage nodes, and generating data cubes, we can achieve spatiotemporal network structure analysis of multi-dimensional data, and determine the environmental status and pollution source location.
It improves data storage efficiency, shortens the response time for retrieving key data, enables accurate positioning of environmental status and precise tracing of pollution sources, and supports intuitive decision-making in environmental management.
Smart Images

Figure CN120583122B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of data processing, and more particularly relates to an environment monitoring method and system based on big data analysis. BACKGROUND
[0002] In the field of intelligent environmental protection technology driven by industrial big data, environmental monitoring is the core foundation of building an ecological protection system. Currently, large-scale deployment of sensor networks continuously generates massive industrial environmental data, which presents the characteristics of multi-source heterogeneity and spatio-temporal coupling.
[0003] The existing technology has double technical bottlenecks in dealing with such data: on the one hand, there is a lack of hierarchical storage strategy for the characteristics of industrial environmental data, resulting in low storage efficiency of massive data and high delay in retrieval response of monitoring data; on the other hand, the traditional data processing framework cannot realize multi-dimensional correlation analysis — for example, it cannot effectively fuse multi-dimensional data such as pollutant concentration data, temperature, humidity, etc., making it difficult to capture the implicit correlation between data in the environmental state recognition process.
[0004] The technical problems directly caused by the above technical defects are: the processing timeliness of industrial environmental data is insufficient, which cannot support real-time environmental state monitoring requirements; the analysis capability of multi-dimensional data is weak, resulting in low recognition accuracy of environmental abnormal events, and further affecting the response speed and traceability accuracy of environmental pollution event warning. How to break through the technical bottlenecks of efficient storage and multi-dimensional correlation analysis of industrial big data, and realize real-time and accurate environmental state monitoring, has become a core technical problem that needs to be solved. SUMMARY
[0005] The present application aims to provide an environment monitoring method and system based on big data analysis to improve large-scale data processing capability and realize real-time and accurate monitoring of the environment.
[0006] The first aspect of the embodiment of the present application provides an environment monitoring method based on big data analysis, comprising:
[0007] Obtaining an original data stream corresponding to at least one monitoring node in an environment monitoring area, and determining a target storage node of each data stream according to the data type of the monitoring data in each data stream of the original data stream; wherein one data stream represents the monitoring data collected by one collection device in one monitoring node at one collection timestamp; the monitoring data in the data stream includes physical index data, image data or video data; the physical index data includes at least one index parameter data of temperature, humidity and pollutant concentration;
[0008] Processing the monitoring data in all target storage nodes to obtain a data cube corresponding to the original data stream;
[0009] determine monitoring results of each environment state in the environment monitoring area according to the spatio-temporal grid data in the data cube; the monitoring results include distribution information corresponding to each environment state and pollution source positioning information, and the environment state includes a normal state and an abnormal state.
[0010] In a second aspect, the embodiment of the application provides an environment monitoring system based on big data analysis, which comprises:
[0011] The acquisition module is used to acquire original data streams corresponding to at least one monitoring node in an environment monitoring area, and determine target storage nodes of each data stream according to data types of monitoring data in each data stream in the original data streams; wherein one data stream represents monitoring data collected by one collection device in one monitoring node at one collection timestamp; the monitoring data in the data stream includes physical index data, image data or video data; the physical index data includes data of at least one index parameter such as temperature, humidity and pollutant concentration;
[0012] The processing module is used to process monitoring data in all target storage nodes to obtain a data cube corresponding to the original data streams.
[0013] The determining module is used to determine monitoring results of each environment state in the environment monitoring area according to spatio-temporal grid data in the data cube; the monitoring results include distribution information corresponding to each environment state and pollution source positioning information, and the environment state includes a normal state and an abnormal state.
[0014] The environment monitoring method and system based on big data analysis provided by the embodiment of the application have the following advantages:
[0015] (1) The application assigns corresponding target storage nodes to different types of data streams by identifying data types of monitoring data in each data stream of the original data streams, solves the unified storage problem of multi-source heterogeneous data in industrial environment data, improves storage efficiency, and shortens the key data retrieval response time.
[0016] (2) The application generates a data cube by processing monitoring data, can integrate scattered data into spatio-temporal network structure data, forms a three-dimensional spatio-temporal network structure that integrates time, space and index parameters, and provides a standardized data model basis for multi-dimensional analysis of multi-dimensional data;
[0017] (3) The application determines monitoring results of each environment state by spatio-temporal network data, i.e. determines distribution information corresponding to each environment state and pollution source positioning information, realizes accurate positioning of each environment state in the environment monitoring area and accurate tracing of pollution sources, and provides intuitive decision support for environmental management. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating an environmental monitoring method based on big data analysis provided in an embodiment of this application;
[0020] Figure 2 This is a structural block diagram of an environmental monitoring system based on big data analysis provided in an embodiment of this application;
[0021] Figure 3 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0022] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0023] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.
[0024] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of an environmental monitoring method based on big data analysis provided in this application. The method includes:
[0025] S101. Obtain the raw data stream corresponding to at least one monitoring node in the environmental monitoring area, determine the target storage node for each data segment based on the data type of the monitoring data in each data stream of the raw data stream, and store each data segment based on each target storage node.
[0026] In this embodiment, the environmental monitoring area can be an industrial plant, a city functional area, or the like, and original data streams corresponding to at least one monitoring node in the environmental monitoring area can be acquired. The distributed sensor network can be deployed in the environmental monitoring area, and each monitoring node can be configured with multiple types of collection devices, such as sensors for collecting physical index data of index parameters such as temperature, humidity, and pollutant concentration, cameras for collecting image data, or cameras for collecting video data, and the like. The monitoring data collected by each collection device at a collection timestamp constitutes a data stream in the original data stream, and each data stream carries monitoring data, a timestamp corresponding to the monitoring data, a spatial position of the monitoring node, and identification information of the collection device; the monitoring data in the data stream includes physical index data, image data, or video data; and one data slice includes at least one data stream. The original data stream can include all monitoring data collected in a preset time period.
[0027] In this embodiment, after the original data stream is acquired, for each data stream in the original data stream, the monitoring data in the data stream can be acquired, and the data type of the monitoring data can be determined. If the data type of the monitoring data is structured data such as physical index data of temperature, humidity, and pollutant concentration, it is classified as a relational storage category; if the data type of the monitoring data is unstructured data such as image or video, it is marked as an object storage category, thereby obtaining an identification result corresponding to each data stream.
[0028] In this embodiment, after the identification result of each data stream is obtained, each data stream can be transmitted to a corresponding storage architecture according to the identification result of each data stream. The storage architecture can include a relational database cluster and a distributed file storage unit. For example, a data stream marked as a relational storage category can be transmitted to the relational database cluster; and a data stream marked as an object storage category can be transmitted to the distributed file storage unit and the target storage node corresponding to each data slice can be determined based on a consistent hashing algorithm.
[0029] In this embodiment, data slicing processing can be performed on data in different storage architectures to obtain multiple data slices, and each target storage node of each data slice in each storage architecture can be determined, and each data slice can be stored based on each target storage node.
[0030] In this embodiment, by classifying each data stream in the original data stream according to the data type of the monitoring data in each data stream and storing it in the corresponding target storage node, the targeted management of different types of monitoring data can be realized, the efficiency and standardization of data storage are improved, and a good data foundation is provided for subsequent analysis and processing of the monitoring data. At the same time, this classification storage method also facilitates the differential management and application of different types of data.
[0031] S102, processing the monitoring data in all target storage nodes to obtain a data cube corresponding to the original data stream.
[0032] In this embodiment, the monitoring data in each target storage node corresponding to the original data stream can be obtained, and the monitoring data is processed to obtain a data cube corresponding to the original data stream. Specifically, structured physical index data such as temperature, humidity, and pollutant concentration can be extracted from the database storage unit, and image data or video data can be obtained from the distributed file storage unit. Align and associate the two types of data according to the monitoring node identifier and timestamp to form a unified data set.
[0033] In this embodiment, the integrated data set can be standardized in the time and space dimensions. For example, the time dimension can be divided into time slices according to a preset granularity, such as 1 hour, and the space dimension can be divided into geographical grids according to spatial positions, such as 1km x 1km, each spatio-temporal grid corresponds to a unique spatio-temporal coordinate. For physical index data, it can be aggregated according to the spatio-temporal grid to calculate the mean, extreme value, and other statistics of each index parameter; for image data or video data, key frame features such as smoke probability can be extracted and mapped to the corresponding spatio-temporal grid.
[0034] In this embodiment, a multi-dimensional data cube can be constructed based on a spatiotemporal grid. Using time, space, and indicator parameters as the three basic dimensions, the time dimension includes all time slices, the space dimension includes all geographic grids, and the indicator parameter dimension includes physical indicators such as temperature, humidity, and pollutant concentration, as well as image feature parameters. Monitoring data from each spatiotemporal grid can be filled into the corresponding dimensional intersections according to the type of indicator parameter, forming a three-dimensional data array. For example, pollutant concentration can refer to PM2.5 concentration. In a certain spatiotemporal grid, the average PM2.5 concentration from 08:00:00 to 09:00:00 on 2025-06-08T is 50 μg / m³, corresponding to an image feature value of 0.7 and a smoke probability. These values are then stored in the corresponding coordinates of the data cube. During the construction of the data cube, missing data needs to be interpolated. For missing values in the time dimension, linear interpolation can be used to fill in the missing values; for missing values in the spatial dimension, the inverse distance weighted method of adjacent grid data can be used. Then, the interpolated data is standardized to convert physical index data with different dimensions into dimensionless values, such as Z-Score standardization, to ensure the consistency of data for each index parameter in the data cube.
[0035] The data cube obtained in this embodiment supports multi-dimensional slicing and drill-down analysis. Slicing along the time dimension reveals the trend of indicator parameters changing over time, slicing along the spatial dimension presents the spatial distribution of indicator parameters at a certain moment, and slicing along the indicator parameter dimension allows for comparison of the correlation between different indicator parameters. By processing the raw data stream to obtain the data cube, scattered monitoring data can be transformed into a unified spatiotemporal parameter model, providing a standardized data foundation for subsequent determination of environmental status and pollution source location, and improving the efficiency and accuracy of multi-dimensional analysis.
[0036] S103. Determine the monitoring results of each environmental state in the environmental monitoring area based on the spatiotemporal grid data in the data cube.
[0037] In this embodiment, the cross data of time dimension, spatial dimension and indicator parameter dimension can be extracted from the data cube to form spatiotemporal grid data; wherein, each spatiotemporal grid data can be used to represent the physical indicator statistical value within the spatiotemporal range corresponding to each grid cell in the data cube, such as the average pollutant concentration, temperature extreme value and image feature parameters.
[0038] In this embodiment, a sliding window mechanism can be used to determine the parameter dataset corresponding to each preset time window of the spatiotemporal grid data, and to determine the average pollutant concentration in the parameter dataset corresponding to each time window. Specifically, the average pollutant concentration corresponding to a time window can be the average pollutant concentration at each time point in the parameter dataset within that preset time window.
[0039] If the average is greater than a preset quality standard threshold, the average change rate of each monitoring data in the parameter data set in the time window relative to historical data is further determined, if there is a monitoring data corresponding to a change rate greater than a preset fluctuation threshold, the spatial position of the parameter data set corresponding to the time window is determined, and the environmental state of the region corresponding to the spatial position is determined as an abnormal state; otherwise, the environmental state of the spatial position is determined as a normal state; all parameter data sets corresponding to abnormal states can be classified based on a random forest algorithm to obtain a mild data set corresponding to a mild abnormal state and a severe data set corresponding to a severe abnormal state. The preset quality standard threshold can be determined based on the current environmental quality standard, for example, the pollutant concentration can refer to the PM2.5 concentration, and the PM2.5 concentration corresponding to the quality standard threshold can be set to 35 μg / m³, which can be set according to actual conditions, and the present application is not limited.
[0040] In the present embodiment, pollution source positioning analysis can be performed on the parameter data set of each abnormal state. The abnormal state can include a mild abnormal state and a severe abnormal state. The pollution source positioning information of the region corresponding to the mild abnormal state and the region corresponding to the severe abnormal state can be determined respectively, and the pollution source positioning information can include the coordinate range of the pollution source.
[0041] In the present embodiment, the distribution information of the key influence factor corresponding to the mild abnormal state and the distribution information of the key influence factor corresponding to the severe abnormal state can be determined according to the parameter data set of each abnormal state, and the distribution information of each index parameter corresponding to the normal state can be determined according to the parameter data set of each normal state. Through the multi-dimensional analysis method, accurate derivation from the space-time grid data to the pollution source positioning information can be realized, and reliable pollution source tracing basis for environmental management decision-making can be provided
[0042] In the embodiment, the current access right can be determined, and different access rights correspond to different data display intensities. For example, the ordinary user only displays the environment state overview and the pollution area distribution, and the administrator can view the pollution source accurate coordinates and the pollution contribution degree details. The data field matched with the current access right can be filtered through the permission verification tool, and the field is matched with the preset report template. The report template includes title, environment state summary, pollution source analysis, trend prediction, and suggestion measures, etc. Each block corresponds to different data filling rules. For example, the environment state summary block automatically extracts the area proportion of each state area and the typical pollutant concentration value; the pollution source analysis block displays the main pollution source according to the pollution contribution degree, and inserts the pollution diffusion path diagram. For image or video data, the key frame or video clip of the abnormal area is retrieved through the storage path index, and the corresponding analysis block of the report is embedded, such as inserting the monitoring image showing obvious pollution characteristics in the severe abnormal area analysis.
[0043] In an embodiment of the present application, the target storage node of each data slice is determined according to the data type of the monitoring data in each data stream of the original data stream, comprising:
[0044] For each data stream in the original data stream, the data type of the monitoring data in the data stream is obtained. If the data type of the monitoring data is structured data, the storage category of the data stream is determined as a relational storage category, the data stream is transmitted to a relational database cluster, and the data stream is subjected to field mapping processing to obtain structured records. If the data type of the monitoring data is unstructured data, the storage category of the data stream is determined as an object storage category, the data stream is transmitted to a distributed file storage unit to obtain a preliminary storage path.
[0045] The structured records corresponding to the original data stream and the data in the preliminary storage path corresponding to the original data stream are subjected to slicing processing respectively to obtain a plurality of data slices.
[0046] The target storage node corresponding to each data slice is determined based on a consistent hash algorithm.
[0047] In the embodiment, for each data stream in the original data stream, the data type of the monitoring data in the data stream can be obtained by analyzing the content of the data stream. If the monitoring data is structured data such as temperature and humidity, the storage category of the data stream is identified as a relational storage category; if the monitoring data is unstructured data such as image data or video data, the storage category of the data stream is identified as an object storage category, and the corresponding identification result is generated.
[0048] In this embodiment, each data stream can be distributed according to a preset distribution rule through a data routing distribution mechanism. The preset distribution rule is that, for each data stream in the original data stream, if the storage category of the monitoring data of the data stream is a relational storage category, the data stream is transmitted to a relational database cluster; at the same time, the data stream can be subjected to field mapping processing to obtain a structured record corresponding to the data stream. The fields in the data stream can include a timestamp, a spatial position of a collection device, a number of the data stream, an index parameter corresponding to the monitoring data, and a value of the monitoring data; if the storage category of the monitoring data of the data stream is an object storage category, the data stream is transmitted to a distributed file storage unit to obtain a preliminary storage path. For example, in a distributed sensor network in the field of environmental monitoring, the identification results of the respective storage categories of the structured data and the unstructured data can be subjected to data distribution according to the preset distribution rule. For structured data such as numerical information of temperature, humidity, etc., the data routing distribution mechanism can be used to transmit the data to the relational database cluster; assuming that a sensor node in an industrial park collects a temperature of 36.5 degrees Celsius and a humidity of 62.8%, these data can be directly mapped as fields in the database to form a structured record. For unstructured data such as on-site monitoring video, the system transmits the data to the distributed file storage unit to obtain a preliminary storage path such as “ / video / 2023 / 10 / 02 / vid001”.
[0049] In this embodiment, the data in the structured record corresponding to the original data stream in the relational database cluster and the preliminary storage path corresponding to the original data stream in the distributed file storage unit can be subjected to sharding processing respectively to obtain a plurality of data shards. The sharding processing can be based on a preset shard size. The preset shard size can be set in combination with the respective different storage architectures of the relational database cluster and the distributed file storage unit, and the specific value and setting logic are as follows: for the relational database cluster, according to the page size of the database therein, the preset shard size can be an integer multiple of the page size to avoid cross-page reading and writing, and the specific preset shard size can be determined according to actual conditions; for the distributed file storage unit, according to the block size therein, the preset shard size can be set to 1 / 2 to 2 times the block size to balance I / O efficiency and storage overhead, and the specific preset shard size can be determined according to actual conditions.
[0050] In this embodiment, a consistent hashing algorithm can be used to determine the target storage nodes corresponding to each data shard. The target storage nodes corresponding to each data shard can be determined based on the distribution information of the storage nodes in the storage architecture corresponding to the storage category of each data shard, and the storage architecture can include the relational database cluster and the distributed file storage unit.
[0051] In this embodiment, after obtaining the target storage nodes of each data shard, the storage node mapping table can also be determined according to the correspondence between each data shard and each target storage node. As an example, if data shard A1 is structured data and is allocated to target storage node B1 of the relational database cluster, and data shard A2 is unstructured data and is allocated to target storage node B2 of the distributed file storage unit, the correspondence between A1 and B1 and the correspondence between A2 and B2 can be added to the storage node mapping table. This storage node mapping table can ensure the efficiency of subsequent retrieval.
[0052] In an embodiment of the present application, after determining the target storage nodes of each data shard, the path index corresponding to each target node can also be determined.
[0053] For the processing of structured data records, the trend of the numerical value of each monitoring data index parameter over time can be analyzed. If the corresponding numerical value of any index parameter in the monitoring data exceeds the preset alarm threshold of the index parameter, the monitoring data can be marked with an abnormal alarm identifier, and the time and spatial location of the abnormal occurrence can be recorded. As an example, the temperature data gradually rises from 25.0 degrees Celsius to 35.2 degrees Celsius in the past 24 hours, exceeding the preset threshold range of 30.0 degrees Celsius, an abnormal alarm identifier can be generated, and the time and spatial location of the abnormal occurrence can be recorded, such as "October 1, 2023, 14:00, node A3". This helps to quickly locate the problem area and improve monitoring efficiency.
[0054] After the abnormal alarm is triggered, the corresponding monitoring data can be obtained according to the storage path index. Assuming that the image data retrieved through the path " / image / 2023 / 10 / 01 / img001" shows that there are smoke marks near node A3, image processing tools can be used to extract features to determine whether there are environmental abnormal features such as smoke density or color abnormalities. Combined with the time and spatial location of the abnormal data, it is finally confirmed that there is a potential fire risk at this spatial location. This correlation analysis effectively improves the accuracy of abnormal confirmation and avoids misjudgment that may be caused by relying solely on numerical data.
[0055] Fast retrieval through the path index corresponding to the target storage node can ensure that relevant data can be retrieved in time when the abnormality is confirmed. This storage allocation rule not only optimizes the data management efficiency, but also provides reliable data support for subsequent analysis. Overall, the combination of classified storage, abnormal judgment, and image feature extraction enables the environmental monitoring system to achieve accurate early warning and rapid response in complex scenarios, significantly improving the environmental safety guarantee capability.
[0056] In an embodiment of the present application, according to the preliminary storage path, the target storage node of the data stream is determined based on a consistent hashing algorithm, comprising:
[0057] For each data shard, the data shard is subjected to a hash calculation, and the calculated hash value is mapped to a preset virtual hash ring.
[0058] A first storage node closest to the hash value along the preset virtual hash ring is found clockwise, and if the load rate of the first storage node is less than a preset load threshold, the first storage node is determined as the target storage node of the data shard.
[0059] In the present application, for each data shard, the data shard can be subjected to a hash calculation. As an example, a key field in the data shard can be extracted, and an MD5 or SHA-256 hash function can be used to calculate the key field to generate a 28-bit or 256-bit hash value.
[0060] In the present application, the hash value obtained by hashing the data shard can be mapped to a preset virtual hash ring. The preset virtual hash ring can include a virtual hash ring corresponding to a relational database cluster and a virtual hash ring corresponding to a distributed file storage unit. The preset hash ring corresponding to the data shard can be determined according to the storage category of the data shard. If the storage category of the data shard is a relational storage category, mapping processing is performed based on the virtual hash ring corresponding to the relational database cluster, and if the storage category of the data shard is an object storage category, mapping processing is performed based on the virtual hash ring corresponding to the distributed file storage unit.
[0061] In the present embodiment, the first storage node closest to the hash value along the preset virtual hash ring can be found clockwise. After finding the first storage node, the load rate of the first storage node can be verified. As an example, the load rate can be calculated based on the CPU utilization rate, disk read / write rate, and remaining storage space of the first storage node, such as the weighted sum of the CPU utilization rate, disk read / write rate, and remaining storage space, and if the load rate is less than a preset load threshold, the node is determined as the target storage node; if the load rate is greater than or equal to the preset load threshold, the next node closest to the hash value is found clockwise, and the node is determined as the target storage node until the load rate of the node is less than the preset load threshold. The preset load threshold can be set according to the actual storage situation, such as 60%, or a dynamic adjustment mechanism can be set to automatically adjust the preset load threshold according to the data traffic in different time periods, for example, the threshold is reduced by 10% in peak period and increased by 10% in valley period, to better adapt to the load demand in different time periods.
[0062] In this embodiment, after determining the target storage node of the data shard, the data shard can be stored based on the target storage node, and the load rate of the target storage node in the storage architecture corresponding to the data shard is updated, and the storage rate of each storage node in the storage architecture to the data shard corresponding to the original data stream is determined, and if there is a difference between the storage rate of a storage node and the average storage rate of all nodes greater than a preset difference threshold, the data shards stored by each storage node in the storage architecture can be dynamically adjusted based on the load balancing tool. As an example, the preset difference threshold can be set according to the actual situation of each storage architecture, such as 15%, and during data transmission, if there is uneven storage of data shards in the storage nodes of the structured data in the relational database cluster, storage node B1 stores 70% of the data shards, and node B2 only stores 30% of the data shards, the average storage rate of the two nodes is 50%, and if the storage rate of node B1 is 70%, the difference between the average storage rate is 20%, which exceeds the preset difference threshold of 15%, dynamic adjustment is performed, and the storage node mapping table between the target storage node and the data shard is updated after adjustment. It can effectively avoid single node overload and improve the stability of data storage.
[0063] In this embodiment, for the adjusted target storage location, if a storage path conflict of unstructured data in the distributed file storage unit is detected, such as two video data being simultaneously allocated to the path " / video / 2023 / 10 / 02 / vid001", path reassignment can be performed. For example, one of the data paths is adjusted to " / video / 2023 / 10 / 02 / vid002", and the final physical storage node distribution result is determined. This ensures the uniqueness of data storage and reduces access conflicts.
[0064] The above storage allocation and adjustment process closely combines environmental monitoring business needs. Whether it is accurate recording of data or fast retrieval, it provides a reliable foundation for subsequent environmental anomaly analysis. For example, when temperature data is abnormal, the system can quickly locate the relevant records according to the storage node mapping table, and simultaneously retrieve the video data at the corresponding time point for comparison. This multi-dimensional data collaborative management approach significantly improves the response capability of the monitoring system, providing stronger protection for environmental safety.
[0065] In an embodiment of the present application, the monitoring data in all target storage nodes is processed to obtain a data cube corresponding to the original data stream, including:
[0066] The monitoring data in all target storage nodes is standardized to obtain a standardized monitoring data set;
[0067] Based on a distributed computing framework, the standardized monitoring dataset is divided into blocks according to a preset time window or a preset spatial region, to obtain a plurality of data blocks.
[0068] The data conversion processing is performed on each data block through parallel computing, the monitoring data corresponding to each data block is mapped into a grid unit corresponding to a pre-established space-time network framework, to obtain a converted gridded dataset.
[0069] The gridded dataset is subjected to data integration processing according to a preset dimension, to obtain a data cube; the preset dimension includes at least one of a time dimension, a space dimension and an index parameter dimension.
[0070] In this embodiment, the monitoring data can be subjected to standardized processing. For example, for physical index data such as temperature and humidity, the sliding window is used to detect abnormal values, and the forward filling method is used to complete the missing values; for image data or video data, the file integrity can be checked, such as automatically marking damaged pictures, and compressing to a unified resolution, such as 280x720. At the same time, the data unit is unified, such as converting concentration to pg / m3, the time format is ISO 8601, to form a standardized monitoring dataset.
[0071] In this embodiment, the standardized dataset can be subjected to block processing based on a distributed computing framework, such as Apache Spark. The data blocks are divided according to a preset time window, such as 1 hour, or a spatial region, such as 1kmx1km. When the data blocks are divided, the time stamp, spatial coordinates and other information of the data can be retained, to facilitate subsequent mapping.
[0072] In this embodiment, parallel computing can be performed through the distributed computing framework to map the data blocks to the space-time grid framework. For example, a space-time grid model can be pre-established, such as dividing the time dimension into time slices according to a 1-hour granularity, dividing the space dimension into equidistant grids according to the longitude and latitude, and each grid unit corresponds to a unique space-time coordinate. For each data block, the time stamp and spatial coordinates of the data are extracted and mapped into the corresponding grid unit, to form a gridded dataset. In the specific implementation process, if there is a lack of data in a certain grid region in a certain time period, the grid can be marked as a to-be-processed state, to facilitate subsequent adjustment.
[0073] In the embodiment, the gridded data set can be integrated according to a preset dimension to generate a data cube. A time dimension, a space dimension, and an index parameter dimension are taken as a three-dimensional framework, the time dimension includes all time slices, the space dimension includes all space grids, and the index parameter dimension includes physical indexes such as temperature, humidity, and pollutant concentration and image feature parameters such as smoke probability. The aggregated data of each grid unit is filled into a corresponding dimension intersection point. For example, the pollutant concentration can refer to the PM2.5 concentration, the time slice "08:00-09:00", and the space grid "116.40°E-116.41°E, 39.90°N-39.91°N" correspond to the PM2.5 concentration average value 50 μg / m³, which is stored in the corresponding coordinates of the data cube. During the integration, the missing data can be complemented in the space dimension by using the inverse distance weighting method or complemented in the time dimension by using the linear interpolation method, and Z-Score standardization processing is performed to ensure the dimensional consistency of different indexes, and finally a data cube supporting multi-dimensional slice analysis is formed.
[0074] In the embodiment, it can also be judged whether the final data cube meets the preset grid structure requirement. It can be checked whether each grid unit includes complete data dimensions. If the data dimensions in a grid unit are missing, such as missing data at a certain time point, the grid unit can be marked as unqualified, and a subsequent complement mechanism is triggered. This way ensures the integrity of the data structure and provides a reliable basis for subsequent analysis.
[0075] In an embodiment of the present application, the monitoring data in all target storage nodes is standardized to obtain a standardized monitoring data set, including:
[0076] The monitoring data in all target storage nodes is classified based on index parameters to obtain an index data set corresponding to each data type;
[0077] For each index data set, a missing value proportion of the index data set is determined. If the missing value proportion exceeds a preset proportion threshold, the index data set is interpolated based on a time sequence to obtain a processed first monitoring data set;
[0078] Data exceeding a preset normal range in the first monitoring data set is determined as abnormal data, and the abnormal data in the first monitoring data set is corrected to obtain a second monitoring data set;
[0079] The second monitoring data set is standardized to obtain a standardized monitoring data set.
[0080] In this embodiment, the monitoring data in all target storage nodes can be classified based on the index parameters, so as to obtain index data sets corresponding to each index parameter. For example, a temperature data set corresponding to temperature and a humidity data set corresponding to humidity are obtained.
[0081] In this embodiment, for each index data set, the proportion of missing values can be determined. If the proportion of missing values exceeds a preset proportion threshold, such as 10%, the monitoring data set is interpolated and completed based on time series to obtain a first processed monitoring data set. Taking PM2.5 data in a certain time period as an example, if the missing proportion reaches 20%, which exceeds the threshold, the missing values can be estimated and completed by linear interpolation and other methods according to the numerical trend of the previous and subsequent time points. As an example, taking an air quality data set as an example, assuming that the PM2.5 data missing proportion reaches 20%, which exceeds the preset 10% threshold. At this time, the system will start the time series interpolation tool, estimate the value of the missing time point according to the PM2.5 numerical trend of the previous and subsequent time points, such as completing the missing PM2.5 value to 35 micrograms per cubic meter, to form the completed data set.
[0082] In this embodiment, the data in the first monitoring data set that exceeds the preset normal range after the completion processing can be determined as abnormal data, and the abnormal data is corrected to obtain a second monitoring data set. If the air quality data set is completed, the noise value at a certain time point is 20 decibels, which is far beyond the normal range of 50-70 decibels, the historical data and surrounding node data can be combined to correct it to a reasonable value of 65 decibels by using statistical tools. Specifically, for the classified monitoring data set, the data preprocessing tool scans the missing values. This way ensures the continuity of the data and provides a complete basis for subsequent analysis.
[0083] In this embodiment, the second monitoring data set can be standardized to obtain a standardized monitoring data set. For example, the data unit is unified, the PM2.5 unit is unified to micrograms per cubic meter, the field name is standardized, such as “PM25” is unified to “PM2.5”, and the consistency and standardization of the data format are ensured, so as to facilitate subsequent data sharding management and analysis application in distributed storage.
[0084] Through classification, missing value processing, abnormal value correction and standardization processing of the monitoring data in the target storage node, the data quality can be effectively improved, and a reliable data basis is provided for subsequent analysis and application of environmental monitoring data.
[0085] In an embodiment of the present application, the monitoring result of the environmental state in the environmental monitoring area is determined according to the spatio-temporal grid data in the data cube, comprising:
[0086] Obtain the space-time grid data in the data cube according to the time dimension and the space dimension;
[0087] The space-time grid data is segmented and processed based on the preset time window by using a sliding window mechanism, to obtain a parameter data set corresponding to each time window.
[0088] The pollution source positioning information is standardized, and the mean value of the pollutant concentration in the parameter data set is determined; if the mean value is greater than a preset quality standard threshold, the state corresponding to the time window is determined as an abnormal state.
[0089] The average change rate of each monitoring data in the parameter data set of the abnormal state is determined, and if the average change rate of a monitoring data is greater than a preset fluctuation threshold, the parameter data set of the abnormal state is determined as a risk parameter set.
[0090] The pollution source positioning information in the monitoring result is determined according to the risk parameter set.
[0091] In this embodiment, the space-time grid data can be extracted from the data cube. By specifying the time dimension and the space dimension for slicing operation, all grid data in the corresponding space-time range can be obtained, and each grid data contains a data set corresponding to index parameters such as temperature, humidity, and PM2.5 concentration.
[0092] In this embodiment, the space-time grid data can be segmented and processed by using a sliding window mechanism. For example, the time window size can be set to 1 hour, and the sliding step is 30 minutes. The first time window can cover 00:00-01:00, the second time window covers 00:30-01:30, and so on, to obtain a parameter data set corresponding to each time window. For each time window, the pollutant concentration (such as PM2.5, SO2) can be extracted to form a parameter data set, and the mean value of all pollutant concentrations contained in each time window can be calculated. The mean value of the pollutant concentration can be compared with the quality standard threshold to determine whether the time window is in an abnormal state. If the mean value is greater than the preset quality standard threshold, the parameter data set of the time window is marked as an abnormal state. If the mean value is less than or equal to the preset quality standard threshold, the parameter data set of the time window is marked as a normal state. The preset quality standard threshold can be determined based on the current environmental quality standard. For example, the pollutant concentration can refer to the PM2.5 concentration, and the quality standard threshold corresponding to the PM2.5 concentration can be set to 35 μg / m³. The specific value can be set according to the actual situation, and the present application is not limited.
[0093] In the embodiment, the change rate of each monitoring data in the parameter data set in all time windows corresponding to each abnormal state can be further determined. If the change rate of one monitoring data is greater than the preset fluctuation threshold, such as the PM2.5 concentration of one monitoring data suddenly increases from 40 μg / m3 to 80 μg / m3, the change rate reaches 100%, and the preset fluctuation range is 30%, it indicates that the sudden pollution source may exist due to the sharp change of the pollutant concentration. The parameter data set of the abnormal state can be marked as a risk parameter set, an early warning signal is generated based on the trend early warning tool, and it is prompted that there may be an environmental risk. Alternatively, the early warning mechanism can remind the relevant personnel to pay attention to the potential problem in time, and enhance the response capability of the monitoring system. In addition, the generation of the early warning signal can also be based on the trend analysis combined with the historical data, and whether to generate the early warning signal is determined according to the analysis result. For example, by comparing the change rate of the pollutant concentration in the same time period in the past preset time period, such as one week, if the change rate shows a continuous upward trend, the early warning signal is generated. Through the above processing, not only the risk parameter set corresponding to the abnormal data can be effectively identified, but also the potential risk can be found in advance through the early warning mechanism, which provides an important reference for environmental monitoring and management.
[0094] In the embodiment, the mode recognition and category division of the risk parameter set can be performed based on the classification tool, such as being divided into two categories of mild abnormality and severe abnormality, and the key influence factors corresponding to each category are determined. The pollution source positioning information is determined according to the distribution information of the key influence factors in each category.
[0095] In one embodiment of the present application, the pollution source positioning information in the monitoring result is determined according to the risk parameter set, comprising:
[0096] The risk parameter set is divided into categories to obtain a mild risk parameter set corresponding to a mild abnormal environment state and a severe risk parameter set corresponding to a severe abnormal environment state;
[0097] The mode recognition is performed on each category obtained after the division to obtain a potential mode set corresponding to each category. Each mode in the potential mode set corresponding to one category is used to represent the association relationship between the multiple index parameters under the condition of the category;
[0098] For each potential mode set corresponding to each category, the spatial distribution and weight of each index parameter are determined. The index parameter with a weight greater than a preset weight threshold is determined as a key influence factor of the mode;
[0099] According to the spatial distribution of the key influence factors corresponding to each category, the current wind direction data, and the preset pollution diffusion model, the pollution source positioning information corresponding to the category is determined.
[0100] In this embodiment, based on the risk parameter set, a classification tool such as a random forest algorithm can be used for pattern recognition and classification. For example, the risk parameter set corresponding to an abnormal state can be divided into two categories, i.e., mild abnormality and severe abnormality. For example, when the PM2.5 concentration exceeds 50 μg / m3 and the temperature is higher than 25℃, it is a "severe abnormality" category corresponding to the mode.
[0101] In this embodiment, the potential mode set corresponding to each category can be determined, and the potential mode set can include one or more modes. Each mode in the potential mode set corresponding to a category is used to represent at least one association relationship between multiple index parameters under the condition of the category.
[0102] In this embodiment, for the potential mode set corresponding to each category, the height and weight of each index corresponding to the spatial distribution of each index parameter in the category can be determined. The index parameter with a weight greater than a preset weight threshold in the mode can be determined as a key influencing factor of the mode. The index parameters corresponding to each category can be determined according to each mode in the potential mode set corresponding to the category, and the spatial distribution of each index parameter can be determined according to the spatial position of each index parameter corresponding to the category. The weight of each index parameter can be determined based on the contribution of each index parameter to the category. The greater the weight of the index parameter, the higher the contribution of the index parameter to the classification. The index parameter with a weight greater than a preset weight threshold can be determined as a key influencing factor. As an example, the weight of each index parameter can be determined by determining the information gain corresponding to each index parameter according to the proportion of each index parameter in each category, and normalizing the information gain corresponding to each index parameter to obtain the weight of each index parameter. For example, the weight of PM2.5 concentration is 0.6, the weight of temperature is 0.2, and the weight of humidity is 0.1. If the preset weight threshold is 0.5, PM2.5 concentration is determined as the key influencing factor of the mode. The preset weight threshold can be determined according to actual conditions, which is not limited in the present application.
[0103] In this embodiment, the key influencing factors in each mode corresponding to each category can be determined, and the spatial distribution of all key influencing factors corresponding to each category can be determined. Based on the current wind direction data and a preset diffusion model such as a Gaussian diffusion model, a spatial correlation tool is used to obtain the pollution source distribution information corresponding to each category, i.e., the pollution source distribution information corresponding to the region of mild abnormality in the abnormal state, and the pollution source distribution information corresponding to the region of severe abnormality in the abnormal state. The pollution source distribution information can include the coordinate range of the pollution source distribution.
[0104] In the embodiment, through the whole process of data standardization, pattern recognition, key factor extraction and spatial correlation analysis, accurate derivation from multi-dimensional environmental data to pollution source positioning is realized. With the help of the feature analysis capability of algorithms such as random forest, combined with spatial distribution and pollution diffusion model, the accuracy of environmental monitoring and the efficiency of decision support are effectively improved, which provides a scientific basis for pollution control.
[0105] In an embodiment of the present application, the corresponding distribution information of each environmental state is determined, including:
[0106] In the embodiment, after obtaining the parameter data set corresponding to each time window according to the space-time grid data, and determining the mean value of the pollutant concentration in the parameter data set corresponding to each time window, if the mean value is less than or equal to the preset quality standard threshold, the parameter data set of the time window is marked as normal state. The parameter data set in the time window corresponding to all normal states can be standardized to obtain a standard normal parameter set, and the spatial distribution of each parameter index in the standard normal parameter set can be further determined. The spatial distribution of each parameter index can be determined as the spatial distribution information corresponding to the region with normal state of environmental state.
[0107] In the embodiment, after obtaining the risk parameter set and classifying the risk parameter set, the risk parameter set corresponding to the mild abnormal environmental state and the risk parameter set corresponding to the severe abnormal environmental state are obtained. The spatial distribution of each parameter index can be further determined according to the risk parameter set corresponding to the mild abnormal environmental state, and the spatial distribution of the parameter index can be determined as the spatial distribution of the environmental state in the mild abnormal state. The spatial distribution of each parameter index can be determined according to the risk parameter set corresponding to the severe abnormal environmental state, and the spatial distribution of the parameter index can be determined as the spatial distribution information corresponding to the region with severe abnormal state of environmental state.
[0108] In an embodiment of the present application, the environmental monitoring report of the environmental monitoring area can also be determined according to the pollution source positioning information corresponding to each environmental state, including:
[0109] According to the distribution information corresponding to each environmental state and the pollution source positioning information corresponding to the abnormal state, an initial data set corresponding to each environmental state is determined;
[0110] A preset output result and a preset output template corresponding to the current access permission level are determined;
[0111] Target data matching the preset output result is obtained from the initial data set;
[0112] A monitoring report is determined according to the target data and the preset output template;
[0113] If the current access permission level is greater than the preset permission threshold, a visual distribution view corresponding to the target data is determined.
[0114] In this embodiment, the initial data set corresponding to each environmental state can be determined according to the distribution information and the pollution source positioning information corresponding to each environmental state. The initial data set can include the distribution information corresponding to the normal state region, the distribution information corresponding to the abnormal state region, and the pollution source positioning information corresponding to the abnormal state region.
[0115] In this embodiment, after obtaining the initial data set, the preset output result and the preset output template corresponding to the current access permission level can be determined. Different access permission levels correspond to different output contents and format requirements. As an example, the normal user permission can only obtain the environmental state overview and the key influence factor information, and the corresponding output template is a concise report format; while the administrator permission can obtain more detailed pollution source positioning information and analysis data, and the corresponding output template is a format containing more data details and analysis contents.
[0116] In this embodiment, the target data corresponding to the preset output result can be obtained from the initial data set. According to the preset output result determined according to the current access permission level, each field of the initial data set is matched with the preset output template to obtain the target field presented in the monitoring report, and the target data corresponding to the field is extracted from the initial data set. For example, under the normal user permission, the environmental state summary data such as the average PM2.5 of each region and the pollution trend overview data such as the concentration change curve in the past week can be screened from the initial data set according to the preset output result corresponding to its access permission, and the detailed pollution source positioning information and emission data are ignored.
[0117] In this embodiment, if the current access permission level is greater than the preset permission threshold, such as access permission level 3 and above, a visual distribution view corresponding to the target data is determined. For example, the administrator user has an access permission level of 4, and after obtaining the detailed pollution source positioning data, an interactive visual distribution view can also be generated by superimposing the pollution source positioning information on the heat map corresponding to each environmental state, supporting zooming, querying and other operations, and intuitively displaying the monitoring results.
[0118] Through the standardized processing of the distribution information and the pollution source positioning information corresponding to each environmental state, and in combination with the user permission grading to generate differentiated monitoring reports, data security is guaranteed, and the decision-making needs of different users are met.
[0119] In an embodiment of the present application, if the current access permission is greater than the preset permission threshold, a visual distribution view corresponding to the target data is determined, comprising:
[0120] determine heat map data corresponding to the target data;
[0121] superimpose and fuse the heat map data with a map base corresponding to the environmental monitoring area to obtain a visual distribution view;
[0122] obtain at least one preset index parameter;
[0123] perform hierarchical division on the visual distribution view according to preset hierarchical thresholds corresponding to each preset index parameter to obtain a hierarchical distribution view corresponding to each preset hierarchical threshold.
[0124] In the embodiment, if the current access permission is greater than the preset permission threshold, the target data can be extracted in a structured manner, such as extracting index information such as pollutant concentration and spatial coordinates from the target data, associating and matching the extracted data with a pre-established spatial coordinate library to obtain a monitoring data view corresponding to a spatial position. A heat map generation tool is used to render the data view to obtain corresponding heat map data. The color mapping rule in the heat map data can be set according to the concentration value range, such as a higher concentration and a more red color, to generate a heat map data reflecting the spatial distribution of pollutants.
[0125] In the embodiment, the heat map data can be superimposed and fused with a map base corresponding to the environmental monitoring area to obtain a visual distribution view. For example, a WebGIS engine can be called to load a preset map base, such as a city vector map containing roads and buildings, and the heat map data can be superimposed on the base in the form of a layer, and the heat map can be accurately matched with the actual geographical area through transparency adjustment and coordinate matching.
[0126] In the embodiment, at least one preset index parameter can be obtained. The preset index parameter can be set by the user or randomly set in advance. The visual distribution view can be divided into hierarchical distribution views corresponding to preset hierarchical thresholds according to each preset index parameter. For example, suppose that the preset hierarchical threshold corresponding to the pollutant concentration is set, such as a concentration exceeding 50 micrograms per cubic meter being divided into a high-risk level and a concentration being less than 20 micrograms per cubic meter being divided into a low-risk level. A user with an access permission level greater than the preset permission threshold can view the detailed distribution of different risk levels by switching the preset hierarchical threshold, such as displaying not only the red heat map area but also the position icon and emission data label of the industrial pollution source in the high-risk level view.
[0127] In the embodiment, through the superimposed fusion of the heat map and the map base map, combined with the hierarchical division of the preset index parameters, the high-privilege user can obtain a multi-dimensional and refined visual distribution view, quickly drill from the macro pollution distribution to the specific pollution source influence range, and provide intuitive and accurate data support for environmental decision-making.
[0128] In an embodiment of the present application, after obtaining the hierarchical distribution view corresponding to each preset hierarchical threshold, the layered hierarchical distribution view can be matched with the display template of the decision dashboard by using a dashboard construction tool to obtain the decision dashboard content. Based on the dynamic monitoring demand, the decision dashboard content is updated in linkage with the interactive map to obtain a real-time decision visualization interface. As an example, assuming that the air quality data of a certain area, the data of the hierarchical distribution view corresponding to the high-risk level can be matched with the display template of the decision dashboard to form an interface containing the trend graph and the index parameters corresponding to the display template. At the same time, through the linkage update with the interactive map, the change of the pollutant concentration can be further displayed in real time. This dynamic monitoring method can help the decision maker to timely grasp the environmental state. In the construction process of the decision dashboard, matching the hierarchical distribution view with the display template can effectively improve the information presentation efficiency. For the presentation interface of real-time decision support, if the pollutant concentration of a certain area suddenly increases during the monitoring process, the content of the decision dashboard is refreshed in real time, and the color of the area corresponding to the heat map deepens, prompting the area corresponding to the high-risk level. This can ensure that the decision maker can quickly respond to potential environmental problems and improve management efficiency.
[0129] The environmental monitoring method based on big data analysis corresponding to the above embodiment, Figure 2 The structure block diagram of the environmental monitoring system based on big data analysis provided by an embodiment of the present application is shown in FIG. 1. For ease of illustration, only the parts related to the embodiments of the present application are shown. For reference Figure 2 The environmental monitoring system based on big data analysis 20 includes an acquisition module 21, a processing module 22, and a determination module 23.
[0130] The acquisition module 21 is configured to acquire the original data stream corresponding to at least one monitoring node in the environmental monitoring area, and determine the target storage node of each data stream according to the data type of the monitoring data in each data stream in the original data stream. One data stream represents the monitoring data collected by one collection device in one monitoring node at one collection timestamp. The monitoring data in the data stream includes physical index data, image data, or video data. The physical index data includes data of at least one index parameter such as temperature, humidity, and pollutant concentration.
[0131] The processing module 22 is configured to process the monitoring data in all target storage nodes to obtain a data cube corresponding to the original data stream.
[0132] The determining module 23 is configured to determine a monitoring result of each environment state in the environment monitoring area according to the spatio-temporal grid data in the data cube; the monitoring result includes distribution information corresponding to each environment state and pollution source positioning information, and the environment state includes a normal state and an abnormal state.
[0133] In an embodiment of the present application, the obtaining module 21 is further configured to: for each data stream in the original data stream, obtain a data type of the monitoring data in the data stream, if the data type of the monitoring data is structured data, determine that the storage category of the data stream is a relational storage category, transmit the data stream to a relational database cluster, perform field mapping processing on the data stream to obtain structured records; if the data type of the monitoring data is unstructured data, determine that the storage category of the data stream is an object storage category, transmit the data stream to a distributed file storage unit to obtain a preliminary storage path; perform sharding processing on the structured records corresponding to the original data stream and the data in the preliminary storage path corresponding to the original data stream respectively to obtain a plurality of data shards; and determine a target storage node corresponding to each data shard based on a consistent hashing algorithm.
[0134] The obtaining module 21 is further configured to: for each data shard, perform hash calculation on the data shard, and map the calculated hash value to a preset virtual hash ring; find a first storage node closest to the hash value along the preset virtual hash ring in a clockwise direction, and if the load rate of the first storage node is less than a preset load threshold, determine the first storage node as the target storage node of the data shard.
[0135] The processing module 22 is further configured to: perform standardization processing on the monitoring data in all target storage nodes to obtain a standardized monitoring data set; perform block processing on the standardized monitoring data set according to a preset time window or a preset space region based on a distributed computing framework to obtain a plurality of data blocks; perform data conversion processing on each data block through parallel computing, map the monitoring data corresponding to each data block to a grid cell corresponding to a pre-established spatio-temporal network framework, and obtain a converted gridded data set; and perform data integration processing on the gridded data set according to a preset dimension to obtain a data cube; the preset dimension includes at least one of a time dimension, a space dimension, and an index parameter dimension.
[0136] The processing module 22 is further configured to: classify the monitoring data in all target storage nodes based on the index parameters to obtain an index data set corresponding to each index parameter; determine, for each index data set, a proportion of missing values of the index data set; if the proportion of missing values exceeds a preset proportion threshold, perform interpolation completion processing on the index data set based on a time sequence to obtain a first processed monitoring data set; determine data in the first monitoring data set that exceeds a preset normal range as abnormal data, and obtain a second monitoring data set after performing correction processing on the abnormal data in the first monitoring data set; and perform standardization processing on the second monitoring data set to obtain a standardized monitoring data set.
[0137] The determining module 23 is further configured to: obtain space-time grid data in the data cube according to a time dimension and a space dimension; perform segmentation processing on the space-time grid data based on a preset time window by using a sliding window mechanism to obtain a parameter data set corresponding to each time window; determine, for each parameter data set corresponding to each time window, a mean value of the pollutant concentration in the parameter data set; if the mean value is greater than a preset quality standard threshold, determine a state of the parameter data set corresponding to the time window as an abnormal state; determine an average change rate corresponding to each monitoring data in the parameter data set in the abnormal state, and if the average change rate of a monitoring data is greater than a preset fluctuation threshold, determine the parameter data set in the abnormal state as a risk parameter set; and determine, according to the risk parameter set, pollution source positioning information in the monitoring result.
[0138] The determining module 23 is further configured to: determine an initial data set corresponding to each environmental state according to distribution information corresponding to each environmental state and pollution source positioning information corresponding to an abnormal state; determine a preset output result and a preset output template corresponding to a current access permission level; obtain target data matching the preset output result from the initial data set; determine a monitoring report according to the target data and the preset output template; and if the current access permission level is greater than a preset permission threshold, determine a visual distribution view corresponding to the target data.
[0139] The determining module 23 is further configured to: determine heat map data corresponding to the target data; perform superimposition and fusion processing on the heat map data and a map base map corresponding to the environmental monitoring area to obtain the visual distribution view; obtain at least one preset index parameter; perform hierarchical division on the visual distribution view according to each preset hierarchical threshold corresponding to each preset index parameter to obtain a hierarchical distribution view corresponding to each preset hierarchical threshold.
[0140] Referring to Figure 3 , Figure 3 A schematic block diagram of an electronic device is provided in an embodiment of the present application. As shown in Figure 3The electronic device 300 in the embodiment shown can include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 complete communication with each other through a communication bus 305. The memory 304 is configured to store a computer program, and the computer program includes program instructions. The processor 301 is configured to execute the program instructions stored in the memory 304. Specifically, the processor 301 is configured to invoke the program instructions to perform the functions of various modules / units in the above-mentioned device embodiments, for example Figure 2 The functions of the acquisition module 21, the processing module 22, and the determination module 23 shown.
[0141] It should be understood that, in the embodiments of the present application, the processor 301 can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0142] The input device 302 can include a touchpad, a fingerprint collection sensor (used to collect fingerprint information and direction information of a fingerprint of a user), a microphone, etc., and the output device 303 can include a display (LCD, etc.), a speaker, etc.
[0143] The memory 304 can include read-only memory and random access memory, and provide instructions and data for the processor 301. A portion of the memory 304 can also include non-volatile random access memory. For example, the memory 304 can also store device type information.
[0144] In specific implementations, the processor 301, the input device 302, and the output device 303 described in the embodiments of the present application can perform the implementation manners described in the first and second embodiments of the environment monitoring method based on big data analysis provided by the embodiments of the present application, and can also perform the implementation manners of the electronic device described in the embodiments of the present application, which will not be described here.
[0145] In another embodiment of the present application, a computer readable storage medium is provided, which stores a computer program. The computer program includes program instructions, which, when executed by a processor, implement all or part of the processes of the above-mentioned embodiment methods. The computer program can also instruct related hardware to complete the implementation. The computer program can be stored in a computer readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate form. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0146] The computer readable storage medium can be an internal storage unit of the electronic device of any of the preceding embodiments, such as a hard disk or a memory of the electronic device. The computer readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the computer readable storage medium can include both the internal storage unit and the external storage device of the electronic device. The computer readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer readable storage medium can also be used to temporarily store data that has been output or will be output.
[0147] Those skilled in the art can appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0148] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the electronic device and the units described above can refer to the corresponding processes in the above-mentioned method embodiments, which will not be described here.
[0149] In several embodiments provided in the present application, it should be understood that the disclosed electronic device and method can be implemented in other manners. For example, the embodiments of the apparatus described above are merely schematic, and the division of units is merely logical function division, and there can be other division manners in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between units can be indirect coupling or communication connection through some interfaces, or can be electrical, mechanical or other forms of connection.
[0150] The units described as separated components may or may not be physically separated, and the components displayed as units may or may not be physical units, i.e., may be located in one place, or may be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.
[0151] In addition, each functional unit in the various embodiments of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0152] The above is merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any skilled person in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An environmental monitoring method based on big data analysis, characterized in that, include: The process involves acquiring raw data streams corresponding to at least one monitoring node in an environmental monitoring area, determining the target storage node for each data segment based on the data type of the monitoring data in each data stream, and storing each data segment based on the target storage node. A data stream represents the monitoring data collected by a data acquisition device at a monitoring node at a given timestamp. A data segment includes at least one data stream, and the monitoring data in the data stream includes physical indicator data, image data, or video data. The physical indicator data includes data for at least one of the following parameters: temperature, humidity, and pollutant concentration. The monitoring data in all target storage nodes are processed to obtain the data cube corresponding to the original data stream; The monitoring results of each environmental state in the environmental monitoring area are determined based on the spatiotemporal grid data in the data cube; the monitoring results include the distribution information corresponding to each environmental state and the pollution source location information, and the environmental state includes normal state and abnormal state; The step of determining the monitoring results of each environmental state in the environmental monitoring area based on the spatiotemporal grid data in the data cube includes: The spatiotemporal grid data in the data cube is obtained based on the time and space dimensions; A sliding window mechanism is used to segment the spatiotemporal grid data based on a preset time window to obtain the parameter dataset corresponding to each time window; For each time window, the mean concentration of pollutants in the parameter dataset is determined; if the mean concentration is greater than a preset quality standard threshold, the state of the parameter dataset corresponding to the time window is determined to be an abnormal state. The average rate of change of each monitoring data in the parameter dataset of the abnormal state is determined. If the average rate of change of a monitoring data is greater than a preset fluctuation threshold, the parameter dataset of the abnormal state is determined as the risk parameter set. The pollution source location information in the monitoring results is determined based on the risk parameter set; The step of determining the pollution source location information in the monitoring results based on the risk parameter set includes: The risk parameter set is categorized to obtain a mild risk parameter set corresponding to a mildly abnormal environmental state and a severe risk parameter set corresponding to a severely abnormal environmental state. Pattern recognition is performed on each category after classification to obtain a set of latent patterns for each category; each pattern in the latent pattern set corresponding to a category is used to characterize the correlation between multiple index parameters under the condition of that category. For each category's potential pattern set, determine the spatial distribution and weight of each indicator parameter in the potential pattern set; identify the indicator parameters whose weights exceed a preset weight threshold as the key influencing factors of the category. Based on the spatial distribution of key influencing factors corresponding to each category, current wind direction data, and a preset pollution diffusion model, the location information of pollution sources corresponding to each category is determined.
2. The method according to claim 1, characterized in that, The step of determining the target storage node corresponding to each data shard based on the data type of the monitored data in each data stream of the original data stream includes: For each data stream in the original data stream, the data type of the monitored data in the data stream is obtained. If the data type of the monitored data is structured data, the storage category of the data stream is determined to be relational storage category, and the data stream is transmitted to a relational database cluster. Field mapping processing is performed on the data stream to obtain structured records. If the data type of the monitored data is unstructured data, the storage category of the data stream is determined to be object storage category, and the data stream is transmitted to a distributed file storage unit to obtain a preliminary storage path. The structured records corresponding to the original data stream and the data in the initial storage path corresponding to the original data stream are respectively fragmented to obtain multiple data fragments; The target storage node corresponding to each data shard is determined based on the consistent hashing algorithm.
3. The method according to claim 2, characterized in that, The step of determining the target storage node corresponding to each data shard based on the consistent hashing algorithm includes: For each data shard, a hash calculation is performed on the data shard, and the calculated hash value is mapped to a preset virtual hash ring; The first storage node closest to the hash value is found clockwise along the preset virtual hash ring. If the load rate of the first storage node is less than the preset load threshold, the first storage node is determined as the target storage node for the data shard.
4. The environmental monitoring method based on big data analysis according to claim 1, characterized in that, The process of processing the monitoring data in all target storage nodes to obtain the data cube corresponding to the original data stream includes: The monitoring data in all target storage nodes are standardized to obtain a standardized monitoring dataset. Based on a distributed computing framework, the standardized monitoring dataset is divided into blocks according to a preset time window or a preset spatial region to obtain multiple data blocks; By performing data transformation processing on each data block through parallel computing, the monitoring data corresponding to each data block is mapped to the grid cell corresponding to the pre-established spatiotemporal network framework to obtain the transformed gridded dataset. The data cube is obtained by integrating the gridded dataset according to a preset dimension; the preset dimension includes at least one of the following: time dimension, spatial dimension, and indicator parameter dimension.
5. The environmental monitoring method based on big data analysis according to claim 4, characterized in that, The standardization process is performed on the monitoring data in all target storage nodes to obtain a standardized monitoring dataset, including: The monitoring data in all target storage nodes are classified based on the indicator parameters to obtain the indicator dataset corresponding to each indicator parameter. Interpolation and completion processing is performed on each indicator dataset to obtain the processed first monitoring dataset. Data in the first monitoring dataset that exceeds the preset normal range is identified as abnormal data. The abnormal data is then corrected to obtain the second monitoring dataset. The second monitoring dataset is standardized to obtain a standardized monitoring dataset.
6. The environmental monitoring method based on big data analysis according to claim 1, characterized in that, The method further includes: Based on the distribution information corresponding to each environmental state and the pollution source location information corresponding to the abnormal state, the initial data set corresponding to each environmental state is determined. Determine the preset output results and preset output templates corresponding to the current access permission level; Obtain target data that matches the preset output result from the initial data set; A monitoring report is determined based on the target data and the preset output template; If the current access permission level is greater than the preset permission threshold, then a visualization distribution view corresponding to the target data is determined.
7. The method according to claim 6, characterized in that, If the current access permission is greater than a preset permission threshold, then determining the visualization distribution view corresponding to the target data includes: Determine the heat map data corresponding to the target data; The heat map data is overlaid and fused with the base map corresponding to the environmental monitoring area to obtain a visual distribution view; Obtain at least one preset indicator parameter; Based on the preset level thresholds corresponding to each preset index parameter, the visualization distribution view is divided into levels to obtain the level distribution view corresponding to each preset level threshold.
8. An environmental monitoring system based on big data analysis, characterized in that, An environmental monitoring method based on big data analysis as described in any one of claims 1 to 7, comprising: The acquisition module is used to acquire the raw data stream corresponding to at least one monitoring node in the environmental monitoring area, and determine the target storage node for each data stream based on the data type of the monitoring data in each data stream. A data stream represents the monitoring data collected by a data acquisition device in a monitoring node at a single acquisition timestamp. The monitoring data in the data stream includes physical indicator data, image data, or video data. The physical indicator data includes data for at least one of the following parameters: temperature, humidity, and pollutant concentration. The processing module is used to process the monitoring data in all target storage nodes to obtain the data cube corresponding to the original data stream; The determination module is used to determine the monitoring results of each environmental state in the environmental monitoring area based on the spatiotemporal grid data in the data cube; the monitoring results include the distribution information corresponding to each environmental state and the pollution source location information, and the environmental state includes normal state and abnormal state.
Citation Information
Patent Citations
Distributed storage method and system for mass monitoring data of ecological and natural disasters
CN118364013A
Land space planning environment influence monitoring method and device
CN119848703A