A coal bunker environment intelligent monitoring method and system based on data fusion
Patent Information
- Application Number
- CN202610814154.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-08
- Publication Date
- 2026-09-11
AI Technical Summary
首先,多数方案采用简单的加权融合或基于规则引擎的决策驱动方式,仅实现数据层面的物理叠加,缺乏对不同模态数据之间深层语义关联的挖掘能力
Smart Images

Figure CN122736308A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of industrial intelligent monitoring technology, specifically relating to a method and system for intelligent monitoring of coal bunker environment based on data fusion. Background Technology
[0002] With the rapid development of Industrial Internet of Things (IIoT) technology, coal bunkers, as key storage nodes in the coal supply chain, are directly affected by the safety monitoring and intelligent management of their internal environment, impacting production efficiency and operational safety. Coal bunker environmental monitoring typically involves the real-time acquisition and comprehensive analysis of heterogeneous data from multiple sources, including temperature, gas concentration, dust content, video surveillance, and audio signals. Due to the diverse data types, varying acquisition frequencies, and complex spatial distribution, effectively integrating this heterogeneous data to comprehensively and accurately perceive the environmental status of the coal bunker has become a pressing technical challenge in this field.
[0003] Several multi-sensor data fusion solutions have emerged for coal bunkers or similar industrial scenarios. However, these solutions generally suffer from the following shortcomings. First, most solutions employ simple weighted fusion or rule-based decision-driven approaches, achieving only physical overlay at the data level and lacking the ability to mine deep semantic relationships between different modalities of data.
[0004] Secondly, the fusion weights often remain unchanged during system operation, constituting a static weighting strategy. This makes it impossible to dynamically adjust the weight coefficients based on the historical trends of each sensor's data, hindering its adaptation to the gradual evolution of hazards in coal bunker environments. Furthermore, existing technologies generally do not fully consider the topological characteristics of sensor spatial distribution and lack analysis of correlations between sensor data from different spatial locations, resulting in a lack of effective data support for hotspot identification and hazard diffusion path prediction. More critically, existing fusion analyses are mostly limited to a single dimension, either time or space, failing to synergistically integrate time and spatial weights, thus limiting the accuracy and reliability of comprehensive risk assessment.
[0005] To address the aforementioned technical deficiencies, this invention provides a data fusion-based intelligent monitoring method and system for coal bunker environments. The aim is to achieve accurate assessment and early warning of coal bunker environmental risks through cross-modal feature alignment, temporal dynamic weighting, spatial topology correlation analysis, and spatiotemporal weight synergistic fusion. Summary of the Invention
[0006] The purpose of this invention is to provide a method and system for intelligent monitoring of coal bunker environment based on data fusion, which can effectively solve the problems in the background art mentioned above.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0008] A method for intelligent monitoring of coal bunker environment based on data fusion includes the following specific steps:
[0009] Cross-modal feature extraction is performed on the collected coal bunker environmental data, which includes temperature data, gas concentration data, dust content data, video data, and audio data.
[0010] The extracted modal feature vectors are uniformly mapped to a feature space of a preset dimension to obtain the standardized feature vectors corresponding to each data source.
[0011] The historical temporal evolution features of each sensor data are extracted using a temporal convolutional network, and the temporal dimension fusion weights of each data source are dynamically calculated based on the historical temporal evolution features using a multi-head attention mechanism.
[0012] Based on the spatial location information of each sensor node, a sensor spatial distribution topology map is constructed. A graph neural network is used to analyze the spatial correlation between sensor data at different spatial locations. Combined with the hotspot area identification and hazard diffusion path prediction results, the spatial dimension fusion weight of each data source is determined.
[0013] The time dimension fusion weight and the spatial dimension fusion weight are synergistically fused to calculate the comprehensive fusion coefficient of each data source;
[0014] Based on the comprehensive fusion coefficient, the standardized feature vectors corresponding to each data source are weighted and fused to obtain the fused feature vector;
[0015] Based on the fused feature vector, a comprehensive risk index is output through a risk assessment model.
[0016] Furthermore, the preset dimension feature space is a 256-dimensional feature space.
[0017] The step of uniformly mapping the extracted modal feature vectors to a feature space of a preset dimension includes:
[0018] The 32-dimensional feature vectors of temperature data, 64-dimensional feature vectors of gas concentration data, 32-dimensional feature vectors of dust content data, 96-dimensional feature vectors of video data, and 32-dimensional feature vectors of audio data are mapped to 256 dimensions through their respective fully connected mapping layers. Then, L2 norm normalization is performed on each of the mapped 256-dimensional intermediate feature vectors to obtain the standardized feature vectors corresponding to each data source.
[0019] Furthermore, the temporal convolutional network adopts a 4-layer causal convolutional structure, with the number of convolutional kernels in the 1st to 4th layers being 32, 64, 128 and 256 respectively, the kernel size being 3, and the dilation factors being 1, 2, 4 and 8 respectively.
[0020] The step involves extracting historical temporal evolution features of data from each sensor using a temporal convolutional network, and dynamically calculating the temporal dimension fusion weights of each data source based on these historical temporal evolution features using a multi-head attention mechanism. This includes:
[0021] The temporal feature vectors output from each data source after processing by the temporal convolutional network are combined into a feature matrix. The feature matrix is then linearly transformed to generate a query matrix, a key matrix, and a value matrix. The scaled dot product attention is calculated to obtain an attention weight matrix. The attention weight matrix is then averaged row by row and normalized using the softmax function to obtain the temporal dimension fusion weights of each data source.
[0022] Furthermore, the construction of the sensor spatial distribution topology map includes:
[0023] Using each sensor node as a vertex of the graph, when the Euclidean distance between two sensor nodes is less than a preset distance threshold, an edge is established between the two sensor nodes. The weight of the edge between the two sensor nodes is inversely proportional to the Euclidean distance between the two sensor nodes.
[0024] The method of using graph neural networks to analyze the spatial correlation between sensor data from different spatial locations includes:
[0025] The graph neural network is configured as a 3-layer graph convolutional structure, with each layer containing 128 hidden units. The first layer first maps the feature vectors of each sensor node to 128 dimensions. Then, the stacked graph convolutional layers aggregate the features of the central sensor node and the features of the adjacent sensor nodes of the central sensor node, and output the spatial correlation feature vector of each sensor node.
[0026] Furthermore, the hotspot area identification and hazard spread path prediction include:
[0027] When the current measurement value of a sensor node exceeds 1.5 times the historical average value of the sensor node, the sensor node is determined to be a potential hotspot node, and multiple adjacent potential hotspot nodes in the sensor spatial distribution topology map form a hotspot group.
[0028] Starting from the hotspot area node, the probability of danger spreading to adjacent sensor nodes is calculated based on the weight of the edges in the sensor spatial distribution topology graph. Iterative diffusion is performed through breadth-first search until the diffusion probability is lower than a preset termination threshold. The sensor nodes and edges involved in the diffusion process are organized into a tree structure according to the diffusion order to obtain the danger diffusion path tree.
[0029] Furthermore, the time-dimension fusion weights and spatial-dimension fusion weights are synergistically fused to calculate the comprehensive fusion coefficient of each data source. The calculation formula is as follows:
[0030] ,
[0031] in, For the first The comprehensive fusion coefficient of individual sensor data sources, For the first Time-dimensional fusion weights of individual sensor data sources For the first Spatial dimension fusion weights of individual sensor data sources The fusion ratio for the time dimension weight is 0.6. The fusion ratio of the spatial dimension weight is 0.4.
[0032] Furthermore, the temporal convolutional network, the multi-head attention mechanism, the graph neural network, and the risk assessment model are jointly trained end-to-end. The training objective of the end-to-end joint training is to minimize the deviation between the predicted comprehensive risk index output by the risk assessment model and the calibrated comprehensive risk index.
[0033] During end-to-end joint training, all learnable parameters in the temporal convolutional network, the multi-head attention mechanism, the graph neural network, and the risk assessment model are updated synchronously through the backpropagation algorithm.
[0034] Furthermore, the cross-modal feature extraction of the collected coal bunker environmental data includes:
[0035] The audio data is preprocessed sequentially by pre-emphasis filtering, framing, windowing, and Mel-frequency conversion. The audio data features are then extracted using a two-dimensional convolutional network and a bidirectional long short-term memory network to obtain the modal feature vector corresponding to the audio data.
[0036] Furthermore, it also includes:
[0037] Based on the comprehensive risk index sequence of the most recent 10 monitoring periods, the sequence prediction model is used to predict the direction of risk index change in the next monitoring period. When the difference between the predicted comprehensive risk index and the comprehensive risk index of the current monitoring period is greater than the preset change threshold of 5, a risk warning is generated in advance.
[0038] The sequence prediction model employs a long short-term memory network with a two-layer stacked structure, and the number of hidden units in the sequence prediction model is 64.
[0039] Furthermore, this invention also provides a data fusion-based intelligent monitoring system for coal bunker environments, used to implement the aforementioned data fusion-based intelligent monitoring method for coal bunker environments, comprising:
[0040] The multi-dimensional feature vector space construction module is used to extract cross-modal features from collected temperature data, gas concentration data, dust content data, video data, and audio data, and to uniformly map the feature vectors of each modality to a feature space of a preset dimension to obtain the standardized feature vectors corresponding to each data source.
[0041] The time dimension weight allocation module is used to extract the historical time-series evolution features of each sensor data using a temporal convolutional network, and dynamically calculate the time dimension fusion weight of each data source based on the historical time-series evolution features through a multi-head attention mechanism.
[0042] The spatial dimension weight allocation module is used to construct a sensor spatial distribution topology map based on the spatial location information of each sensor node, use graph neural networks to analyze the spatial correlation between sensor data at different spatial locations, and combine the hotspot area identification and hazard diffusion path prediction results to determine the spatial dimension fusion weight of each data source.
[0043] The spatiotemporal collaborative fusion assessment module is used to collaboratively fuse the fusion weights of the time dimension and the fusion weights of the spatial dimension, calculate the comprehensive fusion coefficient of each data source, and perform weighted fusion of the standardized feature vectors corresponding to each data source according to the comprehensive fusion coefficient to obtain the fusion feature vector, and then output the comprehensive risk index through the risk assessment model.
[0044] In summary, the beneficial technical effects of the present invention are as follows:
[0045] 1. This invention constructs a multi-dimensional feature vector space, which uniformly maps multi-source heterogeneous data such as temperature, gas, dust, video, and audio to a high-dimensional feature space, realizing semantic alignment and standardized representation of cross-modal data. This solves the problem that different types of sensor data cannot be directly fused and analyzed in the prior art, and lays a data foundation for subsequent spatiotemporal correlation analysis.
[0046] 2. This invention extracts the temporal evolution features of each sensor data through a temporal convolutional network, and combines an attention mechanism to dynamically calculate the time dimension fusion weight based on historical change trends, thereby realizing dynamic adjustment of the fusion weight. It can adaptively perceive the early warning timeliness of each sensor data according to the progressive evolution characteristics of dangerous factors in the coal bunker environment. Compared with the static weighted fusion of the prior art, it has stronger adaptability and accuracy.
[0047] 3. This invention constructs a sensor spatial distribution topology map and uses graph neural networks to analyze the correlation between sensor data at different spatial locations, thereby achieving automatic identification of hotspot areas and effective prediction of hazard diffusion paths. It solves the problems of existing technologies that do not consider the spatial distribution characteristics of sensors and cannot analyze the regional correlation and diffusion trends of hazardous factors, providing more comprehensive decision support for coal bunker safety management.
[0048] 4. This invention integrates the weights of the time dimension and the weights of the spatial dimension, calculates the comprehensive fusion coefficient of the data from each sensor, and outputs the comprehensive risk assessment result. This achieves collaborative analysis of the time and spatial dimensions, significantly improves the accuracy of the comprehensive risk assessment, and controls the early warning response time within the preset time range, effectively enhancing the intelligence level and safety assurance capability of coal bunker environmental monitoring. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0050] Figure 1 This is a schematic diagram of the overall scheme for a data fusion-based intelligent monitoring method for coal bunker environments.
[0051] Figure 2 This is a schematic diagram of the overall scheme of a coal bunker environment intelligent monitoring system based on data fusion. Detailed Implementation
[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0053] like Figure 1 As shown, this invention provides a data fusion-based intelligent monitoring method for coal bunker environments, covering the entire technical chain from data acquisition to comprehensive risk assessment output. In the coal bunker environmental monitoring scenario, temperature sensors, gas sensors, dust sensors, video monitoring equipment, and audio acquisition equipment acquire environmental parameter information inside the coal bunker from different dimensions. Each sensor node synchronously records its spatial location information during the data acquisition process, providing a location basis for subsequent spatial correlation analysis.
[0054] This invention achieves unified representation of cross-modal data by constructing a multi-dimensional feature vector space, uses temporal convolutional networks and attention mechanisms to achieve dynamic adaptive allocation of temporal dimension weights, utilizes graph neural networks to achieve deep mining of spatial correlations, and finally outputs a comprehensive risk assessment result through the synergistic fusion of spatiotemporal weights. The specific implementation steps are as follows:
[0055] Step S1: Collect environmental data from the coal bunker. This involves deploying various sensors and acquisition devices to obtain multi-source heterogeneous environmental data within the coal bunker, and simultaneously recording the spatial location and temporal reference of the sensors. This lays the data foundation for subsequent cross-modal feature extraction and spatiotemporal fusion analysis. The specific operation is performed according to the following sub-steps:
[0056] Step S101: Deploy temperature sensors, gas sensors, dust sensors, video monitoring equipment, and audio acquisition equipment inside the coal bunker to acquire temperature data, gas concentration data, dust content data, video data, and audio data.
[0057] The temperature sensor has a measurement range of -40℃ to 85℃, a resolution of 0.1℃, and a measurement accuracy of ±0.5℃, in order to accurately capture temperature changes inside the coal bunker.
[0058] Gas sensors include three types: methane sensors, carbon monoxide sensors, and hydrogen sulfide sensors. All utilize electrochemical detection principles and feature rapid response and high selectivity. The methane sensor has a detection range of 0 to 100% volume percentage and a detection sensitivity of 0.01% volume percentage; the carbon monoxide sensor has a detection range of 0 to 1000 ppm and a detection sensitivity of 1 ppm; and the hydrogen sulfide sensor has a detection range of 0 to 100 ppm and a detection sensitivity of 0.1 ppm.
[0059] The dust sensor uses the principle of light scattering to measure dust concentration, with a detection range of 0 to 1000 mg / m³. 3 The detection sensitivity is 0.1 mg / m³. 3 .
[0060] The video surveillance equipment uses wide dynamic range cameras with a dynamic range of 120dB and is equipped with infrared supplementary lighting to adapt to changes in lighting inside the coal bunker. The camera resolution is no less than 2 megapixels, and the frame rate is 25fps to obtain clear video footage of the coal bunker's interior.
[0061] The audio acquisition equipment uses a wideband microphone with a frequency response range of 20Hz to 20kHz, a sampling rate of 44.1kHz, and a sampling accuracy of 16bit, to completely acquire the sound signals inside the coal bunker.
[0062] Each sensor node records its spatial location information synchronously during deployment.
[0063] In step S102, each sensor node transmits the collected data to the data aggregation unit. The sensor nodes complete the data transmission via wired or wireless communication. Wired communication uses an RS485 bus or Ethernet interface, while wireless communication uses the LoRa or WiFi protocol.
[0064] The data transmission frequency is set according to the different sensor types: the data transmission frequency of temperature sensors and gas sensors is 1Hz; the data transmission frequency of dust sensors is 0.5Hz; the data transmission frequency of video surveillance equipment is 25fps; and the data transmission frequency of audio acquisition equipment is 44.1kHz.
[0065] In step S103, each sensor node synchronously reports its spatial location information while transmitting data. The spatial location information consists of the sensor node's three-dimensional coordinates inside the coal bunker. The coordinate system uses the coal bunker entrance as the origin, with the horizontal east and north directions as the X and Y axes respectively, and the vertical upward direction as the Z axis. The spatial coordinates of the sensor nodes are determined by a GPS module or an indoor positioning base station, with an accuracy better than 0.1m. For areas inside the coal bunker where GPS signals cannot be received, an indoor positioning base station is used to achieve precise positioning.
[0066] In step S104, after receiving data from different sensor nodes, the data aggregation unit performs timestamp synchronization processing. Using the local clock of the data aggregation unit as a reference, and employing a network time protocol or hardware synchronization signal, the timestamps in each sensor data packet are uniformly calibrated to the same time base. After synchronization processing, the time synchronization accuracy between data from different sensor nodes is better than 10ms.
[0067] Step S105 involves periodically calibrating or standardizing various sensors to ensure the accuracy and stability of the collected data. The temperature sensor has a calibration cycle of 90 days, using a standard temperature source and comparing values at 0℃, 25℃, and 50℃. The required calibration accuracy is ±0.2℃. The gas sensor is equipped with an automatic calibration function, performing zero-point and full-scale calibration every 7 days. Zero-point calibration is performed in a clean air environment, and full-scale calibration is completed using the corresponding standard concentration gas.
[0068] Step S1 completes the comprehensive collection of five types of data in the coal bunker environment: temperature, gas, dust, video, and audio, as well as the recording of spatial coordinates of each sensor and the unification of time reference. The raw data stream with precise spatiotemporal markers serves as the direct information source for constructing the multidimensional feature vector space in step S2, supporting the unified mapping of heterogeneous data to a high-dimensional semantic space and standardized representation.
[0069] For step S2, a multi-dimensional feature vector space is constructed. The temperature data, gas concentration data, dust content data, video data, and audio data collected and spatiotemporally aligned in step S1 are preprocessed and cross-modal feature extracted respectively. The original features of each modality are uniformly mapped to the feature space of the preset dimension, and the output is a standardized feature representation that can be directly used for subsequent spatiotemporal fusion analysis.
[0070] The total dimension of the multidimensional feature vector space is set to 256 dimensions, of which 32 dimensions are allocated to temperature data, 64 dimensions to gas concentration data, 32 dimensions to dust content data, 96 dimensions to video data, and 32 dimensions to audio data. The specific steps are as follows:
[0071] Step S201, the preprocessing of temperature data includes outlier filtering and smoothing filtering. The outlier filtering stage uses a statistical method based on the interquartile range to calculate the first quartile of the temperature data within the current sliding window. With the 3rd quartile It will exceed Scope or Data points within the specified range are identified as outliers and removed.
[0072] The smoothing filtering stage uses the Kalman filter algorithm. The process noise covariance matrix is adaptively adjusted according to the mean of the temperature difference. When the mean exceeds the threshold, the covariance is increased to speed up tracking, and when it is below the threshold, the covariance is decreased to enhance suppression.
[0073] The feature extraction of temperature data is accomplished by a one-dimensional convolutional neural network with three convolutional layers and a kernel size of 3. The number of kernels in the first to third layers are 16, 32 and 64, respectively. Each layer is followed by a ReLU activation function, and the final layer is subjected to global average pooling to obtain a 32-dimensional temperature feature vector.
[0074] Step S202, the preprocessing of gas concentration data includes baseline correction and multi-sensor fusion. Baseline correction uses a moving average method with a window width of 120 sampling points, subtracting the mean within the moving window from the original sensor readings to obtain the baseline value. Multi-sensor fusion uses a weighted average method, where the weights of multiple sensors for the same gas are dynamically adjusted based on the historical measurement accuracy over the past 30 days, with sensors having higher accuracy receiving higher weights.
[0075] Feature extraction of gas concentration data employs a combination of temporal and statistical features. Temporal features are extracted using a Long Short-Term Memory (LSTM) network with 64 hidden units; statistical features include 18 parameters: maximum, minimum, average, standard deviation, peak factor, and slope. The two features are concatenated and mapped through a fully connected layer to obtain a 64-dimensional gas concentration data feature vector.
[0076] Step S203, the preprocessing of dust content data includes environmental factor compensation and time series smoothing. Environmental factor compensation uses contemporaneous temperature and humidity readings, and corrects the original dust readings using a multiple linear regression model established by least squares fitting of historical data. Time series smoothing uses the exponentially weighted moving average method, with a smoothing coefficient of 0.3.
[0077] Feature extraction of dust content data is accomplished by a one-dimensional convolutional neural network and a single-head attention mechanism. The one-dimensional convolutional neural network structure is the same as that of S201, and the single-head attention dimension is set to 32, finally outputting a 32-dimensional dust content data feature vector.
[0078] Step S204, the preprocessing of the video data includes video frame sampling, image enhancement, and object detection. Frame sampling downscales the video stream from 25fps to 1fps. Image enhancement performs histogram equalization and contrast stretching sequentially. Object detection uses the YOLO algorithm to identify abnormal event regions such as human activity, smoke, and flames, and outputs detection boxes and categories.
[0079] Video feature extraction is performed jointly by a pre-trained ResNet50 and a Transformer encoder. The ResNet50 extracts 2048-dimensional image space features within the detection box, which are then reduced to 512 dimensions by linear mapping and fed into a 4-layer Transformer encoder with 8 heads per layer. The output is then mapped through a fully connected layer to obtain a 96-dimensional video data feature vector.
[0080] Step S205, the preprocessing of the audio data includes pre-emphasis filtering, framing, windowing, and Mel spectrum conversion. The pre-emphasis filter coefficients are 0.97; the framing window is 25ms, and the frame shift is 10ms; the windowing uses a Hamming window; and a Mel spectrum is generated using 64 Mel filter banks.
[0081] Audio feature extraction is accomplished jointly by a two-dimensional convolutional network and a bidirectional long short-term memory network. The two-dimensional convolutional network extracts local time-frequency features from the Mel spectrogram, while the bidirectional long short-term memory network extracts temporal context. After global average pooling, a 32-dimensional audio data feature vector is obtained.
[0082] Step S206: Map the modal feature vectors extracted in steps S201 to S205 to a 256-dimensional feature space and normalize them.
[0083] The 32-dimensional vectors of temperature data, 64-dimensional vectors of gas concentration data, 32-dimensional vectors of dust content data, 96-dimensional vectors of video data, and 32-dimensional vectors of audio data are each mapped to 256 dimensions through their respective fully connected mapping layers. The fully connected mapping layers use the ReLU activation function and a dropout rate of 0.5. After mapping, five 256-dimensional intermediate feature vectors are obtained, denoted as... , , , and .
[0084] L2 norm normalization is performed on all these intermediate feature vectors, calculated as follows:
[0085]
[0086] in, , For vectors The Each feature vector has a magnitude of 1 after normalization, resulting in a uniform scale.
[0087] The training data for the one-dimensional convolutional neural network and long short-term memory network in steps S201, S202, and S203 are all taken from corresponding sensor sequences of no less than 6 months from the historical monitoring database of coal bunkers, and undergo the same preprocessing as in each step. The input of the training samples are sequences of the past 60 sampling points, the prediction target is the preprocessed features of the next sampling point, the loss function is mean squared error, the optimizer is Adam, the initial learning rate is 0.001, and training terminates when the loss on the validation set converges. The attention mechanism parameters in S203 are jointly trained with this one-dimensional convolutional neural network.
[0088] The Transformer encoder in step S204 is trained on historical video clips from inside the coal bunker, labeled with event categories such as normal working conditions, personnel activities, smoke, and flames. It takes the image spatial feature sequence extracted by ResNet50 after preprocessing consecutive frames in S204 as input, the event category as the prediction target, and cross-entropy loss as the loss function. During training, the ResNet50 pre-trained weights are frozen, and only the Transformer encoder and fully connected mapping layers are updated. The optimizer is Adam, the initial learning rate is 0.0001, and training terminates when the accuracy on the validation set stabilizes.
[0089] In step S205, the two-dimensional convolutional network and the bidirectional long short-term memory network are trained using historical audio clips from within the coal bunker, labeled with event categories such as equipment operation sounds, coal falling sounds, and abnormal impact sounds. Using fixed-duration Mel-spectrum image segments as input and event categories as prediction targets, the network employs cross-entropy loss as the loss function. Both networks are trained jointly using Adam as the optimizer, with an initial learning rate of 0.001, and training terminates when the validation set loss converges.
[0090] After the feature extraction network is trained, each fully connected mapping layer can fix the feature extraction network parameters and fine-tune them with the objective function of the final risk assessment task in step S5 as a guide, so that the mapped features are adapted to the downstream fusion assessment.
[0091] Step S2, through the above sub-steps, processes the five heterogeneous data types—temperature, gas, dust, video, and audio—through adapted preprocessing and feature extraction networks to obtain semantically aligned and scale-uniform 256-dimensional feature vectors. , , , and .
[0092] These feature vectors serve as the normalized inputs for steps S3 and S4, enabling the temporal convolutional network to extract temporal evolution features to compute temporal dimension weights. to This also enables graph neural networks to aggregate spatial neighborhood information at the sensor node level to output spatial dimension weights. to .
[0093] Step S3: Assign weights to the time dimension, which shall be carried out in accordance with the following steps;
[0094] Step S301: Construct a temporal convolutional network to extract the historical temporal evolution features of each sensor data. The temporal convolutional network adopts a multi-layer causal convolutional structure with 4 convolutional layers. The number of filters in each layer is 32, 64, 128, and 256 respectively, and the kernel size is 3 for all layers. The dilation factor increases exponentially with each layer, with dilation factors of 1, 2, 4, and 8 for layers 1 to 4, respectively, to expand the temporal receptive field. The preset time window length is 60 seconds, and temporal feature extraction is performed based on the historical sampling data within this time window.
[0095] For temperature and gas sensors, the initial sampling frequency is 1 Hz, and the input sequence length is 60 sampling points. For dust sensors, the sampling frequency is 0.5 Hz, and the input sequence length is 30 sampling points.
[0096] For video and audio data, since temporal modeling has been completed through the corresponding network in step S2, the 96-dimensional video feature vector and 32-dimensional audio feature vector output from step S2 are directly used as the input of the temporal convolutional network in this step, and the original sampling points are no longer used.
[0097] The video and audio feature vectors each correspond to abstract time points within a time window. Temporal convolutional networks extract the changing features in the time dimension through one-dimensional convolution operations.
[0098] After processing by a temporal convolutional network, the temperature sensor data outputs 32-dimensional temporal features, the gas sensor data outputs 64-dimensional temporal features, the dust sensor data outputs 32-dimensional temporal features, the video data outputs 96-dimensional temporal features, and the audio data outputs 32-dimensional temporal features.
[0099] Step S302: A multi-head attention mechanism is used to dynamically calculate the time dimension fusion weights based on the temporal characteristics of each data source. The attention mechanism employs a multi-head structure with four attention heads, each with a 64-dimensional feature dimension.
[0100] The time-series feature vectors obtained from the five sensor data sources in step S301 are combined into a feature matrix. The maximum value of the time-series feature dimension of each sensor, 64 dimensions, is taken as the comprehensive feature dimension. The 32-dimensional feature vectors of temperature, dust, and audio are zero-padded to reach 64 dimensions. The 96-dimensional feature vector of video is linearly mapped to 64 dimensions, while the 64-dimensional feature vector of gas remains unchanged, resulting in a feature matrix with a dimension of 5 rows and 64 columns. Each row corresponds to a time-series feature representation of a sensor data source.
[0101] For the characteristic matrix Perform a linear transformation to generate the query matrix. Key matrix Sum matrix All dimensions The transformation formula is:
[0102]
[0103] in, To query the weight matrix, The key weight matrix, The weight matrix consists of values, with each weight matrix having a dimension of 1. .
[0104] Calculate the scaled dot product attention to obtain the attention weight matrix. , dimension The calculation formula is:
[0105]
[0106] matrix The Middle Line number The elements of the column represent the first... The sensor for the first Attention weight values for each sensor.
[0107] Multiplying the attention weight matrix by the value matrix yields the attention-weighted feature matrix. :
[0108]
[0109] The attention-weighted feature matrix Data feature representations incorporating temporal context information are obtained through aggregation using fully connected layers and residual connections. To obtain the temporal dimension fusion weights for each data source, the attention weight matrix is... The average is taken row by row, and then normalized using the softmax function to obtain a 5-dimensional vector. Each component of this vector is the time dimension fusion weight of the five sensor data sources, denoted as follows: , , , , Satisfying the normalization condition And all weights are positive numbers.
[0110] The weights obtained through adaptive learning via the attention mechanism capture the historical drastic changes in the data from each sensor. Data sources with more drastic changes are assigned higher time dimension weights during fusion.
[0111] In step S303, the learnable parameters of the temporal convolutional network and the multi-head attention mechanism are jointly optimized during the training of the comprehensive risk assessment model in the subsequent step S5.
[0112] The training data was taken from the historical monitoring database of the coal bunker, including historical sequences of each sensor and a comprehensive risk index calibrated by experts based on historical accident records. The training samples were constructed as follows: sensor data within a continuous time window was used as input, and the comprehensive risk index corresponding to the end of the window was used as the prediction target.
[0113] The loss function used is mean squared error, the optimizer is Adam, the initial learning rate is 0.0001, and the training objective is to minimize the deviation between the predicted risk index and the calibrated risk index. The linear transformation weight matrix in the temporal convolutional network and the attention mechanism is updated synchronously using the backpropagation algorithm. , , And all parameters of the risk assessment model in step S5, until the validation set loss converges.
[0114] Step S3 adaptively extracts the temporal evolution patterns from the historical time-series features of each sensor and outputs dynamic time-dimensional fusion weights. to These weights are related to the spatial dimension weights output in step S4. to Enter the data together in step S5 to calculate the overall fusion coefficient of the data from each sensor. , driving weighted fusion and risk assessment of spatiotemporal collaboration.
[0115] For step S4, the step of analyzing spatial relationships, the following implementation method shall be followed.
[0116] Step S401: Construct a sensor spatial distribution topology graph based on the spatial location information of each sensor node. The sensor spatial distribution topology graph uses each sensor node as a vertex, and determines the edges and edge weights of the graph based on the Euclidean distance between the sensor nodes. A preset distance threshold of 5m is used; when the Euclidean distance between two sensor nodes is less than 5m, an edge connecting these two sensor nodes is established in the sensor spatial distribution topology graph.
[0117] The edge weight is inversely proportional to the Euclidean distance between sensor nodes. The calculation formula is:
[0118]
[0119] in, Indicates the first The sensor node and the first Edge weights between sensor nodes The preset distance threshold is 5m. For the first The sensor node and the first The Euclidean distance between sensor nodes, in meters. Approaching distance threshold At that time, edge weight Close to 1.
[0120] The sensor spatial distribution topology map is stored and represented using an adjacency matrix. The dimension of the adjacency matrix is... ,in Let be the total number of sensor nodes, and be the in the adjacency matrix. Line number The elements of the column are the edge weights. If there is no edge between two sensor nodes, the corresponding element takes the value of 0.
[0121] The number of sensor nodes inside the coal bunker is determined based on the bunker's volume and monitoring accuracy requirements. For a volume of 1000 m³... 3 The coal bunker has 20 to 30 sensor nodes, including 8 to 10 temperature sensors, 6 to 8 gas sensors, 4 to 6 dust sensors, 2 to 4 video monitoring devices, and 2 to 3 audio acquisition devices.
[0122] The spatial distribution of each sensor node adopts a combination of uniform grid layout and dense layout in key areas. While maintaining basic uniform coverage in the overall coal bunker space, the sensor deployment density is increased in key areas prone to danger, such as the coal pile surface, bunker wall corners, and ventilation openings.
[0123] Step S402: Analyze the spatial correlation between sensor data from different spatial locations using a graph neural network. The graph neural network employs a multi-layer graph convolutional structure with three layers, each containing 128 hidden units.
[0124] The first layer input of the graph neural network is the feature vectors of each sensor extracted in step S2. Specifically, the feature vectors of the temperature sensor are 32-dimensional, the gas sensor is 64-dimensional, the dust sensor is 32-dimensional, the video surveillance equipment is 96-dimensional, and the audio acquisition equipment is 32-dimensional.
[0125] Since different types of sensors have different feature dimensions, the first layer of graph convolution first maps the feature vectors of each sensor to a unified 128-dimensional space through a fully connected layer, resulting in a sensor feature matrix with unified dimensions. , dimension , This represents the total number of sensor nodes.
[0126] Graph convolution operations aggregate the features of the central sensor node with the features of its neighboring sensor nodes. Output of layer graph convolution The calculation formula is:
[0127]
[0128] in, For the first The input feature matrix of layer graph convolution, The sensor feature matrix after unified mapping; To add the adjacency matrix after adding self-loops, The adjacency matrix constructed in step S401 for A 3D identity matrix is used, with self-loops added to preserve the unique characteristics of the central sensor node. for The degree matrix, ; For the first The learnable weight matrix of the layer graph convolution has dimensions of , For the first Layer feature dimension For the first Layer feature dimension; ReLU is the activation function.
[0129] The forward propagation process of the graph convolutional network consists of three stacked graph convolutional layers, with batch normalization and ReLU activation performed after each graph convolution. The third graph convolutional layer outputs a spatial correlation feature vector for each sensor node, with a dimension of 128.
[0130] Step S403: Identify potential hotspot regions based on the spatial correlation feature vector output by the graph neural network. The threshold for hotspot identification is set to 1.5 times the historical average value of the corresponding parameter.
[0131] For temperature sensor nodes, a node is identified as a potential hotspot if its current measured temperature exceeds 1.5 times its historical 30-day average temperature. Similarly, for gas sensor nodes, a node is identified as a potential hotspot if its current measured concentration exceeds 1.5 times its historical 30-day average concentration.
[0132] For dust sensor nodes, when the current measured content of a dust sensor node exceeds 1.5 times the historical 30-day average content of that dust sensor node, the dust sensor node is determined to be a potential hotspot node.
[0133] Multiple potentially hotspot nodes that are adjacent in the sensor's spatial distribution topology map form a hotspot cluster. The boundary of the hotspot cluster is determined by the clustering analysis results of the spatial correlation feature vectors output by the graph neural network. Specifically, this can be achieved by performing connected component analysis on the spatial correlation feature vectors, classifying nodes with feature vector similarity higher than a preset similarity threshold and topologically adjacent into the same hotspot cluster.
[0134] Step S404: Starting from the identified hotspot area, calculate the probability of danger spreading to adjacent sensor nodes based on the weight of the edges in the sensor spatial distribution topology map, and output the danger spread path tree.
[0135] Let the number of a sensor node in the current hotspot area be . Sensor nodes The set of adjacent sensor nodes is denoted as For adjacent sensor node sets Any sensor node Danger originates from sensor nodes Diffusion to sensor nodes probability The calculation formula is:
[0136]
[0137] in, For sensor nodes With sensor nodes The edge weights between them are calculated in step S401, and the denominator is... For sensor nodes The sum of edge weights between it and all its adjacent sensor nodes.
[0138] When danger originates from sensor nodes Diffusion to sensor nodes After that, sensor nodes Become a new source of diffusion, continue computing from the sensor node. The probability of diffusion to its neighboring sensor nodes. This process is iterated until the diffusion probability falls below a preset termination threshold of 0.05, at which point diffusion stops.
[0139] The hazard propagation path tree is constructed using a breadth-first search algorithm, starting from all hotspot nodes and propagating simultaneously. The sensor nodes and edges involved in the propagation process are organized into a tree structure according to the propagation order. The root node is the hotspot node, and the tree levels represent the propagation depth. Each level's nodes represent the locations of sensors that may be affected by the hazard at the corresponding propagation depth.
[0140] Step S405: Calculate the spatial dimension weights of the five sensor data sources. For the five sensor data sources—temperature, gas, dust, video, and audio—their respective spatial dimension weights are denoted as follows: , , , , .
[0141] The spatial dimension weight is calculated as follows: For multiple sensor nodes of the same type of sensor, the spatial correlation feature vector of each sensor node is correlated with the hotspot area identification result in step S403 and the hazard diffusion path tree in step S404.
[0142] Sensor nodes located within hotspot areas or on dangerous diffusion path trees receive higher spatial dimension weights. The specific weight values are determined based on the spatial distance between the sensor node and the center of the hotspot area, as well as the diffusion probability at the diffusion path tree level.
[0143] For sensor nodes that are not in hotspot areas and are not on the danger spread path tree, a low but non-zero spatial dimension weight is assigned based on the shortest path distance between the sensor node and the nearest hotspot area in the sensor spatial distribution topology map.
[0144] The weights of all sensor nodes within the same type of sensor are aggregated to obtain the spatial dimension weights corresponding to the sensor data source. The spatial dimension weights of the five sensor data sources satisfy the normalization condition, i.e. .
[0145] In step S406, the learnable parameters of the graph neural network are jointly optimized during the training of the comprehensive risk assessment model in step S5. The training data is taken from the historical monitoring database of the coal bunker, which includes historical feature sequences and historical spatial location information of each sensor node, as well as hotspot area labels and comprehensive risk indices identified by experts based on historical accident records.
[0146] The training objective of the graph neural network consists of two parts: the first part is the node-level hotspot prediction loss, which uses the binary classification labels of whether each sensor node belongs to a hotspot area as the supervision signal and adopts the binary cross-entropy loss function; the second part is the graph-level risk prediction loss, which uses the comprehensive risk index as the supervision signal and adopts the mean squared error loss function. The total loss function is the weighted sum of the two losses, with the weight ratio set to 1:1.
[0147] The optimizer used is Adam, with an initial learning rate of 0.001. The learnable weight matrix of the graph convolution is updated using the backpropagation algorithm. The training terminates when the validation set loss converges, along with the parameters of the first fully connected mapping layer.
[0148] Step S4, through the above sub-steps, mines spatial relationships from the sensor spatial distribution topology and outputs spatial dimension weights. to Hotspot area identification results and hazard diffusion path tree. Spatial dimension weights and time dimension weights output from step S3. to Enter the data together in step S5 to calculate the overall fusion coefficient of the data from each sensor. The hotspot area identification results and the danger spread path tree are used to determine the scope of the affected areas in the early warning information.
[0149] Finally, step S5 involves fusing the spatiotemporal weights and outputting the evaluation results, which is carried out according to the following steps:
[0150] Step S501: The time dimension weights obtained in step S3 and the spatial dimension weights obtained in step S4 are fused together to calculate the comprehensive fusion coefficient of each sensor data source.
[0151] Time dimension weight to Spatial Dimension Weights to Each element had already met the normalization condition before fusion. The weighted summation method was used to calculate the... The overall fusion coefficient of individual sensor data sources The calculation formula is:
[0152]
[0153] in, Indicates the first The comprehensive fusion coefficient of individual sensor data sources, The values range from 1 to 5, corresponding to the five sensor data sources: temperature sensor, gas sensor, dust sensor, video surveillance equipment, and audio acquisition equipment, respectively. The output of step S3 Time dimension weighting of each sensor data source; The output of step S4 Spatial dimension weights of each sensor data source; The fusion ratio for the time dimension weight is set to 0.6; The fusion ratio of the spatial dimension weight is 0.4.
[0154] The fusion coefficients of the five sensor data sources are denoted as follows: , , , , The overall fusion coefficient satisfies the normalization condition. .
[0155] Step S502: The feature vectors of each sensor data source are weighted and fused according to the comprehensive fusion coefficient to obtain the fused feature vector.
[0156] The five 256-dimensional feature vectors extracted and normalized in step S2 are then weighted and summed according to the comprehensive fusion coefficients calculated in step S501. The fused feature vectors are then... The calculation formula is:
[0157]
[0158] in, The fused data feature vector has 256 dimensions. For the first The overall fusion coefficient of individual sensor data sources; For the first The 256-dimensional feature vectors extracted and normalized from each sensor data source in step S2 are as follows: Corresponding temperature data feature vector , Corresponding gas concentration data feature vector , Corresponding dust content data feature vector , Corresponding video data feature vector , Corresponding audio data feature vector .
[0159] Fusion feature vectors The data status of the coal bunker environment at the current monitoring time and its correlation characteristics in the time and spatial dimensions are comprehensively characterized.
[0160] Step S503: Based on the fused feature vector, a risk assessment model is used to calculate the comprehensive risk index of the coal bunker environment.
[0161] The risk assessment model employs a deep neural network structure, comprising three fully connected hidden layers with 256 hidden units per layer, using the ReLU activation function, and an output layer containing one neuron. The comprehensive risk index ranges from 0 to 100.
[0162] The fused feature vector obtained in step S502 The input risk assessment model is propagated forward, and the activation values of the output layer neurons are used as the raw values of the comprehensive risk index. The raw values are mapped to the range of 0 to 100 through linear mapping to obtain the comprehensive risk index.
[0163] Based on the comprehensive risk index, the environmental risk of coal bunkers is divided into four levels: a comprehensive risk index range of 0 to 25 indicates a low risk level, a range of 25 to 50 indicates a medium risk level, a range of 50 to 75 indicates a high risk level, and a range of 75 to 100 indicates an extremely high risk level.
[0164] Step S504: Generate early warning decision information based on the comprehensive risk level. The trigger condition for the early warning decision is that the comprehensive risk level reaches the medium risk level or above.
[0165] The warning information includes the risk level, the affected area, and recommended response measures. The affected area is determined based on the hotspot area group identified in step S403 and the hazard diffusion path tree output in step S404, and is described by the spatial coordinates of sensor nodes and the coverage of diffusion paths.
[0166] When the overall risk level is medium, the warning information recommends increasing the monitoring frequency and initiating manual inspections. When the overall risk level is high, the warning information recommends immediately notifying relevant personnel and taking on-site verification measures. When the overall risk level is extremely high, the warning information recommends activating the emergency response procedure and automatically triggering the audible and visual alarm devices.
[0167] Step S505: The risk assessment model is trained as follows. The training data is taken from the historical monitoring database of the coal bunker, including the fused feature vectors generated in step S502 for each monitoring period and the comprehensive risk index calibrated by experts based on historical accident records as labels.
[0168] Training samples are used to fuse feature vectors The input is the calibrated comprehensive risk index, and the prediction target is the standard comprehensive risk index. The loss function is the mean squared error, and the training objective is to minimize the deviation between the predicted comprehensive risk index output by the risk assessment model and the standard comprehensive risk index.
[0169] The optimizer used is Adam, with an initial learning rate of 0.0001. As described in step S303, the learnable parameters of this risk assessment model are jointly optimized with the parameters of the temporal convolutional network and multi-head attention mechanism in step S3; as described in step S406, the learnable parameters of this risk assessment model are also jointly optimized with the parameters of the graph neural network in step S4.
[0170] During training, all learnable parameters in the risk assessment model, temporal convolutional network, attention mechanism, and graph neural network are updated synchronously through backpropagation until the validation set loss converges.
[0171] Step S506: The risk evolution prediction function uses a sequence prediction model to predict the direction of risk index change in the next monitoring period based on the trend of the comprehensive risk index change over the most recent 10 monitoring periods.
[0172] The sequence prediction model employs a Long Short-Term Memory (LSTM) network with 64 hidden units and a two-layer stacked structure. The input to the sequence prediction model is the composite risk index sequence for the most recent 10 monitoring periods, and the output is the predicted composite risk index for the next monitoring period.
[0173] When the difference between the predicted comprehensive risk index and the comprehensive risk index for the current monitoring period exceeds a preset change threshold of 5, the risk index is determined to be on an upward trend and the increase exceeds the threshold, thus generating an early risk warning. The risk warning includes the predicted risk index, the predicted risk level, and suggested precautions.
[0174] The training data for the sequence prediction model is taken from the comprehensive risk index sequence of consecutive monitoring periods in the historical monitoring database of coal bunkers. The training samples use the comprehensive risk index sequence of 10 consecutive monitoring periods as input, and the comprehensive risk index of the 11th monitoring period that follows as the prediction target. The loss function is mean squared error, the optimizer is Adam, the initial learning rate is 0.001, and training terminates when the loss on the validation set converges.
[0175] Step S5, through the above sub-steps, weights the time dimension. to Spatial Dimension Weights to Synergistic integration is the comprehensive integration coefficient to It drives the weighted fusion of multi-source heterogeneous feature vectors, and outputs a comprehensive risk index and corresponding risk level and early warning decision through a risk assessment model.
[0176] The temporal convolutional network, attention mechanism, graph neural network and risk assessment model involved in steps S3, S4 and S5 are jointly trained end-to-end to optimize their respective learnable parameters. This allows the four stages of time dimension weight allocation, spatial dimension weight allocation, feature fusion and risk assessment to adapt to each other under a unified training objective, and output a comprehensive risk assessment result that is consistent with the real risk situation.
[0177] To facilitate understanding, specific application examples are provided to illustrate the technical solution of this invention.
[0178] Assume that the coal storage silo volume of a certain coal mine is 2000m³ 3 The warehouse is equipped with 28 sensor nodes, including 10 temperature sensors, 8 gas sensors, 6 dust sensors, 2 video surveillance cameras, and 2 broadband microphones. The spatial coordinates are determined by the indoor positioning system.
[0179] Within monitoring period T, the temperature measured at node 1 of temperature sensor was 45℃, with a historical 30-day average of 28℃; the methane concentration at node 3 of gas sensor was 0.8% by volume, with a historical 30-day average of 0.3% by volume; and the dust content at node 2 of dust sensor was 120 mg / m³. 3 The historical 30-day average was 85 mg / m³. 3 Video surveillance detected slight smoke coming from the surface of the coal pile, but no abnormalities were detected in the audio.
[0180] After feature extraction in step S2 and time dimension weight calculation in step S3, the time dimension weights of each sensor are obtained as follows: temperature α1=0.25, gas α2=0.35, dust α3=0.20, video α4=0.12, and audio α5=0.08.
[0181] After spatial analysis in step S4, the Euclidean distance between temperature sensor node 1 and gas sensor node 3 is 3.2m, and the edge weight W 12 =5 2 / 3.2 2 =2.44; The Euclidean distance between gas sensor node 3 and dust sensor node 2 is 4.1m, and the edge weight W 23 =5 2 / 4.1 2 =1.49. Temperature sensor node 1's current temperature of 45℃ exceeds the historical average of 28℃ by 1.5 times, or 42℃. Gas sensor node 3's methane concentration of 0.8% exceeds the historical average of 0.3% by 1.5 times, or 0.45%. Both are identified as hotspot nodes and form a hotspot cluster. Dust sensor node 2 is located on a hazardous diffusion path. The spatial dimension weights for each sensor are: β1=0.30, β2=0.32, β3=0.22, β4=0.10, β5=0.06.
[0182] After fusion in step S5, the comprehensive fusion coefficients are: γ1 = 0.6 × 0.25 + 0.4 × 0.30 = 0.27, γ2 = 0.6 × 0.35 + 0.4 × 0.32 = 0.338, γ3 = 0.6 × 0.20 + 0.4 × 0.22 = 0.208, γ4 = 0.6 × 0.12 + 0.4 × 0.10 = 0.112, and γ5 = 0.6 × 0.08 + 0.4 × 0.06 = 0.072. After weighted fusion, the input to the risk assessment model outputs a comprehensive risk index of 62.5, classifying it as a high-risk level. Warning information: The risk level is high, affecting the center of the coal pile surface and a 2m radius above it. It is recommended to immediately notify safety management personnel to conduct on-site verification, activate ventilation within the storage area, and prepare fire emergency equipment. Risk evolution prediction shows that the risk index will rise to 68 in the next cycle, exceeding the current value by more than 5, triggering an upward trend warning.
[0183] like Figure 2 As shown, the intelligent monitoring system for coal bunker environment based on data fusion disclosed in this invention includes:
[0184] The multi-dimensional feature vector space construction module is used to extract cross-modal features from collected temperature data, gas concentration data, dust content data, video data, and audio data, and to uniformly map the feature vectors of each modality to a feature space of a preset dimension to obtain the standardized feature vectors corresponding to each data source.
[0185] The time dimension weight allocation module is used to extract the historical time-series evolution features of each sensor data using a temporal convolutional network, and dynamically calculate the time dimension fusion weight of each data source based on the historical time-series evolution features through a multi-head attention mechanism.
[0186] The spatial dimension weight allocation module is used to construct a sensor spatial distribution topology map based on the spatial location information of each sensor node, use graph neural networks to analyze the spatial correlation between sensor data at different spatial locations, and combine the hotspot area identification and hazard diffusion path prediction results to determine the spatial dimension fusion weight of each data source.
[0187] The spatiotemporal collaborative fusion assessment module is used to collaboratively fuse the fusion weights of the time dimension and the fusion weights of the spatial dimension, calculate the comprehensive fusion coefficient of each data source, and then perform weighted fusion of the standardized feature vectors corresponding to each data source based on the comprehensive fusion coefficient to obtain the fusion feature vector, and finally output the comprehensive risk index through the risk assessment model.
[0188] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, the embodiments should be regarded as exemplary and non-limiting in all respects.
[0189] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for intelligent monitoring of coal bunker environment based on data fusion, characterized in that, Includes the following steps: Cross-modal feature extraction is performed on the collected coal bunker environmental data, which includes temperature data, gas concentration data, dust content data, video data, and audio data. The extracted modal feature vectors are uniformly mapped to a feature space of a preset dimension to obtain the standardized feature vectors corresponding to each data source. The historical temporal evolution features of each sensor data are extracted using a temporal convolutional network, and the temporal dimension fusion weights of each data source are dynamically calculated based on the historical temporal evolution features using a multi-head attention mechanism. Based on the spatial location information of each sensor node, a sensor spatial distribution topology map is constructed. A graph neural network is used to analyze the spatial correlation between sensor data at different spatial locations. Combined with the hotspot area identification and hazard diffusion path prediction results, the spatial dimension fusion weight of each data source is determined. The time dimension fusion weight and the spatial dimension fusion weight are synergistically fused to calculate the comprehensive fusion coefficient of each data source; Based on the comprehensive fusion coefficient, the standardized feature vectors corresponding to each data source are weighted and fused to obtain the fused feature vector; Based on the fused feature vector, a comprehensive risk index is output through a risk assessment model.
2. The intelligent monitoring method for coal bunker environment based on data fusion according to claim 1, characterized in that, The preset dimension feature space is a 256-dimensional feature space. The step of uniformly mapping the extracted modal feature vectors to a feature space of a preset dimension includes: The 32-dimensional feature vectors of temperature data, 64-dimensional feature vectors of gas concentration data, 32-dimensional feature vectors of dust content data, 96-dimensional feature vectors of video data, and 32-dimensional feature vectors of audio data are mapped to 256 dimensions through their respective fully connected mapping layers. Then, L2 norm normalization is performed on each of the mapped 256-dimensional intermediate feature vectors to obtain the standardized feature vectors corresponding to each data source.
3. The intelligent monitoring method for coal bunker environment based on data fusion according to claim 1, characterized in that, The temporal convolutional network adopts a 4-layer causal convolutional structure. The number of convolutional kernels in the 1st to 4th layers are 32, 64, 128 and 256 respectively, the kernel size is 3, and the dilation factors are 1, 2, 4 and 8 respectively. The step involves extracting historical temporal evolution features of data from each sensor using a temporal convolutional network, and dynamically calculating the temporal dimension fusion weights of each data source based on these historical temporal evolution features using a multi-head attention mechanism. This includes: The temporal feature vectors output from each data source after processing by the temporal convolutional network are combined into a feature matrix. The feature matrix is then linearly transformed to generate a query matrix, a key matrix, and a value matrix. The scaled dot product attention is calculated to obtain an attention weight matrix. The attention weight matrix is then averaged row by row and normalized using the softmax function to obtain the temporal dimension fusion weights of each data source.
4. The intelligent monitoring method for coal bunker environment based on data fusion according to claim 1, characterized in that, The construction of the sensor spatial distribution topology map includes: Using each sensor node as a vertex of the graph, when the Euclidean distance between two sensor nodes is less than a preset distance threshold, an edge is established between the two sensor nodes. The weight of the edge between the two sensor nodes is inversely proportional to the Euclidean distance between the two sensor nodes. The method of using graph neural networks to analyze the spatial correlation between sensor data from different spatial locations includes: The graph neural network is configured as a 3-layer graph convolutional structure, with each layer containing 128 hidden units. The first layer first maps the feature vectors of each sensor node to 128 dimensions. Then, the stacked graph convolutional layers aggregate the features of the central sensor node and the features of the adjacent sensor nodes of the central sensor node, and output the spatial correlation feature vector of each sensor node.
5. The intelligent monitoring method for coal bunker environment based on data fusion according to claim 1, characterized in that, The hotspot area identification and hazard spread path prediction include: When the current measurement value of a sensor node exceeds 1.5 times the historical average value of the sensor node, the sensor node is determined to be a potential hotspot node, and multiple adjacent potential hotspot nodes in the sensor spatial distribution topology map form a hotspot group. Starting from the hotspot area node, the probability of danger spreading to adjacent sensor nodes is calculated based on the weight of the edges in the sensor spatial distribution topology graph. Iterative diffusion is performed through breadth-first search until the diffusion probability is lower than a preset termination threshold. The sensor nodes and edges involved in the diffusion process are organized into a tree structure according to the diffusion order to obtain the danger diffusion path tree.
6. The intelligent monitoring method for coal bunker environment based on data fusion according to claim 1, characterized in that, The time-dimension fusion weights and spatial-dimension fusion weights are then fused collaboratively to calculate the comprehensive fusion coefficient for each data source. The calculation formula is as follows: , in, For the first The comprehensive fusion coefficient of individual sensor data sources, For the first Time-dimensional fusion weights of individual sensor data sources For the first Spatial dimension fusion weights of individual sensor data sources The fusion ratio for the time dimension weight is 0.
6. The fusion ratio of the spatial dimension weight is 0.
4.
7. The intelligent monitoring method for coal bunker environment based on data fusion according to claim 1, characterized in that, The temporal convolutional network, the multi-head attention mechanism, the graph neural network, and the risk assessment model are jointly trained end-to-end. The training objective of the end-to-end joint training is to minimize the deviation between the predicted comprehensive risk index output by the risk assessment model and the calibrated comprehensive risk index. During end-to-end joint training, all learnable parameters in the temporal convolutional network, the multi-head attention mechanism, the graph neural network, and the risk assessment model are updated synchronously through the backpropagation algorithm.
8. The intelligent monitoring method for coal bunker environment based on data fusion according to claim 1, characterized in that, The cross-modal feature extraction of the collected coal bunker environmental data includes: The audio data is preprocessed sequentially by pre-emphasis filtering, framing, windowing, and Mel-frequency conversion. The audio data features are then extracted using a two-dimensional convolutional network and a bidirectional long short-term memory network to obtain the modal feature vector corresponding to the audio data.
9. The intelligent monitoring method for coal bunker environment based on data fusion according to claim 1, characterized in that, Also includes: Based on the comprehensive risk index sequence of the most recent 10 monitoring periods, the sequence prediction model is used to predict the direction of risk index change in the next monitoring period. When the difference between the predicted comprehensive risk index and the comprehensive risk index of the current monitoring period is greater than the preset change threshold of 5, a risk warning is generated in advance. The sequence prediction model employs a long short-term memory network with a two-layer stacked structure, and the number of hidden units in the sequence prediction model is 64.
10. A coal bunker environment intelligent monitoring system based on data fusion, used to implement the coal bunker environment intelligent monitoring method based on data fusion as described in any one of claims 1-9, characterized in that, include: The multi-dimensional feature vector space construction module is used to extract cross-modal features from collected temperature data, gas concentration data, dust content data, video data, and audio data, and to uniformly map the feature vectors of each modality to a feature space of a preset dimension to obtain the standardized feature vectors corresponding to each data source. The time dimension weight allocation module is used to extract the historical time-series evolution features of each sensor data using a temporal convolutional network, and dynamically calculate the time dimension fusion weight of each data source based on the historical time-series evolution features through a multi-head attention mechanism. The spatial dimension weight allocation module is used to construct a sensor spatial distribution topology map based on the spatial location information of each sensor node, use graph neural networks to analyze the spatial correlation between sensor data at different spatial locations, and combine the hotspot area identification and hazard diffusion path prediction results to determine the spatial dimension fusion weight of each data source. The spatiotemporal collaborative fusion assessment module is used to collaboratively fuse the fusion weights of the time dimension and the fusion weights of the spatial dimension, calculate the comprehensive fusion coefficient of each data source, and perform weighted fusion of the standardized feature vectors corresponding to each data source according to the comprehensive fusion coefficient to obtain the fusion feature vector, and then output the comprehensive risk index through the risk assessment model.