Multi-modal fusion mine microseismic energy evaluation method, device, equipment and storage medium

By combining waveform and audio signals through multimodal fusion and graph convolutional networks, the problems of signal noise and modal differences in mine microseismic monitoring were solved, and more accurate microseismic source energy assessment was achieved.

CN121598328BActive Publication Date: 2026-04-21CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2026-01-30
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies cannot effectively address the problems of high signal noise, significant modal differences, complex propagation paths, and inconsistent observation quality of different geophones in mine microseismic monitoring, resulting in low accuracy of microseismic source energy assessment.

Method used

A multimodal fusion method is adopted, which combines waveform and audio signals, and performs cross-modal feature fusion through graph convolutional networks. The graph convolutional network is constructed to predict energy by utilizing the spatial location and propagation delay of the detector.

Benefits of technology

It improves the accuracy and stability of microseismic source energy assessment, reduces the impact of noise and environmental interference, and comprehensively explores the multimodal characteristics of microseismic events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121598328B_ABST
    Figure CN121598328B_ABST
Patent Text Reader

Abstract

This invention discloses a multimodal fusion method, apparatus, equipment, and storage medium for assessing microseismic energy in mines. The method includes: preprocessing multimodal signals collected from the mine area to obtain initial waveform feature vectors and initial audio feature vectors; fusing the initial waveform feature vectors and initial audio feature vectors across modes based on the spatial location of each detector to output fused mode features; inputting the fused mode features of each detector into a graph convolutional network for aggregation to obtain target features of microseismic events; and predicting the source energy of microseismic events based on the target features. By acquiring multimodal signals and fusing them across modes, the limitations of single signal monitoring are overcome, the multimodal correlation features of microseismic events are fully explored, and the impact of interference factors on the assessment results is reduced. Combined with the spatial feature aggregation capability of the graph convolutional network, the complementary information of multi-source data is effectively integrated, significantly improving the accuracy and stability of source energy assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent mining technology, and in particular to a multimodal fusion method, apparatus, equipment, and storage medium for assessing microseismic energy in mines. Background Technology

[0002] During mining operations, microseismic events are prone to occur due to factors such as mining disturbances, stress redistribution, and surrounding rock fracturing. Accurate assessment of the source energy of microseismic events can provide quantitative evidence for identifying the risks of dynamic disasters such as rockbursts, roof collapses, and rock bursts, and is also a key link in the mine safety monitoring and early warning system.

[0003] In current engineering applications, microseismic energy estimation largely relies on inversion of waveform signal characteristics recorded by geophones or empirical formulas. However, the underground mining environment is complex, and signal propagation is significantly affected by medium inhomogeneity, tunnel structure, attenuation, and scattering. Furthermore, different geophones are affected by installation orientation, distance, and local noise interference, leading to fluctuations and systematic biases in energy estimation results. Meanwhile, some monitoring systems can simultaneously acquire accompanying acoustic information in addition to waveforms, but existing methods often employ simple splicing or fixed-weight fusion, lacking modeling of cross-modal time alignment, complementary relationships, and reliability differences, making it difficult to fully leverage the advantages of multi-source observation. Moreover, multi-geophone data fusion typically uses averaging or manual point selection, failing to adaptively allocate the contributions of each sensor based on the geophone's spatial structure and propagation priors, further limiting the assessment accuracy and robustness. Therefore, a new method is urgently needed that can simultaneously utilize multi-modal information such as waveforms and audio, combined with physical factors such as the spatial location, distance, and propagation delay of the seismic source and geophone, to achieve a more stable and accurate energy assessment. Summary of the Invention

[0004] The main objective of this invention is to provide a multi-modal fusion method, device, equipment, and storage medium for assessing the energy of mine microseismic sources. This invention aims to solve the technical problems of low accuracy in assessing the energy of mine microseismic sources caused by the inability of existing technologies to effectively address the difficulties of high noise in mine microseismic monitoring signals, significant modal differences, complex propagation paths, and inconsistent observation quality of different detectors.

[0005] To achieve the above objectives, this invention provides a multimodal fusion method for assessing microseismic energy in mines, the method comprising the following steps:

[0006] Acquire microseismic signal data, which includes multi-mode signals of microseismic events collected by multiple detectors and the spatial location of each detector in the mining area. The multi-mode signals include waveform image signals and audio signals.

[0007] The multimodal signal is preprocessed to obtain an initial waveform feature vector and an initial audio feature vector;

[0008] Based on the spatial location of each detector, the initial waveform feature vector and the initial audio feature vector of each detector are fused across modes to output the fused mode features of each detector;

[0009] A graph convolutional network is constructed based on the microseismic events and the spatial locations of each detector. The fused modal features of each detector are then input into the graph convolutional network for aggregation to obtain the target features of the microseismic events. The graph convolutional network includes detector nodes and source nodes of the microseismic events.

[0010] Predicting source energy of microseismic events based on target characteristics.

[0011] Optionally, the preprocessing of the multimodal signal to obtain an initial waveform feature vector and an initial audio feature vector includes:

[0012] A time window clipping function is constructed based on the acquisition time window of the multimode signals of each detector, and the multimode signals are clipped based on the time window clipping function to obtain clipped waveform signals and clipped audio signals.

[0013] The cropped waveform signal and the cropped audio signal are converted to the time-frequency domain using the short-time Fourier transform function to obtain the time-frequency diagram of the waveform signal and the time-frequency diagram of the audio signal.

[0014] The waveform signal time-frequency diagram and the audio signal time-frequency diagram are standardized to obtain the initial waveform feature vector and the initial audio feature vector.

[0015] Optionally, the step of performing cross-modal fusion of the initial waveform feature vector and the initial audio feature vector of each detector based on the spatial position of each detector, and outputting the fused modal features of each detector, includes:

[0016] The initial waveform feature vector and initial audio feature vector of each detector are linearly converted into waveform attention intermediate vector and audio attention intermediate vector, respectively. The waveform attention intermediate vector includes waveform query vector, waveform key vector and waveform value vector, and the audio attention intermediate vector includes audio query vector, audio key vector and audio value vector.

[0017] Self-attention analysis is performed based on the waveform attention intermediate vector and the audio attention intermediate vector to obtain waveform mode self-attention output results and audio mode self-attention output results;

[0018] Based on the spatial position of each detector, the waveform mode self-attention output result, and the audio mode self-attention output result, the initial waveform feature vector and the initial audio feature vector are fused across modes to output the fused mode features of each detector.

[0019] Optionally, the step of performing cross-modal fusion of the initial waveform feature vector and the initial audio feature vector based on the waveform mode self-attention output result and the audio mode self-attention output result, and outputting the fused mode features of each detector, includes:

[0020] The initial waveform feature vector is linearly converted into a first query vector, and the initial audio feature vector is linearly converted into a first key vector and a first value vector;

[0021] Based on the first query vector, the first key vector, and the first value vector, cross-modal co-attention analysis is performed to obtain the first cross-modal co-attention feature of waveform mode focusing on audio mode;

[0022] The initial audio feature vector is linearly transformed into a second query vector, and the initial waveform feature vector is linearly transformed into a second key vector and a second value vector;

[0023] Based on the second query vector, the second key vector, and the second value vector, cross-modal co-attention analysis is performed to obtain the second cross-modal co-attention feature of the audio modality-focused waveform modality;

[0024] The waveform modal self-attention output result, the first cross-modal co-attention feature, and the initial waveform feature vector are residually connected to obtain the waveform cross-modal interaction feature;

[0025] The audio modal self-attention output, the second cross-modal co-attention feature, and the initial audio feature vector are residually concatenated to obtain the audio cross-modal interaction feature;

[0026] Based on the spatial position of each detector, the waveform cross-modal interaction features and the audio cross-modal interaction features are fused together to output the fused modal features of each detector.

[0027] Optionally, the step of fusing the waveform cross-modal interaction features and the audio cross-modal interaction features based on the spatial position of each detector to output the fused modal features of each detector includes:

[0028] The initial predicted distance between each geophone and the seismic source is determined based on the spatial location and preliminary positioning coordinates of each geophone. The preliminary positioning coordinates are determined based on the three-dimensional velocity model of the mining area and the travel time difference between P-waves and S-waves, referring to the following formula:

[0029]

[0030] in, For detector coordinate, Microseismic events Preliminary positioning coordinates Let be the Euclidean distance function. For detector Microseismic events The initial predicted distance between the earthquake sources;

[0031] The propagation delay of each detector is calculated based on the initial predicted distance, referring to the following formula:

[0032]

[0033] in, For detector With the epicenter The propagation delay The equivalent wave velocity in the medium;

[0034] The spatial position codes of each detector are obtained by encoding the three-dimensional coordinates of the spatial position of each detector based on multiple frequencies.

[0035] The spatial location codes are mapped to the required dimensions of the model to obtain the location code vectors of each detector.

[0036] The input vector of the gated network is constructed based on the location encoding vector, the initial prediction distance, the propagation delay, the waveform cross-modal interaction features, and the audio cross-modal interaction features;

[0037] The input vector is fed into a gating network, which is configured to output a gating vector based on the input vector through a multilayer perceptron network. The formula for calculating the gating vector is as follows:

[0038]

[0039]

[0040] in, and These represent the waveform cross-modal interaction features and the audio cross-modal interaction features, respectively. For detector Position encoding vector, Indicates the first The first microseismic event The gating vector corresponding to each detector A multilayer perceptron network representing a gated network. This represents the Sigmoid activation function;

[0041] Based on the gate vector, the waveform cross-modal interaction features and the audio cross-modal interaction features are fused across modally to output the fused modal features of each detector, as shown in the following formula:

[0042]

[0043] in, For the first The first microseismic event Fusion mode characteristics of individual detectors This is for element-wise multiplication.

[0044] Optionally, the graph convolutional network is configured to calculate the position-aware edge weights between the detector node and the source node using a position-aware edge weight function, referring to the following formula:

[0045]

[0046] in, For detector nodes Source nodes corresponding to microseismic events Position-aware edge weights between them For the Sigmoid function, , and This represents learnable hyperparameters, which control the weights of distance, latency, and spatial location information in the graph convolutional network. For detector nodes With the focal node The initial predicted distance between them For detector nodes With the focal node The propagation delay and These are the detector nodes. and epicenter node Position encoding vector, It is a detector node and epicenter node The inner product of the position encoding vectors is used to measure the spatial angular relationship between the detector node and the source node;

[0047] The graph convolutional network is further configured to normalize the position-aware edge weights using a softmax function to obtain the target edge weights, as shown in the following formula:

[0048]

[0049] in, The standardized target edge weights, and These are the detector nodes. Detector node The focal node of microseismic events Position-aware edge weights between them To collect microseismic events The set of all detector nodes for the information;

[0050] The graph convolutional network is further configured to perform graph convolution operations on the fused modal features of each detector based on the target edge weights, and output the target features of the microseismic event, as shown in the following formula:

[0051]

[0052] in, For microseismic events that aggregate information from all geophone nodes Target characteristics, Microseismic events Intermediate detector node The fusion modal features.

[0053] Optionally, the prediction of source energy of microseismic events based on target features includes:

[0054] Based on target characteristics, the source energy of microseismic events can be predicted using the following formula:

[0055]

[0056] in, Microseismic events Predicted energy, For learnable regression weight vectors, This is the transpose of the regression weight vector. For microseismic events that aggregate information from all geophone nodes Target characteristics, This is a learnable bias term.

[0057] Furthermore, to achieve the above objectives, the present invention also proposes a multi-modal fusion mine microseismic energy assessment device, wherein the multi-modal fusion mine microseismic energy assessment device comprises:

[0058] The data acquisition module is used to acquire microseismic signal data, which includes multi-mode signals of microseismic events collected by multiple detectors and the spatial position of each detector in the mining area. The multi-mode signals include waveform image signals and audio signals.

[0059] The preprocessing module is used to preprocess the multimodal signal to obtain an initial waveform feature vector and an initial audio feature vector;

[0060] The cross-modal fusion module is used to perform cross-modal fusion of the initial waveform feature vector and the initial audio feature vector of each detector based on the spatial position of each detector, and output the fused modal features of each detector;

[0061] The feature aggregation module is used to construct a graph convolutional network based on the microseismic event and the spatial location of each detector, and input the fused modal features of each detector into the graph convolutional network for aggregation to obtain the target features of the microseismic event. The graph convolutional network includes detector nodes and source nodes of the microseismic event.

[0062] The energy prediction module is used to predict the source energy of microseismic events based on target characteristics.

[0063] Furthermore, to achieve the above objectives, this application also proposes a multimodal fusion mine microseismic energy assessment device, the device comprising: a memory, a processor, and a multimodal fusion mine microseismic energy assessment program stored in the memory, the processor being used to run the multimodal fusion mine microseismic energy assessment program, the computer program being configured to implement the steps of the multimodal fusion mine microseismic energy assessment method as described above.

[0064] In addition, to achieve the above objectives, this application also proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the multimodal fusion mine microseismic energy assessment method described above.

[0065] This invention acquires microseismic signal data, including multimodal signals of microseismic events collected by multiple geophones and the spatial location of each geophone in a mining area. The multimodal signals include waveform image signals and audio signals. The multimodal signals are preprocessed to obtain initial waveform feature vectors and initial audio feature vectors. Based on the spatial location of each geophone, the initial waveform feature vectors and initial audio feature vectors of each geophone are fused across modes to output the fused mode features of each geophone. A graph convolutional network is constructed based on the microseismic events and the spatial location of each geophone, and the fused mode features of each geophone are input into the network. Graph convolutional networks are used to aggregate and obtain target features of microseismic events. The graph convolutional network includes detector nodes and source nodes of microseismic events. Based on the target features, the source energy of microseismic events is predicted. Because this invention overcomes the limitations of traditional single signal monitoring by acquiring multimodal signals and fusing them across modes, it comprehensively mines the waveform, audio and spatial correlation features of microseismic events, improves the completeness of feature representation, and effectively integrates complementary information from multiple sources by combining the spatial feature aggregation capability of graph convolutional networks. This reduces the impact of noise, environmental interference and other factors on the evaluation results, and significantly improves the accuracy and stability of source energy assessment. Attached Figure Description

[0066] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0067] Figure 1 This is a schematic diagram of the structure of a multimodal fusion mine microseismic energy assessment device for hardware operation environment involved in the embodiments of the present invention;

[0068] Figure 2 This is a flowchart illustrating the first embodiment of the multimodal fusion method for assessing microseismic energy in mines according to the present invention.

[0069] Figure 3 This is a flowchart illustrating the second embodiment of the multimodal fusion method for assessing microseismic energy in mines according to the present invention.

[0070] Figure 4 This is a flowchart illustrating the third embodiment of the multimodal fusion method for assessing microseismic energy in mines according to the present invention.

[0071] Figure 5 This is a structural block diagram of the first embodiment of the multimodal fusion mine microseismic energy assessment device of the present invention.

[0072] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0073] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0074] Reference Figure 1 , Figure 1 This is a schematic diagram of the structure of a multimodal fusion mine microseismic energy assessment device for hardware operation environment involved in the embodiments of the present invention.

[0075] like Figure 1As shown, the multimodal fusion mine microseismic energy assessment device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be high-speed random access memory (RAM) or stable non-volatile memory (NVM), such as a disk storage device. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0076] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the multimodal fusion mine microseismic energy assessment device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0077] like Figure 1 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a multimodal fusion mine microseismic energy assessment program.

[0078] exist Figure 1 In the multimodal fusion mine microseismic energy assessment device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and memory 1005 in the multimodal fusion mine microseismic energy assessment device of the present invention can be set in the multimodal fusion mine microseismic energy assessment device. The multimodal fusion mine microseismic energy assessment device calls the multimodal fusion mine microseismic energy assessment program stored in the memory 1005 through the processor 1001 and executes the multimodal fusion mine microseismic energy assessment method provided in the embodiment of the present invention.

[0079] This invention provides a multimodal fusion method for assessing microseismic energy in mines, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the multimodal fusion method for assessing microseismic energy in mines according to the present invention.

[0080] In this embodiment, the multimodal fusion method for assessing microseismic energy in mines includes the following steps:

[0081] Step S10: Acquire microseismic signal data.

[0082] It should be noted that this embodiment is applied to mine microseismic energy assessment. Addressing the problems of high noise, significant modal differences, complex propagation paths, and inconsistent observation quality among different detectors in mine microseismic monitoring signals, a multi-modal fusion method for mine microseismic energy assessment is designed. This method achieves effective alignment and complementary fusion of microseismic waveform modes and audio modes. Modal adaptive weighting is achieved by combining the spatial location, distance, and propagation delay of the source-detector system. Furthermore, the contribution of each detector to energy estimation is learned using the spatial structure of multiple detectors, resulting in a more accurate assessment of microseismic source energy.

[0083] It should be understood that the executing entity of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or a terminal electronic device capable of performing the above functions. The following uses a multimodal fusion mine microseismic energy assessment device (hereinafter referred to as the assessment device) as an example to illustrate this embodiment and the following embodiments.

[0084] It should be noted that microseismic signal data refers to the set of relevant signals collected by monitoring equipment when microseismic events occur in the strata rock due to stress release during mining, blasting and other operations in a mine. The microseismic signal data includes multi-mode signals of microseismic events collected by multiple geophones and the spatial location of each geophone in the mining area.

[0085] Multimodal signals refer to two or more types of signals originating from the same microseismic event but exhibiting different forms. In this embodiment, they include waveform image signals and audio signals, which reflect the characteristics of the microseismic event from different dimensions. Waveform image signals can be image data transformed from the vibration signals generated by the microseismic event according to the time-series variation law, which can intuitively present the time-domain and frequency-domain characteristics of the vibration, such as amplitude, period, and duration. Audio signals can be the sound wave signals that accompany the occurrence of the microseismic event, captured by a detector or dedicated audio acquisition equipment. Their frequency, intensity, and other characteristics are potentially related to the energy of the microseismic event.

[0086] It should be noted that a geophone can be a sensor capable of converting ground vibration (mechanical vibration) into a measurable electrical signal. It is the core acquisition device of a mine microseismic monitoring system, capable of capturing vibration-related signals generated by microseismic events. Spatial location can be the three-dimensional spatial coordinates of the geophone within the mining area.

[0087] In practice, the assessment equipment acquires microseismic signal data from the mine safety monitoring system, specifically including:

[0088] Detector signal: for each microseismic event There are N detectors in total. Once the corresponding signals are acquired, the signals acquired by each detector can be divided into two main modes: waveform image signals and audio signals.

[0089] Spatial coordinates of the detector: for each detector Its spatial location in the mine is These are the spatial coordinates of the detector.

[0090] Step S20: Preprocess the multimodal signal to obtain the initial waveform feature vector and the initial audio feature vector.

[0091] It should be noted that the initial waveform feature vector can be a one-dimensional numerical vector obtained by transforming the preprocessed waveform image signal through a feature extraction algorithm. It contains key feature information of the waveform and can be used for subsequent feature fusion and model input.

[0092] The initial audio feature vector can be a one-dimensional numerical vector obtained by transforming the preprocessed audio signal through an audio feature extraction algorithm. It contains key audio feature information and forms a complementary initial feature set with the initial waveform feature vector.

[0093] Understandably, given a microseismic event and the detector that acquired the event Its raw input includes waveform image signals. With audio signals Because different modes have different sampling rates and lengths, it is necessary to perform time alignment, time-frequency transformation, and mode alignment on the waveform and audio signal.

[0094] In some embodiments, the evaluation device can use a wavelet threshold denoising algorithm to remove noise such as environmental vibrations (e.g., equipment operation vibrations, personnel activity vibrations) introduced during the acquisition process for waveform image signals, while retaining the effective vibration waveforms of micro-vibration events; for audio signals, an adaptive filtering algorithm can be used to remove mining operation noise (e.g., blasting residual noise, mechanical operation noise), while retaining audio segments of the micro-vibration event period and eliminating invalid period signals through signal truncation operations.

[0095] In some embodiments, the evaluation device can standardize the multimodal signal, and then perform standardization operations on the waveform image signal and audio signal after data cleaning, for example, using the Z-score standardization method to unify the numerical range of the two types of signals to the same order of magnitude (e.g., mean of 0 and variance of 1), to avoid the effective features being masked during subsequent feature extraction due to excessive differences in signal amplitude, thereby effectively removing noise and invalid information in the original signal, improving the signal-to-noise ratio of the signal, and reducing the interference of noise on subsequent feature fusion and model evaluation.

[0096] Step S30: Based on the spatial position of each detector, perform cross-modal fusion of the initial waveform feature vector and the initial audio feature vector of each detector, and output the fused modal features of each detector.

[0097] It should be noted that the fused modal features can be feature vectors obtained by processing through cross-modal fusion algorithms. They integrate the effective information of the initial waveform feature vector and the initial audio feature vector, and combine the correlation information of the detector's spatial position. They can more comprehensively reflect the local features of microseismic events than single modal features.

[0098] In practical implementation, the evaluation device can evaluate the standardized modal feature vectors. and Mapped to a unified time axis and encoded into a unified latent coding space to ensure cross-modal vector space consistency, cross-modal self-attention adaptive fusion is performed. Specifically, a position-aware energy-adaptive cross-modal self-attention fusion mechanism is constructed. This mechanism first performs bidirectional cross-modal self-attention modeling of the temporal dependencies and complementarities of waveform and audio modes in a unified latent space. Then, based on event energy estimation and detector spatial location information, a gating function is constructed to perform fine-grained adaptive weighting of the two modes in the channel dimension, thereby achieving a multi-modal fusion representation strongly coupled with the energy assessment task.

[0099] Step S40: Construct a graph convolutional network based on the microseismic event and the spatial location of each detector, and input the fused modal features of each detector into the graph convolutional network for aggregation to obtain the target features of the microseismic event.

[0100] It should be noted that Graph Convolutional Networks (GCNs) are deep learning models based on graph-structured data. Their core principle is to aggregate and update node features through the adjacency relationships of nodes in the graph, effectively mining spatial correlation features in the data and used to process spatial distribution structure data of geophones. The graph convolutional network includes geophone nodes and source nodes for microseismic events.

[0101] Detector nodes can be constructed by treating each detector as a node in the graph during the graph convolutional network construction process. The initial features of the nodes are the fused modal features output in step three. Source nodes can be virtual nodes added to the graph convolutional network, representing the source location of the microseismic event. Their initial features can be initialized using information such as the center coordinates of the spatial location of each detector and the approximate time of the microseismic event.

[0102] In practical implementation, the evaluation device can use graph convolutional networks to model the spatial structure of multiple detectors. For microseismic events... The set of nodes for constructing a graph convolutional network is Corresponding to the collected microseismic events N detectors, and microseismic events The epicenter itself. The edge set consists of all connected nodes. For microseismic events. All sensor nodes that have collected the relevant information Input feature vector The edge weights between the source node and the detector node are determined by a graph convolutional network. After normalization, the importance of each sensor node in determining the energy of the source node is obtained, and the final microseismic source node features are output. By taking advantage of the graph convolutional network's ability to process spatially correlated data, the spatial dependencies between the detectors are effectively mined, and the collaborative aggregation of features from multiple detectors is achieved, avoiding information loss caused by processing individual detector features in isolation.

[0103] In some embodiments, the evaluation device constructs a graph structure with each detector as a detector node and adds one source node; it calculates the adjacency weights between detector nodes based on the spatial coordinates of the detectors (the closer the distance, the higher the weight), and simultaneously calculates the adjacency weights between each detector node and the source node (based on the approximate distance from the detector to the source), thus constructing the adjacency matrix of the graph; the network structure consists of 2-3 graph convolutional layers, each convolutional layer is followed by a batch normalization layer and an activation function (such as the ReLU function), and finally a global pooling layer is set for feature aggregation.

[0104] The fused modal features of each detector are used as the initial features of the corresponding detector node; the initial features of the source node are initialized in the following way: the approximate location of the source is estimated based on the arrival time difference of the signals acquired by each detector, and the location coordinates, the occurrence time of the microseismic event, the mean of the fused modal features of all detectors, and other information are combined into the initial feature vector of the source node.

[0105] The initialized graph structure is input into a graph convolutional network. Neighborhood aggregation of node features is achieved through convolutional operations in each layer (i.e., each node feature is updated to a weighted sum of its own features and the features of its neighboring nodes). After multiple convolutional updates, the final features of all nodes are summarized through a global pooling layer, outputting microseismic event target features with fixed dimensions. During training, gradient descent algorithm is used to optimize network parameters, and the mapping error between target features and real seismic source energy is used as the loss function to improve the effectiveness of feature aggregation.

[0106] Furthermore, in order to effectively aggregate the fusion mode features of all geophones related to microseismic events in the mining space, in one embodiment, the graph convolutional network is configured to calculate the position-aware edge weights between the geophone nodes and the source nodes through a position-aware edge weight function.

[0107] It should be noted that, firstly, considering the influence of the distance between the seismic source and the detector, the propagation delay, and the geometric position on energy estimation, a function for the position-aware edge weights is designed, referring to the following formula:

[0108]

[0109] in, For detector nodes Source nodes corresponding to microseismic events Position-aware edge weights between them For the Sigmoid function, , and This represents learnable hyperparameters, which control the weights of distance, latency, and spatial location information in the graph convolutional network. For detector nodes With the focal node The initial predicted distance between them For detector nodes With the focal node The propagation delay and These are the detector nodes. and epicenter node Position encoding vector, It is a detector node and epicenter node The inner product of the position encoding vectors is used to measure the spatial angular relationship between the detector node and the source node;

[0110] The graph convolutional network is further configured to normalize the position-aware edge weights using a softmax function to obtain the target edge weights, as shown in the following formula:

[0111]

[0112] in, The standardized target edge weights, and These are the detector nodes. Detector node The focal node of microseismic events Position-aware edge weights between them To collect microseismic events The set of all detector nodes for the information;

[0113] The graph convolutional network is further configured to perform graph convolution operations on the fused modal features of each detector based on the target edge weights, and output the target features of the microseismic event.

[0114] Understandably, in the end, the evaluation device can utilize edge weights. This is used to regulate the propagation of information between the detector node and the source node, specifically through graph convolution operations, as shown in the following formula:

[0115]

[0116] in, For microseismic events that aggregate information from all geophone nodes Target characteristics, Microseismic events Intermediate detector node The fusion modal features.

[0117] Step S50: Predict the source energy of microseismic events based on target characteristics.

[0118] It should be noted that source energy prediction refers to establishing a mapping relationship between target characteristics and the actual source energy of microseismic events.

[0119] In some embodiments, the evaluation device may use a deep learning regression model (such as a multilayer perceptron MLP or a gradient boosting tree GBRT) as the source energy prediction model; the input layer dimension of the model is consistent with the target feature dimension, the hidden layer is set to 2-4 layers, the number of neurons in each layer is adaptively adjusted according to the feature dimension, and the output layer is a single neuron (corresponding to the estimated source energy value); before model training, the dataset is divided into a training set, a validation set, and a test set. The training set is used for model parameter fitting, the validation set is used to adjust model hyperparameters (such as learning rate and number of hidden layers), and the test set is used to evaluate the model prediction performance.

[0120] Furthermore, in order to accurately assess the energy of the microseismic source, step S50 above may include:

[0121] Step S501: Predict the source energy of microseismic events based on target characteristics.

[0122] In the specific implementation, the source node features are output through a graph convolutional network. Microseismic events Prediction of earthquake source energy:

[0123]

[0124] in, Microseismic events Predicted energy; For learnable regression weight vectors, To transpose it; For microseismic event source nodes that aggregate information from all geophone nodes The features are output by the graph convolutional network; This is a learnable bias term.

[0125] This embodiment acquires microseismic signal data, which includes multimodal signals of microseismic events collected by multiple geophones and the spatial location of each geophone in the mining area. The multimodal signals include waveform image signals and audio signals. The multimodal signals are preprocessed to obtain initial waveform feature vectors and initial audio feature vectors. Based on the spatial location of each geophone, the initial waveform feature vectors and initial audio feature vectors of each geophone are fused across modes to output the fused mode features of each geophone. A graph convolutional network is constructed based on the microseismic events and the spatial location of each geophone, and the fused mode features of each geophone are input into the network. Graph convolutional networks are used to aggregate data and obtain target features of microseismic events. The graph convolutional network includes detector nodes and source nodes of microseismic events. Based on the target features, the source energy of microseismic events is predicted. Because this embodiment overcomes the limitations of traditional single signal monitoring by acquiring multimodal signals and fusing them across modes, it comprehensively explores the waveform, audio, and spatial correlation features of microseismic events, improving the completeness of feature representation. Combined with the spatial feature aggregation capability of graph convolutional networks, it effectively integrates complementary information from multiple sources, reduces the impact of noise, environmental interference, and other factors on the evaluation results, and significantly improves the accuracy and stability of source energy assessment.

[0126] refer to Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the multimodal fusion method for assessing microseismic energy in mines according to the present invention.

[0127] Based on the first embodiment described above, in this embodiment, step S20 further includes:

[0128] Step S201: Construct a time window clipping function based on the acquisition time window of the multimode signals of each detector, and perform time window clipping on the multimode signals based on the time window clipping function to obtain clipped waveform signals and clipped audio signals.

[0129] Understandably, since each detector may acquire signals of different lengths, and the signals acquired by different detectors will also have different start times, time window trimming and alignment are necessary. Assume the... Signal data of microseismic events are detected by the detector. The collection time is A unified time window needs to be selected. The signal from the detector is clipped. The time window clipping function is defined as follows:

[0130]

[0131]

[0132] in, For the event Middle The signal obtained by trimming the waveform signal detected by each detector; For the event , by the The signal obtained by trimming the audio signal detected by each detector; , These are the waveform and audio signal initially acquired by the detector, respectively. Indicates the signal Extracted time range .

[0133] In some embodiments, the evaluation device can combine the typical duration of mine microseismic events (e.g., common microseismic events last from several seconds to tens of seconds) and the signal triggering mechanism of the geophones (e.g., triggering acquisition when the vibration amplitude exceeds a preset threshold) to define the range of the acquisition time window. For example, it can be set to "0.5-1 seconds before triggering to 5-10 seconds after triggering" to ensure coverage of the complete occurrence process of the microseismic event; at the same time, it can perform minor calibration on the start time of the time window of each geophone to take into account the acquisition delay differences of geophones in different areas, so as to ensure that the acquisition time windows of multiple geophones for the same microseismic event are synchronized.

[0134] Step S202: Convert the cropped waveform signal and the cropped audio signal to the time-frequency domain using the short-time Fourier transform function to obtain the time-frequency diagram of the waveform signal and the time-frequency diagram of the audio signal.

[0135] It should be noted that the short-time Fourier transform function is a mathematical transform function that converts a time-domain signal into a time-frequency domain signal. The core of it is to perform a piecewise Fourier transform on the time-domain signal through a sliding time window, which can preserve the time information of the signal and obtain the frequency characteristics at each time point, and is used for time-frequency analysis of non-stationary signals (such as microseismic signals).

[0136] In its implementation, the evaluation device converts the waveform and audio signals from each detector to the time-frequency domain, using a short-time Fourier transform (STFT) to transform the signals into a time-frequency image. This operation maps the waveform and audio signals from one-dimensional signals to a two-dimensional time-frequency image.

[0137] For waveforms and audio signals, the first Microseismic events were detected by the detector. The time-frequency images on the above are as follows:

[0138]

[0139]

[0140] in, , These are events Middle The waveform and time-frequency diagram of the audio signal from each detector. Represents the frequency dimension. It represents the time dimension, and the waveform has the same dimension as the processed audio signal.

[0141] Step S203: Standardize the time-frequency diagram of the waveform signal and the time-frequency diagram of the audio signal to obtain the initial waveform feature vector and the initial audio feature vector.

[0142] It should be understood that for multimodal (waveform signals, audio signals) data, cross-modal alignment is used to enable features from different modalities to be compared and fused in the same space. Specifically, the feature vector of each modality is standardized to ensure that features from different modalities can be compared at the same scale.

[0143]

[0144]

[0145] in, and events Middle The standardized waveform and audio feature vector of each detector; and events Middle The feature vector is composed of the time-frequency domain features of the waveform and audio extracted by each detector; and These are the mean and standard deviation of the waveform signal characteristics, respectively; and These are the mean and standard deviation of the audio signal characteristics, respectively.

[0146] This embodiment effectively removes redundant information and improves processing efficiency through time window pruning. It fully mines the time-frequency joint features of the signal by leveraging short-time Fourier transform to compensate for the shortcomings of single-domain analysis. Standardized processing achieves unified quantification of features and ensures the effectiveness of subsequent processing. The overall process not only enhances the extraction quality of effective features of microseismic signals, but also reduces the complexity of data processing and computational costs through the collaborative design of each step. It provides high-quality and highly adaptable basic feature data for core links such as cross-modal feature fusion and microseismic event energy assessment, effectively improving the intelligence level and feature mining capabilities of the entire microseismic monitoring and assessment system.

[0147] refer to Figure 4 , Figure 4 This is a flowchart illustrating the third embodiment of the multimodal fusion method for assessing microseismic energy in mines according to the present invention.

[0148] Based on the above embodiments, in this embodiment, step S30 further includes:

[0149] Step S301: Linearly convert the initial waveform feature vector and initial audio feature vector of each detector into waveform attention intermediate vector and audio attention intermediate vector, respectively.

[0150] It should be noted that the waveform attention intermediate vector includes a waveform query vector, a waveform key vector, and a waveform value vector, and the audio attention intermediate vector includes an audio query vector, an audio key vector, and an audio value vector.

[0151] It is understandable that in this embodiment, the self-attention intermediate vector of multimodal features is extracted through a cross-modal self-attention layer. On the one hand, the cross-modal self-attention layer models the temporal correlation within a single modality. On the other hand, it explicitly aligns and fuses the complementary information of waveform and audio modalities in the time dimension through a bidirectional co-attention mechanism, thereby obtaining a joint representation that is more sensitive to and more robust to microseismic information.

[0152] Step S302: Perform self-attention analysis based on the waveform attention intermediate vector and the audio attention intermediate vector to obtain the waveform mode self-attention output result and the audio mode self-attention output result.

[0153] For waveforms and audio modes, respectively:

[0154]

[0155]

[0156]

[0157]

[0158]

[0159] in, , events Middle The standardized waveform and audio feature vector of each detector; , events Middle Waveform and audio query vectors for each detector. , events Middle Waveform and audio key vector of each detector; , events Middle Waveform and audio value vector (Value) of each detector; , events Middle Each detector projects waveform and audio features into a linear transformation of the query vector; , events Middle Each detector projects waveform and audio features into a linear transformation of the key vector; , events Middle Each detector projects waveform and audio features into a linear transformation of a value vector; , events Middle Self-attention output of each detector waveform mode and audio mode; The standard scaling factor is used to avoid excessively large dot products that could lead to unstable gradients. To normalize the exponential function, attention weights are provided so that the sum of the weights is 1; , For the event Middle The waveform and audio key vector transpose of each detector.

[0160] Step S303: Based on the spatial position of each detector, the waveform mode self-attention output result, and the audio mode self-attention output result, the initial waveform feature vector and the initial audio feature vector are fused across modes to output the fused mode features of each detector.

[0161] In practical implementation, the evaluation device can construct a cross-modal fusion weight matrix by combining the feature importance and spatial weight coefficients of the self-attention output results. For the same detector, the basic fusion weights are first determined based on the information entropy of the waveform and audio self-attention output results (the higher the information entropy, the larger the basic weight, indicating that the modal feature contains more effective information); then the basic weights are multiplied by the spatial weight coefficients of the detector to obtain the final waveform modal fusion weights and audio modal fusion weights, ensuring that the fusion process simultaneously considers feature quality and spatial correlation.

[0162] Furthermore, in order to effectively extract cross-modal interaction features and effectively mine the feature association information within each modality, step S303 above may include:

[0163] Step S3031: Linearly convert the initial waveform feature vector into a first query vector, and linearly convert the initial audio feature vector into a first key vector and a first value vector;

[0164] Step S3032: Perform cross-modal co-attention analysis based on the first query vector, the first key vector, and the first value vector to obtain the first cross-modal co-attention feature of waveform mode focusing on audio mode;

[0165] Step S3033: Linearly convert the initial audio feature vector into a second query vector, and linearly convert the initial waveform feature vector into a second key vector and a second value vector;

[0166] Step S3034: Perform cross-modal co-attention analysis based on the second query vector, the second key vector, and the second value vector to obtain the second cross-modal co-attention feature of the audio modality of interest waveform;

[0167] Step S3035: Perform residual connection between the waveform modal self-attention output result, the first cross-modal co-attention feature and the initial waveform feature vector to obtain the waveform cross-modal interaction feature;

[0168] Step S3036: Perform a residual connection between the audio modal self-attention output result, the second cross-modal co-attention feature, and the initial audio feature vector to obtain the audio cross-modal interaction feature;

[0169] Step S3037: Based on the spatial position of each detector, perform cross-modal fusion of the waveform cross-modal interaction feature and the audio cross-modal interaction feature, and output the fused modal feature of each detector.

[0170] Understandably, this is to ensure that the two modalities are aligned and complementary. Let the waveform be the query and the audio be the key / value pair; that is, the waveform modality supplements information from the audio modality.

[0171]

[0172] in, For the event Middle Each detector waveform mode focuses on the cross-modal co-attention characteristics of the audio mode; The information flows from the audio mode to the waveform mode.

[0173] Let the audio be the query and the waveform be the key / value:

[0174]

[0175] in, For the event Middle The cross-modal co-attention characteristics of the audio mode of the detector are of interest to the waveform mode; This represents the flow of information from waveform mode to audio mode.

[0176] Finally, the original features, self-attention results, and cross-modal results are combined using residual linking to obtain the cross-modal representation after mutual perception:

[0177]

[0178]

[0179] in, , Representing events respectively Middle The waveforms and audio signals of each detector are fused with cross-modal interaction information to form their characteristics. , events Middle The standardized waveform and audio feature vector of each detector; , events Middle Self-attention output of each detector waveform mode and audio mode; , events Middle Cross-modal co-attention features of detector waveform modes and audio modes.

[0180] Furthermore, in order to accurately fuse cross-modal features and achieve complementarity and interaction of information between modalities, step S3037 above may include:

[0181] Step S30371: Determine the initial predicted distance between each detector and the seismic source based on the spatial position and preliminary positioning coordinates of each detector. The preliminary positioning coordinates are determined based on the three-dimensional velocity model of the mining area and the wave travel time difference between P-waves and S-waves, referring to the following formula:

[0182]

[0183] in, For detector coordinate, Microseismic events Preliminary positioning coordinates Let be the Euclidean distance function. For detector Microseismic events The initial predicted distance between the earthquake sources;

[0184] Step S30372: Calculate the propagation delay of each detector based on the initial predicted distance, referring to the following formula:

[0185]

[0186] in, For detector With the epicenter The propagation delay The equivalent wave velocity in the medium;

[0187] Step S30373: Encode the three-dimensional coordinates of the spatial position of each detector according to multiple frequencies to obtain the spatial position code of each detector;

[0188] Step S30374: Map the spatial location codes to the required dimensions of the model to obtain the location code vectors of each detector;

[0189] Step S30375: Construct the input vector of the gating network based on the position encoding vector, the initial prediction distance, the propagation delay, the waveform cross-modal interaction features, and the audio cross-modal interaction features;

[0190] Step S30376: Input the input vector into a gating network, the gating network being configured to output a gating vector based on the input vector through a multilayer perceptron network, the calculation formula of the gating vector being as follows:

[0191]

[0192]

[0193] in, and These represent the waveform cross-modal interaction features and the audio cross-modal interaction features, respectively. For detector Position encoding vector, Indicates the first The first microseismic event The gating vector corresponding to each detector A multilayer perceptron network representing a gated network. This represents the Sigmoid activation function;

[0194] Step S30377: Based on the gate vector, perform cross-modal fusion of the waveform cross-modal interaction features and the audio cross-modal interaction features, and output the fused modal features of each detector, referring to the following formula:

[0195]

[0196] in, For the first The first microseismic event Fusion mode characteristics of individual detectors This is for element-wise multiplication.

[0197] Understandably, the reliability of waveform and audio for energy estimation varies under different geometric conditions (distance, orientation, propagation path), hence the need to construct position-driven modal fusion weights.

[0198] First, distance and time delay estimation:

[0199] Preliminary estimate of the distance from the detector to the seismic source:

[0200]

[0201] in, For detector coordinate; For events given by the system Preliminary coordinate location; To calculate the Euclidean distance; For detector With the epicenter The distance.

[0202] Calculate the propagation delay prior:

[0203]

[0204] in, For detector With the epicenter The propagation delay; It is the equivalent wave velocity in the medium.

[0205] Second, spatial location coding:

[0206] Since the original coordinates are continuous real numbers with limited expressive power, and spatial structures at different scales are not easy to capture, spatial encoding combined with linear projection is used to map the coordinates to learnable embeddings for use by gating networks.

[0207] First, each coordinate dimension is encoded using multiple frequencies:

[0208]

[0209] in, For detector Spatial location encoding; It is a multi-frequency set; , , Detector x, y, z coordinates in the three-dimensional space of the mine; , These are the sine and cosine functions, respectively.

[0210] Secondly, map the high-dimensional sinusoidal encoding to the dimension required by the model:

[0211]

[0212] in, For detector Position encoding vector, It is a learnable matrix; For detector Spatial location encoding.

[0213] Third, position-aware modal gating:

[0214] The effectiveness of waveform and audio modes in energy estimation varies depending on the propagation distance and orientation. Therefore, modal weights are generated using location priors and then fine-grained weighting is applied at the channel dimension.

[0215] First, construct the input vector of the gated network:

[0216]

[0217] in, This indicates vector concatenation; These are cross-modal interaction information features of waveforms and audio, respectively; For detector Location encoding vector; For detector With the epicenter The distance; For detector With the epicenter The propagation delay; The input vector to the constructed gating network integrates waveform information, audio information, distance between the detector and the source, and propagation path information.

[0218] Subsequently, the gated network uses a multilayer perceptron (MLP) for its output:

[0219]

[0220] Finally, modal adaptive fusion is performed:

[0221]

[0222] in, This is the channel-level gating vector output by the multilayer perceptron (MLP). For the Sigmoid function, compress the output to... interval; It is a multilayer perceptron network; The modal features for final fusion; These are cross-modal interaction information features of waveforms and audio, respectively; This is for element-wise multiplication.

[0223] This embodiment adapts the initial features to the attention mechanism through linear transformation, strengthens the core features of each mode and suppresses redundant information by leveraging self-attention analysis, and achieves dynamic weighted fusion of multimodal features by combining the spatial position of the detector. It effectively mines the complementary information and intramodal correlation information of multimodal features, and also incorporates the physical correlation characteristics of spatial position, which significantly improves the representation ability and robustness of the fused features, reduces the feature learning difficulty of subsequent models, provides core technical support for the accurate assessment of microseismic events, and improves the intelligence and accuracy of the entire microseismic monitoring and analysis system.

[0224] Furthermore, embodiments of the present invention also propose a computer-readable storage medium storing a multimodal fusion mine microseismic energy assessment program, wherein when the multimodal fusion mine microseismic energy assessment program is executed by a processor, it implements the steps of the multimodal fusion mine microseismic energy assessment method described above.

[0225] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0226] The aforementioned computer-readable storage medium may be included in a multimodal fusion mine microseismic energy assessment device; or it may exist independently and not assembled into a multimodal fusion mine microseismic energy assessment device.

[0227] Furthermore, this invention also proposes a computer program product, including a multimodal fusion mine microseismic energy assessment program, which, when executed by a processor, implements the steps of the multimodal fusion mine microseismic energy assessment method as described above.

[0228] The specific implementation of the computer program product of this invention is basically the same as the embodiments of the above-mentioned multimodal fusion mine microseismic energy assessment method, and will not be repeated here.

[0229] Reference Figure 5 , Figure 5 This is a structural block diagram of the first embodiment of the multimodal fusion mine microseismic energy assessment device of the present invention.

[0230] like Figure 5As shown in the embodiment of the present invention, the multimodal fusion mine microseismic energy assessment device includes:

[0231] The data acquisition module 10 is used to acquire microseismic signal data, which includes multi-mode signals of microseismic events collected by multiple detectors and the spatial position of each detector in the mining area. The multi-mode signals include waveform image signals and audio signals.

[0232] Preprocessing module 20 is used to preprocess the multimodal signal to obtain initial waveform feature vector and initial audio feature vector;

[0233] The cross-modal fusion module 30 is used to perform cross-modal fusion of the initial waveform feature vector and the initial audio feature vector of each detector based on the spatial position of each detector, and output the fused modal features of each detector;

[0234] The feature aggregation module 40 is used to construct a graph convolutional network based on the microseismic event and the spatial location of each detector, and input the fused modal features of each detector into the graph convolutional network for aggregation to obtain the target features of the microseismic event. The graph convolutional network includes detector nodes and source nodes of the microseismic event.

[0235] Energy prediction module 50 is used to predict the source energy of microseismic events based on target characteristics.

[0236] This embodiment acquires microseismic signal data, which includes multimodal signals of microseismic events collected by multiple geophones and the spatial location of each geophone in the mining area. The multimodal signals include waveform image signals and audio signals. The multimodal signals are preprocessed to obtain initial waveform feature vectors and initial audio feature vectors. Based on the spatial location of each geophone, the initial waveform feature vectors and initial audio feature vectors of each geophone are fused across modes to output the fused mode features of each geophone. A graph convolutional network is constructed based on the microseismic events and the spatial location of each geophone, and the fused mode features of each geophone are input into the network. Graph convolutional networks are used to aggregate data and obtain target features of microseismic events. The graph convolutional network includes detector nodes and source nodes of microseismic events. Based on the target features, the source energy of microseismic events is predicted. Because this embodiment overcomes the limitations of traditional single signal monitoring by acquiring multimodal signals and fusing them across modes, it comprehensively explores the waveform, audio, and spatial correlation features of microseismic events, improving the completeness of feature representation. Combined with the spatial feature aggregation capability of graph convolutional networks, it effectively integrates complementary information from multiple sources, reduces the impact of noise, environmental interference, and other factors on the evaluation results, and significantly improves the accuracy and stability of source energy assessment.

[0237] The multimodal fusion mine microseismic energy assessment device provided in this application, employing the multimodal fusion mine microseismic energy assessment method described in the above embodiments, can solve the technical problems of multimodal fusion mine microseismic energy assessment. Compared with the prior art, the beneficial effects of the multimodal fusion mine microseismic energy assessment device provided in this application are the same as those of the multimodal fusion mine microseismic energy assessment method described in the above embodiments, and other technical features in the multimodal fusion mine microseismic energy assessment device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0238] It should be understood that the above are merely illustrative examples and do not constitute any limitation on the technical solution of the present invention. In specific applications, those skilled in the art can make settings as needed, and the present invention does not impose any restrictions on this.

[0239] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this invention. In practical applications, those skilled in the art can select some or all of the workflow to achieve the purpose of this embodiment according to actual needs, and no restrictions are imposed here.

[0240] In addition, for technical details not described in detail in this embodiment, please refer to the multimodal fusion mine microseismic energy assessment method provided in any embodiment of the present invention, which will not be repeated here.

[0241] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0242] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0243] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory / random access memory, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0244] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. A multimodal fusion method for assessing microseismic energy in mines, characterized in that, The multimodal fusion method for assessing microseismic energy in mines includes: Acquire microseismic signal data, which includes multi-mode signals of microseismic events collected by multiple detectors and the spatial location of each detector in the mining area. The multi-mode signals include waveform image signals and audio signals. The multimodal signal is preprocessed to obtain an initial waveform feature vector and an initial audio feature vector; Based on the spatial location of each detector, the initial waveform feature vector and the initial audio feature vector of each detector are fused across modes to output the fused mode features of each detector; A graph convolutional network is constructed based on the microseismic events and the spatial locations of each detector. The fused modal features of each detector are then input into the graph convolutional network for aggregation to obtain the target features of the microseismic events. The graph convolutional network includes detector nodes and source nodes of the microseismic events. Predicting source energy of microseismic events based on target characteristics; The method involves cross-modal fusion of the initial waveform feature vector and the initial audio feature vector of each detector based on their spatial location, outputting the fused modal features of each detector, including: The initial predicted distance between each geophone and the seismic source is determined based on the spatial location and preliminary positioning coordinates of each geophone. The preliminary positioning coordinates are determined based on the three-dimensional velocity model of the mining area and the wave travel time difference between P-waves and S-waves, referring to the following formula: in, For detector coordinate, Microseismic events Preliminary positioning coordinates Let be the Euclidean distance function. For detector Microseismic events The initial predicted distance between the earthquake sources; The propagation delay of each detector is calculated based on the initial predicted distance, referring to the following formula: in, For detector With the epicenter The propagation delay The equivalent wave velocity in the medium; The spatial position codes of each detector are obtained by encoding the three-dimensional coordinates of the spatial position of each detector based on multiple frequencies. The spatial location codes are mapped to the required dimensions of the model to obtain the location code vectors of each detector. The input vector of the gated network is constructed based on the waveform cross-modal interaction features, the audio cross-modal interaction features, the position encoding vector, the initial prediction distance, and the propagation delay. The waveform cross-modal interaction features and the audio cross-modal interaction features are obtained by performing linear transformation, cross-modal co-attention analysis, and residual connection on the initial waveform feature vector and the initial audio feature vector. The input vector is fed into a gating network, which is configured to output a gating vector based on the input vector through a multilayer perceptron network. The formula for calculating the gating vector is as follows: in, and These represent the waveform cross-modal interaction features and the audio cross-modal interaction features, respectively. For detector Position encoding vector, Indicates the first The first microseismic event The gate vector corresponding to each detector A multilayer perceptron network representing a gated network. This represents the Sigmoid activation function; Based on the gate vector, the waveform cross-modal interaction features and the audio cross-modal interaction features are fused across modally to output the fused modal features of each detector, as shown in the following formula: in, For the first The first microseismic event Fusion mode characteristics of individual detectors This is for element-wise multiplication.

2. The multimodal fusion method for mine microseismic energy assessment as described in claim 1, characterized in that, The preprocessing of the multimodal signal to obtain initial waveform feature vectors and initial audio feature vectors includes: A time window clipping function is constructed based on the acquisition time window of the multimode signals of each detector, and the multimode signals are clipped based on the time window clipping function to obtain clipped waveform signals and clipped audio signals. The cropped waveform signal and the cropped audio signal are converted to the time-frequency domain using the short-time Fourier transform function to obtain the time-frequency diagram of the waveform signal and the time-frequency diagram of the audio signal. The waveform signal time-frequency diagram and the audio signal time-frequency diagram are standardized to obtain the initial waveform feature vector and the initial audio feature vector.

3. The multimodal fusion method for mine microseismic energy assessment as described in claim 1, characterized in that, The method involves cross-modal fusion of the initial waveform feature vector and the initial audio feature vector of each detector based on their spatial location, outputting the fused modal features of each detector, including: The initial waveform feature vector and initial audio feature vector of each detector are linearly converted into waveform attention intermediate vector and audio attention intermediate vector, respectively. The waveform attention intermediate vector includes waveform query vector, waveform key vector and waveform value vector, and the audio attention intermediate vector includes audio query vector, audio key vector and audio value vector. Self-attention analysis is performed based on the waveform attention intermediate vector and the audio attention intermediate vector to obtain waveform mode self-attention output results and audio mode self-attention output results; Based on the spatial position of each detector, the waveform mode self-attention output result, and the audio mode self-attention output result, the initial waveform feature vector and the initial audio feature vector are fused across modes to output the fused mode features of each detector.

4. The multimodal fusion method for mine microseismic energy assessment as described in claim 3, characterized in that, The method of fusing the initial waveform feature vector and the initial audio feature vector across modes based on the waveform mode self-attention output result and the audio mode self-attention output result, and outputting the fused mode features of each detector, includes: The initial waveform feature vector is linearly converted into a first query vector, and the initial audio feature vector is linearly converted into a first key vector and a first value vector; Based on the first query vector, the first key vector, and the first value vector, cross-modal co-attention analysis is performed to obtain the first cross-modal co-attention feature of waveform mode focusing on audio mode; The initial audio feature vector is linearly transformed into a second query vector, and the initial waveform feature vector is linearly transformed into a second key vector and a second value vector; Based on the second query vector, the second key vector, and the second value vector, cross-modal co-attention analysis is performed to obtain the second cross-modal co-attention feature of the audio modality-focused waveform modality; The waveform modal self-attention output result, the first cross-modal co-attention feature, and the initial waveform feature vector are residually connected to obtain the waveform cross-modal interaction feature; The audio modal self-attention output, the second cross-modal co-attention feature, and the initial audio feature vector are residually concatenated to obtain the audio cross-modal interaction feature; Based on the spatial position of each detector, the waveform cross-modal interaction features and the audio cross-modal interaction features are fused together to output the fused modal features of each detector.

5. The multimodal fusion method for mine microseismic energy assessment as described in claim 1, characterized in that, The graph convolutional network is configured to calculate the position-aware edge weights between the detector node and the source node using a position-aware edge weight function, as shown in the following formula: in, For detector nodes Source nodes corresponding to microseismic events Position-aware edge weights between them For the Sigmoid function, , and This represents learnable hyperparameters, which control the weights of distance, latency, and spatial location information in the graph convolutional network. For detector nodes With the focal node The initial predicted distance between them For detector nodes With the focal node The propagation delay and These are the detector nodes. and epicenter node Position encoding vector, It is a detector node and epicenter node The inner product of the position encoding vectors is used to measure the spatial angular relationship between the detector node and the source node; The graph convolutional network is further configured to normalize the position-aware edge weights using a softmax function to obtain the target edge weights, as shown in the following formula: in, The standardized target edge weights, and These are the detector nodes. Detector node The focal node of microseismic events Position-aware edge weights between them To collect microseismic events The set of all detector nodes for the information; The graph convolutional network is further configured to perform graph convolution operations on the fused modal features of each detector based on the target edge weights, and output the target features of the microseismic event, as shown in the following formula: in, For microseismic events that aggregate information from all geophone nodes Target characteristics, Microseismic events Intermediate detector node The fusion modal features.

6. The multimodal fusion method for mine microseismic energy assessment as described in claim 1, characterized in that, The prediction of source energy for microseismic events based on target features includes: Based on target characteristics, the source energy of microseismic events can be predicted using the following formula: in, Microseismic events Predicted energy, For learnable regression weight vectors, This is the transpose of the regression weight vector. For microseismic events that aggregate information from all geophone nodes Target characteristics, This is a learnable bias term.

7. A multimodal fusion mine microseismic energy assessment device, characterized in that, The multimodal fusion mine microseismic energy assessment device includes: The data acquisition module is used to acquire microseismic signal data, which includes multi-mode signals of microseismic events collected by multiple detectors and the spatial position of each detector in the mining area. The multi-mode signals include waveform image signals and audio signals. The preprocessing module is used to preprocess the multimodal signal to obtain an initial waveform feature vector and an initial audio feature vector; The cross-modal fusion module is used to perform cross-modal fusion of the initial waveform feature vector and the initial audio feature vector of each detector based on the spatial position of each detector, and output the fused modal features of each detector; The feature aggregation module is used to construct a graph convolutional network based on the microseismic event and the spatial location of each detector, and input the fused modal features of each detector into the graph convolutional network for aggregation to obtain the target features of the microseismic event. The graph convolutional network includes detector nodes and source nodes of the microseismic event. The energy prediction module is used to predict the source energy of microseismic events based on target characteristics. The cross-modal fusion module is also used to determine the initial predicted distance between each detector and the seismic source based on the spatial position of each detector and the preliminary positioning coordinates. The preliminary positioning coordinates are determined based on the three-dimensional velocity model of the mining area and the wave travel time difference between P-waves and S-waves, referring to the following formula: in, For detector coordinate, Microseismic events Preliminary positioning coordinates Let be the Euclidean distance function. For detector Microseismic events The initial predicted distance between the earthquake sources; The propagation delay of each detector is calculated based on the initial predicted distance, referring to the following formula: in, For detector With the epicenter The propagation delay The equivalent wave velocity in the medium; The spatial position codes of each detector are obtained by encoding the three-dimensional coordinates of their spatial locations using multiple frequencies. These spatial position codes are then mapped to the required dimensions of the model to obtain the position code vectors for each detector. An input vector for a gating network is constructed based on waveform cross-modal interaction features, audio cross-modal interaction features, the position code vectors, the initial prediction distance, and the propagation delay. The waveform cross-modal interaction features and audio cross-modal interaction features are obtained through linear transformation, cross-modal co-attention analysis, and residual connections of the initial waveform feature vector and the initial audio feature vector. The input vector is then input to the gating network, which is configured to output a gating vector based on the input vector using a multilayer perceptron network. The calculation formula for the gating vector is as follows: in, and These represent the waveform cross-modal interaction features and the audio cross-modal interaction features, respectively. For detector Position encoding vector, Indicates the first The first microseismic event The gate vector corresponding to each detector A multilayer perceptron network representing a gated network. This represents the Sigmoid activation function; Based on the gate vector, the waveform cross-modal interaction features and the audio cross-modal interaction features are fused across modally to output the fused modal features of each detector, as shown in the following formula: in, For the first The first microseismic event Fusion mode characteristics of individual detectors This is for element-wise multiplication.

8. A multimodal fusion mine microseismic energy assessment device, characterized in that, The multimodal fusion mine microseismic energy assessment device includes: a memory, a processor, and a multimodal fusion mine microseismic energy assessment program stored in the memory. The processor is used to run the multimodal fusion mine microseismic energy assessment program, which is configured to implement the multimodal fusion mine microseismic energy assessment method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a multimodal fusion mine microseismic energy assessment program, which, when executed by a processor, implements the multimodal fusion mine microseismic energy assessment method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Mine micro-seismic early warning method, device and equipment based on space-time diagram neural network

    CN118915143A

  • Microseismic event identification method, device and equipment and storage medium

    CN120316728A