Mine monitoring-oriented mine environment super-resolution image reconstruction method based on vibration induction
By combining a two-layer LSTM network with blueprint separable convolution, the problems of image blur and noise in mine working environments are solved, and high-frequency details are restored and computational efficiency is improved. It is suitable for scenarios such as mine monitoring and drone imaging.
Patent Information
- Application Number
- CN202510655741.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The complex vibrations in the mine operating environment lead to blurred and noisy image acquisition. Existing technologies are difficult to adapt to dynamic changes, have poor ability to restore high-frequency details of images, and have high computational complexity, which cannot meet the real-time requirements of mine monitoring.
A two-layer LSTM network is used to predict vibration characteristics, combined with blueprint separable convolution and coordinate attention mechanism to perform multi-scale feature extraction and dynamic compensation, and residual learning is used to achieve image super-resolution reconstruction to generate high-resolution images.
It effectively compensates for blur caused by vibration, improves the ability to restore image details, reduces computational complexity, and adapts to application scenarios with high real-time requirements such as mine monitoring.
Smart Images

Figure CN120725870A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image enhancement technology, and in particular to a mine environment super-resolution image reconstruction method based on vibration sensing for mine monitoring. Background Art
[0002] The dynamic and complex nature of mine operating environments exposes surveillance image acquisition to multiple interference sources. Mechanical and geological vibrations combine to create multi-directional, non-stationary disturbances, causing continuous displacement of monitoring equipment, resulting in image motion blur and misalignment of features between adjacent frames. Furthermore, the low illumination and high dust levels underground further exacerbate image noise and detail loss. Existing static super-resolution methods (such as SRCNN) rely on fixed blur model assumptions and are unable to adapt to the dynamic changes in vibration scenarios. Traditional video super-resolution techniques, however, suffer from motion estimation errors due to the nonlinear characteristics of vibration direction and amplitude, and the high computational complexity of multi-frame processing, making them difficult to meet the real-time requirements of underground edge devices.
[0003] Existing image reconstruction techniques often focus on optimizing single-modal data. Static super-resolution utilizes only the spatial features of a single frame, ignoring the dynamic relationship between vibration characteristics and blur patterns. While video super-resolution incorporates inter-frame temporal information, it relies on the assumption of linear optical flow, making it difficult to model the discontinuous motion trajectories caused by multi-source vibration in mines. Furthermore, traditional methods generally employ fixed-parameter blur correction and are unable to dynamically adjust reconstruction strategies based on external vibration sensor data (direction, amplitude, and frequency). This limits the ability to recover high-frequency image details in complex vibration scenarios.
[0004] Although super-resolution algorithms based on deep learning (such as convolutional neural networks) have improved reconstruction effects through feature extraction, their computational frameworks still have significant defects: the existing network models do not have a multimodal fusion mechanism designed for vibration data, and image features and vibration parameters are difficult to effectively combine due to differences in representation scales; at the same time, traditional sliding window convolution operations have low efficiency in utilizing the spatiotemporal information of the reference frame, and do not dynamically optimize the convolution kernel weights according to the vibration characteristics, resulting in high resource utilization of edge devices and insufficient real-time processing performance, which restricts the reliability improvement of mine safety monitoring systems. Summary of the Invention
[0005] The technical problem to be solved by the present invention is that the ability to restore high-frequency details of images in complex vibration scenes is poor.
[0006] To this end, the present invention provides a mine environment super-resolution image reconstruction method based on vibration sensing for mine monitoring.
[0007] The technical solution adopted by the present invention to solve its technical problem is:
[0008] A mine environment super-resolution image reconstruction method based on vibration sensing for mine monitoring, comprising:
[0009] S1, collecting vibration data and image data, and preprocessing the collected vibration data and image data;
[0010] S2, inputs the preprocessed vibration data sequence into a two-layer LSTM network, uses LSTM to analyze the vibration data and predict the vibration characteristics at the next moment;
[0011] S3, the preprocessed image is input into a convolutional neural network constructed using blueprint separable convolution, as well as a residual attention module and a coordinate attention mechanism for multi-scale feature extraction;
[0012] S4 uses the vibration parameters predicted by LSTM to generate dynamic compensation parameters, which are used to adjust the convolution kernel parameters of the super-resolution network;
[0013] S5, performing temporal alignment of the vibration data and the image data to generate corresponding compensation parameters, and simultaneously performing spatial alignment and fusion of the two to obtain fusion features;
[0014] S6, based on fusion features, adopts a residual network structure, realizes image magnification through multi-layer convolution and upsampling modules, and uses residual learning to reduce information loss, performs super-resolution reconstruction on the image, and generates high-resolution images.
[0015] Furthermore, in step 2, the specific steps of analyzing the vibration data using LSTM include:
[0016] S2.1, package vibration data into batches;
[0017] S2.2, batch i The kth group of features X under ik (θ k ,α k ,A k ,ω k ) performs forward propagation;
[0018] S2.3, based on the threshold value obtained during the forward propagation process, the calculation unit outputs y k and long-term memory;
[0019] S2.4, output y to the unit k Perform full connection output;
[0020] S2.5, perform steps S2.2 to S2.4 on all data in the batch, calculate the error, calculate the partial differential of the error with respect to the fully connected weights and unit output, and get the output gradient step output and input gradient δ0;
[0021] S2.7, gradient descent updates the weight matrix and backpropagates the input gradient.
[0022] Furthermore, in the step three, the blueprint separation convolution first performs a weighted combination of the input in the depth direction, and then performs multi-layer depth convolution on multiple channels at the same time. The residual attention module first adds the input and the output of the multi-layer convolution layer to realize residual learning, and then sends the added output features to the coordinate attention module, and then adds the output features after the coordinate attention module to the original input to obtain the final output features of the lightweight residual attention module.
[0023] Furthermore, the step three specifically includes the following steps:
[0024] S3.1, uses pooling kernels of size (H, 1) and (1, W) to perform pooling operations in the horizontal and vertical directions and obtain a set of perceptual feature maps with sizes of C×H×1 and C×1×W respectively;
[0025] S3.2, the perceptual feature maps of the above two directions are fused and spliced to obtain: X′=δ(f 1×1 ([z h ,z w ])), Where δ(·) represents the h_swish activation function; [z h ,z w ] represents the fusion operation of the output of the cth channel with a height of h and the output of the cth channel with a width of w; f 1×1 Represents 1×1 convolution; represents the intermediate feature map encoding spatial information, and r represents the downsampling step size;
[0026] S3.3, split F' into two independent tensors X' h and X' w , and change the feature map X' through 1×1 convolution layer h and X' w The number of channels is: g h =σ(f 1×1 (X' h )), g w =σ(f 1×1 (X' w )), Where σ(·) represents the sigmoid activation function; g h and g w As the attention weight is used to adjust the module's attention to the input image, the final output of the coordinate attention module can be expressed as: Y CA =X c(i,j)×g h ×g w .
[0027] Furthermore, in step 4, the vibration prediction parameters obtained after the hidden state output by the LSTM network is mapped by the fully connected layer contain multiple dimensions, which are used to guide the dynamic compensation operations of different areas in the image respectively, thereby achieving an organic combination of local and global compensation.
[0028] Furthermore, in step 4, the process of generating the dynamic fuzzy compensation matrix can be expressed as: K comp =g(d t+1 ,m t+1 ), where g(.) function is based on (d t+1 ,m t+1 )Adjust the direction and strength of the convolution kernel to adapt to the current environmental vibration state.
[0029] Furthermore, in step 5, the dynamic compensation parameter is K, and its calculation formula is: Among them, θ t is the vibration direction at the current moment, θ t+1 is the vibration direction of the next moment predicted by LSTM, and σ is the fuzzy adjustment parameter; the fusion of vibration data and image data adopts the weighted fusion method, and the fusion weight of vibration features and image features is calculated through the self-attention mechanism.
[0030] Furthermore, in step five, the fusion feature F fused Expressed as: F fused =α vib ·F vib +α img ·F img , where F vib is the vibration compensation feature obtained by mapping the LSTM output through the fully connected layer, and the weight satisfies α vib +α img =1.
[0031] A computer device comprising:
[0032] processor;
[0033] a memory for storing executable instructions;
[0034] The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the random walk-based mine image detail enhancement algorithm as described above.
[0035] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor implements the random walk-based mine image detail enhancement algorithm as described above.
[0036] The beneficial effects of the present invention are:
[0037] This application utilizes vibration data and employs low-pass filtering and Kalman filtering for denoising. It then performs time-series modeling on the vibration data through a two-layer LSTM network, predicts the vibration parameters at the next moment, and generates dynamic compensation parameters. Simultaneously, the feature extraction module extracts multi-scale features of the image through blueprint separable convolution and a coordinate attention mechanism. Combined with inter-frame motion compensation information, the feature alignment module performs a nonlinear transformation on the features of the target and reference frames to achieve precise alignment. Next, the fusion module combines the vibration compensation features with the image features through adaptive weight calculation to generate a fused high-dimensional feature representation. Finally, the super-resolution reconstruction module upsamples and enhances the fused features based on residual learning and a coordinate attention mechanism, thereby restoring a high-resolution image. Compared to existing methods, this invention not only effectively compensates for vibration-induced blur but also improves the detail restoration capability of super-resolution reconstruction through precise alignment and multi-scale fusion. Furthermore, the lightweight network structure reduces computational complexity and improves inference efficiency, making it widely applicable in real-time applications such as mine monitoring, drone imaging, and robot navigation. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The present invention will be further described below with reference to the accompanying drawings and examples.
[0039] Figure 1 This is a flow chart of the implementation of the mine environment super-resolution image reconstruction method based on vibration sensing for mine monitoring in the present invention.
[0040] Figure 2 This is a structural block diagram of the mine environment super-resolution image reconstruction method based on vibration sensing for mine monitoring in the present invention.
[0041] Figure 3 It is a structural diagram of the blueprint separable convolution module in the present invention.
[0042] Figure 4 This is a comparison chart of the detail effects of using the method of the present invention and the blueprint separable convolution method to recognize different images. DETAILED DESCRIPTION
[0043] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0044] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", "axial", "radial", "circumferential" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, features defined as "first" or "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.
[0045] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0046] A mine environment super-resolution image reconstruction method based on vibration sensing for mine monitoring.
[0047] S1, data acquisition and preprocessing
[0048] S1.1 In a mine monitoring environment, vibration data, including vibration direction, acceleration, amplitude, and frequency, is collected through an IMU sensor array; image data is collected through a camera.
[0049] It should be noted that in the IMU sensor array, the range of the three-axis accelerometer is ±10g and the sampling frequency is not less than 50Hz, the range of the three-axis gyroscope is ±250° / s and the sampling frequency is not less than 50Hz, and the range of the seismic wave sensor is ±1000μm / s and the sampling frequency is not less than 10Hz.
[0050] S1.2 Preprocess the collected vibration data. First, use low-pass filtering to remove high-frequency noise, then use Kalman filtering to smooth the data, and finally standardize it using the mean and standard deviation.
[0051] S1.3 preprocesses low-resolution images collected from mine monitoring. First, optical flow is used to align the images between frames. Non-local mean filtering is then used to perform preliminary denoising to suppress global noise and repetitive interference. Bilateral filtering is then used to further enhance image edge and texture details, preserving important structural information while denoising. This results in a clear, preprocessed image. Adaptive histogram equalization is then used to enhance image contrast, yielding the final image.
[0052] S2, inputs the preprocessed vibration data sequence into a two-layer LSTM network, uses LSTM to analyze the vibration data, and predicts the vibration characteristics at the next moment.
[0053] S2.1. Package the vibration data, integrate the four dimensional features of vibration direction θ, acceleration α, amplitude A, and frequency ω, and divide the data into batches. The size of each batch is between (min, max):
[0054]
[0055] Where n represents the batch size, i represents the batch number, and u and r are the mean and root mean square of the distribution, respectively.
[0056] S2.2, batch i The kth group of features X under ik (θ k ,α k ,A k ,ω k ) to perform forward propagation:
[0057]
[0058] Among them, forgetgate k Represents the forget gate, inputgate k Represents input gate, outputgate k represents the output gate, Indicates that each gate is connected to the input X k The weight of Indicates that each gate is connected to the short-term memory h k-1 The weight of (b f ,b g ,b j ,b o ) represents the bias term of each layer.
[0059] S2.3, based on the threshold value obtained in step S2.2, calculate the output and long-term memory:
[0060]
[0061] Among them, c k ,c k-1 Indicates long-term memory, y k Represents the unit output, h k Represents short-term memory.
[0062] S2.4, y for S2.3 k Perform full connection output:
[0063] y k+1 =W0·y k
[0064] Among them, W0 represents the full connection weight, y k+1 Represents the prediction of the next moment (θ k+1 ,α k+1 ,A k+1 ,ω k+1 ).
[0065] S2.5. Perform steps S2.2 to S2.4 on all data in each batch and calculate the error:
[0066]
[0067] Among them, MSE(.) represents the square error, X i(k+1) Represents the k+1th group of features in the i-th batch of data, and makes an error on W0,y k The partial differential of , the output gradient step output and input gradient δ0.
[0068] S2.6, gradient descent updates the weight matrix and backpropagates the input gradient:
[0069] (w x ,w h ,b) new =Adam((w x ,w h ,b),backward(δ0))
[0070] in, Represents the updated weight matrix.
[0071] S3, the preprocessed image is input into a convolutional neural network constructed using blueprint separable convolution, as well as a residual attention module and a coordinate attention mechanism for multi-scale feature extraction.
[0072] Blueprint separation convolution is another separation method of standard convolution. Its principle is mainly to first perform weighted combination in the depth direction and then perform depth convolution. The characteristic of this convolution method is that it introduces a cross-channel bridging mechanism, which can process the features of multiple channels at the same time, improve the feature expression ability, and thus improve the accuracy of the model; the residual attention module first adds the input and the output of the two convolution layers to realize residual learning, and then sends the added output features to the coordinate attention module, and then adds the output features after the coordinate attention module to the initial input to obtain the final output features of the lightweight residual attention module. The specific calculation formula is:
[0073] F mid =F n-1 +f BSConv (ReLu(f BSConv (F n-1 )))
[0074] F n =F n-1 +H CA (F mid )
[0075] Among them, Fi (i = 1, ..., n) is the input of the i-th residual attention module; Fmid is the output after adding the input and the output of the two convolutional layers; fBSConv is the blueprint separation convolution operation; HCA is the coordinate attention mechanism to extract features; ReLu is the ReLu activation function.
[0076] The coordinate attention mechanism adds additional coordinate information to the neural network, enabling the model to better understand the spatial relationship between pixels. It typically adds a coordinate attention module after the original convolutional layer. This module weights features based on pixel location information, allowing the model to focus more on certain important features, thereby improving reconstruction quality. The main steps are as follows:
[0077] In S3.1, we use pooling kernels of size (H, 1) and (1, W) to perform pooling operations in both horizontal and vertical directions and obtain a set of perceptual feature maps with sizes C×H×1 and C×1×W, respectively. Therefore, the output of the c-th channel with a height of h and a width of w is:
[0078]
[0079] S3.2, we fuse and splice the perceptual feature maps of the above two directions to obtain:
[0080]
[0081] In this expression, δ(·) represents the h_swish activation function; [z h ,z w ] indicates a fusion operation; f 1×1 Represents 1×1 convolution; represents the intermediate feature map encoding spatial information, and r represents the downsampling step size.
[0082] S3.3, we split F' into two independent tensors X' h and X' w , and change the feature map X' through 1×1 convolution layer h and X' w The number of channels, then we get:
[0083]
[0084] In this expression, σ(·) represents the sigmoid activation function; g h and g w As the attention weight is used to adjust the module's attention to the input image, the final output of the coordinate attention module can be expressed as:
[0085] Y CA =X c (i,j)×g h ×g w
[0086] S4 uses the vibration parameters predicted by the LSTM to generate dynamic compensation parameters, which are used to adjust the convolution kernel parameters of the super-resolution network. The vibration prediction parameters obtained after the hidden state output of the LSTM network is mapped through the fully connected layer contain multiple dimensions, which are used to guide the dynamic compensation operation of different areas in the image, achieving an organic combination of local and global compensation.
[0087] Perform forward propagation through the LSTM network to obtain the vibration parameter P at the next moment t+1 as follows:
[0088] P t+1 ={a t+1 ,d t+1 ,f t+1 ,m t+1}
[0089] The vibration parameters contain global and local correction information for image blur. The dynamic compensation parameters are not only used to correct image blur, but also to guide the parameter setting of the upsampling module during image reconstruction to adapt to the dynamic changes of image features under different vibration environments. The process of generating the dynamic blur compensation matrix can be expressed as:
[0090] K comp =g(dt+1 ,m t+1 )
[0091] Among them, the g(.) function is based on (d t+1 ,m t+1 )Adjust the direction and strength of the convolution kernel to adapt to the current environmental vibration state.
[0092] S5, performing temporal alignment of the vibration data and the image data to generate corresponding dynamic compensation parameters, and simultaneously performing spatial alignment and fusion of the two to obtain fusion features.
[0093] The dynamic compensation parameter is K, and its calculation formula is:
[0094]
[0095] Among them, θ t is the vibration direction at the current moment, θ t+1 is the vibration direction of the next moment predicted by LSTM, σ is the fuzzy adjustment parameter, and this formula is used to determine the directionality of fuzzy compensation in the image; the fusion step of the vibration data and the image data adopts the weighted fusion method, and the fusion weight of the vibration feature and the image feature is calculated through the self-attention mechanism. The fused feature F fused Given by: F fused =α vib ·F vib +α img ·F img , where F vib is the vibration compensation feature obtained by mapping the LSTM output through the fully connected layer, and the weight satisfies α vib +α img =1.
[0096] S6, based on the fused features, uses multi-layer convolution, upsampling, and residual learning modules to perform super-resolution reconstruction of the image, generating a high-resolution image. The super-resolution image reconstruction step uses a residual network structure, achieving image magnification through multi-layer convolution and upsampling modules, and using residual learning to reduce information loss, thereby improving the restoration of image details.
[0097] Example 1
[0098] Step 1: Multimodal sensors, including triaxial accelerometers, gyroscopes, and seismic wave sensors, are installed around key mine equipment (such as mine car tracks and around drilling rigs) to capture environmental vibration characteristics. The sensors are deployed at a distance of 0.5 meters from the cameras to ensure synchronized data collection and minimize viewing angle deviation.
[0099] The sensor collects vibration data at a sampling frequency of 50 Hz and obtains the following characteristics: acceleration a t : reflects the vibration intensity; direction dt : Main vibration direction, calculated by fusion with gyroscope data; frequency f t : Record the vibration frequency range and calculate it using FFT; amplitude m t : Quantify the vibration magnitude, obtained by calculating the signal peak value.
[0100] The collected data needs to be denoised and normalized to improve the prediction accuracy.
[0101] Denoising: Use low-pass filter (Butterworth filter) to eliminate high-frequency noise
[0102]
[0103] Among them, w i is the filter weight, and x(t) is the input signal.
[0104] Normalization: Min-Max normalization of data
[0105]
[0106] Map the data to [0,1] to reduce the dimension effect.
[0107] Step 2: Time Series Prediction (LSTM)
[0108] Input data: Take the vibration characteristic data of the latest 20 frames to form a time series input:
[0109]
[0110] Network structure: Use an LSTM network with two hidden layers, 128 units per layer, to output the vibration prediction characteristics for the next moment:
[0111]
[0112] The 2-layer LSTM network structure is as follows:
[0113] Input layer: dimension is 20×4 (time steps 20, number of features 4)
[0114] Hidden layer: 2 layers of LSTM, 128 units per layer
[0115] Output layer: predict the vibration characteristics of the next moment
[0116] Training process: The model is trained using 10,000 mine vibration data, the optimizer uses Adam, the loss function is mean square error (MSE), and the learning rate is 0.001
[0117]
[0118] Step 3: Prediction results
[0119] The trained LSTM model predicts vibration data for use in the dynamic compensation module. For example, if the prediction results show a vibration direction of 45° and an amplitude of 0.8, the fuzzy compensation parameters will be adjusted accordingly.
[0120] Step 4: Fuzzy compensation parameter generation
[0121] According to the predicted vibration characteristics (direction, amplitude), the dynamic fuzzy compensation matrix K is generated comp :
[0122]
[0123] Among them, the function g adjusts the direction and strength of the convolution kernel, for example:
[0124] when Deflect the convolution kernel to 45°;
[0125] when Amplify the weights to enhance the compensation effect.
[0126] Dynamic adjustment example:
[0127] Assuming the current vibration direction D = 45° and the amplitude M = 0.7, the compensation matrix is:
[0128]
[0129] This matrix is used to adjust the direction of image blur and perform dynamic compensation for areas of high vibration.
[0130] Step 5: Super-resolution reconstruction module
[0131] Blueprint separable convolution design: The convolution structure is divided into two stages: local convolution and global convolution. Local convolution is used to extract detailed features and adopts 3×3 convolution kernel; global convolution integrates vibration compensation information and full-image features through 5×5 convolution.
[0132] Dynamic Fusion: Vibration Compensation Parameter K comp Dynamically adjust the weights of the convolution kernel to make the model adaptive to vibration blur characteristics.
[0133] F'=Conv 5×5 (Conv 3×3 (I)·K comp )
[0134] Among them, K comp is the aforementioned fuzzy compensation matrix.
[0135] Network structure:
[0136] Input: low-resolution image ILR and the fuzzy compensation matrix K comp
[0137] Blueprint-based Separate Convolution Technology
[0138] Introducing residual modules and coordinate attention mechanisms
[0139] Output: High-resolution image I HR
[0140] Network process:
[0141] I HR =Reconstruct(I LR ,K comp )
[0142] Evaluation indicators
[0143] The super-resolution effect is evaluated using PSNR and SSIM:
[0144]
[0145] in:
[0146] PSNR (Peak Signal-to-Noise Ratio): A measure of image quality
[0147] SSIM (structural similarity): evaluates the structural consistency between the reconstructed image and the original image
[0148] Step 6: Multimodal Fusion Module
[0149] 1. Module Function
[0150] Integrate vibration sensing data and image features to improve reconstruction effects, and calculate the fusion weights of vibration features and image features through the self-attention mechanism.
[0151] 2. Weighted Fusion Design
[0152] The vibration feature branch provides blur correction parameters;
[0153] The image feature branch extracts spatial texture information;
[0154] Through the weighted fusion formula:
[0155] F final =αF image +βF vibration
[0156] Among them, α and β are fusion weight parameters.
[0157] 3. Integration process
[0158] Through the above LSTM time prediction, the moment when the vibration sensing is about to occur is fused and matched with the image, and the corresponding timestamps are aligned to enrich the image features.
[0159] The fused features are the features originally extracted by preprocessing and are appropriately fused and weighted.
[0160] Example 2
[0161] This embodiment conducts a mine monitoring scenario test in a mine operating area based on a mine environment super-resolution image reconstruction method based on vibration sensing for mine monitoring.
[0162] In a mine operation area, track vibrations range from 10 to 50 Hz, and drilling rig operation causes unstable vibrations. A surveillance camera equipped with a vibration sensor is installed within the mine operation area to capture low-resolution images.
[0163] The vibration sensor collects the vibration characteristics of the environment in real time; the vibration characteristics at the next moment are predicted through LSTM to generate a dynamic fuzzy compensation matrix; the vibration characteristics are fused with the features extracted by preprocessing through the fusion module; the low-resolution image is input into the super-resolution reconstruction module, and the super-resolution reconstruction is completed by combining the compensation information.
[0164] The experimental results are shown in Table 1:
[0165] Table 1 Comparison of experimental data between the method adopted by the present invention and the blueprint separable convolution method
[0166]
[0167] The original image PSNR value is 22.5. The present invention was compared with Blueprint Separable Convolution (BSC). The comparative data shows that the present invention has significant advantages in image quality: PSNR reaches 30.5dB, higher than BSC's 28.7dB, indicating that it is more effective in enhancing pixel-domain detail; the SSIM value is 0.92, surpassing BSC's 0.89, confirming that it better protects the structural integrity of images in vibration scenarios. Although the computational complexity (1.2G FLOPS) is slightly higher than BSC (0.8G FLOPS), the present invention achieves breakthroughs in vibration compensation and inter-frame alignment capabilities by leveraging LSTM vibration information prediction and adaptive compensation mechanisms, making it suitable for complex vibration environments such as mine monitoring and drone imaging. In contrast, BSC lacks a vibration compensation mechanism and is only suitable for general super-resolution tasks.
[0168] In order to more intuitively illustrate the technical effect of the present invention, we compared the enhancement results of different algorithms through experiments. Figure 4 As shown, Figure 4Figures a, b, and c are three actual images of the underground mine environment. The red frame in Figure a marks the wall image under the mine, and the red frames in Figures b and c mark two building structure images respectively. Figure 4 Small figures a1, b1, and c1 are respectively enlarged images of the images in the red frames of small figures a, b, and c; small figures a1, b1, and c1 are respectively schematic diagrams of the results of processing the images in the red frames of small figures a, b, and c using blueprint separable convolution; small figures a1, b1, and c1 are respectively schematic diagrams of the results of processing the images in the red frames of small figures a, b, and c using the algorithm of this application.
[0169] Observing the processing results of the blueprint separable convolution, it can be seen that it has obvious defects in restoring image details: when facing detailed areas such as wall texture and building structure, the processed image presents a fuzzy texture, and the original detail features are not accurately restored. For example, in parts such as wall hole texture and building lines, the problem of detail loss is prominent and the visual effect is poor. In contrast, the method of the present invention has significant advantages in detail processing. For complex texture areas, the present invention can accurately capture and enhance details, clearly restore features such as wall hole texture and building structure lines, so that the processed image details are rich and natural. Whether it is the complex texture of the mine wall or the fine part of the building structure, the method of the present invention can maintain the integrity of the details during the enhancement process, and the visual effect is significantly better than the blueprint separable convolution method, which fully demonstrates its technical superiority in image detail restoration and enhancement.
[0170] In summary, the experimental results above demonstrate the algorithm's applicability in the complex environments of underground mines. Through technological integration and innovation, this invention significantly enhances adaptability to complex vibration scenarios while improving image quality. Compared to the Blueprint Separable Convolution method, it demonstrates more comprehensive technical competitiveness and provides more efficient and reliable technical support for scenarios with stringent vibration compensation requirements, such as mine monitoring and robotic vision.
[0171] Example 3
[0172] An embodiment of the present application provides a computer device, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement a vibration sensing-based mine environment super-resolution image reconstruction method for mine monitoring as provided in the above-mentioned method embodiment.
[0173] Example 4
[0174] The embodiment of the present application also provides a computer-readable storage medium, which can be set in a server to store at least one instruction or at least one program related to a method for implementing a super-resolution image reconstruction method of a mine environment based on vibration sensing for mine monitoring in a method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the method embodiment provided by the above method embodiment. A super-resolution image reconstruction method of a mine environment based on vibration sensing for mine monitoring. Optionally, in this embodiment, the above storage medium can be located in at least one network server among multiple network servers of a computer network. Optionally, in this embodiment, the above storage medium can include but is not limited to: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store program codes.
[0175] With the above-described preferred embodiments of the present invention as a guide, and with reference to the above description, relevant personnel are fully capable of making various changes and modifications without departing from the technical spirit of this invention. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.
Claims
1. A mine environment super-resolution image reconstruction method based on vibration sensing for mine monitoring, characterized in that: include, S1, collecting vibration data and image data, and preprocessing the collected vibration data and image data; S2, inputs the preprocessed vibration data sequence into a two-layer LSTM network, uses LSTM to analyze the vibration data and predict the vibration characteristics at the next moment; S3, the preprocessed image is input into a convolutional neural network constructed using blueprint separable convolution, as well as a residual attention module and a coordinate attention mechanism for multi-scale feature extraction; S4 uses the vibration parameters predicted by LSTM to generate dynamic compensation parameters, which are used to adjust the convolution kernel parameters of the super-resolution network; S5, performing temporal alignment of the vibration data and the image data to generate corresponding compensation parameters, and simultaneously performing spatial alignment and fusion of the two to obtain fusion features; S6, based on fusion features, adopts a residual network structure, realizes image magnification through multi-layer convolution and upsampling modules, and uses residual learning to reduce information loss, performs super-resolution reconstruction on the image, and generates high-resolution images.
2. The mine environment super-resolution image reconstruction method based on vibration sensing for mine monitoring according to claim 1 is characterized in that: In step 2, the specific steps of using LSTM to analyze vibration data include: S2.1, package vibration data into batches; S2.2, batch i The kth group of features X under ik (θ k ,α k ,A k ,ω k ) performs forward propagation; S2.3, based on the threshold value obtained during the forward propagation process, the calculation unit outputs y k and long-term memory; S2.4, output y to the unit k Perform full connection output; S2.5, perform steps S2.2 to S2.4 on all data in the batch, calculate the error, calculate the partial differential of the error with respect to the fully connected weights and unit output, and get the output gradient step output and input gradient δ0; S2.7, gradient descent updates the weight matrix and backpropagates the input gradient.
3. The mine environment super-resolution image reconstruction method based on vibration sensing for mine monitoring according to claim 1 is characterized in that: In the step three, the blueprint separation convolution first performs a weighted combination of the input in the depth direction, and then performs multi-layer depth convolution on multiple channels at the same time. The residual attention module first adds the input and the output of the multi-layer convolution layer to realize residual learning, and then sends the added output features to the coordinate attention module, and then adds the output features after the coordinate attention module to the original input to obtain the final output features of the lightweight residual attention module.
4. The mine environment super-resolution image reconstruction method based on vibration sensing for mine monitoring according to claim 3 is characterized in that: The step three specifically includes the following steps: S3.1, uses pooling kernels of size (H, 1) and (1, W) to perform pooling operations in the horizontal and vertical directions and obtain a set of perceptual feature maps with sizes of C×H×1 and C×1×W respectively; S3.2, the perceptual feature maps of the above two directions are fused and spliced to obtain: X′=δ(f 1×1 ([z h ,z w ])), Where δ(·) represents the h_swish activation function; [z h ,z w ] represents the fusion operation of the output of the cth channel with a height of h and the output of the cth channel with a width of w; f 1×1 Represents 1×1 convolution; represents the intermediate feature map encoding spatial information, and r represents the downsampling step size; S3.3, split F' into two independent tensors X ’h and X ’w , and change the feature map X' through 1×1 convolution layer h and X' w The number of channels is: g h =σ(f 1×1 (X' h )), g w =σ(f 1×1 (X' w )), Where σ(·) represents the sigmoid activation function; g h and g w As the attention weight is used to adjust the module's attention to the input image, the final output of the coordinate attention module can be expressed as: Y CA =X c (i,j)×g h ×g w .
5. The mine environment super-resolution image reconstruction method based on vibration sensing for mine monitoring according to claim 1 is characterized in that: In step 4, the vibration prediction parameters obtained after the hidden state output by the LSTM network is mapped by the fully connected layer include multiple dimensions, which are used to guide the dynamic compensation operations of different areas in the image respectively, thereby achieving an organic combination of local and global compensation.
6. The mine environment super-resolution image reconstruction method based on vibration sensing for mine monitoring according to claim 5 is characterized in that: In step 4, the process of generating the dynamic fuzzy compensation matrix can be expressed as: K comp =g(d t+1 ,m t+1 ), where g(.) function is based on (d t+1 ,m t+1 )Adjust the direction and strength of the convolution kernel to adapt to the current environmental vibration state.
7. The mine environment super-resolution image reconstruction method based on vibration sensing for mine monitoring according to claim 1 is characterized in that: In step 5, the dynamic compensation parameter is K, and its calculation formula is: Among them, θ t is the vibration direction at the current moment, θ t+1 is the vibration direction of the next moment predicted by LSTM, and σ is the fuzzy adjustment parameter; the fusion of vibration data and image data adopts the weighted fusion method, and the fusion weight of vibration features and image features is calculated through the self-attention mechanism.
8. The mine environment super-resolution image reconstruction method based on vibration sensing for mine monitoring according to claim 7 is characterized in that: In step 5, the fusion feature F fused Expressed as: F fused =α vib ·F vib +α img ·F img , where F vib is the vibration compensation feature obtained by mapping the LSTM output through the fully connected layer, and the weight satisfies α vib +α img =1.
9. A computer device, characterized in that: include: processor; a memory for storing executable instructions; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the random walk-based mine image detail enhancement algorithm according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the random walk-based underground mine image detail enhancement algorithm according to any one of claims 1 to 8.
Citation Information
Patent Citations
Mine image super-resolution reconstruction method and system based on multi-scale residual network
CN113592718A
Mine blurred image super-resolution reconstruction method based on MLP improved model
CN115496665A
Image super-resolution reconstruction method for mine fuzzy environment
CN115797181A
Image deblurring network model and rotating body vibration displacement visual measurement method
CN116309123A
Low-resolution vibration image target detection network structure, training method and application
CN116863332A