Method and system for denoising geomagnetic data using dense residual shuffle attention network
By constructing a dense residual shuffling attention network DRSANet, the residual connections and feature interactions are improved, which solves the problem of insufficient denoising accuracy of geomagnetic data, achieves efficient and stable noise suppression, and improves the reliability of earthquake prediction and space weather monitoring.
Patent Information
- Application Number
- CN202511405403.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-04-07
- Estimated Expiration
- 2045-09-29
AI Technical Summary
Existing methods for denoising geomagnetic data suffer from insufficient accuracy, high computational resource consumption, and poor training results for deep learning networks. These methods are unable to effectively suppress human noise interference, thus affecting the reliability of earthquake prediction and space weather monitoring.
A Dense Residual Shuffle Attention Network (DRSANet) is adopted. By improving the residual connection mechanism of the RDSAB module and optimizing the cross-layer feature interaction of the HDRDSAB module, and combining it with the Shuffle Attention module, an end-to-end geomagnetic data denoising framework is constructed, and the network is trained using the noise types in the training sample library.
It improves the denoising accuracy of geomagnetic data, reduces signal loss, enhances the robustness and adaptability of the model, significantly improves the signal-to-noise ratio, and solves the problems of high signal loss and poor adaptability in traditional methods.
Smart Images

Figure CN121210849B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary field of electronic information, artificial intelligence and geophysics, and specifically to a method and system for denoising geomagnetic data using a dense residual shuffling attention network. Background Technology
[0002] Geomagnetic sounding is a deep geophysical method that uses changes in the magnetic fields of the magnetosphere and ionosphere as field sources to acquire response curves from global geomagnetic observatories, in order to explore the mantle transition zone and the conductive structure of the upper lower mantle. Geomagnetic field observation data contains rich information about the Earth's interior and exterior, and is of great value in many fields such as earthquake prediction and space weather monitoring. However, with the advancement of urbanization, the spatial and temporal range of human-caused noise has expanded and interference has intensified, seriously polluting the observation signals of geomagnetic observatories, affecting the reliability of earthquake prediction and space weather monitoring, and limiting the accuracy of deep Earth exploration. Therefore, noise suppression of geomagnetic observatory data is crucial.
[0003] Faced with severe human-caused noise pollution, various methods such as Fourier transform and wavelet transform are commonly used for denoising geomagnetic station data. However, these methods have limitations: frequency domain analysis is prone to losing high-frequency effective signals; Hilbert-Huang transform suffers from mode aliasing; principal component analysis requires orthogonal statistical modes; and independent component decomposition is only suitable for single-channel processing. Due to the complex and varied forms of human-caused noise, existing denoising methods are insufficient to meet practical needs.
[0004] In recent years, deep learning algorithms have developed rapidly, attracting attention in the field of geophysics, and their application in geomagnetic signal processing has gradually increased. Examples include LSTM-based geoelectric signal denoising methods (Wang Kaixiang et al., 2020), a multi-source noise removal method for airborne transient electromagnetic data based on a denoising autoencoder (DAE) (Wue et al., 2019), magnetotelluric power frequency interference suppression based on recurrent neural networks (Xu Taotao et al., 2020), and magnetotelluric signal denoising based on convolutional neural networks and long short-term memory neural networks (CNN-LSTM) (patent application number: 202110320241.4). However, the deep learning networks used in existing methods still suffer from problems such as gradient vanishing or exploding, excessive parameter count, excessive memory consumption, and the need to improve denoising accuracy. Summary of the Invention
[0005] To address the technical problems of insufficient accuracy, high computational resource consumption, and poor training effect of deep learning networks in existing geomagnetic data denoising technologies, this invention provides a geomagnetic data denoising method and system that utilizes a dense residual shuffling attention network, which has high training efficiency, good feature extraction effect, and stable performance.
[0006] To achieve the above-mentioned technical objectives, the technical solution of the present invention is as follows:
[0007] A method for denoising geomagnetic data using a dense residual shuffling attention network includes the following steps:
[0008] The geomagnetic data to be denoised is input into the geomagnetic data denoising model for noise suppression, thereby obtaining the denoised geomagnetic data;
[0009] The geomagnetic data denoising model is a Dense Residual Washing Attention Network (DRSANet) after training. DRSANet includes an input header, a backbone module, and an output layer connected in sequence.
[0010] The input header expands the original input to the required number of layers and then uses it as the input to the backbone module;
[0011] The main module includes an odd number of Block modules, with the Block module in the middle position serving as the boundary between the first and second halves. Starting from the first Block module, the output of each Block module in the first half is processed by a downsampling module and used as the input of the next Block module, until the Block module in the middle position is reached.
[0012] Starting from the middle block, the output of each block in the latter half is processed by an upsampling module, and then added to the output of the first block (which was of the same size before downsampling) to serve as the input for the next block, until the last block. Each block includes a sequentially connected dual-branch convolutional module, an RDSAB module, an HDRDSAB module, and a CA module to extract and fuse the detailed and macroscopic features of the input data.
[0013] The output layer receives the output of the last block module of the backbone module, and the output of the output layer is added to the original input of DRSANet to obtain the final output of DRSANet.
[0014] Furthermore, the input header includes 64 convolutional layers with a kernel size of 1×3 and a stride of 1×1, and a ReLU activation function connected after the convolutional layers. The output layer includes one convolutional layer with a kernel size of 1×3 and a stride of 1×1.
[0015] Furthermore, the dual-branch convolutional module in the Block module includes:
[0016] Two parallel branch paths: one branch path includes a 1×3 convolutional layer and a 1×3 dilated convolutional layer with a dilation rate of 2, with a ReLU activation function connected after each convolutional layer; the other branch path includes a 1×3 dilated convolutional layer with a dilation rate of 3 and a 1×3 dilated convolutional layer with a dilation rate of 4, with a ReLU activation function connected after each convolutional layer.
[0017] The connection processing unit connects the outputs of the first branch path and the second branch path, then feeds them into a 1×3 convolution, and finally performs a residual connection with the original input of the dual-branch convolution module to produce the output.
[0018] Furthermore, the RDSAB module in the Block module includes multiple sequentially connected convolutional layers. Each convolutional layer first performs a convolution operation with a kernel size of 1×3, and then is processed by the ReLU activation function. The output of the ReLU activation function is then concatenated with the original input of the convolutional layer based on the channel dimension to obtain the output of the convolutional layer. The output of the last convolutional layer is fed into the Shuffle Attention module for spatial and channel shuffling attention operations, and then after passing through a 1×1 convolution, it is residually connected with the original input of the RDSAB module before being output.
[0019] The HDRDSAB module in the Block module has the same structure as the RDSAB module, but the expansion rate of the convolutional layers is larger than that of the convolutional layers in the RDSAB module, in order to expand the receptive field of the model.
[0020] Furthermore, the Shuffle Attention module in the Block module includes the following units connected in sequence:
[0021] The channel segmentation unit is used to attenuate the channel dimension of the data and then split the data into two parts;
[0022] The channel attention calculation unit performs average pooling on a portion of the segmented input data, calculates the channel attention parameters, and multiplies them element-wise with the input data of the channel attention calculation unit to output the channel attention feature results.
[0023] The spatial attention computation unit groups and normalizes the input data after segmentation, then calculates the spatial attention parameters, and multiplies them element by element with the input data of the spatial attention computation unit to output the spatial attention feature results.
[0024] The channel shuffling calculation unit concatenates the channel attention feature results with the spatial attention feature results, then reshapes the concatenated result into an array shape identical to the input data, performs channel shuffling again, and finally outputs the shuffling attention result as the output of the Shuffle Attention module.
[0025] Furthermore, the CA module in the Block module first performs average pooling on the input, then processes it through a convolutional layer with a channel decay rate of 16 and a kernel of 1×3 and a ReLU nonlinear layer, then through a convolutional layer with a channel expansion rate of 16 and a kernel of 1×3 and a Sigmoid nonlinear layer to obtain channel attention, and finally multiplies it with the input of the CA module as the output.
[0026] Furthermore, during training, the Dense Residual Washing Attention Network (DRSANet) selects clean, noise-free data from existing observation data and then adds noise to create a training sample library.
[0027] Furthermore, the method for creating the training sample library is as follows: the selected clean data is segmented using an equal-length segmentation method, each data segment is used as a label sample, and then three types of noise with different amplitudes and durations are added to the label samples respectively. The noise types include square wave noise, impulse noise and Gaussian noise, so that each label sample forms multiple different noisy data, thereby forming the training sample library.
[0028] An electronic device, comprising:
[0029] One or more processors;
[0030] Storage device for storing one or more programs.
[0031] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.
[0032] A computer-readable medium storing a computer program that, when executed by a processor, implements the aforementioned method.
[0033] The technical advantages of this invention lie in its construction of a novel Dense Residual Shuffle Attention Network (DRSANet). Specifically, this invention first improves the RDSAB (Residual Dense Attention Block) module, strengthening its residual connection mechanism to enhance feature propagation capabilities and reduce the loss of effective signals in deep networks. Secondly, it optimizes the HDRDSAB (High-Density Residual Dense Attention Block) module, increasing the breadth of cross-layer feature interaction through dilated convolutions to improve the ability to capture subtle geomagnetic signals. Simultaneously, it innovatively integrates the ShuffleAttention module into the network architecture, leveraging its efficient cross-channel feature filtering advantages to accurately distinguish effective signals from noise components in geomagnetic data, thus solving the problem of low signal-to-noise recognition accuracy in complex geomagnetic environments using traditional attention mechanisms.
[0034] The network structure of this invention utilizes the concept of dense residual structure, which enables gradients to be transferred more directly from deeper layers to shallower layers, thereby effectively alleviating the gradient vanishing problem and learning more complex feature representations.
[0035] The backbone network of this invention fully utilizes residual structures, which effectively transmits shallow features to deep networks, avoiding feature loss or attenuation during transmission. Simultaneously, features learned from different layers can be reused by subsequent layers, improving feature utilization and enabling the network to more fully extract useful information from the data.
[0036] The network structure of this invention, through dense connections and a shuffled attention structure, enables the DRSANet network to learn richer and more representative features, making the model more adaptable to changes in input data and thus enhancing its robustness. DRSANet can better maintain performance stability when facing noise and data augmentation.
[0037] Through innovative reorganization and collaborative design of the aforementioned modules, the DRSANet of this invention organically combines the advantages of feature reuse in residual networks, fine-grained feature extraction from high-density connections, and adaptive feature selection from shuffling attention, ultimately forming an end-to-end geomagnetic data denoising framework. This invention eliminates the need for manually setting complex parameters, avoiding the subjective bias caused by manually setting thresholds in traditional methods. It significantly improves the signal-to-noise ratio while effectively preserving the valid geomagnetic signal, solving the pain points of existing methods such as significant signal loss, poor adaptability, and reliance on manual parameter tuning, providing a novel and efficient solution for geomagnetic data denoising.
[0038] The present invention will now be further described in conjunction with the accompanying drawings and embodiments. Attached Figure Description
[0039] Figure 1 This is a diagram of the DRSANet network structure of the present invention.
[0040] Figure 2 This is a structural diagram of the Block module in the main module of the present invention.
[0041] Figure 3 This is a structural diagram of the dual-branch convolutional DBCM module in the Block module of this invention.
[0042] Figure 4 This is a structural diagram of the RDSAB module in the Block module of this invention.
[0043] Figure 5 This is a structural diagram of the HDRDSAB module in the Block module of this invention.
[0044] Figure 6 This is a structural diagram of the Shuffle Attention module in the Block module of this invention.
[0045] Figure 7 This is a structural diagram of the CA module in the Block module of this invention.
[0046] Figure 8 This is a graph showing the change in loss value during the training process of the DRSANet network model of this invention.
[0047] Figure 9 This is a diagram showing the Gaussian noise denoising effect in an embodiment of the present invention.
[0048] Figure 10 This is a diagram showing the effect of pulse noise denoising in an embodiment of the present invention.
[0049] Figure 11 This is a diagram showing the square wave noise denoising effect in an embodiment of the present invention.
[0050] Figure 12 This is a diagram showing the denoising effect of measured JGU data in an embodiment of the present invention.
[0051] Figure 13 The above are geomagnetic transformation function diagrams of measured JGU data in this embodiment of the invention, where (a) is the transformation function and coherence data diagram of high-quality data, (b) is the transformation function and coherence data diagram of noisy data, and (c) is the transformation function and coherence data diagram of denoised data. Detailed Implementation
[0052] The geomagnetic data denoising method using a dense residual shuffle attention network provided in this embodiment addresses the noise suppression problem of geomagnetic data by introducing the concepts of residual shuffle attention and dense residual structures. The constructed DRSANet (Dense Residual Shuffle Attention Network) network structure is as follows: Figure 1 As shown.
[0053] The DRSANet provided in this embodiment includes an input header, a backbone module, and an output layer connected in sequence.
[0054] The input header of this embodiment includes 64 convolutional layers with 1×3 kernels and 1×1 stride, and a ReLU activation function connected after the convolutional layers, which is used to expand the original input of DRSANet to the required number of layers and then use it as the input of the backbone module.
[0055] The main module of this embodiment includes a total of 5 Block modules, with the 3rd Block module as the dividing middle Block module. Starting from the 1st Block module, the outputs of the 1st and 2nd Block modules are processed by a downsampling module before being used as the input of the next Block module.
[0056] Starting from the third block, the outputs of the subsequent third and fourth blocks are processed by an upsampling module, and then added to the outputs of the first half of the blocks (which were of the same size before downsampling) as the input to the next block, until the final fifth block. In this embodiment, the downsampling module in the backbone uses convolution to perform 1 / 2 downsampling, while the upsampling module uses deconvolution plus a non-linear ReLU operation to perform 2x upsampling. Therefore, in this embodiment, the output added to the third block is the output of the second block, and the output added to the fourth block is the output of the first block. Meanwhile, the number of Block modules in the main module can be adjusted according to actual needs, as long as the total number of Block modules is odd. The Block module in the middle position serves as the boundary between the first and second halves. The outputs of the Block modules in the first half are processed by the downsampling module and then used as the input of the next Block module. The outputs of the Block modules in the second half are added to the outputs of the Block modules in the first half that are the same size before being processed by the downsampling module, and then used as the input of the next Block module. The scaling factor of the downsampling module and the scaling factor of the upsampling module should be reciprocals of each other.
[0057] The output layer receives the output of the last block module of the backbone module, and the output of the output layer is added to the original input of DRSANet to obtain the final output of DRSANet.
[0058] Taking the DRSANet network structure in this embodiment as an example, the data processing flow is as follows:
[0059] 1) The original input data is processed by the input header (64 convolutional layers) of the DRSANet network to obtain the output array h.
[0060] 2) h passes through the first block module Block1 to obtain the output array b1.
[0061] 3) Downsample b1 to obtain array b_1.
[0062] 4) b_1 passes through the second block module Block2 to obtain array b2.
[0063] 5) Downsample b2 to obtain array b_2.
[0064] 6) b_2 passes through the third block, Block3, to obtain array b3;
[0065] 7) After b3 is processed by the upsampling module, the array b_3 is obtained.
[0066] 8) Add arrays b_3 and b2 together, input the 4th Block module Block4, and get array b4.
[0067] 9) After b4 is processed by the upsampling module, the array b_4 is obtained.
[0068] 10) Add arrays b_4 and b1 together, input the 5th block block Block5, and get array b_out.
[0069] 11) After b_out is convolved by the output layer, it is added to the original input data of DRSANet to obtain the final denoised data.
[0070] The Block module in this embodiment includes a DBCM dual-branch convolutional module, an RDSAB module, an HDRDSAB module, and a CA module connected in sequence, thereby extracting and fusing the detailed and macroscopic features of the input data. The structure of a single Block module is as follows: Figure 2 As shown. The data processing flow includes:
[0071] 1) The input data is processed by the dual-branch network module (DBCM module) to obtain a 64-layer data structure.
[0072] 2) The RDSAB module is then used to further extract detailed features from the data.
[0073] 3) The macroscopic features of the data are then obtained through long-order modeling using the HDRDSAB module.
[0074] 4) Finally, the CA module performs feature fusion on the data features.
[0075] Specifically, the dual-branch convolutional module (DBCM) structure in the Block module of this embodiment is as follows: Figure 3 As shown, the diagram includes two parallel branch paths and connection processing units following each branch path. One branch path consists of a 1×3 convolutional layer and a 1×3 dilated convolutional layer with a dilation rate of 2, with a ReLU activation function connected after each convolutional layer. The other branch path consists of a 1×3 dilated convolutional layer with a dilation rate of 3 and a 1×3 dilated convolutional layer with a dilation rate of 4, with a ReLU activation function also connected after each convolutional layer.
[0076] The connection processing unit concatenates the outputs of the first and second branch paths, then feeds them into a 1×3 convolution, and finally performs a residual concatenation with the original input of the dual-branch convolution module to produce the output. The two branch paths of the DBCM have different fields of view to obtain feature information at different scales. Finally, the feature information from the two branch outputs is combined and residually concatenated with the original input of the DBCM to maximize the preservation of detailed information in the input data.
[0077] In this embodiment, the specific network parameters and data processing flow in the dual-branch convolutional module (DBCM) of the Block module are as follows:
[0078] 1) Let the input data of DBCM be x. After x is convolved by 64 layers with a kernel size of 1×3 and a stride of 1×1, the array x1 is obtained.
[0079] 2) x1 passes through the ReLU nonlinear layer with activation function to obtain x1_1.
[0080] 3) Data x1_1 is processed by 64 layers of convolution with kernel size of 1×3, stride of 1×1 and expansion rate of 2 to obtain array x1_2.
[0081] 4) x1_2 is passed through the ReLU nonlinear layer to obtain x1_3.
[0082] 5) Data x is convolved through 64 layers with a kernel size of 1×3, a stride of 1×1, and a spread of 3 to obtain array x2.
[0083] 6) x2 passes through the ReLU nonlinear layer to obtain x2_1.
[0084] 7) x2_1 is convolved through 64 layers of convolution with kernel size of 1×3, stride of 1×1, and expansion rate of 4 to obtain array x2_2.
[0085] 8) x2_2 is passed through a ReLU nonlinear layer to obtain x2_3.
[0086] 9) Concatenate x1_3 and x2_3 to obtain array x3.
[0087] 10) x3 is convolved with 64 layers of 1×3 kernels and 1×1 stride to obtain array x4.
[0088] 11) x4 is passed through a ReLU nonlinear layer to obtain x5.
[0089] 12) Finally, x5 is added to the input DBCM data x to obtain the output array Out.
[0090] The structure of the RDSAB module in the Block module of this embodiment is as follows: Figure 4 As shown, the module comprises eight sequentially connected convolutional layers. Each convolutional layer first performs a 1×3 kernel convolution, followed by a ReLU activation function. The output of the ReLU activation function is then concatenated with the original input of the convolutional layer along the channel dimension to obtain the output of the convolutional layer. The output of the last convolutional layer is fed into the Shuffle Attention module for spatial and channel-wise shuffling attention operations, and then passed through a 1×1 convolution before being residually concatenated with the original input of the RDSAB module before output. In this embodiment, the RDSAB module uses eight convolutional layers. The number of convolutional layers can be adjusted according to specific needs. Generally, if the data being processed is relatively simple, the number of convolutional layers can be reduced to speed up computation, while if the data is more complex, the number of convolutional layers can be increased to better extract features.
[0091] The specific network parameters and data processing flow of the RDSAB module in this embodiment are as follows:
[0092] 1) Let the input data of RDSAB be x. Then x first goes through 64 layers of convolution with a kernel of 1×3 and a stride of 1×1 to obtain the array x1.
[0093] 2) x1 passes through the ReLU nonlinear layer to obtain x1_1.
[0094] 3) Concatenate x and x1_1 to obtain the array x1_2.
[0095] 4) Data x1_2 is convolved through 64 layers with a kernel size of 1×3 and a stride of 1×1 to obtain array x2.
[0096] 5) x2 passes through the ReLU nonlinear layer to obtain x2_1.
[0097] 6) Concatenate x1_2 and x2_1 to obtain the array x2_2.
[0098] 7) Data x2_2 is convolved through 64 layers with a kernel size of 1×3 and a stride of 1×1 to obtain array x3.
[0099] 8) x3 passes through the ReLU nonlinear layer to obtain x3_1.
[0100] 9) Concatenate x2_2 and x3_1 to obtain the array x3_2.
[0101] 10) Data x3_2 is convolved through 64 layers with a kernel size of 1×3 and a stride of 1×1 to obtain array x4.
[0102] 11) x4 passes through a ReLU nonlinear layer to obtain x4_1.
[0103] 12) Concatenate x3_2 and x4_1 to obtain the array x4_2.
[0104] 13) Data x4_2 is convolved through 64 layers with a kernel size of 1×3 and a stride of 1×1 to obtain array x5.
[0105] 14) x5 passes through a ReLU nonlinear layer to obtain x5_1.
[0106] 15) Concatenate x4_2 and x5_1 to obtain the array x5_2.
[0107] 16) Data x5_2 is convolved through 64 layers with a kernel size of 1×3 and a stride of 1×1 to obtain array x6.
[0108] 17) x6 passes through a ReLU nonlinear layer to obtain x6_1.
[0109] 18) Concatenate x5_2 and x6_1 to obtain the array x6_2.
[0110] 19) Data x6_2 is convolved through 64 layers with a kernel size of 1×3 and a stride of 1×1 to obtain array x7.
[0111] 20) x7 passes through the ReLU nonlinear layer with activation function to obtain x7_1.
[0112] 21) Concatenate x6_2 and x7_1 to obtain the array x7_2.
[0113] 22) Data x7_2 is convolved through 64 layers with a kernel size of 1×3 and a stride of 1×1 to obtain array x8.
[0114] 23) x8 is passed through a ReLU nonlinear layer to obtain x8_1.
[0115] 24) Concatenate x7_2 and x8_1 to obtain the array x8_2.
[0116] 25) After the x8_2 is shuffled by the attention module, the array x9 is obtained.
[0117] 26) x9 is convolved through 64 layers of convolution with a kernel size of 1×3 and a stride of 1×1 to obtain the array x10.
[0118] 27) x10 is finally added to the input x of RDSAB to obtain the output array out.
[0119] The structure of the HDRDSAB module in the Block module of this embodiment is as follows: Figure 5 As shown, the specific module structure is similar to the RDSAB module, except that the convolutional operations in each layer are replaced with dilation rates of 1, 2, 3, 4, 3, 2, 1, and 1, thereby expanding the receptive field of the model. In contrast, all convolutional layers in the RDSAB module are standard convolutional layers, meaning they all have a dilation rate of 1. Furthermore, the number of convolutional layers and the corresponding dilation rates in the HDRDSAB module can be adjusted according to actual needs.
[0120] The network structure and data processing flow of the HDRDSAB module in this embodiment are as follows:
[0121] 1) Let the input data of HDRDSAB be x. Then x first goes through 64 layers of convolution with kernel size of 1×3, stride of 1×1 and expansion rate of 1 to obtain array x1.
[0122] 2) x1 passes through the ReLU nonlinear layer to obtain x1_1.
[0123] 3) Concatenate x and x1_1 to obtain the array x1_2.
[0124] 4) Data x1_2 is convolved through 64 layers with a kernel size of 1×3, a stride of 1×1, and a spread of 2 to obtain array x2.
[0125] 5) x2 passes through the ReLU nonlinear layer to obtain x2_1.
[0126] 6) Concatenate x1_2 and x2_1 to obtain the array x2_2.
[0127] 7) Data x2_2 is convolved through 64 layers with a kernel size of 1×3, a stride of 1×1, and a spread of 3 to obtain array x3.
[0128] 8) x3 passes through the ReLU nonlinear layer to obtain x3_1.
[0129] 9) Concatenate x2_2 and x3_1 to obtain the array x3_2.
[0130] 10) Data x3_2 is convolved through 64 layers with a kernel size of 1×3, a stride of 1×1, and a spread of 4 to obtain array x4.
[0131] 11) x4 passes through a ReLU nonlinear layer to obtain x4_1.
[0132] 12) Concatenate x3_2 and x4_1 to obtain the array x4_2.
[0133] 13) Data x4_2 is convolved through 64 layers with a kernel size of 1×3, a stride of 1×1, and a spread of 3 to obtain array x5.
[0134] 14) x5 passes through a ReLU nonlinear layer to obtain x5_1.
[0135] 15) Concatenate x4_2 and x5_1 to obtain the array x5_2.
[0136] 16) Data x5_2 is convolved through 64 layers with a kernel size of 1×3, a stride of 1×1, and a spread of 2 to obtain array x6.
[0137] 17) x6 passes through a ReLU nonlinear layer to obtain x6_1.
[0138] 18) Concatenate x5_2 and x6_1 to obtain the array x6_2.
[0139] 19) Data x6_2 is convolved through 64 layers with a kernel size of 1×3, a stride of 1×1, and a spread of 1 to obtain array x7.
[0140] 20) x7 passes through the ReLU nonlinear layer with activation function to obtain x7_1.
[0141] 21) Concatenate x6_2 and x7_1 to obtain the array x7_2.
[0142] 22) Data x7_2 is convolved through 64 layers with a kernel size of 1×3, a stride of 1×1, and a spread of 1 to obtain array x8.
[0143] 23) x8 is passed through a ReLU nonlinear layer to obtain x8_1.
[0144] 24) Concatenate x7_2 and x8_1 to obtain the array x8_2.
[0145] 25) After the x8_2 is shuffled by the attention module, the array x9 is obtained.
[0146] 26) x9 is convolved through 64 layers of convolution with a kernel size of 1×3 and a stride of 1×1 to obtain the array x10.
[0147] 27) Finally, x10 is added to the input HDRDSAB data x to obtain the output array out.
[0148] The structure of the Shuffle Attention module in the Block module of this embodiment is as follows: Figure 6 As shown, it includes the following units connected in sequence:
[0149] A channel splitting unit is used to divide data into two parts along the channel dimension;
[0150] The channel attention calculation unit performs average pooling on a portion of the segmented data, calculates the channel attention parameters, and multiplies them element-wise with the input data of the channel attention calculation unit to output the channel attention feature results.
[0151] The spatial attention calculation unit groups and normalizes the other part of the segmented data, then calculates the spatial attention parameters, and multiplies them element by element with the input data of the spatial attention calculation unit to output the spatial attention feature results.
[0152] The channel shuffling calculation unit concatenates the channel attention feature results with the spatial attention feature results, then reshapes the concatenated result into an array shape identical to the input data, performs channel shuffling again, and finally outputs the shuffling attention result as the output of the Shuffle Attention module.
[0153] In this embodiment, the network structure and data processing flow of the Shuffle Attention module are as follows:
[0154] 1) Use the reshape function to attenuate the number of channels in the input data by a factor of 8. Figure 6 In the example, the input data is 80×576×1440, where 80 represents the batch size of the input data, 576 represents the number of channels in the input data, and 1440 represents the length of the input data. After processing by the reshape function, the number of channels in the input data decreases by a factor of 8 to 72. At the same time, in order to maintain the consistency of the data volume, the batch size is increased by a factor of 8, from 80 to 640. Therefore, the resulting data is 640×72×1440.
[0155] 2) The output of 1) is further divided into two data points based on the channel, i.e. Figure 6The grouping process yielded two datasets of 640×36×1440.
[0156] 3) One of the data x obtained in step 2) is average pooled and then the channel attention coefficient is calculated by the function f(x)=w×x+a. Finally, the channel attention feature data is obtained by element-wise multiplication with the input data with a broadcast mechanism. Here, f(x) represents the channel attention coefficient function, w is the weight parameter of the channel dimension, and a is the bias parameter of the channel dimension.
[0157] 4) Another data y is grouped and normalized, and then the spatial attention coefficient is calculated by the function f(y)=v×y+b. Finally, the spatial attention feature data is obtained by performing element-wise multiplication with the input data with a broadcast mechanism. Here, f(y) represents the spatial attention coefficient function, v is the weight parameter of the spatial dimension, and b is the bias parameter of the spatial dimension.
[0158] 5) Then, concatenate the output data from 3) and 4) based on the channel dimension to obtain a 640×72×1440 data.
[0159] 6) Reshape and flatten the output data from step 5), including using the `.view` function to reshape the output data to the same shape as the input. Then perform Channel Shuffle, which involves adjusting the data shape by grouping the data to obtain the desired result. Figure 6 The data is 80×288×2×1440, where "×2" represents the grouping result. Then the data dimensions are swapped, and finally the data is transformed into the same shape as the input to obtain the output data out.
[0160] The structure of the CA module in the Block module of this embodiment is as follows: Figure 7 As shown, the CA module first performs average pooling on the input, then processes it through a convolutional layer with a channel decay rate of 16 and a kernel of 1×3, and a ReLU nonlinear layer. Next, it passes through a convolutional layer with a channel expansion rate of 16 and a kernel of 1×3, and a Sigmoid nonlinear layer to obtain channel attention. Finally, it multiplies the input of the CA module with the original input to obtain the output.
[0161] The network structure and data processing flow of the CA module are as follows:
[0162] 1) x is average pooled to generate an array x1 of length 1;
[0163] 2) x1 is convolved by 4 layers of convolution with kernel size of 1×3 and stride of 1×1 to obtain array x2.
[0164] 3) x2 passes through the ReLU nonlinear layer to obtain x3.
[0165] 4) x3 is convolved through 64 layers of convolution with a kernel size of 1×3 and a stride of 1×1 to obtain the array x4.
[0166] 5) x4 is passed through the Sigmoid nonlinear layer to obtain x8.
[0167] 6) Perform element-wise broadcast multiplication of x8 and x to obtain the output Out.
[0168] It should be noted that the number and size of the Block module, DBCM module, RDSAB module, HDRDSAB module, and CA module mentioned above in this embodiment are set based on the model training effect. Therefore, the above example is only for illustration. Without departing from the concept of this invention, the number of modules, the number of residual blocks, and the size and number of convolution kernels can be adjusted.
[0169] In this embodiment, the Dense Residual Washing Attention Network (DRSANet) selects clean, noise-free data from existing observation data during training, and then adds noise to create a training sample library. The method for creating the training sample library is as follows: the selected clean data is segmented using equal-length segments, with each segment serving as a label sample. Then, three types of noise with different amplitudes and durations are added to each label sample. These noise types include square wave noise, impulse noise, and Gaussian noise, resulting in multiple noisy data sets for each label sample. The specific number of noisy data sets can be adjusted according to requirements, ranging from hundreds to tens of thousands, thus forming the training sample library.
[0170] In this embodiment, the Adam optimizer is used during model training, with a batch size of 80 and an initial learning rate of 0.0001. The ReduceLROnPlateau decay strategy is used to adjust the learning rate: first, the changes in validation loss are monitored during training. If the validation loss shows little change over 10 epochs, the strategy adjusts the learning rate to 95% of its current value. If the validation loss continues to decrease, the learning rate remains unchanged until the entire training process is complete. A total of 100 epochs are trained. In other feasible embodiments, no specific limitations are imposed, and other optimizers can be selected.
[0171] like Figure 8This chart shows the changes in accuracy and loss during the training process of the DRSANet network model. The red solid line represents the changes in accuracy and loss on the validation set during model training, while the blue solid line represents the changes in accuracy and loss on the training set after the model has been trained on the validation set. The curves show that as the number of training iterations increases, the model's accuracy gradually increases, while the model's loss gradually decreases and eventually stabilizes. This indicates a trend of convergence in the data features learned by the model, suggesting that the model's adaptability to the data is increasing, and the model's error is decreasing.
[0172] like Figure 9-11 This is a graph showing the denoising effect of the DRSANet network model. It's evident that regardless of whether it's Gaussian white noise, impulse noise, or square wave noise, the denoised data (yellow curve) closely matches the overall trend of the original data (blue curve). It accurately reflects the changes in the original data at its peaks and troughs. Compared to the red curve after adding noise, the denoised data effectively suppresses noise interference, making the data more reflective of the true situation. Therefore, denoising is effective in preserving the characteristics of the original data and suppressing noise, which is beneficial for more accurate subsequent analysis and modeling.
[0173] Figure 12 This image shows the denoising effect of data from the JGU station in Jinggu area, spanning 50 days from January 8th to February 26th, 2018. The data type is segmented data, with three components: X, Y, and Z. As can be seen from the image, the DRSANet network can completely remove large impulse noise from all three components, making the time-domain curve smoother and achieving a high degree of fit with the effective components of the original data.
[0174] Figure 13 This diagram illustrates the calculation of the geomagnetic transfer function (GJT) using measured JGU data before and after noise reduction, with a comparison made using GJT calculated from high-quality data. Figure 13 As can be seen from (a), the period of the geomagnetic transformation function in high-quality data is 3×10. 2 s-4×10 4 The values at each frequency point between s are continuous, the curve is smooth, and the error bars are small, demonstrating the good characteristics of high-quality data. Furthermore, the real and imaginary parts of the transfer function for data containing large pulses are within 3×10... 2 s-1×10 3 The error bars of the transformation function are longer during the s-week period, and the shape of its curve is also significantly different compared to high-quality data, as shown in Figure (b). After denoising using the DRSANet network, the transformation function is shown in Figure (c). The period of the geomagnetic transformation function calculated from the denoised data is 3 × 10⁻⁶. 2 s-1×10 3The continuity of the values at each frequency point between s is greatly improved. In summary, this invention can effectively remove noise in geomagnetic data and significantly improve data quality.
[0175] According to embodiments of the present invention, the present invention also provides an electronic device and a computer-readable medium.
[0176] Electronic devices include:
[0177] One or more processors;
[0178] Storage device for storing one or more programs.
[0179] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.
[0180] In practical use, users can interact with servers, which are also electronic devices, via a network to receive or send messages. Terminal devices are generally various electronic devices equipped with a display and used through a human-computer interface, including but not limited to smartphones, tablets, laptops, and desktop computers. Various specific application software can be installed on these terminal devices as needed, including but not limited to web browsers, instant messaging software, social media platforms, and shopping apps.
[0181] Similarly, the computer-readable medium of the present invention stores a computer program thereon, which, when executed by a processor, implements the geomagnetic data denoising method of the embodiments of the present invention.
[0182] It should be emphasized that the examples described in this invention are illustrative rather than limiting. Therefore, this invention is not limited to the examples described in the specific embodiments. Any other embodiments derived by those skilled in the art based on the technical solutions of this invention, without departing from the spirit and scope of this invention, whether modifications or substitutions, are also within the protection scope of this invention.
Claims
1. A method for denoising geomagnetic data using a dense residual shuffling attention network, characterized in that, Includes the following steps: The geomagnetic data to be denoised is input into the geomagnetic data denoising model for noise suppression, thereby obtaining the denoised geomagnetic data; The geomagnetic data denoising model is a Dense Residual Washing Attention Network (DRSANet) after training. DRSANet includes an input header, a backbone module, and an output layer connected in sequence. The input header expands the original input to the required number of layers and then uses it as the input to the backbone module; The main module includes an odd number of Block modules, with the Block module in the middle position serving as the boundary between the first and second halves. Starting from the first Block module, the output of each Block module in the first half is processed by a downsampling module and used as the input of the next Block module, until the Block module in the middle position is reached. Starting from the middle block, the output of each block in the latter half is processed by an upsampling module, and then added to the output of the first block (which was of the same size before downsampling) to serve as the input for the next block, until the last block. Each block includes a sequentially connected dual-branch convolutional module, an RDSAB module, an HDRDSAB module, and a CA module to extract and fuse the detailed and macroscopic features of the input data. The output layer receives the output of the last block module of the backbone module, and the output of the output layer is added to the original input of DRSANet to obtain the final output of DRSANet. The RDSAB module in the Block module includes multiple sequentially connected convolutional layers. Each convolutional layer first performs a convolution operation with a kernel size of 1×3, and then is processed by the ReLU activation function. The output of the ReLU activation function is then concatenated with the original input of the convolutional layer based on the channel dimension to obtain the output of the convolutional layer. The output of the last convolutional layer is fed into the Shuffle Attention module for spatial and channel shuffling attention operations, and then after passing through a 1×1 convolution, it is residually concatenated with the original input of the RDSAB module before being output. The HDRDSAB module in the Block module has the same structure as the RDSAB module, but the expansion rate of the convolutional layer is larger than that of the convolutional layer in the RDSAB module, in order to expand the receptive field of the model. The Shuffle Attention module in the Block module includes the following units connected in sequence: The channel segmentation unit is used to attenuate the channel dimension of the data and then split the data into two parts; The channel attention calculation unit performs average pooling on a portion of the segmented input data, calculates the channel attention parameters, and multiplies them element-wise with the input data of the channel attention calculation unit to output the channel attention feature results. The spatial attention computation unit groups and normalizes the input data after segmentation, then calculates the spatial attention parameters, and multiplies them element by element with the input data of the spatial attention computation unit to output the spatial attention feature results. The channel shuffling calculation unit concatenates the channel attention feature results with the spatial attention feature results, then reshapes the concatenated result into an array shape identical to the input data, performs channel shuffling again, and finally outputs the shuffling attention result as the output of the Shuffle Attention module.
2. The method for denoising geomagnetic data using a dense residual shuffling attention network according to claim 1, characterized in that, The input header includes 64 convolutional layers with a kernel size of 1×3 and a stride of 1×1, and ReLU activation functions connected after the convolutional layers. The output layer includes one convolutional layer with a kernel size of 1×3 and a stride of 1×1.
3. The method for denoising geomagnetic data using a dense residual shuffling attention network according to claim 1, characterized in that, The dual-branch convolutional module in the Block module includes: Two parallel branch paths: one branch path includes a 1×3 convolutional layer and a 1×3 dilated convolutional layer with a dilation rate of 2, with a ReLU activation function connected after each convolutional layer; the other branch path includes a 1×3 dilated convolutional layer with a dilation rate of 3 and a 1×3 dilated convolutional layer with a dilation rate of 4, with a ReLU activation function connected after each convolutional layer. The connection processing unit connects the outputs of the two parallel branch paths, then feeds them into a 1×3 convolution, and finally performs a residual connection with the original input of the dual-branch convolution module to produce the output.
4. The method for denoising geomagnetic data using a dense residual shuffling attention network according to claim 1, characterized in that, The CA module in the Block module first performs average pooling on the input, then processes it through a convolutional layer with a channel decay rate of 16 and a kernel of 1×3 and a ReLU nonlinear layer, then through a convolutional layer with a channel expansion rate of 16 and a kernel of 1×3 and a Sigmoid nonlinear layer to obtain channel attention, and finally multiplies it with the input of the CA module to obtain the output.
5. The method for denoising geomagnetic data using a dense residual shuffling attention network according to claim 1, characterized in that, During training, the Dense Residual Washing Attention Network (DRSANet) selects clean, noise-free data from existing observation data and then adds noise to create a training sample library.
6. The method for denoising geomagnetic data using a dense residual shuffling attention network according to claim 5, characterized in that, The training sample library is constructed as follows: the selected clean data is segmented using an equal-length segmentation method, and each data segment is used as a label sample. Then, three types of noise with different amplitudes and durations are added to the label samples. The noise types include square wave noise, impulse noise, and Gaussian noise, so that each label sample forms multiple different noisy data, thereby forming the training sample library.
7. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-6.
8. A computer-readable medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.
Citation Information
Patent Citations
CNN-LSTM-based magnetotelluric signal noise suppression method and system
CN113158553A
Geomagnetic signal noise suppression method and system
CN116953808A
Geomagnetic data denoising method and system based on reference trace data constraint and deep learning
CN117272138A