Signal denoising method and device, equipment and storage medium
Through the conditional generation adversarial network, the synthetic signal is generated and combined with the convolutional neural network and the proxy attention mechanism, the problem of insufficient clean data in accelerometer signal denoising is solved, and the efficient signal denoising effect is achieved, maintaining the time continuity and characteristic information of the signal.
Patent Information
- Application Number
- CN202510661056.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-05
AI Technical Summary
The prior art is difficult to effectively remove noise in the accelerometer signal without clean data, especially due to filter design difficulties and unsatisfactory denoising effects due to noise and signal frequency similarity, and existing methods may destroy the time continuity and dependencies of the signal.
The conditional generation adversarial network is used to generate synthetic signals as training samples, and combined with the convolutional neural network and the proxy attention mechanism, a signal denoising model is constructed, and the synthetic signals with small differences from the original signal are denoised.
In the absence of clean data, the denoising effect of the accelerometer signal is significantly improved, maintaining the time continuity and characteristic information of the signal, which is better than traditional methods.
Smart Images

Figure CN120429644A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of artificial intelligence and signal processing, and in particular to a signal denoising method, apparatus, device and storage medium. Background Art
[0002] Inertial sensors are one of the key components commonly used in navigation systems. Micro-electromechanical system inertial sensors (MEMS) are popular due to their small size, high cost-effectiveness, low power consumption and low price. They are widely used in various applications, platforms and environments, such as in the air, on the ground and at sea. Among them, MEMS accelerometers are a commonly used sensor type in MEMS inertial sensors. If the motion information (accelerometer signal) recorded by the MEMS accelerometer is combined with other information such as speed and displacement, important characteristics such as the performance and attitude of the carrier can be estimated and analyzed, thereby achieving high-precision navigation. However, due to the complexity of the real environment, the above-mentioned accelerometer signals often contain noise. If the noise measurement is propagated into the navigation solution, it will cause the solution to drift.
[0003] Traditional signal denoising methods generally use noise modeling, subtraction, data smoothing, etc. to suppress noise and improve the signal-to-noise ratio of the signal. However, noisy signals often show adverse frequency fluctuations similar to the underlying scientific signals, which makes it difficult to achieve the best noise filter, resulting in unsatisfactory denoising effects. With the development of neural networks, they have achieved remarkable results in application fields such as image denoising and speech denoising. However, most denoising neural network training requires clean signals as label data, but due to the complexity of the environment and defects in the equipment, it is very difficult to collect clean signals. Therefore, in actual scenarios, such methods are difficult to apply and play a role. Based on the above problems, someone proposed an image denoising method that can be used without using clean images. Training a denoising neural network with labeled samples requires paired noisy data. This paired noisy data is obtained by downsampling randomly selected data points from the original image data to generate a pair of noisy sub-images. However, when processing accelerometer signals, directly randomly sampling the sub-signals from the original signal results in a sequence length that is much shorter than the original signal, resulting in loss of key information. Furthermore, because accelerometer signals are highly temporally correlated, the state at a given moment is closely linked to its preceding and following states. Random downsampling disrupts this temporal continuity and dependency, impairing the signal's preceding and following information, making it impossible to accurately reflect the true characteristics of the original signal. Therefore, this method is not suitable for accelerometer signal processing. Summary of the Invention
[0004] This application provides a signal denoising method, apparatus, device and storage medium, which can construct a training sample in the absence of clean data, and combine convolution with proxy attention mechanism to improve the denoising effect of accelerometer signals.
[0005] In a first aspect, the present application provides a signal denoising method, comprising:
[0006] Acquire a training sample, where the training sample includes an original signal and a synthesized signal generated based on the original signal;
[0007] Taking the original signal as input and the synthesized signal as output, training a convolutional neural network that introduces a proxy attention module to obtain a signal denoising model;
[0008] The noisy signal is input into the signal denoising model to obtain a denoised clean signal.
[0009] In one or more possible embodiments, the synthetic signal that meets the preset conditions is generated based on a conditional generative adversarial network, wherein the conditional generative adversarial network includes a generator and a discriminator, the generator is used to generate a synthetic signal based on the original signal and a random vector, and the discriminator uses a soft label method to classify the original signal and the synthetic signal.
[0010] In one or more possible embodiments, the conditional generative adversarial network is trained in the following manner:
[0011] The mean square error loss function is used to determine the original signal loss and the synthetic signal loss;
[0012] Using the synthetic signal loss as the generator loss, and determining the discriminator loss according to the original signal loss and the synthetic signal loss;
[0013] When it is determined that the generator loss and the discriminator loss meet the preset condition, a trained conditional generative adversarial network is obtained.
[0014] In one or more possible embodiments, the preset condition is that both the generator loss and the discriminator loss approach a preset threshold.
[0015] In one or more possible embodiments, the convolutional neural network introducing the proxy attention module further includes an encoder and a decoder.
[0016] The encoder is used to compress the original signal layer by layer through multi-layer convolution operations to obtain feature information of different time scales; the decoder is used to restore the feature information of different time scales through upsampling and output the synthesized signal.
[0017] In one or more possible embodiments, the training of the convolutional neural network introducing the proxy attention module includes:
[0018] Inputting the original signal into the encoder to obtain feature information of different time scales;
[0019] Based on a preset proxy token, the feature information of the different time scales is processed to obtain global context feature information;
[0020] The global context feature information is restored according to an upsampling method, and the synthesized signal is output.
[0021] In one or more possible embodiments, processing the feature information of different time scales based on a preset proxy token to obtain global context feature information includes:
[0022] Determine a key matrix K and a value matrix V according to the characteristic information of the different time scales;
[0023] Taking the preset proxy token as a query Q, calculating the similarity between the query Q and the key matrix K, and normalizing the similarity according to a Softmax function;
[0024] According to the normalized similarity and the value matrix V, global context feature information is determined.
[0025] In a second aspect, the present application provides a signal denoising device, the device comprising:
[0026] A sample acquisition module is used to acquire a training sample, wherein the training sample includes an original signal and a synthetic signal generated according to the original signal;
[0027] a model training module, configured to take the original signal as input and the synthesized signal as output, and train a convolutional neural network that introduces a proxy attention module to obtain a signal denoising model;
[0028] The signal denoising module is used to input the noisy signal into the signal denoising model to obtain a denoised clean signal.
[0029] In a third aspect, the present application provides an electronic device, comprising:
[0030] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any one of the methods in the first aspect.
[0031] In a fourth aspect, the present application provides a computer storage medium storing a computer program, wherein the computer program is used to enable a computer to execute any one of the methods in the first aspect.
[0032] According to a signal denoising method, apparatus, device and storage medium provided in this application, a training sample can be constructed in the absence of clean data, and the denoising effect of the accelerometer signal can be improved by combining convolution and proxy attention mechanisms. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification are used to explain the principles of the present application, and do not constitute an improper limitation on the present application.
[0034] Figure 1 A flowchart of a signal denoising method provided according to an embodiment of the present application;
[0035] Figure 2 A schematic diagram of a conditional generative adversarial network provided according to an embodiment of the present application;
[0036] Figure 3 A schematic diagram of a convolutional neural network provided according to an embodiment of the present application;
[0037] Figure 4 A fluctuation comparison diagram of a synthetic signal and an original signal provided according to an embodiment of the present application;
[0038] Figure 5 A visualization diagram comparing a synthesized signal with an original signal using a PCA method according to an embodiment of the present application is provided;
[0039] Figure 6 Schematic diagram of a signal denoising device provided according to an embodiment of the present application;
[0040] Figure 7 A schematic diagram of an electronic device provided according to an embodiment of the present application;
[0041] Figure 8 A schematic diagram of a computer storage medium provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0042] To make the objectives, technical solutions, and advantages of the present invention more apparent, the present invention will be further described in detail below with reference to the accompanying drawings. It is apparent that the embodiments described are only some, not all, of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are intended to fall within the scope of protection of the present invention.
[0043] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.
[0044] Moreover, in the description of the embodiments of the present application, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; “and / or” in the text is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, “multiple” refers to two or more than two.
[0045] It should be noted that the embodiments of the present disclosure may include multiple steps. For the convenience of description, these steps are numbered, but these numbers do not limit the execution time slots or execution order between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not limit this.
[0046] For ease of understanding, the terms involved in the embodiments of the present invention are explained below:
[0047] Supervised learning: refers to the method of using labeled training data to train models for learning and prediction. In the supervised learning process, the model receives a labeled training data set and learns from this data to identify the best association between input features and output labels; in this way, the model can accurately predict or classify unseen samples.
[0048] Unsupervised learning: This method does not rely on labeled datasets to train models. In unsupervised learning, algorithms attempt to extract information and identify hidden patterns from raw data. It is primarily used for tasks such as data clustering, discovering association rules, and performing feature dimensionality reduction. It aims to uncover the inherent structure of data for effective data organization, anomaly detection, or information extraction.
[0049] Generative Adversarial Networks: Generative Adversarial Networks (GANs) are a new type of generative model based on deep learning and adversarial learning, proposed by Ian Goodfellow and others. GANs mainly consist of two networks: a discriminator and a generator. Its basic idea is to let the generator and the discriminator compete with each other to find the optimal balance. The generator is responsible for generating new data samples. It receives random noise as input and converts it into data samples with a distribution similar to real data through a neural network. The goal of the generator is to create increasingly realistic data to the point where the discriminator cannot distinguish. The discriminator is responsible for distinguishing whether the input data is real or generated by the generator. It improves its judgment ability by learning the characteristics of real data. The goal of the discriminator is to correctly identify real data and fake data generated by the generator.
[0050] U-Net: A fully convolutional neural network, it can be considered to consist of three parts: an encoder, skip connections, and a decoder. The encoder primarily extracts data features and downsamples the data. It is generally composed of convolutional layers and pooling layers. Each layer gradually reduces the spatial size of the data while increasing the depth of its feature map. Skip connections connect the feature maps of corresponding layers in the encoder and decoder. The feature maps output by each encoder layer are saved and then concatenated with the upsampled feature maps during the upsampling stage. The decoder is responsible for converting the features extracted by the encoder to the size of the original data and upsampling the features. This is generally achieved through transposed convolutional layers or upsampling layers. During each upsampling process, the decoder concatenates the feature maps output by the skip connection stage. The final output of the decoder is a feature map of the same size as the input data.
[0051] Before describing the embodiments of the present application, a theoretical proof is provided that a similar noise signal can be generated based on only a single noise signal, and a neural network can learn to denoise by using noise-to-noise difference mapping, as follows:
[0052] Existing unsupervised learning methods require paired noisy data from the same research object in the same scenario as training data. However, in real-world environments, ensuring that the environmental conditions for data collection are exactly the same is extremely difficult to achieve. This is because slight changes in factors such as device angle and ambient temperature may lead to differences between the data, and these factors are often uncontrollable. Therefore, if an algorithm can be used to construct a noisy signal that meets the conditions based on only a single noisy signal to form a noisy signal pair, the conditions for using unsupervised learning methods can be greatly reduced.
[0053] For example, consider two very close and independent noisy signals Y1 and Y2 that are the same research object. The clean signal corresponding to the noisy signal Y1 is X, and the clean signal corresponding to the noisy signal Y2 is X+W, where W approaches 0 to represent the small error between the two clean signals. Therefore, the following two formulas can be obtained:
[0054] E(Y1)=X; E(Y2)=X+W
[0055] Among them, E represents the expectation operator, and f is used for the neural network algorithm that needs to be trained. θ (g) is used to represent that θ is the parameter of the neural network algorithm. In the original supervised learning training method, the expected minimization formula (1) is:
[0056] E‖f θ (Y1)-X‖ 2 (1)
[0057] Substituting the noisy signal Y2 into the above minimization formula, we get formula (2):
[0058] E‖f θ (Y1)-X‖ 2
[0059] =E‖(f θ (Y1)-Y2)+(Y2-X)‖ 2
[0060] =E‖f θ (Y1)-Y2‖ 2 +E‖Y2-X‖ 2 +2E[(f θ (Y1)-Y2) T (Y2-X)] (2)
[0061] Similarly, bring the clean signal x into 2E[(f θ (Y1)-Y2) T (Y2-X)] to obtain formula (3):
[0062] 2E[(f θ (Y1)-Y2) T (Y2-X)]=2E[(f θ (Y1)-X) T (Y2-X)]-2E[‖Y2-X‖ 2 ] (3)
[0064] Substituting the above formula (3) into formula (2) yields formula (4):
[0065] E‖f θ (Y1)-X‖2
[0066] =E‖f θ (Y1)-Y2‖ 2 +E‖Y2-X‖ 2 +2E[(f θ (Y1)-Y2) T (Y2-X)]
[0067] =E‖f θ (Y1)-Y2‖ 2 +E‖Y2-X‖ 2 +2E[(f θ (Y1)-X) T (Y2-X)]-2E[‖Y2-X‖ 2 ]
[0068] =E‖f θ (Y1)-Y2‖ 2 -σ 2 +2E[(f θ (Y1)-X) T (Y2-X)] (4)
[0069] Among them, σ 2 represents the mean square error of Y2 and X, and given the condition that the noisy signals Y1 and Y2 are independent of each other, formula (4) is adjusted to obtain formula (5)
[0070] E‖f θ (Y1)-Y2‖ 2 -σ 2 +2E[(f θ (Y1)-X) T (Y2-X)]
[0071] =E‖f θ (Y1)-Y2‖ 2 -σ 2 +2WE[f θ (Y1)-X](5)
[0072] According to the relationship between two very close and independent noisy signals Y1 and Y2, and W approaches 0, we can get 2WE[f θ (Y1)-X] approaches 0, and simplifying formula (5) we can get:
[0073] E‖f θ (Y1)-X‖ 2 =E‖f θ (Y1)-Y2‖ 2 -σ 2
[0074] Due to σ2 is a constant, so E‖f θ (Y1)-X‖ 2 Equivalent to E‖f θ (Y1)-Y2‖ 2 ;
[0075] Finally, we can conclude that E‖f θ (Y1)-X‖ 2 Equivalent to E‖f θ (Y1)-Y2‖ 2 , without the need for signal frequency characteristics and noise characteristics, as long as we have two noisy signal pairs with small enough differences based on an original noisy signal, we can use E‖f θ (Y1)-Y2‖ 2 To train the denoising neural network, it is proved that a similar noise signal can be generated based on only a single noise signal, and the neural network can learn to denoise by mapping the difference between noise and noise.
[0076] Example 1
[0077] This application provides a signal denoising method, such as Figure 1 As shown, including:
[0078] Step 101: obtaining a training sample, wherein the training sample includes an original signal and a synthetic signal generated based on the original signal;
[0079] In one or more possible embodiments, the signal in this application refers to an accelerometer signal, the original signal in the training sample is obtained based on a MEMS accelerometer, and the synthetic signal is generated based on a conditional generative adversarial network (cGAN); the original noisy signal is input as conditional information into the generator of the cGAN. After receiving this conditional information and random noise, the generator learns how to generate a new noisy signal under the supervision of a discriminator. This newly generated signal, that is, the synthetic signal, together with the original signal, can form a pair of noisy signals with small differences. Ultimately, the pair of noisy signals with small differences is used as training samples for a denoising neural network model; the conditional generative adversarial network includes a generator and a discriminator. The generator is used to generate a synthetic signal based on the original signal and a random vector. The discriminator classifies the original signal and the synthetic signal using a soft label method.
[0080] In one or more possible embodiments, Figure 2The following is a schematic diagram of cGAN. Both the generator and the discriminator are built based on the Transformer encoder architecture. The encoder consists of two composite blocks. The first block consists of an attention module, and the second block consists of a feedforward MLP with a GELU activation function. A normalization layer is applied before the two blocks, and a Dropout layer is added after each block. Both blocks use residual connections. Specifically, an original signal (noisy signal Y) and a random vector (random noise Z) are input into the generator. The original signal is the collected noisy signal. The random vector has the same dimension as the original signal, and N uniformly distributed random values are in (0 ,1) range, N represents the potential dimension of the synthetic signal, which is an adjustable hyperparameter. The above random vector is then mapped to a sequence with the same actual signal length and M embedding dimension, which is also an adjustable hyperparameter; the above sequence is divided into multiple patch blocks (patch), and a position encoding value is added to each patch block. Each patch block is then input into the transformer encoder block, and the output of the encoder block is passed through the Conv2D layer (convolutional layer) to reduce the dimension of the synthetic data. The Conv2D layer is set to the kernel size (1,1), which does not change the width and height of the synthetic data. The filter size is set to the same dimension size as the actual data sequence.
[0081] The above discriminator is similar to the ViT model, which is a model for binary classification used to distinguish whether the input sequence is an actual signal or a generated signal. In the ViT model, an image is evenly divided into multiple small blocks of the same width and height, each of which is called a patch; each sequence input to the above discriminator is regarded as an image with a height of 1, where the time step of the input sequence represents the width of the image. Therefore, if you want to add position encoding to the time series input, you only need to evenly divide the width into multiple segments, keep the height of each segment unchanged, and regard the sequence signal output by the sensor as an image with a height of 1. The length of the sequence is the width D of the image; the above sequence can have a single channel or multiple channels, which can be regarded as the number of channels of the image (RGB). Select a patch of size N and divide the sequence into D / N subsequences. Add a soft position encoding value at the end of each subsequence to learn the position value during model training.
[0082] In one or more possible embodiments, the conditional generative adversarial network cGAN also needs to be trained, and the parameters of the generator and discriminator are updated. The mean square error loss function is used to determine the original signal loss and the synthetic signal loss; the synthetic signal loss is used as the generator loss, and the discriminator loss is determined based on the original signal loss and the synthetic signal loss; when it is determined that the generator loss and the discriminator loss meet the above preset conditions, a trained conditional generative adversarial network is obtained; first, the input of the generator consists of two vectors: Y and Z, Y represents a noisy signal (original signal), and Z represents a random vector. The generator generates a synthetic signal based on Y and Z, expressed as G(Y, Z); the discriminator is used to determine whether the received signal is a real signal or a synthetic signal synthesized by the generator. For example, D(g) is used to represent the classification output of the discriminator, where g can be a real signal or a synthetic signal. During the training process, Real_l is set to represent the label of the real signal, with a value of 1; Fake_l represents the label of the synthetic signal, with a value of 0. In order to enhance the stability of the cGAN model and improve the training quality, the label value can also be set in the following way; for example, the above original signal and the above synthetic signal can be classified using a soft label method, Real_l is a floating point number close to 1, and Fake_l is a floating point number close to 0. This adjustment helps to reduce the gradient vanishing problem and model vibration; the values of Real_l and Fake_l can also be flipped in a certain proportion to reduce the mode collapse caused by the excessive competition between the generator and the discriminator. In this application, the loss function of the generator and the loss function of the discriminator are both calculated using the mean square error loss. The use of the mean square error loss MSELoss can effectively reduce the gradient vanishing problem and improve the stability and convergence speed of the cGAN model. The discriminator loss D_Loss is determined according to the real signal loss D_RealLoss and the synthetic signal loss D_FakeLoss, and the generator loss G_Loss is determined by the loss of the synthetic signal, specifically using the following formula:
[0083] D_FakeLoss=MSELoss(D(G(Y,Z)),Fake_l)
[0084] D_RakeLoss=MSELoss(D(Y),Real_l)
[0085] D_Loss=D_FakeLoss+D_RakeLoss
[0086] G_Loss = D_FakeLoss
[0087] Among them, D_Loss represents the discriminator loss, G_Loss represents the generator loss, D_FakeLoss represents the synthetic signal loss, and D_RealLoss represents the real signal loss. The above real signal loss D_RealLoss and synthetic signal loss D_FakeLoss are both determined by the MSELoss function. When it is determined that the generator loss G_Loss and the discriminator loss D_Loss are both close to 0, it means that the game between the generator and the discriminator has reached equilibrium, and the discriminator has mistaken the synthetic signal generated by the generator for the real signal. When it is determined that the generator loss and the discriminator loss meet the above preset conditions, that is, when it is determined that the generator loss G_Loss and the discriminator loss D_Loss are both close to 0, the training of the conditional generative adversarial network is determined to be complete. Then, the trained conditional generative adversarial network can be used to generate a synthetic signal. The original signal and the synthetic signal generated based on the original signal are used as a pair of training samples to train the convolutional neural network introduced with the proxy attention module.
[0088] Step 102: Using the original signal as input and the synthesized signal as output, a convolutional neural network with a proxy attention module is trained to obtain a signal denoising model.
[0089] In one or more possible embodiments, the above original signal is used as input, and the above original signal is expressed as Y=X+N, where X represents a clean signal (that is, a signal to be denoised), and N represents noise. The convolutional neural network introduced with the agent attention module in this application can learn to extract the noise signal (Y1, Y2, ...Y n ) to the clean denoised signal (X1, X2, ... X n ) mapping, unlike traditional denoising methods, does not require the establishment of an error analysis model. Furthermore, since this application has previously demonstrated that the expected minimization result of two independent noisy signals with very small differences is equivalent to the expected minimization result of the noisy signal and the corresponding clean signal, the convolutional neural network with the proxy attention module can be trained using the original signal as input and the above-mentioned synthetic signal as output. In other words, the synthetic signal generated by the conditional generative adversarial network is used as the clean signal for training.
[0090] In one or more possible embodiments, Figure 3As shown, this is a module diagram of a convolutional neural network in which a proxy attention module is introduced in this application, including an encoder and a decoder; the encoder (downsampling layer) is used to compress the original signal (noisy signal) layer by layer through multi-layer convolution operations to obtain feature information of different time scales; the decoder (upsampling layer) is used to restore the feature information of the different time scales through upsampling, and output the synthetic signal (denoised signal, clean signal); the upper layer network and the lower layer network in the encoder and decoder are connected through multi-layer convolution operations, which can fully utilize the feature information of each layer for better fitting signal features. A proxy attention module is also included between the encoder and the decoder for obtaining global context feature information.
[0091] In one or more possible embodiments, the proxy attention module in the present application integrates the advantages of Softmax attention and linear attention, which not only preserves the efficiency of linear complexity but also maintains high expressiveness; specifically, the convolutional neural network introducing the proxy attention module is trained, including: inputting the above-mentioned original signal into the above-mentioned encoder to obtain feature information of different time scales; based on the preset proxy token, processing the above-mentioned feature information of different time scales to obtain global context feature information; fusing the above-mentioned feature information of different time scales with the above-mentioned global context feature information to obtain feature fusion global information; restoring the above-mentioned feature fusion global information according to the upsampling method, and outputting the above-mentioned synthetic signal; wherein, the role of the above-mentioned proxy attention module is to obtain corresponding global context feature information based on the feature information of different time scales output by the above-mentioned encoder, and fusing the above-mentioned feature information of different time scales with the above-mentioned global context feature information to obtain feature fusion global information; finally, the decoder restores the above-mentioned feature fusion global information by upsampling and outputs the above-mentioned synthetic signal.
[0092] In one or more possible embodiments, the above-mentioned feature information of different time scales is processed based on the preset proxy token to obtain global context feature information, including: determining the key matrix K and the value matrix V according to the feature information of different time scales; taking the above-mentioned preset proxy token as the query Q, calculating the similarity between the query Q and the above-mentioned key matrix K, and normalizing the above-mentioned similarity according to the Softmax function; determining the global context feature information according to the normalized similarity and the above-mentioned value matrix V; when calculating the similarity between the above-mentioned key matrix K and the query Q, a dot product calculation can be used, and after calculating the similarity corresponding to the preset proxy token, the Softmax function is applied to normalize the similarity, and the normalized similarity is used as the attention weight, and then the value matrix V is weightedly summed to obtain the global context feature information; after obtaining the above-mentioned global context feature information, a second attention calculation is performed, and the above-mentioned global context feature information is used as the new value matrix V ′ , then the query Q and the new value matrix V ′ Apply the Softmax function again and transform the value matrix V ′ The preset proxy token is broadcast back so that the preset proxy token receives the global context feature information, thereby taking all the original input signals into account during the denoising process.
[0093] Step 103: Input the noisy signal into the signal denoising model to obtain a denoised clean signal.
[0094] In one or more possible embodiments, after determining that model training is complete, the noisy signal can be input into the signal denoising model to obtain a denoised clean signal. Prior to this, an experiment is conducted to further verify the effectiveness of the signal denoising model. First, the sampling frequency of a smartphone and an inertial laboratory MRU unit is set to the same 100 Hz. Then, the two units start working synchronously to collect signal data at the same time. This allows for obtaining a relatively high-precision signal (the signal collected by the MRU) and a lower-precision signal (the signal collected by the smartphone). The signal collected by the smartphone is then used as the noisy signal, and the signal collected by the MRU is used as the clean signal to form an experimental dataset.
[0095] The generator and discriminator of the conditional generative adversarial network of the present application are set with initial learning rates of 1e-4 and 3e-4 respectively, and the Adam optimizer is adopted with its momentum parameter set to (0.9, 0.999) to achieve a more stable training process; the above-mentioned noisy signal is input into the trained conditional generative adversarial network to obtain a training data set (that is, training sample), which includes signals collected by smartphones and synthetic signals generated based on the signals collected by smartphones. In order to evaluate the performance of the trained conditional generative adversarial network, principal component analysis (PCA) dimensionality reduction technology is used to visualize the data, which is used to intuitively observe the similarity between the synthetic signal generated by the conditional generative adversarial network and the signal collected by the smartphone; specifically, Figure 4 As shown, compared with the consistency of the shape, fluctuation and overall trend of the above-mentioned synthetic signal with the signal collected by the smartphone, the above-mentioned synthetic signal visually presents a signal pattern similar to the real signal (the signal collected by the smartphone), which shows that the conditional generative adversarial network trained in this application is able to generate data similar to the original noisy signal.
[0096] To further illustrate the similarity between the original signal (the signal collected by the smartphone) and the synthesized signal, PCA can be used to visualize the synthesized signal and the original signal to map the signal data into a two-dimensional space, as shown in the following example: Figure 5 As shown, it can be seen that the data points of the above-mentioned synthetic signal and the original signal are basically clustered together, and there is no very obvious separation, which shows that the two signals are highly similar. That is, the conditional generative adversarial network of this application has successfully learned the characteristics of the original signal and can generate a signal similar to the original signal.
[0097] Subsequently, the trained conditional generative adversarial network with the best similarity is used to construct a training data set (the training data set includes signals collected by the smartphone and synthetic signals generated based on the signals collected by the smartphone), and the training data set is used to train the signal denoising neural network. At the same time, the signal denoising neural network is trained using clean signals (signals collected by the MRU) as label data, and the performance of all denoising neural networks is compared. The signal denoising neural networks used in the experiments of this application that are trained using clean signals as label data include at least one of the following:
[0098] 1. SG (Savitzky-Golay) filter: It mainly reduces the noise in the signal through convolution operation. It is a denoising technology based on local smoothing.
[0099] 2. Moving Average (MA): This method suppresses unstable signal noise by averaging measurements within a rolling window. It is a smoothing-based denoising technique.
[0100] 3. DWT (Discrete Wavelet Transform): Discrete wavelet transform captures signal details at different frequency levels through multi-scale analysis;
[0101] 4. KNN (K-Nearest Neighbors): The K-nearest neighbor algorithm is a supervised classification algorithm that estimates the denoised signal value by finding the average of the k clean data points in the signal that are most similar to the noisy data point.
[0102] 5. RNN, GRU, and LSTM network variants: These are recurrent neural networks that can capture temporal dependencies when processing sequential data.
[0103] 6. U-Net network: network model of encoder-decoder structure, basic network model framework;
[0104] To ensure a fair comparison, all neural network-based methods were trained for the same number of times, 700. The inputs to all neural network models consisted of data pairs consisting of noisy and clean signals. All frameworks were implemented in PyTorch. After training, the denoising performance of the neural network models was evaluated using four quantitative metrics widely recognized in the denoising field. These metrics include:
[0105] RMSE, Root Mean Squared Error, is a measure of the standard deviation of the difference between the predicted value and the actual value. It is expressed by taking the square root of the average of the squared errors to express the accuracy of the prediction.
[0106] MAE, Mean Absolute Error, measures the size of the error by calculating the average of the absolute values of the differences between the predicted value and the actual value;
[0107] PSNR, Peak signal-to-noise ratio, is calculated by comparing the maximum possible power in the signal (i.e., the power of the original clean signal) with the noise power in the model output signal;
[0108] RAE, Relative Absolute Error, this indicator expresses the relative size of the error by calculating the ratio of the predicted error to the average of its true value;
[0109] The specific formula is as follows:
[0110]
[0111] Among them, y cRepresents a real clean signal, y e represents the denoised signal output by the model, i represents the i-th signal, and m represents the number of signal pairs. MAE, RMSE, and RAE mainly focus on the accuracy of the neural network model's predictions. Lower values indicate better denoising effects. PSNR measures the overall quality of the denoised signal. Higher PSNR values indicate that the denoising process successfully reduces noise while preserving the details and structure of the original signal as much as possible. The experimental data of the above signal denoising methods are shown in the following table:
[0112]
[0113] Among them, in the above table, N2N indicates that the training method uses the unsupervised learning training (Noise2Noise) method (that is, the method of training using the above training data set), and N2C indicates that the method of using supervised learning training (that is, using clean data as label data). It can be seen that the signal denoising model trained using the above training data set is equivalent to the neural network model trained using the clean signal as label data, which proves that the training samples generated by the conditional generative adversarial network in this application are effective, that is, as long as a noisy signal pair with sufficiently small dependence can be obtained on the basis of the original noisy signal, the use of such a signal pair can help the denoising neural network learn how to denoise from the noise difference, and finally realize the denoising function. According to the data in the above table, it can also be seen that the denoising effect of the signal denoising method proposed in this application is better than other methods.
[0114] According to a signal denoising method provided in this application, a training sample can be constructed in the absence of clean data, and the convolution and proxy attention mechanisms can be combined to improve the denoising effect of the accelerometer signal.
[0115] Example 2
[0116] Corresponding to the above signal denoising method, the present invention also proposes a signal denoising device, such as Figure 6 As shown; since the device embodiment of the present invention corresponds to the above-mentioned method embodiment, details not disclosed in the device embodiment can be referred to the above-mentioned method embodiment, and will not be repeated in the present invention.
[0117] The sample acquisition module 601 is used to acquire training samples, where the training samples include original signals and synthetic signals generated based on the original signals;
[0118] A model training module 602 is configured to use the original signal as input and the synthesized signal as output to train a convolutional neural network that incorporates a proxy attention module to obtain a signal denoising model;
[0119] The signal denoising module 603 is used to input the noisy signal into the above-mentioned signal denoising model to obtain a denoised clean signal.
[0120] Example 3
[0121] The present application also provides a signal denoising device, comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above-mentioned signal denoising method.
[0122] like Figure 7 As shown, the device includes a processor 701 , a memory 702 , a communication interface 703 and a bus 704 . The processor 701 , the memory 702 and the communication interface 703 are interconnected via the bus 704 .
[0123] The processor 701 is configured to read and execute instructions in the memory 702 , so that at least one processor can execute the signal denoising method provided in the above embodiment.
[0124] The memory 702 is used to store various instructions and programs of the signal denoising method provided in the above embodiment.
[0125] The bus 704 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0126] The processor 701 may be a central processing unit (CPU), a network processor (NP), a graphics processing unit (GPU), or any combination of a CPU, NP, and GPU. It may also be a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0127] Example 4
[0128] In addition, the present application also provides a computer-readable storage medium, such as Figure 8 As shown, the computer storage medium stores a computer program, and the computer program is used to enable a computer to execute any one of the methods in the above embodiments.
[0129] The memory may include readable media in the form of volatile memory, such as random access memory (RAM) 801 and / or cache memory 802 , and may further include read-only memory (ROM) 803 .
[0130] The memory may also include a program / utility 805 having a set (at least one) of program modules 804, such program modules 804 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0131] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0132] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0133] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0134] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0135] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A signal denoising method, characterized in that: include: Acquire a training sample, where the training sample includes an original signal and a synthesized signal generated based on the original signal; Taking the original signal as input and the synthesized signal as output, training a convolutional neural network that introduces a proxy attention module to obtain a signal denoising model; The noisy signal is input into the signal denoising model to obtain a denoised clean signal.
2. The method according to claim 1, characterized in that The synthetic signal that meets the preset conditions is generated according to a conditional generative adversarial network, wherein the conditional generative adversarial network includes a generator and a discriminator, the generator is used to generate a synthetic signal based on the original signal and a random vector, and the discriminator uses a soft label method to classify the original signal and the synthetic signal.
3. The method according to claim 2, characterized in that The conditional generative adversarial network is trained in the following way: The mean square error loss function is used to determine the original signal loss and the synthetic signal loss; Using the synthetic signal loss as the generator loss, and determining the discriminator loss according to the original signal loss and the synthetic signal loss; When it is determined that the generator loss and the discriminator loss meet the preset condition, a trained conditional generative adversarial network is obtained.
4. The method according to claim 3, characterized in that The preset condition is that both the generator loss and the discriminator loss are close to a preset threshold.
5. The method according to claim 1, wherein The convolutional neural network introducing the agent attention module also includes an encoder and a decoder. The encoder is used to compress the original signal layer by layer through multi-layer convolution operations to obtain feature information of different time scales; the decoder is used to restore the feature information of different time scales through upsampling and output the synthesized signal.
6. The method according to claim 5, characterized in that The training of the convolutional neural network introducing the agent attention module includes: Inputting the original signal into the encoder to obtain feature information of different time scales; Based on a preset proxy token, the feature information of the different time scales is processed to obtain global context feature information; The global context feature information is restored according to an upsampling method, and the synthesized signal is output.
7. The method according to claim 6, characterized in that The processing of the feature information of different time scales based on the preset proxy token to obtain global context feature information includes: Determine a key matrix K and a value matrix V according to the characteristic information of the different time scales; Taking the preset proxy token as a query Q, calculating the similarity between the query Q and the key matrix K, and normalizing the similarity according to a Softmax function; According to the normalized similarity and the value matrix V, global context feature information is determined.
8. A signal denoising device, characterized in that: The device comprises: A sample acquisition module is used to acquire a training sample, wherein the training sample includes an original signal and a synthetic signal generated according to the original signal; a model training module, configured to take the original signal as input and the synthesized signal as output, and train a convolutional neural network that introduces a proxy attention module to obtain a signal denoising model; The signal denoising module is used to input the noisy signal into the signal denoising model to obtain a denoised clean signal.
9. An electronic device, characterized in that: The electronic device comprises: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any one of the methods of claims 1-7.
10. A computer storage medium, characterized in that The computer storage medium stores a computer program, and the computer program is used to make a computer execute the method according to any one of claims 1 to 7.