Denoising method and system based on digital filtering and deep learning
By combining digital filtering and deep learning methods, multi-level semantic constraints and interactions are applied to the data to be denoised, solving the problem of poor denoising effect in existing technologies and achieving a more reliable denoising effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-03-13
AI Technical Summary
Existing denoising methods suffer from information loss or incomplete denoising when faced with complex, nonlinear, or non-stationary noise features. Furthermore, deep learning-based methods are prone to overfitting or semantic distortion when there is a discrepancy between the training data distribution and the test data.
By combining digital filtering and deep learning, the first denoised data is formed by digital filtering of the data to be denoised, semantic vectors are formed by latent semantic mining, and semantic constraints and semantic restoration are performed by a deep learning neural network model. By integrating and aggregating semantic interactions at multiple levels, the target denoised data is formed.
It improves the reliability of denoised data, takes into account both the richness and accuracy of semantic information, and addresses the problem of poor denoising performance in existing technologies.
Smart Images

Figure CN121658779A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data analysis technology, and more specifically, to a denoising method and system based on digital filtering and deep learning. Background Technology
[0002] In many technological fields such as data analysis, signal processing, image recognition, speech recognition, and intelligent sensing, data denoising is a fundamental step that significantly impacts the accuracy and robustness of subsequent analysis and modeling. Existing technologies commonly employ denoising methods including traditional digital filtering methods such as median filtering, Kalman filtering, and wavelet transform, as well as emerging deep learning-based denoising methods such as autoencoders and convolutional neural networks. While these methods demonstrate some denoising effectiveness in specific scenarios, they still have limitations. For example, traditional digital filtering methods rely on preset rules or fixed model parameters, making them unable to fully adapt to complex, nonlinear, or non-stationary noise characteristics, easily leading to information loss or incomplete denoising. While deep learning-based denoising methods possess stronger feature extraction capabilities, the lack of explicit prior constraints makes them prone to overfitting, denoising failure, or semantic distortion when there are discrepancies between the training and test data distributions. Therefore, existing technologies suffer from unsatisfactory denoising performance. Summary of the Invention
[0003] In view of this, the purpose of this application is to provide a denoising method and system based on digital filtering and deep learning, so as to improve the problem of poor denoising effect in the prior art.
[0004] To achieve the above objectives, this application adopts the following technical solution: A denoising method based on digital filtering and deep learning includes: The data to be denoised is digitally filtered to form the first denoised data, wherein the data to be denoised belongs to the time domain or the spatial domain. Latent semantic mining is performed on the data to be denoised and the first denoised data respectively to form a semantic vector to be denoised and a first denoised semantic vector. Based on the first denoised semantic vector, semantic constraints are applied to the deep denoising process of the semantic vector to be denoised, forming a target denoised semantic vector; The target denoised semantic vector is semantically restored to form the second denoised data. Latent semantic mining, deep denoising, semantic constraints and semantic restoration are implemented through the target denoising model, and the target denoising model belongs to the neural network model formed by deep learning. The first denoised data and the second denoised data are merged to form the target denoised data.
[0005] In a preferred embodiment of this application, in the aforementioned denoising method based on digital filtering and deep learning, the step of semantically constraining the deep denoising process of the semantic vector to be denoised based on the first denoised semantic vector to form a target denoised semantic vector includes: Based on the first denoised semantic vector, the semantic vector to be denoised is subjected to first association mining at multiple levels to achieve deep denoising and obtain first association denoised vectors at multiple levels. The semantic vector to be denoised is subjected to multiple levels of second association mining to achieve deep denoising and obtain multiple levels of second association denoising vectors. According to the corresponding levels, the first associated denoising vectors and the second associated denoising vectors of the multiple levels are semantically interacted to achieve semantic constraints and form semantically interacted denoising vectors of multiple levels, wherein the semantic interaction includes gating adjustment. The semantic interaction denoising vectors of the multiple levels are semantically aggregated to form the target denoised semantic vector.
[0006] In a preferred embodiment of this application, in the aforementioned denoising method based on digital filtering and deep learning, the step of performing multi-level first association mining on the semantic vector to be denoised based on the first denoised semantic vector to achieve deep denoising and obtain multi-level first association denoised vectors includes: Based on the first denoised semantic vector, cross-attention processing is performed on the semantic vector to be denoised to form the first-level first associated denoised vector; Pooling compression is performed on the first denoised semantic vector and the first association denoised vector of the first level respectively to form the first denoised compressed vector of the second level and the first association compressed vector of the second level. Based on the first denoised compressed vector, cross attention processing is performed on the first association compressed vector to form the first association denoised vector of the second level. Pooling compression is performed on the first denoised compression vector and the first associated denoised vector of the previous level to form the first denoised compression vector and the first associated compression vector of the current level. Based on the first denoised compression vector, cross-attention processing is performed on the first associated compression vector to form the first associated denoised vector of the current level.
[0007] In a preferred embodiment of this application, in the aforementioned denoising method based on digital filtering and deep learning, the step of performing multi-level second association mining on the semantic vector to be denoised to achieve deep denoising and obtain multi-level second association denoised vectors includes: The semantic vector to be denoised is subjected to self-attention processing to form the second associated denoising vector of the first level; Pooling compression is performed on the semantic vector to be denoised and the second association denoising vector of the first level respectively to form the second denoising compressed vector of the second level and the second association compressed vector of the second level. Based on the second denoising compressed vector, cross attention processing is performed on the second association compressed vector to form the second association denoising vector of the second level. Pooling compression is performed on the second denoised compression vector and the second associated denoised vector of the previous level to form the second denoised compression vector and the second associated compression vector of the current level. Based on the second denoised compression vector, cross-attention processing is performed on the second associated compression vector to form the second associated denoised vector of the current level.
[0008] In a preferred embodiment of this application, in the aforementioned denoising method based on digital filtering and deep learning, the step of semantically aggregating the semantic interaction denoising vectors at multiple levels to form a target denoised semantic vector includes: For each level of semantic interaction denoising vector, the first mining index and the second mining index of the first association denoising vector and the second association denoising vector corresponding to the semantic interaction denoising vector are determined in the corresponding first association mining and second association mining. The first association mining and the second association mining are implemented based on the attention mechanism. The first mining index and the second mining index are used to reflect the degree of difference between the vectors before and after attention mining. The degree of difference is related to the denoising effect. The first mining index and the second mining index are combined to form a target mining index, and the aggregation weight parameters of the semantic interaction denoising vector at the corresponding level are determined based on the target mining index. Based on the aggregation weight parameters of the semantic interaction denoising vectors at each level, the semantic interaction denoising vectors at each level are aggregated to form the target denoised semantic vector.
[0009] In a preferred embodiment of this application, in the aforementioned denoising method based on digital filtering and deep learning, the step of performing latent semantic mining on the data to be denoised and the first denoised data respectively to form a semantic vector to be denoised and a first denoised semantic vector includes: Fourier transform is performed on the data to be denoised and the first denoised data respectively to form the corresponding spectrum diagram to be denoised and the first denoised spectrum diagram; The latent semantics of the spectrum to be denoised are mined to form a semantic vector to be denoised. Latent semantic mining is performed on the first denoised spectrogram to form a first denoised semantic vector.
[0010] In a preferred embodiment of this application, in the aforementioned denoising method based on digital filtering and deep learning, the step of performing latent semantic mining on the spectrogram to be denoised to form a semantic vector to be denoised includes: The spectrum image to be denoised is convolved to form a convolution vector to be denoised; The spectrum to be denoised is divided into multiple local spectrums to be denoised according to frequency bands, and each local spectrum to be denoised is convolved to form multiple local convolution vectors to be denoised corresponding to the multiple local spectrums to be denoised. Each of the local convolutional vectors to be denoised is used to perform association mining on the local convolutional vectors to be denoised, thereby forming multiple association vectors to be denoised corresponding to the multiple local convolutional vectors to be denoised. The association mining is used to extract potential semantic information that has an association relationship with the local convolutional vectors to be denoised from the local convolutional vectors to be denoised. The multiple denoised related vectors are fused to form a denoised semantic vector.
[0011] In a preferred embodiment of this application, in the aforementioned denoising method based on digital filtering and deep learning, the step of performing latent semantic mining on the first denoised spectrogram to form a first denoised semantic vector includes: The first denoised spectrogram is convolved to form a first denoised convolution vector; The first denoised spectrum is divided into multiple first denoised local spectrums according to frequency bands, and each of the first denoised local spectrums is convolved to form multiple first denoised local convolution vectors corresponding to the multiple first denoised local spectrums. Each of the first denoised local convolution vectors is used to perform association mining on the first denoised convolution vector to form multiple first denoised association vectors corresponding to the multiple first denoised local convolution vectors. The association mining is used to extract potential semantic information that has an association relationship with the first denoised local convolution vector from the first denoised convolution vector. The multiple first denoised correlation vectors are fused to form a first denoised semantic vector.
[0012] In a preferred embodiment of this application, in the aforementioned denoising method based on digital filtering and deep learning, the step of semantically restoring the target denoised semantic vector to form second denoised data includes: Perform a fully connected mapping on the target denoised semantic vector to form a fully connected denoised semantic vector; The shape of the fully connected denoised semantic vector is transformed to form a denoised semantic feature map; The denoised semantic feature map is subjected to multiple levels of transposed convolution processing to obtain the transposed convolution semantic vector of the last level. The transposed convolution semantic vector of the next level is obtained by transposing the transposed convolution semantic vector of the previous level. The transposed convolutional semantic vector of the last layer is activated and output to form the target spectrogram; The target spectrogram is subjected to inverse Fourier transform to form the second denoised data.
[0013] Based on the above, this application also provides a denoising system based on digital filtering and deep learning, comprising: Memory, used to store computer programs; A processor connected to the memory is used to execute the computer program stored in the memory to implement the above-described denoising method based on digital filtering and deep learning.
[0014] The denoising method and system based on digital filtering and deep learning provided in this application firstly performs digital filtering on the data to be denoised to form first denoised data; secondly, latent semantic mining is performed on the data to be denoised and the first denoised data respectively to form a semantic vector to be denoised and a first denoised semantic vector; then, based on the first denoised semantic vector, semantic constraints are applied to the deep denoising process of the semantic vector to be denoised to form a target denoised semantic vector; further, semantic restoration is performed on the target denoised semantic vector to form second denoised data; finally, the first denoised data and the second denoised data are fused to form the target denoised data. Based on the above, since the data to be denoised is first digitally filtered, the semantic vector of the first denoised data can be used to impose semantic constraints on the deep denoising process of the corresponding semantic vector of the data to be denoised. This ensures that the resulting target denoised semantic vector can balance both the richness of semantic information (if deep learning-based latent semantic mining is directly performed on the first denoised data formed by digital filtering, some effective semantic information may be lost due to digital filtering) and the accuracy of semantic information (without the semantic constraints of the semantic vector of the first denoised data, the denoising effect may be relatively low). Therefore, the reliability of the second denoised data formed based on the target denoised semantic vector can be higher, thus fully guaranteeing the reliability of the resulting target denoised data. This achieves an effective combination of deep learning and digital filtering, thereby improving the problem of poor denoising effect in existing technologies. Attached Figure Description
[0015] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings.
[0016] Figure 1This is a block diagram of a denoising system based on digital filtering and deep learning provided in an embodiment of this application.
[0017] Figure 2 This is a flowchart illustrating the denoising method based on digital filtering and deep learning provided in an embodiment of this application.
[0018] Figure 3 This is a first schematic diagram of potential semantic mining provided for an embodiment of this application.
[0019] Figure 4 A second schematic diagram of potential semantic mining provided for embodiments of this application.
[0020] Figure 5 A schematic diagram illustrating semantic constraints provided in the embodiments of this application.
[0021] Figure 6 This is a schematic diagram of the first association mining provided in an embodiment of this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0023] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0024] like Figure 1 As shown in the figure, this application provides a denoising system based on digital filtering and deep learning, which may include a memory and a processor.
[0025] In detail, the memory and the processor are electrically connected directly or indirectly to enable data transmission or interaction. For example, the memory and the processor can be electrically connected via one or more communication buses or signal lines. The processor is used to execute executable computer programs stored in the memory to implement the denoising method based on digital filtering and deep learning provided in the embodiments of this application.
[0026] Optionally, the memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0027] Optionally, the processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a system on chip (SoC), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0028] Understandable. Figure 1 The structure shown is for illustrative purposes only. The denoising system based on digital filtering and deep learning may also include components that are more efficient than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown may include, for example, a communication unit for interacting with other devices (such as image sensors, sound sensors, etc.).
[0029] Combination Figure 2 This application also provides a denoising method based on digital filtering and deep learning that can be applied to the aforementioned denoising system based on digital filtering and deep learning. The method steps defined in the process of the denoising method based on digital filtering and deep learning can be implemented by the denoising system based on digital filtering and deep learning (hereinafter referred to as the denoising system).
[0030] The following will be about Figure 2 The specific process shown will be explained in detail.
[0031] Step S110: Perform digital filtering on the data to be denoised to form the first denoised data.
[0032] In this embodiment, the denoising system can perform digital filtering on the data to be denoised to form first denoised data. The data to be denoised can be either time-domain data (such as audio data) or spatial-domain data (such as image data).
[0033] Step S120: Perform latent semantic mining on the data to be denoised and the first denoised data respectively to form a semantic vector to be denoised and a semantic vector to be denoised.
[0034] In this embodiment, the denoising system can perform latent semantic mining on the data to be denoised and the first denoised data respectively, forming a semantic vector to be denoised and a first denoised semantic vector. That is, it can mine the latent semantic information in the data to be denoised and represent it in vector form to obtain the semantic vector to be denoised; and it can mine the latent semantic information in the first denoised data and represent it in vector form to obtain the first denoised semantic vector.
[0035] Step S130: Based on the first denoised semantic vector, semantic constraints are applied to the deep denoising process of the semantic vector to be denoised to form a target denoised semantic vector.
[0036] In this embodiment, after obtaining the first denoised semantic vector and the semantic vector to be denoised, the denoising system can apply semantic constraints to the deep denoising process of the semantic vector to be denoised based on the first denoised semantic vector to form a target denoised semantic vector. That is, since the first denoised semantic vector is obtained based on the first denoised data from digital filtering, it has relatively less noise semantic information. Therefore, applying corresponding semantic constraints during the deep denoising process can improve the accuracy of deep denoising.
[0037] Step S140: Semantic restoration is performed on the target denoised semantic vector to form the second denoised data.
[0038] In this embodiment, after obtaining the target denoised semantic vector, the denoising system can perform semantic restoration on the target denoised semantic vector to form second denoised data. Latent semantic mining, deep denoising, semantic constraints, and semantic restoration are implemented through a target denoising model, which is a neural network model formed through deep learning (e.g., learning between the data sample to be denoised and the corresponding target denoised data label; the learning process can refer to relevant existing technologies, i.e., the process of training a neural network model). That is, steps S120, S130, and S140 can be implemented through the target denoising model.
[0039] Step S150: Merge the first denoised data and the second denoised data to form the target denoised data.
[0040] In this embodiment of the application, after obtaining the first denoised data and the second denoised data, the denoising system can fuse the first denoised data and the second denoised data to form target denoised data. In this way, the results of digital filtering and deep learning denoising can be taken into account, making the reliability of the formed target denoised data higher.
[0041] Based on the above, since the data to be denoised is first digitally filtered, the semantic vector of the first denoised data can be used to impose semantic constraints on the deep denoising process of the corresponding semantic vector of the data to be denoised. This ensures that the resulting target denoised semantic vector can balance both the richness of semantic information (if deep learning-based latent semantic mining is directly performed on the first denoised data formed by digital filtering, some effective semantic information may be lost due to digital filtering) and the accuracy of semantic information (without the semantic constraints of the semantic vector of the first denoised data, the denoising effect may be relatively low). Therefore, the reliability of the second denoised data formed based on the target denoised semantic vector can be higher, thus fully guaranteeing the reliability of the resulting target denoised data. This achieves an effective combination of deep learning and digital filtering, thereby improving the problem of poor denoising effect in existing technologies.
[0042] Firstly, regarding step S110, it should be noted that the specific method of digital filtering the data to be denoised is not limited and can be selected according to actual needs.
[0043] For example, in an alternative implementation, at least one of several methods, such as median filtering, Kalman filtering, and wavelet transform, can be used to digitally filter the data to be denoised, obtaining the corresponding first denoised data. It should be noted that when using multiple digital filtering methods, these methods can be implemented in parallel. Thus, the resulting multiple digital filtering results can be averaged or weighted to calculate the first denoised data. Alternatively, the multiple digital filtering methods can be performed sequentially; that is, the data to be denoised can first undergo a first digital filter, and then the result of the first digital filter can be subjected to a second digital filter to obtain the first denoised data.
[0044] Secondly, regarding step S120, it should be noted that the specific methods for performing latent semantic mining on the data to be denoised and the first denoised data are not limited and can be selected according to actual needs.
[0045] For example, in an alternative implementation, in order to fully extract the potential semantic information in the data to be denoised and the first denoised data, so that the formed semantic vector can represent more detailed information, the above step S120 may further include the following steps S121, S122 and S123, the specific contents of each step are as follows.
[0046] Step S121: Perform Fourier transform on the data to be denoised and the first denoised data respectively to form the corresponding spectrum diagram to be denoised and the first denoised spectrum diagram.
[0047] In this embodiment of the application, the data to be denoised and the first denoised data can be subjected to Fourier transforms respectively to form corresponding spectrograms to be denoised and first denoised spectrograms. For example, as Figure 3 As shown, the data to be denoised and the data to be denoised can be time-domain audio data.
[0048] Step S122: Perform latent semantic mining on the spectrum to be denoised to form a semantic vector to be denoised.
[0049] In this embodiment of the application, after obtaining the spectrogram to be denoised, latent semantic mining can be performed on the spectrogram to be denoised to form a semantic vector to be denoised. That is, since the spectrogram can represent more detailed information, performing latent semantic mining on the spectrogram to be denoised can result in more detailed information in the formed semantic vector to be denoised.
[0050] Step S123: Perform latent semantic mining on the first denoised spectrogram to form a first denoised semantic vector.
[0051] In this embodiment of the application, after obtaining the first denoised spectrogram, latent semantic mining can be performed on the first denoised spectrogram to form a first denoised semantic vector. That is, since the spectrogram can represent more detailed information, performing latent semantic mining on the first denoised spectrogram can result in more detailed information in the formed first denoised semantic vector.
[0052] It is understood that the specific method of performing latent semantic mining on the spectrogram to be denoised in step S122 is not limited. For example, in an alternative implementation, in order to achieve the mining of detailed semantic information while also fully mining important semantic information, so that important semantic information can be given priority in the formed semantic vector to be denoised, step S122 may further include steps S122a, S122b, S122c and S122d, wherein the specific contents of each step are as follows.
[0053] Step S122a: Convolve the spectrum image to be denoised to form a convolution vector to be denoised.
[0054] In the embodiments of this application, combined with Figure 4 The spectrogram to be denoised can be convolved to form a convolution vector to be denoised, wherein the convolution can be implemented by a convolutional neural network.
[0055] Step S122b: The spectrum to be denoised is divided into multiple local spectrums to be denoised according to frequency bands, and each local spectrum to be denoised is convolved to form multiple local convolution vectors to be denoised corresponding to the multiple local spectrums to be denoised.
[0056] In this embodiment, the spectrogram to be denoised can be segmented according to frequency bands (e.g., 0-50Hz, 51-100Hz, 101-150Hz, etc.) to form multiple local spectrograms to be denoised. Each of these local spectrograms is then convolved to form multiple local convolution vectors corresponding to the multiple local spectrograms to be denoised. In other words, since noise generally has different frequencies from actual sound, segmentation followed by convolution can reduce the interference of noise on actual sound to a certain extent, thereby improving the accuracy of semantic mining.
[0057] Step S122c: Based on each of the local convolutional vectors to be denoised, perform association mining on the convolutional vectors to be denoised to form multiple association vectors to be denoised corresponding to the multiple local convolutional vectors to be denoised.
[0058] In this embodiment, after obtaining the local convolutional vector to be denoised and the convolutional vector to be denoised, association mining can be performed on each of the local convolutional vectors to be denoised to form multiple association vectors corresponding to the multiple local convolutional vectors to be denoised. Association mining is used to extract potential semantic information that has a correlation relationship with the local convolutional vectors to be denoised from the convolutional vectors to be denoised, such as through cross-attention processing. Furthermore, it should be noted that when the local convolutional vectors to be denoised and the convolutional vector to be denoised have different sizes, upsampling and / or downsampling of the local convolutional vectors to be denoised can be performed to make the processed sizes the same, and then association mining can be performed.
[0059] Step S122d: Merge the multiple denoised association vectors to form a denoised semantic vector.
[0060] In this embodiment of the application, after obtaining the plurality of denoised correlation vectors, the plurality of denoised correlation vectors can be fused to form a semantic vector to be denoised. For example, the plurality of denoised correlation vectors can be fused by splicing, averaging or weighted averaging.
[0061] It is understood that the specific method of performing latent semantic mining on the spectrogram to be denoised in step S123 above is not limited. For example, in an alternative implementation, in order to achieve the mining of detailed semantic information while also fully mining important semantic information, so that important semantic information can be given priority in the formed semantic vector to be denoised, step S123 above may further include steps S123a, S123b, S123c and S123d, wherein the specific contents of each step are as follows.
[0062] Step S123a: Convolve the first denoised spectrogram to form a first denoised convolution vector.
[0063] In this embodiment of the application, the first denoised spectrogram can be convolved to form a first denoised convolution vector, wherein the convolution can be implemented by a convolutional neural network.
[0064] Step S123b: The first denoised spectrum is divided into frequency bands to form multiple first denoised local spectrums, and each of the first denoised local spectrums is convolved to form multiple first denoised local convolution vectors corresponding to the multiple first denoised local spectrums.
[0065] In this embodiment, the first denoised spectrogram can be segmented according to frequency bands to form multiple first denoised local spectrograms. Each of these first denoised local spectrograms is then convolved to form multiple first denoised local convolution vectors corresponding to the multiple first denoised local spectrograms. In other words, since noise generally has different frequencies from actual sound, segmentation followed by convolution can reduce the interference of noise on actual sound to a certain extent, thereby improving the accuracy of semantic mining.
[0066] Step S123c: Based on each of the first denoised local convolution vectors, perform association mining on the first denoised convolution vectors to form multiple first denoised association vectors corresponding to the multiple first denoised local convolution vectors.
[0067] In this embodiment, after obtaining the first denoised local convolutional vector and the first denoised convolutional vector, association mining can be performed on each of the first denoised local convolutional vectors to form multiple first denoised association vectors corresponding to the multiple first denoised local convolutional vectors. The association mining is used to extract potential semantic information that has a correlation relationship with the first denoised local convolutional vectors from the first denoised convolutional vectors, such as through cross-attention processing. Furthermore, it should be noted that when the first denoised local convolutional vectors and the first denoised convolutional vectors have different sizes, upsampling and / or downsampling of the first denoised local convolutional vectors can be performed to make the processed sizes the same. Then, association mining is performed to obtain the corresponding first denoised association vectors.
[0068] Step S123d: Fuse the multiple first denoised correlation vectors to form a first denoised semantic vector.
[0069] In this embodiment of the application, after obtaining the plurality of denoised correlation vectors, the plurality of first denoised correlation vectors can be fused to form a first denoised semantic vector. For example, the plurality of first denoised correlation vectors can be fused by concatenation, averaging, or weighted averaging.
[0070] Thirdly, regarding step S130, it should be noted that the specific method of semantic constraint on the deep denoising process of the semantic vector to be denoised is not restricted and can be selected according to actual needs.
[0071] For example, in an alternative implementation, in order to ensure effective suppression of noise semantic information during deep denoising and to make full use of the first denoised semantic vector for semantic constraints to achieve higher accuracy in deep denoising, the above-mentioned step S130 may further include steps S131, S132, S133 and S134, wherein the specific contents of each step are as follows.
[0072] Step S131: Based on the first denoised semantic vector, perform first association mining at multiple levels on the semantic vector to be denoised to achieve deep denoising and obtain first association denoised vectors at multiple levels.
[0073] In the embodiments of this application, combined with Figure 5Based on the first denoised semantic vector, multiple levels of first association mining can be performed on the semantic vector to be denoised to achieve deep denoising and obtain multiple levels of first association denoised vectors. In other words, since noisy semantic information and the semantic information required in reality generally have no correlation or a low correlation, by performing association mining, the associated semantic information can be mined, thereby ignoring the unrelated noisy semantic information, and thus suppressing the noisy semantic information to complete deep denoising.
[0074] Step S132: Perform multi-level second association mining on the semantic vector to be denoised to achieve deep denoising and obtain multi-level second association denoised vectors.
[0075] In this embodiment, the semantic vector to be denoised can also undergo multi-level second association mining to achieve deep denoising and obtain multi-level second association denoised vectors. That is, the noise semantic information can be suppressed by utilizing the fact that there is no correlation or a low correlation between the noise semantic information and the semantic information required by the actual needs within the semantic vector to be denoised, so as to complete the deep denoising.
[0076] Step S133: According to the corresponding levels, the first associated denoising vectors and the second associated denoising vectors of the multiple levels are semantically interacted to achieve semantic constraints and form semantically interactive denoising vectors of multiple levels.
[0077] In this embodiment of the application, after obtaining the first associated denoising vector and the second associated denoising vector, the first associated denoising vector and the second associated denoising vector of the multiple levels can be semantically interacted according to the corresponding levels to achieve semantic constraints and form a semantically interactive denoising vector of multiple levels. For example, the first associated denoising vector of the first level and the second associated denoising vector of the first level can be semantically interacted to form a semantically interactive denoising vector of the first level. As another example, the first associated denoising vector of the second level and the second associated denoising vector of the second level can be semantically interacted to form a semantically interactive denoising vector of the second level. The semantic interaction includes gating adjustment. It should be noted that since the second association denoising vector is formed internally, it has certain limitations, but the richness of the representation of the original semantic information is relatively sufficient. The first association denoising vector is formed based on external semantic information mining, so it has higher accuracy, but the richness of the representation of the original semantic information is relatively lower. Based on this, the first association denoising vector can be activated by functions such as sigmoid to form a corresponding weight parameter distribution. Then, the weight parameter distribution and the second association denoising vector can be multiplied bitwise, and the result of the bitwise multiplication can be summed, weighted averaged, or averaged to obtain the corresponding semantic interaction denoising vector.
[0078] Step S134: Semantically aggregate the semantic interaction denoising vectors of the multiple levels to form the target denoised semantic vector.
[0079] In this embodiment of the application, after obtaining the semantic interaction denoising vectors of the multiple levels, the semantic interaction denoising vectors of the multiple levels can be semantically aggregated to form a target denoised semantic vector. In this way, the target denoised semantic vector can take into account both the accuracy and richness of semantic representation, that is, it has better semantic representation ability.
[0080] It is understood that in step S131 above, the specific method of performing first association mining at multiple levels on the semantic vector to be denoised is not limited. For example, in an alternative implementation, in order to ensure the sufficiency of the first association mining, that is, to make full use of the association relationship between the first denoised semantic vector and the semantic vector to be denoised, so as to achieve effective suppression of noise semantic information, step S131 above may further include steps S131a, S131b and S131c, wherein the specific contents of each step are as follows.
[0081] Step S131a: Based on the first denoised semantic vector, perform cross-attention processing on the semantic vector to be denoised to form the first association denoised vector of the first level.
[0082] In the embodiments of this application, combined with Figure 6 Based on the first denoised semantic vector, cross-attention processing can be performed on the semantic vector to be denoised to form the first-level first associated denoised vector. That is, in the first level, the association relationship between the first denoised semantic vector and the semantic vector to be denoised can be directly used to achieve the corresponding suppression of noise semantic information.
[0083] Step S131b: Pooling compression is performed on the first denoised semantic vector and the first association denoised vector of the first level to form the first denoised compressed vector of the second level and the first association compressed vector of the second level. Based on the first denoised compressed vector, cross-attention processing is performed on the first association compressed vector to form the first association denoised vector of the second level.
[0084] In this embodiment, after obtaining the first-level first-association denoising vector, the first denoised semantic vector and the first-level first-association denoising vector can be pooled and compressed (the pooling method can be the same) to form a second-level first-denoised compressed vector and a second-level first-association compressed vector (the sizes of the two compressed vectors are known). Based on the first denoised compressed vector, cross-attention processing is performed on the first-association compressed vector to form a second-level first-association denoising vector. In other words, through pooling compression, high-level, abstract semantic features can be extracted. Then, the association relationships within these high-level, abstract semantic features are utilized to further suppress noisy semantic information.
[0085] Step S131c: Pooling compression is performed on the first denoised compression vector and the first associated denoised vector of the previous level to form the first denoised compression vector and the first associated compression vector of the current level. Based on the first denoised compression vector, cross-attention processing is performed on the first associated compression vector to form the first associated denoised vector of the current level.
[0086] In this embodiment, after obtaining the first associated denoising vector of the second level, pooling compression can be performed on the first denoising compression vector and the first associated denoising vector of the previous level to form the first denoising compression vector and the first associated compression vector of the current level. Then, based on the first denoising compression vector, cross-attention processing is performed on the first associated compression vector to form the first associated denoising vector of the current level. For example, pooling compression can be performed on the first denoising compression vector and the first associated denoising vector of the second level to form the first denoising compression vector and the first associated compression vector of the third level. Then, based on the first denoising compression vector, cross-attention processing is performed on the first associated compression vector to form the first associated denoising vector of the third level.
[0087] It is understood that in step S132 above, the specific method of performing the second association mining at multiple levels on the semantic vector to be denoised is not limited. For example, in an alternative implementation, in order to ensure the sufficiency of the second association mining, that is, to make full use of the association relationship between the semantics within the semantic vector to be denoised in order to achieve effective suppression of noise semantic information, step S132 above may further include steps S132a, S132b and S132c, wherein the specific contents of each step are as follows.
[0088] Step S132a: Perform self-attention processing on the semantic vector to be denoised to form the second associated denoised vector of the first level.
[0089] In this embodiment, the semantic vector to be denoised can be subjected to self-attention processing to form a second associated denoised vector at the first level. That is, in the first level, the association relationship between the semantic information within the semantic vector to be denoised can be directly utilized to achieve the corresponding suppression of noise semantic information.
[0090] Step S132b: Pooling compression is performed on the semantic vector to be denoised and the second association denoising vector of the first level to form the second denoising compressed vector of the second level and the second association compressed vector of the second level. Based on the second denoising compressed vector, cross-attention processing is performed on the second association compressed vector to form the second association denoising vector of the second level.
[0091] In this embodiment, after obtaining the second association denoising vector of the first level, pooling compression can be performed on the semantic vector to be denoised and the second association denoising vector of the first level to form a second denoising compressed vector of the second level and a second association compressed vector of the second level. Furthermore, based on the second denoising compressed vector, cross-attention processing is performed on the second association compressed vector to form a second association denoising vector of the second level. This is as described above.
[0092] Step S132c: Pooling compression is performed on the second denoised compression vector and the second associated denoised vector of the previous layer respectively to form the second denoised compression vector and the second associated compression vector of the current layer. Based on the second denoised compression vector, cross-attention processing is performed on the second associated compression vector to form the second associated denoised vector of the current layer.
[0093] In this embodiment, after obtaining the second correlation denoising vector of the second level, pooling compression can be performed on the second denoising compression vector and the second correlation denoising vector of the previous level to form the second denoising compression vector and the second correlation compression vector of the current level. Furthermore, based on the second denoising compression vector, cross-attention processing is performed on the second correlation compression vector to form the second correlation denoising vector of the current level. This is as described above.
[0094] It is understood that in step S134 above, the specific method of semantic aggregation of the semantic interaction denoising vectors of the multiple levels is not limited. For example, in an alternative implementation, in order to ensure the reliability of semantic aggregation, that is, to make full use of the semantic representation effectiveness of the semantic interaction denoising vectors of each level, step S134 above may further include steps S134a, S134b and S134c, wherein the specific contents of each step are as follows.
[0095] Step S134a: For each level of semantic interaction denoising vector, determine the first mining index and the second mining index of the first association denoising vector and the second association denoising vector corresponding to the semantic interaction denoising vector in the corresponding first association mining and second association mining.
[0096] In this embodiment, for each level of semantic interaction denoising vector, a first mining index and a second mining index can be determined in the corresponding first association mining and second association mining processes for the first association mining and second association mining, respectively. The first association mining and second association mining are implemented based on an attention mechanism (as described above), and the first mining index and second mining index are used to reflect the degree of difference between the vectors before and after attention mining. For example, for the first level of semantic interaction denoising vector, the difference between the semantic vector to be denoised and the first association denoising vector of the first level can be calculated to obtain a corresponding difference vector (which can represent the difference before and after attention mining). Then, the dispersion of this difference vector can be calculated. The larger the dispersion, the higher the degree of difference and the better the effect of attention mining. Therefore, this dispersion can be determined as the first mining index. Additionally, the difference between the semantic vector to be denoised and the second association denoising vector of the first level can be calculated to obtain a corresponding difference vector. Then, the dispersion of this difference vector can be calculated, and this dispersion can be determined as the second mining index. For example, for the semantic interaction denoising vector of the second level, the difference between the first association compression vector and the first association denoising vector of the second level can be calculated to obtain the corresponding difference vector. Then, the dispersion of this difference vector can be calculated, and this dispersion can be determined as the first mining index. Similarly, the difference between the second association compression vector and the second association denoising vector of the second level can be calculated to obtain the corresponding difference vector. Then, the dispersion of this difference vector can be calculated, and this dispersion can be determined as the second mining index.
[0097] Step S134b: The first mining index and the second mining index are fused to form a target mining index, and the aggregation weight parameters of the semantic interaction denoising vector at the corresponding level are determined based on the target mining index.
[0098] In this embodiment, after obtaining the first mining index and the second mining index, the first mining index and the second mining index can be fused to form a target mining index, and the aggregation weight parameters of the semantic interaction denoising vector at the corresponding level can be determined based on the target mining index. For example, the first mining index and the second mining index can be fused using methods such as mean or weighted mean to obtain the target mining index. Based on this, after obtaining the target mining index at each level, normalization processing can be performed to obtain the aggregation weight parameters.
[0099] Step S134c: Based on the aggregation weight parameters of the semantic interaction denoising vectors at each level, the semantic interaction denoising vectors at each level are aggregated to form the target denoised semantic vector.
[0100] In this embodiment, after obtaining the aggregation weight parameters, the semantic interaction denoising vectors of each level can be aggregated (e.g., weighted summation) based on the aggregation weight parameters of the semantic interaction denoising vectors of each level to form the target denoised semantic vector. It should be noted that if the dimensions of the semantic interaction denoising vectors at different levels are different, upsampling or downsampling operations can be performed first to unify the dimensions before aggregation.
[0101] Fourthly, regarding step S140, it should be noted that the specific method for semantic restoration of the target denoised semantic vector is not limited and can be selected according to actual needs.
[0102] For example, in an alternative implementation, in order to ensure the reliability of semantic restoration and make the reliability of the obtained second denoised data higher, the above step S140 may further include steps S141, S142, S143, S144 and S145, wherein the specific contents of each step are as follows.
[0103] Step S141: Perform a fully connected mapping on the target denoised semantic vector to form a fully connected denoised semantic vector.
[0104] In this embodiment of the application, the target denoised semantic vector can be fully connected to form a fully connected denoised semantic vector. For example, through fully connected mapping, a fully connected denoised semantic vector of size 1*32768 can be obtained.
[0105] Step S142: Perform shape transformation on the fully connected denoised semantic vector to form a denoised semantic feature map.
[0106] In this embodiment, the fully connected denoised semantic vector can be reshaped to form a denoised semantic feature map. The reshape transformation does not change the content of the semantic information; for example, a fully connected denoised semantic vector of size 1*32768 can be transformed into a denoised semantic feature map of size 16*16*128.
[0107] Step S143: Perform multiple levels of transposed convolution processing on the denoised semantic feature map to obtain the transposed convolution semantic vector of the last level.
[0108] In this embodiment, after obtaining the denoised semantic feature map, multiple levels of transposed convolution processing can be performed on the denoised semantic feature map sequentially to obtain the transposed convolution semantic vector of the last level. Specifically, the transposed convolution semantic vector of the next level is obtained by transposing the transposed convolution semantic vector of the previous level; that is, the low-resolution feature map can be progressively upsampled to gradually restore the original data size. For example: The parameters for the first layer of transpose convolution processing can be: Conv2DTranspose + ReLU, 64 kernels, size 3*3, stride=2, and the size of the output transpose convolution semantic vector is 32*32*64. The parameters for the second-level transposed convolution processing can be: Conv2DTranspose + ReLU, 32 kernels, size 3*3, stride=2, and the size of the output transposed convolution semantic vector is 64*64*32. The parameters for the third-level transposed convolution processing can be: Conv2DTranspose, 1 kernel, size 3*3, stride=2, and the size of the output transposed convolution semantic vector is 128*128*1.
[0109] Step S144: Activate the transposed convolutional semantic vector of the last layer to form the target spectrogram.
[0110] In this embodiment, after obtaining the transposed convolutional semantic vector of the last layer, the transposed convolutional semantic vector of the last layer can be activated and output to form the target spectrogram. For example, the activation output can be achieved through identity mapping or mapping functions such as sigmoid.
[0111] Step S145: Perform an inverse Fourier transform on the target spectrogram to form the second denoised data.
[0112] In this embodiment of the application, after obtaining the target spectrum, an inverse Fourier transform can be performed on the target spectrum to form second denoised data. That is, the target spectrum in the frequency domain can be transformed into the time domain to obtain the second denoised data.
[0113] Fifthly, regarding step S150, it should be noted that the specific method of fusing the first denoised data and the second denoised data is not limited and can be selected according to actual needs.
[0114] For example, in an alternative implementation, the first denoised data and the second denoised data can be averaged or weighted averaged to obtain the corresponding target denoised data.
[0115] In summary, the denoising method and system based on digital filtering and deep learning provided in this application firstly performs digital filtering on the data to be denoised to form first denoised data; secondly, latent semantic mining is performed on the data to be denoised and the first denoised data respectively to form a semantic vector to be denoised and a first denoised semantic vector; then, based on the first denoised semantic vector, semantic constraints are applied to the deep denoising process of the semantic vector to be denoised to form a target denoised semantic vector; further, semantic restoration is performed on the target denoised semantic vector to form second denoised data; finally, the first denoised data and the second denoised data are fused to form the target denoised data. Based on the above, since the data to be denoised is first digitally filtered, the semantic vector of the first denoised data can be used to impose semantic constraints on the deep denoising process of the corresponding semantic vector of the data to be denoised. This ensures that the resulting target denoised semantic vector can balance both the richness of semantic information (if deep learning-based latent semantic mining is directly performed on the first denoised data formed by digital filtering, some effective semantic information may be lost due to digital filtering) and the accuracy of semantic information (without the semantic constraints of the semantic vector of the first denoised data, the denoising effect may be relatively low). Therefore, the reliability of the second denoised data formed based on the target denoised semantic vector can be higher, thus fully guaranteeing the reliability of the resulting target denoised data. This achieves an effective combination of deep learning and digital filtering, thereby improving the problem of poor denoising effect in existing technologies.
[0116] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A denoising method based on digital filtering and deep learning, characterized in that, include: The data to be denoised is digitally filtered to form the first denoised data, wherein the data to be denoised belongs to the time domain or the spatial domain. Latent semantic mining is performed on the data to be denoised and the first denoised data respectively to form a semantic vector to be denoised and a first denoised semantic vector. Based on the first denoised semantic vector, semantic constraints are applied to the deep denoising process of the semantic vector to be denoised, forming a target denoised semantic vector; The target denoised semantic vector is semantically restored to form the second denoised data. Latent semantic mining, deep denoising, semantic constraints and semantic restoration are implemented through the target denoising model, and the target denoising model belongs to the neural network model formed by deep learning. The first denoised data and the second denoised data are merged to form the target denoised data.
2. The denoising method based on digital filtering and deep learning according to claim 1, characterized in that, The step of semantically constraining the deep denoising process of the semantic vector to be denoised based on the first denoised semantic vector to form the target denoised semantic vector includes: Based on the first denoised semantic vector, the semantic vector to be denoised is subjected to first association mining at multiple levels to achieve deep denoising and obtain first association denoised vectors at multiple levels. The semantic vector to be denoised is subjected to multiple levels of second association mining to achieve deep denoising and obtain multiple levels of second association denoising vectors. According to the corresponding levels, the first associated denoising vectors and the second associated denoising vectors of the multiple levels are semantically interacted to achieve semantic constraints and form semantically interacted denoising vectors of multiple levels, wherein the semantic interaction includes gating adjustment. The semantic interaction denoising vectors of the multiple levels are semantically aggregated to form the target denoised semantic vector.
3. The denoising method based on digital filtering and deep learning according to claim 2, characterized in that, The step of performing multi-level first association mining on the semantic vector to be denoised based on the first denoised semantic vector to achieve deep denoising and obtain multi-level first association denoised vectors includes: Based on the first denoised semantic vector, cross-attention processing is performed on the semantic vector to be denoised to form the first-level first associated denoised vector; Pooling compression is performed on the first denoised semantic vector and the first association denoised vector of the first level respectively to form the first denoised compressed vector of the second level and the first association compressed vector of the second level. Based on the first denoised compressed vector, cross attention processing is performed on the first association compressed vector to form the first association denoised vector of the second level. Pooling compression is performed on the first denoised compression vector and the first associated denoised vector of the previous level to form the first denoised compression vector and the first associated compression vector of the current level. Based on the first denoised compression vector, cross-attention processing is performed on the first associated compression vector to form the first associated denoised vector of the current level.
4. The denoising method based on digital filtering and deep learning according to claim 2, characterized in that, The step of performing multi-level second association mining on the semantic vector to be denoised to achieve deep denoising and obtain multi-level second association denoised vectors includes: The semantic vector to be denoised is subjected to self-attention processing to form the second associated denoising vector of the first level; Pooling compression is performed on the semantic vector to be denoised and the second association denoising vector of the first level respectively to form the second denoising compressed vector of the second level and the second association compressed vector of the second level. Based on the second denoising compressed vector, cross attention processing is performed on the second association compressed vector to form the second association denoising vector of the second level. Pooling compression is performed on the second denoised compression vector and the second associated denoised vector of the previous level to form the second denoised compression vector and the second associated compression vector of the current level. Based on the second denoised compression vector, cross-attention processing is performed on the second associated compression vector to form the second associated denoised vector of the current level.
5. The denoising method based on digital filtering and deep learning according to claim 2, characterized in that, The step of semantically aggregating the semantic interaction denoising vectors at multiple levels to form the target denoised semantic vector includes: For each level of semantic interaction denoising vector, determine the first mining index and the second mining index of the first association denoising vector and the second association denoising vector corresponding to the semantic interaction denoising vector in the corresponding first association mining and second association mining. The first association mining and the second association mining are implemented based on the attention mechanism. The first mining index and the second mining index are used to reflect the degree of difference between the vectors before and after attention mining. The first mining index and the second mining index are combined to form a target mining index, and the aggregation weight parameters of the semantic interaction denoising vector at the corresponding level are determined based on the target mining index. Based on the aggregation weight parameters of the semantic interaction denoising vectors at each level, the semantic interaction denoising vectors at each level are aggregated to form the target denoised semantic vector.
6. The denoising method based on digital filtering and deep learning according to claim 1, characterized in that, The step of performing latent semantic mining on the data to be denoised and the first denoised data respectively to form a semantic vector to be denoised and a semantic vector to be denoised, includes: Fourier transform is performed on the data to be denoised and the first denoised data respectively to form the corresponding spectrum diagram to be denoised and the first denoised spectrum diagram; The latent semantics of the spectrum to be denoised are mined to form a semantic vector to be denoised. Latent semantic mining is performed on the first denoised spectrogram to form a first denoised semantic vector.
7. The denoising method based on digital filtering and deep learning according to claim 6, characterized in that, The step of performing latent semantic mining on the spectrogram to be denoised to form a semantic vector to be denoised includes: The spectrum image to be denoised is convolved to form a convolution vector to be denoised; The spectrum to be denoised is divided into multiple local spectrums to be denoised according to frequency bands, and each local spectrum to be denoised is convolved to form multiple local convolution vectors to be denoised corresponding to the multiple local spectrums to be denoised. Each of the local convolutional vectors to be denoised is used to perform association mining on the local convolutional vectors to be denoised, thereby forming multiple association vectors to be denoised corresponding to the multiple local convolutional vectors to be denoised. The association mining is used to extract potential semantic information that has an association relationship with the local convolutional vectors to be denoised from the local convolutional vectors to be denoised. The multiple denoised related vectors are fused to form a denoised semantic vector.
8. The denoising method based on digital filtering and deep learning according to claim 6, characterized in that, The step of performing latent semantic mining on the first denoised spectrogram to form a first denoised semantic vector includes: The first denoised spectrogram is convolved to form a first denoised convolution vector; The first denoised spectrum is divided into multiple first denoised local spectrums according to frequency bands, and each of the first denoised local spectrums is convolved to form multiple first denoised local convolution vectors corresponding to the multiple first denoised local spectrums. Each of the first denoised local convolution vectors is used to perform association mining on the first denoised convolution vector to form multiple first denoised association vectors corresponding to the multiple first denoised local convolution vectors. The association mining is used to extract potential semantic information that has an association relationship with the first denoised local convolution vector from the first denoised convolution vector. The multiple first denoised correlation vectors are fused to form a first denoised semantic vector.
9. The denoising method based on digital filtering and deep learning according to any one of claims 1-8, characterized in that, The step of semantically restoring the target denoised semantic vector to form the second denoised data includes: Perform a fully connected mapping on the target denoised semantic vector to form a fully connected denoised semantic vector; The shape of the fully connected denoised semantic vector is transformed to form a denoised semantic feature map; The denoised semantic feature map is subjected to multiple levels of transposed convolution processing to obtain the transposed convolution semantic vector of the last level. The transposed convolution semantic vector of the next level is obtained by transposing the transposed convolution semantic vector of the previous level. The transposed convolutional semantic vector of the last layer is activated and output to form the target spectrogram; The target spectrogram is subjected to inverse Fourier transform to form the second denoised data.
10. A denoising system based on digital filtering and deep learning, characterized in that, include: Memory, used to store computer programs; A processor connected to the memory is used to execute a computer program stored in the memory to implement the denoising method based on digital filtering and deep learning as described in any one of claims 1-9.