Method, apparatus, electronic device and medium for noise removal
By constructing a codec model for deep residual shrinking network and feature reduction network based on the target soft thresholding function, the threshold is automatically set, and the adaptation problem of soft thresholding function in noise removal is solved, achieving higher noise fault diagnosis accuracy and speech recognition accuracy.
Patent Information
- Application Number
- CN202110639632.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-08
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-06-08
AI Technical Summary
In the prior art, how to use soft thresholding functions to perform noise removal more accurately is a challenge, especially in the field of signal noise reduction, the setting of thresholds is difficult to adapt uniformly.
The codec model composed of a deep residual shrinking network and feature reduction network generated based on the target soft thresholding function is adopted. Through training and optimization, thresholds are automatically set to remove noise, including convolutional operations, batch regular normalization, global average pooling and linear rectifying functional operations, forming an automatic codec.
It improves the accuracy and accuracy of noise fault diagnosis, effectively removes noise in voice data, and improves the accuracy of voice recognition.
Smart Images

Figure CN115457946B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to data processing technologies, in particular to a method, device, electronic device and medium for noise removal. Background Art
[0002] With the development of technology, the proportion of high-quality products in the production of the whole country is increasing. Whether it is production products or living products, as the use of products increases, the products will inevitably show wear and tear. Therefore, an accurate fault diagnosis system directly determines the quality of the products.
[0003] Furthermore, in related technologies, judging the operating sound of many products is an important method for their fault diagnosis. Currently, there is usually a method of denoising using soft thresholding in related technologies. Among them, the soft thresholding function (SoftThreshlding), as a classic method, is very practical especially in the field of signal denoising. The natural non-linear property of the soft threshold is very suitable for use in the calculation and conduction process of deep neural networks.
[0004] However, how to use the soft thresholding function more precisely for noise removal has become a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0005] The embodiments of the present application provide a method, device, electronic device and medium for noise removal, and the embodiments of the present application are used to solve the problem in related technologies of how to generate a model using a soft thresholding function for noise recognition.
[0006] Among them, according to one aspect of the embodiments of the present application, a method for noise removal is provided, which is characterized by including:
[0007] Obtain the speech data to be recognized;
[0008] Input the speech data to be recognized into a preset encoding and decoding model to obtain a noise recognition result, where the encoding and decoding model is composed of a target deep residual shrinkage network and a feature restoration network generated based on a target soft thresholding function;
[0009] Based on the noise recognition result, remove the noise data in the speech data to be recognized.
[0010] Optionally, in another embodiment based on the above method of the present application, before obtaining the speech data to be recognized, it further includes:
[0011] Obtain at least one noise-free speech data;
[0012] Add a noise signal to each noise-free speech data to obtain corresponding noisy speech data;
[0013] Training the target deep residual shrinkage network and the feature restoration network by using the at least one noise-free speech data and the at least one corresponding noisy speech data until the codec model that meets the preset training conditions is obtained.
[0014] Optionally, in another embodiment based on the above method of the present application, obtaining the codec model that meets the preset training conditions includes:
[0015] Taking the target deep residual shrinkage network as the encoding module in the codec model; and taking the feature restoration network as the decoding module in the codec model.
[0016] Optionally, in another embodiment based on the above method of the present application, obtaining the codec model that meets the preset training conditions includes:
[0017] Inputting the noisy speech data into the target deep residual shrinkage network to obtain an encoding end output result;
[0018] Inputting the encoding end output result into the feature restoration network to obtain a decoding end output result;
[0019] Constructing a loss function by using the decoding end output result and the corresponding noise-free speech data;
[0020] Detecting that the loss function meets the preset conditions, and determining that the codec model is trained to meet the preset training conditions.
[0021] Optionally, in another embodiment based on the above method of the present application, the loss function is obtained by the following formula:
[0022]
[0023] where N corresponds to the number of the noise-free speech data, x i corresponds to the i-th component of the noise-free speech data, and y i corresponds to the i-th component of the decoding end output result.
[0024] Optionally, in another embodiment based on the above method of the present application, the inputting the noisy speech data into the target deep residual shrinkage network to obtain an encoding end output result includes:
[0025] Inputting the noisy speech data into the target deep residual shrinkage network to obtain a fourth output result with a dimension of CxWx1;
[0026] Using the target soft thresholding function to remove the noise redundancy of the fourth output result to obtain a fifth output result;
[0027] Perform batch normalization on the fifth output result, perform global average pooling operation, and perform rectified linear unit operation to obtain the output result of the encoding end.
[0028] Optionally, in another embodiment based on the method of the present application, after obtaining the output result of the encoding end, it further includes:
[0029] Use the feature restoration network to perform deconvolution operation on the output result of the encoding end to obtain the first restored feature;
[0030] Perform multiple convolution operations on the first restored feature; and perform multiple deconvolution operations on the first restored feature until the output result of the decoding end with the dimension of CxWx1 is obtained.
[0031] Optionally, in another embodiment based on the method of the present application, before obtaining the speech data to be recognized, it further includes:
[0032] Obtain the first input feature, perform convolution operation at least twice on the first input feature, perform batch normalization operation, and perform rectified linear unit operation to obtain the first output result;
[0033] Perform absolute value algorithm on the first output result, and perform global average pooling operation to obtain the second output result;
[0034] Based on the second output result, obtain the target soft thresholding function.
[0035] Optionally, in another embodiment based on the method of the present application, obtaining the target soft thresholding function based on the second output result includes:
[0036] Perform convolution operation on the second output result, perform batch normalization operation, perform fully connected operation, and perform rectified linear unit operation to obtain the third output result;
[0037] Perform sigmoid function on the third output result to obtain the target soft thresholding function.
[0038] According to another aspect of the embodiments of the present application, a noise removal device is provided, including:
[0039] A determination module, configured to obtain the speech data to be recognized;
[0040] A generation module, configured to input the voice data to be recognized into a preset encoding and decoding model to obtain a noise recognition result, where the encoding and decoding model is composed of a target deep residual shrinkage network generated based on a target soft thresholding function and a feature reduction network;
[0041] A clearing module, configured to remove the noise data in the voice data to be recognized based on the noise recognition result.
[0042] According to another aspect of the embodiments of the present application, an electronic device is provided, including:
[0043] A memory, configured to store executable instructions; and
[0044] A display, configured to display with the memory to execute the executable instructions so as to complete the operations of the noise removal method described in any one of the above.
[0045] According to still another aspect of the embodiments of the present application, a computer-readable storage medium is provided, configured to store computer-readable instructions, and when the instructions are executed, the operations of the noise removal method described in any one of the above are performed.
[0046] In the present application, voice data to be recognized is obtained; the voice data to be recognized is input into a preset encoding and decoding model to obtain a noise recognition result, where the encoding and decoding model is composed of a target deep residual shrinkage network generated based on a target soft thresholding function and a feature reduction network; based on the noise recognition result, the noise data in the voice data to be recognized is removed. By applying the technical solution of the present application, the noise data and noise-free data can be used to train the deep residual network generated by the target soft thresholding function and the feature reduction network until an encoder network that meets the training conditions is obtained. So that the encoder network can be deployed on a noise recognition device to achieve the purpose of noise removal.
[0047] Next, through the drawings and embodiments, the technical solution of the present application will be further described in detail. Description of the Drawings
[0048] The drawings constituting a part of the specification depict the embodiments of the present application and, together with the description, are used to explain the principles of the present application.
[0049] Referring to the drawings, the present application can be more clearly understood according to the following detailed description, where:
[0050] Figure 1 It is a schematic diagram of the noise removal method proposed by the present application;
[0051] Figures 2-3 It is a schematic diagram of the method for generating the target soft thresholding function proposed by the present application;
[0052] Figure 4 Schematic diagram of the principle of the encoding and decoding model proposed in this application;
[0053] Figure 5 Schematic diagram of the principle of the encoding end in the encoding and decoding model proposed in this application;
[0054] Figure 6 Schematic diagram of the principle of the decoding end in the encoding and decoding model proposed in this application;
[0055] Figure 7 Schematic diagram of the overall process of noise removal proposed in this application
[0056] Figure 8 Schematic diagram of the structure of the device for noise removal in this application;
[0057] Figure 9 Schematic diagram of the structure of the electronic device for noise removal in this application. Detailed implementation manners
[0058] Various exemplary embodiments of the present application will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present application.
[0059] Meanwhile, it should be understood that, for the sake of convenience of description, the sizes of the various parts shown in the drawings are not drawn in actual proportional relationship.
[0060] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way a limitation on the present application or its application or use.
[0061] Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods and devices should be regarded as part of the specification.
[0062] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0063] In addition, the technical solutions between the various embodiments of the present application can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.
[0064] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of the present application are only used to explain the relative positional relationship, movement conditions, etc. between components in a specific posture (as shown in the attached drawings). If the specific posture changes, the directional indications will also change accordingly.
[0065] Next, in conjunction with Figures 1-7 a method for noise removal according to an exemplary embodiment of the present application will be described. It should be noted that the following application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard. On the contrary, the embodiments of the present application can be applied to any applicable scenario.
[0066] Furthermore, the present application proposes a method, device, target terminal, and medium for noise removal.
[0067] Figure 1 A schematic flowchart of a method for noise removal according to an embodiment of the present application is schematically shown. As Figure 1 shown, the method includes:
[0068] S101, obtaining the speech data to be recognized.
[0069] Furthermore, with the development of technology, the proportion of mechanization and productization in the production of the whole country is also increasing, whether it is products for industry or products for daily life. High-quality products necessarily rely on high-quality production processes, which will inevitably promote the continuous improvement of mechanical production processes. And whether it is production tools or daily life products, as the use of products increases, the products will inevitably show wear and tear. Therefore, an accurate fault diagnosis system directly determines the quality of the products.
[0070] For example, in many mechanical manufacturing or some products in people's daily life, rotary bearings account for a considerable proportion. Whether in manufacturing or in the wear and tear of daily household appliances, the wear of bearings is also the most common phenomenon. Judging the rotation sound of bearings is an important method for fault diagnosis. However, usually, the rotation of bearings will be mixed with a large amount of noise and redundant signals both in production and in life. Therefore, directly judging the degree of damage of faults by sound will bring great errors.
[0071] Currently, the commonly used traditional methods adopt statistical learning methods to analyze sound signal noise. However, traditional methods usually have many parameters that need to be set manually. The setting itself is a very complex decision-making process, which usually requires a large number of statistical experiments to obtain. However, the different environments will determine different parameters. For example, the internal and external differences of the machine itself, the material of the bearing itself, and even the surrounding temperature and humidity may affect the setting of these hyperparameters. Therefore, it is very difficult to uniformly adapt these manually set parameters. Soft Threshlding, as a classic method, is very practical especially in the field of signal denoising. However, as stated before, the threshold in the soft thresholding function is a hyperparameter, and how to set a reasonable value is a very tricky problem.
[0072] In recent years, the popularization of the Internet of Things, big data, and mobile devices, especially the explosive development of deep learning, has made it possible to realize intelligent detection and recognition technologies based on deep learning. Different from traditional methods, deep learning methods can automatically learn the parameter features of perturbed signals and automatically derive the correct and appropriate parameters. Therefore, they have extremely high practical and application values. Moreover, the natural non-linear property of the soft threshold is very suitable for use in the calculation and conduction process of deep neural networks. As a classic deep learning network, the Deep Residual Network (ResNet) has been successfully applied to many fields. The network combining the deep residual network and the non-linear soft thresholding function, namely the Residual Shrinkage Network, has also proven its practicality in the field of signal denoising. The Residual Shrinkage Network adopts an attention mechanism (similar to the Squeeze-and-Excitation Network) to automatically set the threshold, avoiding the trouble of manually setting the threshold.
[0073] Furthermore, the Residual Shrinkage Network has a good performance in fault diagnosis with noise. The information that the network can extract is immune to noise. Therefore, this application can transform the original Residual Shrinkage Network into a feature extraction network, and then use deconvolution to construct a feature restoration network, thus forming a set of auto encode-decode. This network can effectively remove the noise in the original sound signal, thereby providing a clearer sound signal for the speech recognition module, which will greatly improve the accuracy and precision of speech recognition.
[0074] It should be noted that this application does not specifically limit the speech data to be recognized. In one way, it can be speech data with noisy audio. For this noisy audio, it can include a speech source signal and background noise. Among them, the background noise can be various different types of noise signals. For example, different environmental noises such as vehicle noise, industrial noise, wind noise, and sea wave noise may exist due to different environments where the service is located.
[0075] S102. Input the speech data to be recognized into a preset encoding and decoding model to obtain a noise recognition result. The encoding and decoding model is composed of a target deep residual shrinkage network generated based on a target soft thresholding function and a feature restoration network.
[0076] Furthermore, after obtaining the speech data to be recognized, this application can input this data into a pre-trained encoding and decoding model, and then obtain the corresponding noise recognition result.
[0077] Optionally, in this application, obtaining the target soft thresholding function further includes:
[0078] Obtain a first input feature, perform at least two convolutional operations on the first input feature, perform a batch normalization operation, and perform a rectified linear unit operation to obtain a first output result;
[0079] Perform an absolute value algorithm on the first output result and perform a global average pooling operation to obtain a second output result;
[0080] Based on the second output result, obtain the target soft thresholding function.
[0081] Furthermore, this application first explains the process of obtaining the target soft thresholding function as follows:
[0082] In the Residual Shrinkage Network, this application can automatically derive a threshold through an attention mechanism and construct a Soft Threshlding function based on this. The mathematical expression of this function is
[0083]
[0084] where τ is the threshold, and both x and y are real numbers, representing the input and output respectively. The shape of this function is as Figure 2 shown. Its derivative is
[0085]
[0086] Here τ is the threshold, and the shape of the derivative is as Figure 3 shown.
[0087] Furthermore, since the Soft Threshlding function has a value of 0 within the threshold range and remains a linear function with a slope of 1 outside the threshold, it can suppress noise interference within the threshold range. Additionally, the present application can combine the soft thresholding function with a Residual Shrinkage Network that combines it with a Resnet. By using an attention mechanism, the threshold can be automatically derived, achieving a good noise suppression function. We use this network to extract high-order semantic information for noise suppression and construct an auto encode-decode network. The entire network consists of an encoding module and a decoding module. The encoder is a feature extraction network based on the deep residual shrinkage network, and the decoder is a feature restoration network composed of a series of transposed convolutions. The structure of the entire autoencode-decode network is as Figure 4 shown.
[0088] S103. Based on the noise recognition result, remove the noise data from the speech data to be recognized.
[0089] Furthermore, after determining the noise recognition result, the present application can remove the noise redundancy in the speech data to be recognized by using the noise data in the speech data to be recognized. This can improve the response of the entire network model to external noise signals and enhance the accuracy and precision of noise fault diagnosis.
[0090] In the present application, obtain the speech data to be recognized; input the speech data to be recognized into a preset encoding and decoding model to obtain a noise recognition result. The encoding and decoding model is composed of a target deep residual shrinkage network generated based on a target soft thresholding function and a feature restoration network; based on the noise recognition result, remove the noise data from the speech data to be recognized. By applying the technical solution of the present application, the encoder network can be obtained by training the deep residual network generated by the target soft thresholding function and the feature restoration network with noise data and noise-free data until the training conditions are met. Then, the encoder network is deployed to a noise recognition device to achieve the purpose of noise removal.
[0091] Optionally, in a possible implementation manner of the present application, in S102 (obtaining the cooking data to be recognized when the target cooking device is running), it includes:
[0092] Before obtaining the speech data to be recognized, it further includes:
[0093] Obtain at least one noise-free speech data;
[0094] Add a noise signal to each noise-free speech data to obtain the corresponding noisy speech data;
[0095] Using the at least one noise-free speech data and the at least one corresponding noisy speech data to train the target deep residual shrinkage network and the feature restoration network until the codec model that meets the preset training conditions is obtained.
[0096] Furthermore, the present application can first obtain a set of noise-free signals, where the set includes at least one noise-free speech data. Then, noise needs to be added to each of the noise-free speech data to generate a corresponding set of multiple noisy speech data.
[0097] It can be understood that the number of noise-free speech data needs to be the same as that of the noisy speech data. In addition, the present application needs to mark each noise-free speech data and the corresponding noisy speech data, so as to be able to distinguish the corresponding relationship between each noise-free speech data and the noisy speech data subsequently.
[0098] Furthermore, the present application can randomly select a set of noise-free signals and the corresponding noisy signals, and input them into the codec network to obtain the output result of the codec network. So as to subsequently compare the output result with the corresponding noise-free speech data to form a loss function. And when it is determined that the loss function meets the preset conditions subsequently, it is determined that the codec model has been trained.
[0099] Optionally, in a possible implementation manner of the present application, obtaining the codec model that meets the preset training conditions includes:
[0100] Taking the target deep residual shrinkage network as the encoding module in the codec model; and taking the feature restoration network as the decoding module in the codec model.
[0101] Wherein, the loss function is obtained through the following formula:
[0102]
[0103] Wherein, N corresponds to the number of the noise-free speech data, x i corresponds to the i-th component of the noise-free speech data, and y i corresponds to the i-th component of the output result of the decoding end.
[0104] Further, the loss function in this application is used to guide the entire encoding and decoding network, making the final output continuously approach the noise-free signal. The data processing in this application can be as follows: The noise-free speech data is set as x; then a noise signal is added, and this noisy signal is set as x, and x is the input of the encoding and decoding network; the final output of the network is set as y. Since x and y are tensors with exactly the same dimensions, this application can use the method of point-by-point subtraction to construct the loss function. After training like this, the output of the network can continuously approximate the original noise-free signal.
[0105] Optionally, in a possible implementation manner of this application, obtaining the encoding and decoding model that meets the preset training conditions includes:
[0106] Inputting the noisy speech data into the target deep residual shrinkage network to obtain an output result at the encoding end;
[0107] Inputting the output result at the encoding end into the feature restoration network to obtain an output result at the decoding end;
[0108] Constructing a loss function using the output result at the decoding end and the corresponding noise-free speech data;
[0109] When it is detected that the loss function meets the preset conditions, it is determined that the encoding and decoding model has been trained to meet the preset training conditions.
[0110] Further, in this application, when inputting the noisy speech data into the target deep residual shrinkage network to obtain an output result at the encoding end, it includes:
[0111] Inputting the noisy speech data into the target deep residual shrinkage network to obtain a fourth output result with the dimension of CxWx1;
[0112] Using the target soft thresholding function to remove the noise redundancy of the fourth output result to obtain a fifth output result;
[0113] Performing batch normalization operation, global average pooling operation, and rectified linear unit function operation on the fifth output result to obtain the output result at the encoding end.
[0114] Further, as Figure 5 shown, this application can input the original noisy speech data into the target deep residual shrinkage network (i.e., the encoding end of the model), with the dimension of CxWx1. First, perform a convolution operation to extract semantic information once, and the dimension remains CxWx1. Additionally, pass through several of the previously obtained target soft thresholding functions (RSBU) to remove noise redundancy.
[0115] Finally, an output result of the encoding end can also be obtained through an operation of Batch Normalization + Relu (Rectified Linear Unit operation) + Global Average Pooling (GAP). It can be understood that since this output itself has been transplanted with noise interference through several soft thresholding functions (RSBU), it is very suitable as the feature output of the encoder.
[0116] Optionally, after obtaining the output result of the encoding end, the present application further includes:
[0117] Performing a deconvolution operation on the output result of the encoding end by using the feature restoration network to obtain a first restored feature;
[0118] Performing a plurality of convolution operations on the first restored feature; and performing a plurality of deconvolution operations on the first restored feature until a decoding end output result with a dimension of CxWx1 is obtained.
[0119] Furthermore, for the decoding end of the model, as Figure 6 shown, the specific process is as follows:
[0120] First, the present application can use the output result of the encoding end output by the encoding end as the network input of the decoder, and first perform a deconvolution operation to restore the feature once. And through the operations of convolution and deconvolution, the feature is continuously restored until finally a feature identical to the encoder input signal is obtained, that is, a decoding output result with a dimension of CxWx1.
[0121] Even further, as Figure 7 shown, the method for noise removal proposed by the present application is described:
[0122] 1: Obtain a noise-free speech data set.
[0123] 2: Add noise to the noise-free speech data to form a noisy speech data set.
[0124] 3: Randomly select a group of noise-free speech data and the corresponding noisy speech data.
[0125] 4: Input the randomly selected noisy speech data into the encoding and decoding model.
[0126] 5: Obtain the output result of the encoding and decoding model.
[0127] 6: Compare the output result of the encoding and decoding model with the noise-free speech data selected in step 3 to form a loss function.
[0128] 7: Determine whether the loss function satisfies the termination condition.
[0129] 8: If the termination condition is not met, go back to 3 and continue the iteration.
[0130] 9: If the termination condition is met, terminate the training of the encoding and decoding model.
[0131] 10: Save the trained encoding and decoding model.
[0132] 11: Deploy the trained encoding and decoding model to each voice product or module for subsequent use by the voice recognition module.
[0133] 12: End.
[0134] By applying the technical solution of the present application, the noise data and the noise-free data can be used to train the deep residual network and the feature reduction network generated by the target soft thresholding function until an encoder network that meets the training conditions is obtained. So as to deploy the encoder network to the noise recognition device to achieve the purpose of noise removal.
[0135] In another embodiment of the present application, as Figure 8 shown, the present application also provides a noise removal device. Among them, the device includes a determination module 201, a generation module 202, and a clearing module 203, where
[0136] The determination module 201 is configured to obtain the voice data to be recognized;
[0137] The generation module 202 is configured to input the voice data to be recognized into a preset encoding and decoding model to obtain a noise recognition result, and the encoding and decoding model is composed of a target deep residual shrinkage network and a feature reduction network generated based on a target soft thresholding function;
[0138] The clearing module 203 is configured to remove the noise data in the voice data to be recognized based on the noise recognition result.
[0139] In the present application, when obtaining the voice data to be recognized; inputting the voice data to be recognized into a preset encoding and decoding model to obtain a noise recognition result, and the encoding and decoding model is composed of a target deep residual shrinkage network and a feature reduction network generated based on a target soft thresholding function; removing the noise data in the voice data to be recognized based on the noise recognition result. By applying the technical solution of the present application, the noise data and the noise-free data can be used to train the deep residual network and the feature reduction network generated by the target soft thresholding function until an encoder network that meets the training conditions is obtained. So as to deploy the encoder network to the noise recognition device to achieve the purpose of noise removal.
[0140] In another embodiment of the present application, the determination module 201 further includes:
[0141] The determination module 201 is configured to obtain at least one piece of noise-free speech data;
[0142] The determination module 201 is configured to add a noise signal to each piece of noise-free speech data to obtain corresponding noisy speech data;
[0143] The determination module 201 is configured to use the at least one piece of noise-free speech data and the at least one corresponding piece of noisy speech data to train the target deep residual shrinkage network and train the feature restoration network until the codec model that meets the preset training conditions is obtained.
[0144] In another embodiment of the present application, the determination module 201 further includes:
[0145] The determination module 201 is configured to use the target deep residual shrinkage network as the encoding module in the codec model; and use the feature restoration network as the decoding module in the codec model.
[0146] In another embodiment of the present application, the determination module 201 further includes:
[0147] The determination module 201 is configured to input the noisy speech data into the target deep residual shrinkage network to obtain an encoding end output result;
[0148] The determination module 201 is configured to input the encoding end output result into the feature restoration network to obtain a decoding end output result;
[0149] The determination module 201 is configured to construct a loss function by using the decoding end output result and the corresponding noise-free speech data;
[0150] The determination module 201 is configured to determine that the codec model is trained to meet the preset training conditions when it is detected that the loss function meets the preset conditions.
[0151] In another embodiment of the present application, the loss function is obtained through the following formula:
[0152]
[0153] where N corresponds to the number of pieces of noise-free speech data, x i corresponds to the i-th component of the noise-free speech data, and y i corresponds to the i-th component of the decoding end output result.
[0154] In another embodiment of the present application, the determination module 201 further includes:
[0155] The determination module 201 is configured to input the noisy speech data into the target deep residual shrinkage network to obtain a fourth output result with a dimension of CxWx1;
[0156] The determination module 201 is configured to use the target soft thresholding function to remove the noise redundancy of the fourth output result to obtain a fifth output result;
[0157] The determination module 201 is configured to perform batch normalization operation, global average pooling operation, and rectified linear unit operation on the fifth output result to obtain the encoded end output result.
[0158] In another embodiment of the present application, the determination module 201 further includes:
[0159] The determination module 201 is configured to use the feature reduction network to perform a deconvolution operation on the encoded end output result to obtain a first reduced feature;
[0160] The determination module 201 is configured to perform multiple convolution operations on the first reduced feature; and perform multiple deconvolution operations on the first reduced feature until a decoded end output result with a dimension of CxWx1 is obtained.
[0161] In another embodiment of the present application, the determination module 201 further includes:
[0162] The determination module 201 is configured to obtain a first input feature, perform at least two convolution operations, batch normalization operation, and rectified linear unit operation on the first input feature to obtain a first output result;
[0163] The determination module 201 is configured to perform an absolute value algorithm and a global average pooling operation on the first output result to obtain a second output result;
[0164] The determination module 201 is configured to obtain the target soft thresholding function based on the second output result.
[0165] In another embodiment of the present application, the determination module 201 further includes:
[0166] The determination module 201 is configured to perform a convolution operation, batch normalization operation, fully connected operation, and rectified linear unit operation on the second output result to obtain a third output result;
[0167] The determination module 201 is configured to perform a sigmoid function on the third output result to obtain the target soft thresholding function.
[0168] Figure 9 It is a logic structure block diagram of an electronic device shown according to an exemplary embodiment. For example, the electronic device 300 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0169] Referring to Figure 9 , the electronic device 300 may include one or more of the following components: a processor 301 and a memory 302.
[0170] The processor 301 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 301 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 301 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 301 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 301 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0171] The memory 302 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 302 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 302 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 301 to implement the interactive special effect calibration method provided in the method embodiment of the present application.
[0172] In some embodiments, the electronic device 300 may further optionally include: a peripheral device interface 303 and at least one peripheral device. The processor 301, the memory 302, and the peripheral device interface 303 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 303 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 304, a touch display screen 305, a camera 306, an audio circuit 307, a positioning component 308, and a power supply 309.
[0173] The peripheral device interface 303 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 301 and the memory 302. In some embodiments, the processor 301, the memory 302, and the peripheral device interface 303 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 301, the memory 302, and the peripheral device interface 303 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0174] The radio frequency circuit 304 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 304 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 304 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 304 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency circuit 304 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: a metropolitan area network, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 304 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0175] The display screen 305 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 305 is a touch display screen, the display screen 305 also has the ability to collect touch signals on or above the surface of the display screen 305. The touch signals can be input as control signals to the processor 301 for processing. At this time, the display screen 305 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 305, which is provided on the front panel of the electronic device 300; in other embodiments, there may be at least two display screens 305, which are respectively provided on different surfaces of the electronic device 300 or are in a folded design; in still other embodiments, the display screen 305 may be a flexible display screen, which is provided on a curved surface or a folding surface of the electronic device 300. Even further, the display screen 305 can also be set as an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 305 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0176] The camera module 306 is used to capture images or videos. Optionally, the camera module 306 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are respectively any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fused shooting functions. In some embodiments, the camera module 306 may further include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. A two-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0177] The audio circuit 307 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 301 for processing, or input to the radio frequency circuit 304 to implement voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 300. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 301 or the radio frequency circuit 304 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 307 may further include a headphone jack.
[0178] The positioning component 308 is used to locate the current geographical location of the electronic device 300 to implement navigation or LBS (Location Based Service). The positioning component 308 may be a positioning component based on the GPS (Global Positioning System) of the United States, the Beidou system of China, the GLONASS system of Russia, or the Galileo system of the European Union.
[0179] The power supply 309 is used to supply power to each component in the electronic device 300. The power supply 309 may be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 309 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0180] In some embodiments, the electronic device 300 further includes one or more sensors 410. The one or more sensors 410 include but are not limited to: an acceleration sensor 411, a gyroscope sensor 412, a pressure sensor 413, a fingerprint sensor 414, an optical sensor 415, and a proximity sensor 416.
[0181] The acceleration sensor 411 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the electronic device 300. For example, the acceleration sensor 411 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 301 can control the touch display screen 305 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 411. The acceleration sensor 411 can also be used for game or collection of the user's motion data.
[0182] The gyroscope sensor 412 can detect the body orientation and rotation angle of the electronic device 300. The gyroscope sensor 412 can cooperate with the acceleration sensor 411 to collect the 3D actions of the user on the electronic device 300. Based on the data collected by the gyroscope sensor 412, the processor 301 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0183] The pressure sensor 413 can be disposed on the side frame of the electronic device 300 and / or the lower layer of the touch display screen 305. When the pressure sensor 413 is disposed on the side frame of the electronic device 300, it can detect the holding signal of the user on the electronic device 300, and the processor 301 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 413. When the pressure sensor 413 is disposed on the lower layer of the touch display screen 305, the processor 301 can control the operable controls on the UI interface according to the pressure operation of the user on the touch display screen 305. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0184] The fingerprint sensor 414 is used to collect the fingerprints of the user. The processor 301 can identify the user's identity according to the fingerprints collected by the fingerprint sensor 414, or the fingerprint sensor 414 can identify the user's identity according to the collected fingerprints. When the identified user identity is a trusted identity, the processor 301 authorizes the user to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 414 can be disposed on the front, back, or side of the electronic device 300. When there are physical buttons or manufacturer Logos on the electronic device 300, the fingerprint sensor 414 can be integrated with the physical buttons or manufacturer Logos.
[0185] The optical sensor 415 is used to collect the ambient light intensity. In one embodiment, the processor 301 can control the display brightness of the touch display screen 305 according to the ambient light intensity collected by the optical sensor 415. Specifically, when the ambient light intensity is high, the display brightness of the touch display screen 305 is increased; when the ambient light intensity is low, the display brightness of the touch display screen 305 is decreased. In another embodiment, the processor 301 can also dynamically adjust the shooting parameters of the camera module 306 according to the ambient light intensity collected by the optical sensor 415.
[0186] The proximity sensor 416, also known as the distance sensor, is usually disposed on the front panel of the electronic device 300. The proximity sensor 416 is used to collect the distance between the user and the front of the electronic device 300. In one embodiment, when the proximity sensor 416 detects that the distance between the user and the front of the electronic device 300 is gradually decreasing, the touch display screen 305 is controlled by the processor 301 to switch from the lit state to the off state; when the proximity sensor 416 detects that the distance between the user and the front of the electronic device 300 is gradually increasing, the touch display screen 305 is controlled by the processor 301 to switch from the off state to the lit state.
[0187] Those skilled in the art can understand that Figure 9 the structure shown in does not constitute a limitation on the electronic device 300, and may include more or fewer components than shown, or combine certain components, or adopt a different component arrangement.
[0188] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 304 including instructions. The above instructions can be executed by the processor 420 of the electronic device 300 to complete the above noise removal method. The method includes: obtaining the speech data to be recognized; inputting the speech data to be recognized into a preset codec model to obtain a noise recognition result. The codec model is composed of a target deep residual shrinkage network and a feature reduction network generated based on a target soft thresholding function; based on the noise recognition result, removing the noise data in the speech data to be recognized. Optionally, the above instructions can also be executed by the processor 420 of the electronic device 300 to complete other steps involved in the above exemplary embodiment. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0189] In an exemplary embodiment, an application program / computer program product is also provided, including one or more instructions. The one or more instructions can be executed by the processor 420 of the electronic device 300 to complete the above noise removal method. The method includes: executed by the processor 420 of the electronic device 300 to complete the above noise removal method. The method includes: obtaining the speech data to be recognized; inputting the speech data to be recognized into a preset codec model to obtain a noise recognition result. The codec model is composed of a target deep residual shrinkage network and a feature reduction network generated based on a target soft thresholding function; based on the noise recognition result, removing the noise data in the speech data to be recognized. Optionally, the above instructions can also be executed by the processor 420 of the electronic device 300 to complete other steps involved in the above exemplary embodiment.
[0190] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0191] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A method for noise removal, characterized in that, Including: Obtain the speech data to be recognized; Input the speech data to be recognized into a preset encoding and decoding model to obtain a noise recognition result, where the encoding and decoding model is composed of a target deep residual shrinkage network generated based on a target soft thresholding function and a feature restoration network; Based on the noise recognition result, remove the noise data in the speech data to be recognized; Before obtaining the speech data to be recognized, it further includes: obtaining at least one noise-free speech data; adding a noise signal to each noise-free speech data to obtain a corresponding noisy speech data; using the at least one noise-free speech data and the at least one corresponding noisy speech data to train the target deep residual shrinkage network and train the feature restoration network until the encoding and decoding model that meets the preset training conditions is obtained; using the target deep residual shrinkage network as the encoding module in the encoding and decoding model; and using the feature restoration network as the decoding module in the encoding and decoding model; Among them, obtaining the encoding and decoding model that meets the preset training conditions includes: inputting the noisy speech data into the target deep residual shrinkage network to obtain an encoding end output result; inputting the encoding end output result into the feature restoration network to obtain a decoding end output result; constructing a loss function using the decoding end output result and the corresponding noise-free speech data; detecting that the loss function meets the preset conditions, and determining that the encoding and decoding model is trained to meet the preset training conditions; The step of inputting the noisy speech data into the target deep residual shrinkage network to obtain an encoding end output result includes: inputting the noisy speech data into the target deep residual shrinkage network to obtain a fourth output result of CxWx1 dimension; using the target soft thresholding function to remove the noise redundancy of the fourth output result to obtain a fifth output result; performing a batch normalization operation, a global average pooling operation, and a rectified linear unit operation on the fifth output result to obtain the encoding end output result.
2. The method according to claim 1, wherein The loss function is obtained through the following formula: ; where N corresponds to the number of the noise-free speech data, corresponding to the i-th component of the noise-free speech data, corresponding to the i-th component of the output result at the decoding end.
3. The method according to claim 1, characterized in that, After obtaining the encoding end output result, it further includes: Using the feature restoration network to perform a deconvolution operation on the encoding end output result to obtain a first restored feature; Performing multiple convolution operations on the first restored feature; and performing multiple deconvolution operations on the first restored feature until a decoding end output result of CxWx1 dimension is obtained.
4. The method according to claim 1, wherein Before obtaining the speech data to be recognized, it further includes: Obtain a first input feature, and perform at least two convolution operations, a batch normalization operation, and a rectified linear unit operation on the first input feature to obtain a first output result; Perform an absolute value algorithm and a global average pooling operation on the first output result to obtain a second output result; Based on the second output result, obtain the target soft thresholding function.
5. The method according to claim 4, wherein Based on the second output result, obtaining the target soft thresholding function includes: Perform a convolution operation on the second output result, perform a batch normalization operation, perform a fully connected operation, and perform a rectified linear unit operation to obtain a third output result; Perform a sigmoid function on the third output result to obtain the target soft thresholding function.
6. A noise removal device, characterized in that, Comprising: A determination module, configured to obtain speech data to be recognized; A generation module, configured to input the speech data to be recognized into a preset encoding and decoding model to obtain a noise recognition result, where the encoding and decoding model is composed of a target deep residual shrinkage network and a feature restoration network generated based on the target soft thresholding function; A cleaning module, configured to remove noise data in the speech data to be recognized based on the noise recognition result; The determination module is further configured to: before obtaining the speech data to be recognized, further comprising: obtaining at least one noise-free speech data; adding a noise signal to each noise-free speech data to obtain a corresponding noisy speech data; using the at least one noise-free speech data and the at least one corresponding noisy speech data to train the target deep residual shrinkage network and train the feature restoration network until the encoding and decoding model that meets the preset training conditions is obtained; using the target deep residual shrinkage network as the encoding module in the encoding and decoding model; and using the feature restoration network as the decoding module in the encoding and decoding model; where obtaining the encoding and decoding model that meets the preset training conditions includes: inputting the noisy speech data into the target deep residual shrinkage network to obtain an encoding end output result; inputting the encoding end output result into the feature restoration network to obtain a decoding end output result; constructing a loss function using the decoding end output result and the corresponding noise-free speech data; detecting that the loss function meets the preset conditions, and determining that the encoding and decoding model is trained to meet the preset training conditions; inputting the noisy speech data into the target deep residual shrinkage network to obtain an encoding end output result, including: inputting the noisy speech data into the target deep residual shrinkage network to obtain a fourth output result with a dimension of CxWx1; using the target soft thresholding function to remove the noise redundancy of the fourth output result to obtain a fifth output result; performing a batch normalization operation, a global average pooling operation, and a rectified linear unit operation on the fifth output result to obtain the encoding end output result.
7. An electronic device, characterized in that, Comprising: A memory, for storing executable instructions; And, A processor, configured to display with the memory to execute the executable instructions to complete the operations of the noise removal method according to any one of claims 1-5.
8. A computer-readable storage medium for storing computer-readable instructions, characterized in that, When the instructions are executed, the operations of the noise removal method according to any one of claims 1-5 are executed.
Citation Information
Patent Citations
Combined model training method and system
CN109712611A
Universal steganalysis method and system of audio based on spectrograms and deep residual network
CN110120228A