Audio signal processing method, device, electronic device and storage medium
By performing feature extraction and howling suppression model processing on audio signals, the problem of low suppression efficiency of howling signal in the prior art is solved, efficient howling signal suppression is achieved, and the audio signal quality is improved.
Patent Information
- Application Number
- CN202310127570.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-02
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2043-02-02
AI Technical Summary
The existing audio signal processing methods are computationally large and inefficient when suppressing howling signals, making it difficult to adapt to the diversity of different songs and user sound signals.
By extracting the first audio signal and the second audio signal, the pre-trained howling suppression model is used for howling signal detection and suppression, the target audio signal is obtained by combining the feature reduction algorithm to simplify the model complexity and improve processing efficiency.
The calculation amount of the howling suppression model is simplified, the efficiency of howling signal suppression processing is improved, and the audio signal quality is ensured.
Smart Images

Figure CN116229998B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of communication technologies, and in particular to an audio signal processing method, device, electronic device, and storage medium. Background Art
[0002] With the development of audio technology and the entertainment needs of people's daily lives, more and more users are using clients to order songs, sing karaoke, and perform other entertainment activities. During the singing process, users can also use microphones to record the sound signals of singing, and use ear-return devices to mix the sound signals recorded by the microphones with the accompaniment music, and then feed the mixed audio signals back to the users, forming a positive feedback transmission path. In this positive feedback process, the mixed audio signals fed back by the ear-return device are easily recorded by the microphone again, so that the microphone signal recorded by the microphone contains the user's voice signal and the mixed audio signal recorded again (the mixed audio signal is also called a howling signal), thereby affecting the recording quality of the human voice signal. Therefore, a method for suppressing howling signals has emerged.
[0003] Current methods for suppressing howling signals typically perform frequency domain analysis on the microphone signal. These analysis then uses characteristics such as spectral flatness, autocorrelation, and peak harmonic power ratio to generate a detection result. Based on these detection results, notch suppression is then used to eliminate the howling signal from the recorded microphone signal.
[0004] However, in the current methods of suppressing howling, spectrum analysis is time-consuming and computationally intensive due to the diversity of sound signals produced by different songs and different users. Howling signal suppression is performed only based on the spectral characteristics of the sound signal, resulting in low processing efficiency. Summary of the Invention
[0005] The present disclosure provides an audio signal processing method, device, electronic device, and storage medium to at least address the problem of low audio signal processing efficiency in related technologies. The technical solutions of the present disclosure are as follows:
[0006] According to a first aspect of an embodiment of the present disclosure, there is provided a method for processing an audio signal, the method comprising:
[0007] During the audio signal recording process, a first audio signal and a second audio signal are obtained; the first audio signal includes a sound signal of a target object and a howling signal, and the second audio signal is a background audio signal corresponding to the sound signal;
[0008] performing feature extraction on the first audio signal and the second audio signal according to a preset feature extraction algorithm to obtain a first audio feature and a second audio feature;
[0009] Inputting the first audio feature and the second audio feature into a pre-trained howling suppression model for processing to obtain audio features after howling suppression processing;
[0010] The audio features after the howling suppression process are restored according to a preset feature restoration algorithm to obtain a target audio signal.
[0011] In an exemplary embodiment, the extracting features of the first audio signal and the second audio signal according to a preset feature extraction algorithm to obtain the first audio feature and the second audio feature includes:
[0012] Performing short-time Fourier transform processing on the first audio signal and the second audio signal to obtain a first converted signal and a second converted signal in the frequency domain after processing;
[0013] Sampling the first conversion signal and the second conversion signal in the frequency domain according to a preset sampling strategy to obtain a plurality of sampling point data corresponding to the first conversion signal and a plurality of sampling point data corresponding to the second conversion signal, respectively;
[0014] Compression and band division processing is performed on multiple sampling point data corresponding to the first conversion signal and multiple sampling points corresponding to the second conversion signal to obtain a first audio feature corresponding to the first audio signal and a second audio feature corresponding to the second audio signal.
[0015] In an exemplary embodiment, the howling suppression model includes a convolutional layer, a gated recurrent unit, a fully connected layer, and an activation layer. Inputting the first audio feature and the second audio feature into the pre-trained howling suppression model for processing to obtain the audio feature after howling suppression processing includes:
[0016] inputting the first audio feature and the second audio feature into the howling suppression model;
[0017] The first audio feature and the second audio feature are processed sequentially through a convolutional layer, a gated recurrent unit, a fully connected layer, and an activation layer in the howling suppression model, and the audio feature after the howling suppression processing is output.
[0018] In an exemplary embodiment, the restoring the audio features after the howling suppression processing according to a preset feature restoration algorithm to obtain a target audio signal includes:
[0019] Performing restoration processing on the audio features after the howling suppression processing to obtain a restored audio signal in the frequency domain;
[0020] According to an inverse short-time Fourier transform algorithm, an inverse transform is performed on the restored audio signal in the frequency domain to obtain a restored target audio signal in the time domain.
[0021] In an exemplary embodiment, during the audio signal recording process, obtaining the first audio signal and the second audio signal includes:
[0022] During the audio signal collection process, the audio signal collection component collects the sound signal and the howling signal of the target object singing, and obtains a first audio signal including the sound signal and the howling signal;
[0023] A background audio signal corresponding to the sound signal is obtained through a mixing feedback component as a second audio signal.
[0024] In an exemplary embodiment, during the audio signal recording process, before obtaining the first audio signal and the second audio signal, the method further includes:
[0025] Obtain an audio training sample; the audio training sample includes a third audio feature extracted from a third audio signal feature and a fourth audio feature extracted from a fourth audio signal feature; the third audio signal includes a sound signal of a target object and a howling signal, and the fourth audio signal is a background audio signal corresponding to the sound signal;
[0026] Performing model training on a preset howling suppression model according to the audio training sample, and outputting training audio features;
[0027] Performing feature restoration processing on the training audio features according to a preset feature restoration algorithm to obtain a training audio signal;
[0028] The loss calculation of the training audio signal is performed according to a preset standard audio signal until the loss result corresponding to the training audio signal meets a preset loss condition, determining that the howling suppression model training is completed; the standard audio signal is the sound signal of the target object.
[0029] In an exemplary embodiment, performing loss calculation on the training audio signal according to a preset standard audio signal until a loss result corresponding to the training audio signal satisfies a preset loss condition, determining that the howling suppression model training is completed, includes:
[0030] Compressing the training audio signal and the preset standard audio signal by a preset multiple to obtain a compressed training audio signal and a compressed standard audio signal;
[0031] determining an amplitude spectrum distance between each sampling point in the compressed training audio signal and the compressed standard audio signal, and using the amplitude spectrum distance as a loss result corresponding to the training audio signal;
[0032] When the loss result corresponding to the training audio signal meets a preset loss condition, it is determined that the howling suppression model training is completed.
[0033] According to a second aspect of an embodiment of the present disclosure, there is provided an audio signal processing apparatus, the apparatus comprising:
[0034] An acquisition unit is configured to acquire, during the audio signal recording process, a first audio signal and a second audio signal; the first audio signal includes a sound signal of a target object and a howling signal; and the second audio signal is a background audio signal corresponding to the sound signal;
[0035] a feature extraction unit configured to perform feature extraction on the first audio signal and the second audio signal according to a preset feature extraction algorithm to obtain a first audio feature and a second audio feature;
[0036] a processing unit configured to input the first audio feature and the second audio feature into a pre-trained howling suppression model for processing to obtain audio features after howling suppression processing;
[0037] The restoration unit is configured to execute a preset feature restoration algorithm to restore the audio features after the howling suppression process to obtain a target audio signal.
[0038] In an exemplary embodiment, the feature extraction unit includes:
[0039] a first processing subunit, configured to perform short-time Fourier transform processing on the first audio signal and the second audio signal to obtain a first converted signal and a second converted signal in the frequency domain after processing;
[0040] a sampling subunit, configured to sample the first conversion signal and the second conversion signal in the frequency domain according to a preset sampling strategy, and obtain a plurality of sampling point data corresponding to the first conversion signal and a plurality of sampling point data corresponding to the second conversion signal, respectively;
[0041] The second processing sub-unit is configured to perform compression and banding processing on multiple sampling point data corresponding to the first conversion signal and multiple sampling points corresponding to the second conversion signal, respectively, to obtain a first audio feature corresponding to the first audio signal and a second audio feature corresponding to the second audio signal.
[0042] In an exemplary embodiment, the howling suppression model includes a convolutional layer, a gated recurrent unit, a fully connected layer, and an activation layer, and the processing unit includes:
[0043] an input subunit, configured to input the first audio feature and the second audio feature into the howling suppression model;
[0044] The output subunit is configured to process the first audio feature and the second audio feature in sequence through the convolution layer, the gated recurrent unit, the fully connected layer, and the activation layer in the howling suppression model, and output the audio feature after the howling suppression processing.
[0045] In an exemplary embodiment, the reduction unit comprises:
[0046] a conversion subunit, configured to perform restoration processing on the audio features after the howling suppression processing to obtain a restored audio signal in the frequency domain;
[0047] The restoration subunit is configured to perform an inverse transformation on the restored audio signal in the frequency domain according to an inverse short-time Fourier transform algorithm to obtain a restored target audio signal in the time domain.
[0048] In an exemplary embodiment, the acquiring unit includes:
[0049] The first acquisition subunit is configured to collect the sound signal and the howling signal of the target object singing through the audio signal collecting component during the audio signal collection process, and obtain a first audio signal including the sound signal and the howling signal;
[0050] The second acquisition subunit is configured to acquire a background audio signal corresponding to the sound signal as a second audio signal through a mixing feedback component.
[0051] In an exemplary embodiment, the apparatus further comprises:
[0052] a training acquisition unit configured to acquire an audio training sample; the audio training sample includes a third audio feature extracted from a third audio signal feature and a fourth audio feature extracted from a fourth audio signal feature; the third audio signal includes a sound signal of a target object and a howling signal, and the fourth audio signal is a background audio signal corresponding to the sound signal;
[0053] A model training unit is configured to perform model training on a preset howling suppression model according to the audio training sample and output training audio features;
[0054] a training restoration unit configured to perform feature restoration processing on the training audio features according to a preset feature restoration algorithm to obtain a training audio signal;
[0055] The training discrimination unit is configured to perform loss calculation on the training audio signal according to a preset standard audio signal until the loss result corresponding to the training audio signal meets a preset loss condition, thereby determining that the howling suppression model training is completed; the standard audio signal is the sound signal of the target object.
[0056] In an exemplary embodiment, the training discrimination unit includes:
[0057] a loss calculation subunit configured to compress the training audio signal and a preset standard audio signal by a preset multiple to obtain a compressed training audio signal and a compressed standard audio signal; determine an amplitude spectrum distance between each sampling point in the compressed training audio signal and the compressed standard audio signal, and use the amplitude spectrum distance as a loss result corresponding to the training audio signal;
[0058] The determining subunit is configured to determine that the howling suppression model training is completed when the loss result corresponding to the training audio signal meets a preset loss condition.
[0059] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:
[0060] processor;
[0061] a memory for storing instructions executable by the processor;
[0062] The processor is configured to execute the instructions to implement the audio signal processing and display method as described in any one of the first aspects above.
[0063] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the audio signal processing method as described in any one of the first aspects above.
[0064] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided. When the instructions are executed by a processor of an electronic device, the electronic device is capable of executing the audio signal processing method described in any one of the first aspects above.
[0065] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects:
[0066] This method simplifies the complexity of the howling suppression model and reduces the amount of model calculation by pre-extracting features from the first and second audio signals and restoring the audio features output by the model after howling suppression processing. Furthermore, howling signals in the first audio signal are suppressed using the pre-trained howling suppression model, thereby improving the efficiency of howling signal suppression processing.
[0067] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0069] Figure 1 The figure is a flowchart of a method for processing an audio signal according to an exemplary embodiment.
[0070] Figure 2 The figure is a flowchart of a method for extracting audio signal features according to an exemplary embodiment.
[0071] Figure 3 The figure is a schematic diagram showing the internal structure of a howling suppression model according to an exemplary embodiment.
[0072] Figure 4 The figure is a flowchart showing steps of howling suppression processing of a howling suppression model according to an exemplary embodiment.
[0073] Figure 5 The figure is a flowchart of an audio feature restoration method according to an exemplary embodiment.
[0074] Figure 6 The figure is a flowchart of a method for obtaining an audio signal according to an exemplary embodiment.
[0075] Figure 7 The figure is a flowchart of a howling suppression model training method according to an exemplary embodiment.
[0076] Figure 8 The figure is a flowchart showing a method for determining loss of a howling suppression model according to an exemplary embodiment.
[0077] Figure 9 The figure is a block diagram of an audio signal processing apparatus according to an exemplary embodiment.
[0078] Figure 10 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0079] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0080] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.
[0081] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0082] Figure 1 This is a flowchart of an audio signal processing method according to an exemplary embodiment. The audio signal processing method can be applied to electronic devices, for example, earphones, wireless headphones, etc. The embodiment of the present disclosure does not limit the type of specific execution device of the audio data triggering method. The embodiment of the present disclosure uses electronic devices as a general term to describe the technical solution. Figure 1 As shown, the audio signal processing method includes the following steps.
[0083] In step S110 , during the audio signal recording process, a first audio signal and a second audio signal are acquired.
[0084] The first audio signal includes the target object's sound signal and the howling signal, and the second audio signal is a background audio signal corresponding to the sound signal. The background audio signal can be background music used to accompany the target object's sound signal, or other types of audio signals, which are not limited in the present embodiment.
[0085] In implementation, during the audio signal recording process, the electronic device obtains the first audio signal and the second audio signal.
[0086] Specifically, when the target object is singing a song, the target object's voice can be collected by the microphone and converted into a sound signal corresponding to the target object (a sound signal in a preset microphone format). Then, the sound signal is mixed with the background music in an ear return device (which can be an electronic device that executes the audio signal processing method disclosed herein). The ear return device then feeds back the mixture of the sound signal and the background music to the target object, so that the target object can hear the sound signal with the accompaniment of background music, thereby forming a positive feedback process of the sound signal. However, in this positive feedback process of the sound signal, when the ear return device feeds back the mixture to the target object, the microphone can collect not only the target object sound signal, but also other signals, such as other background noise signals generated in the environment and the mixed signal collected by the microphone when the ear return feedback mixes, etc., so that the microphone signal collected by the microphone (i.e., the first audio signal) contains the sound signal of the target object and the howling signal. The howling signal contained in the first audio signal affects the sound clarity of the subsequent mixing of the first audio signal and the background music. In particular, the background music (i.e., the background audio signal) contained in the mixing of the howling signal has a relatively high volume, which greatly interferes with the sound signal of the target object, thereby affecting the effect of the target object singing the song. Therefore, it is necessary to perform howling suppression processing on the first audio signal containing the target object sound signal and the howling signal. In order to obtain an audio signal after howling suppression processing. In order to suppress the howling signal in the first audio signal, the electronic device also collects background music as a second audio signal, which is used to form a reference for comparison with the first audio signal, so as to suppress the background audio signal contained in the first audio signal.
[0087] In step S120, feature extraction is performed on the first audio signal and the second audio signal respectively according to a preset feature extraction algorithm to obtain a first audio feature and a second audio feature.
[0088] During implementation, the electronic device performs feature extraction on the first audio signal and the second audio signal, respectively, based on a preset feature extraction algorithm, to obtain the first audio feature and the second audio feature. Specifically, the electronic device converts the first audio signal and the second audio signal into the frequency domain based on the preset feature extraction algorithm, and then performs feature extraction on the first audio signal and the second audio signal in the converted frequency domain, to obtain the first audio feature corresponding to the feature of the first audio signal and the second audio feature corresponding to the second audio signal, respectively.
[0089] In step S130 , the first audio feature and the second audio feature are input into a pre-trained howling suppression model for processing to obtain audio features after howling suppression processing.
[0090] Among them, the pre-trained howling suppression model can perform howling signal detection and howling signal suppression on the input audio features containing howling signals. The howling suppression model only includes basic convolutional layers, gated recurrent units, fully connected layers and activation layers.
[0091] During implementation, the electronic device inputs the first audio feature and the second audio feature into a pre-trained howling suppression model for processing to obtain the audio feature after howling suppression processing. Specifically, the howling suppression model performs howling detection on each audio feature point in the first audio feature based on the reference correspondence between the first audio feature and the second audio feature during the audio feature processing process, and determines whether each audio feature point included in the first audio feature is an audio feature point of a howling signal. If the corresponding point of the first audio feature is an audio feature point of a howling signal, the audio feature point is eliminated; if the audio feature point is an audio feature point of a non-howling signal, the audio feature point is retained, thereby obtaining the audio feature after howling suppression processing.
[0092] In step S140 , the audio features after the howling suppression process are restored according to a preset feature restoration algorithm to obtain a target audio signal.
[0093] In implementation, the electronic device can restore the audio features after howling suppression processing output by the howling suppression model according to a preset feature restoration algorithm to obtain a target audio signal. The target audio signal is an audio signal that does not contain a howling signal and can be mixed and played.
[0094] Optionally, the preset feature restoration algorithm can be but is not limited to the inverse algorithm of the preset feature extraction algorithm. For example, if the preset feature extraction algorithm is the short-time Fourier algorithm, then the feature restoration algorithm is the inverse of the short-time Fourier algorithm. The embodiments of the present disclosure do not limit the specific feature extraction algorithm.
[0095] In the above-mentioned audio signal processing method, during the audio signal collection process, a first audio signal and a second audio signal are obtained, and features are extracted from the first audio signal and the second audio signal to obtain first audio features and second audio features. Then, with the second audio features as a reference, a howling suppression process is performed on the first audio features using a pre-trained howling suppression model to obtain an audio feature after howling suppression processing. The audio feature does not contain howling signal information. Thus, according to a preset feature restoration algorithm, the audio feature is restored to obtain a restored target audio signal. This method simplifies the complexity of the howling suppression model and reduces the amount of model calculation by pre-extracting features from the first audio signal and the second audio signal, and restoring the audio features after howling suppression processing output by the model. Furthermore, by using the pre-trained howling suppression model to suppress the howling signal in the first audio signal, the complex processing of frequency domain analysis and spectrum calculation of the audio signal is simplified, thereby improving the efficiency of howling signal suppression processing.
[0096] In an exemplary embodiment, Figure 2 As shown, in step S120, according to a preset feature extraction algorithm, feature extraction is performed on the first audio signal and the second audio signal respectively, and obtaining the first audio feature and the second audio feature can be specifically achieved by the following steps:
[0097] In step S211, short-time Fourier transform processing is performed on the first audio signal and the second audio signal to obtain a first converted signal and a second converted signal in the frequency domain after processing.
[0098] During implementation, the electronic device performs short-time Fourier transform processing on the first audio signal and the second audio signal to obtain first and second converted signals in the frequency domain. Specifically, the electronic device converts the audio data contained in the first and second audio signals from the time domain to the frequency domain based on a preset short-time Fourier transform algorithm, thereby determining the first and second converted signals in the frequency domain.
[0099] In step S212, the first conversion signal and the second conversion signal in the frequency domain are sampled according to a preset sampling strategy to obtain a plurality of sampling point data corresponding to the first conversion signal and a plurality of sampling point data corresponding to the second conversion signal.
[0100] During implementation, the electronic device samples the first converted signal and the second converted signal in the frequency domain according to a preset sampling strategy, obtaining multiple sampling point data corresponding to the first converted signal and multiple sampling point data corresponding to the second converted signal. Specifically, the electronic device samples the first converted signal and the second converted signal at a sampling rate of 48 kHz, with a frame length of 20 ms and a frame shift of 10 ms, obtaining 960 sampling point data (i.e., FFT, short-time Fourier transform data) corresponding to the first converted signal and 960 sampling point data corresponding to the second converted signal.
[0101] In step S213, compression and band division processing is performed on the multiple sampling point data corresponding to the first conversion signal and the multiple sampling points corresponding to the second conversion signal to obtain a first audio feature corresponding to the first audio signal and a second audio feature corresponding to the second audio signal.
[0102] In implementation, the electronic device performs compression and banding processing on multiple sampling point data corresponding to the first conversion signal and multiple sampling points corresponding to the second conversion signal, respectively, to obtain a first audio feature corresponding to the first audio signal and a second audio feature corresponding to the second audio signal. Specifically, the 960 sampling point data corresponding to the first conversion signal are halved, for example, the first 481 points are taken. The electronic device then performs compression and banding processing on the first 481 points, compressing them into 64 ERB (equivalent rectangular bandwidth) bands based on human hearing. The summed energy corresponding to each sampling point data in each ERB band (i.e., the sum of the number of sampling point data in each ERB band) is then determined. The logarithm of the energy sum in each ERB band is then calculated, i.e., a logarithm operation is performed on the energy sum in each ERB band, to obtain the first audio feature corresponding to the first conversion signal (i.e., the first audio signal). Similarly, the second conversion signal is sampled and feature extracted in the same manner to obtain the second audio feature corresponding to the second conversion signal (i.e., the second audio signal). The specific processing process is similar to that of the first audio signal and will not be further described in the present embodiment.
[0103] In this embodiment, feature extraction is performed on the first and second audio signals based on a preset feature extraction algorithm to obtain first audio features corresponding to the first audio signal and second audio features corresponding to the second audio signal, respectively. This feature extraction process preprocesses the first and second audio signals, simplifies the internal processing logic of the howling suppression model, reduces the complexity of the howling suppression model, and improves the efficiency of the howling suppression model's howling signal suppression processing.
[0104] In an exemplary embodiment, Figure 3As shown, the howling suppression model includes a convolution layer, a gated recurrent unit, a fully connected layer and an activation layer. Specifically, the howling suppression model may include, but is not limited to, 5 convolution layers (Conv), a gated recurrent unit (GRU, Gated Recurrent Units), a fully connected layer (Dense) and an activation layer (Sigmoid). Among them, each of the 5 convolution layers can be followed by a two-dimensional Batch Normalization (batch normalization layer) and a RelU (activation layer). The first audio signal and the second audio signal are input into the howling suppression model at the same time, and the howling suppression model processes the first audio signal and the second audio signal layer by layer. As shown Figure 4 As shown, in step S130, the first audio feature and the second audio feature are input into a pre-trained howling suppression model for processing to obtain the audio feature after howling suppression processing, such as Figure 4 As shown, this can be achieved through the following steps:
[0105] In step S402, the first audio feature and the second audio feature are input into a howling suppression model.
[0106] In implementation, the electronic device inputs the obtained first audio feature and second audio feature after band division processing into a howling suppression model.
[0107] In step S404, howling suppression processing is performed on the first audio feature and the second audio feature through the convolution layer, the gated recurrent unit, the fully connected layer and the activation layer in the howling suppression model in sequence, and the audio feature after the howling suppression processing is output.
[0108] During implementation, the electronic device sequentially processes the first audio feature and the second audio feature layer by layer through the convolutional layer, gated recurrent unit, fully connected layer, and activation layer in the howling suppression model, and outputs the audio feature after the howling suppression process. Specifically, the five convolutional layers extract features from the first audio feature and the second audio feature layer by layer, and determine the correlation between the first audio feature and the second audio feature. If the audio feature point contained in the first audio feature has a high correlation with the audio feature point in the second audio feature, then the audio feature point is a howling signal feature point. If the audio feature point contained in the first audio feature has a low correlation with the audio feature point in the second audio feature, then the audio feature point is not a howling signal feature point. Then, the GRU layer of the howling suppression model also stores information such as the temporal relationship and temporal correlation of the audio signal features. Based on this GRU layer, further howling signal feature detection and howling suppression can be performed on the first audio feature and the second audio feature. Then, the audio feature after the howling suppression process is output through mapping in the fully connected layer and limiting the range of the audio feature after the howling suppression process through the activation layer. After the audio features undergo howling suppression processing for the corresponding 64 ERB bands, 64 amplitude masks corresponding to the 64 ERB bands are output. Each amplitude mask represents the result of determining whether a point on the current ERB band is a non-howling signal sampling point, i.e., 0 represents a howling signal, and 1 represents a non-howling signal (i.e., a sound signal). Based on the output discrimination results, the non-howling signal feature points in the audio features to be processed are retained, the howling signal feature points are eliminated, and the audio features after howling suppression processing are determined.
[0109] In this embodiment, the first audio feature and the second audio feature are processed by a simplified howling suppression model. Based on the reference role of the second audio feature, the howling suppression model detects the howling signal information in the first audio feature and performs howling suppression processing on the first audio feature to obtain the audio feature after howling suppression processing, thereby improving the audio feature processing efficiency and the audio feature clarity, thereby achieving the elimination of the howling signal.
[0110] In an exemplary embodiment, Figure 5 As shown, in step S140, the audio features after the howling suppression processing are restored according to a preset feature restoration algorithm to obtain a target audio signal, which can be specifically achieved by the following steps:
[0111] In step S502, the audio features after the howling suppression processing are restored to obtain a restored audio signal in the frequency domain.
[0112] In practice, a feature restoration algorithm is pre-stored in the electronic device. This feature restoration algorithm may correspond to a preset feature extraction algorithm, i.e., the feature restoration algorithm is the inverse algorithm (or inverse algorithm) of the feature extraction algorithm. Then, after the howling suppression model outputs the audio features after the howling suppression processing, the electronic device restores the audio features after the howling suppression processing based on the preset feature restoration algorithm to obtain target audio data in the frequency domain.
[0113] In step S504, an inverse transform is performed on the restored audio signal in the frequency domain according to an inverse short-time Fourier transform algorithm to obtain a restored target audio signal in the time domain.
[0114] During implementation, the electronic device performs an inverse transform on the restored audio signal converted to the frequency domain according to a pre-stored inverse short-time Fourier transform algorithm to obtain a restored target audio signal in the time domain. The target audio signal is a signal that can be mixed with the background audio signal in the electronic device and played.
[0115] In this embodiment, a preset feature restoration algorithm is used to restore the audio features after the howling suppression process, thereby reducing the model complexity of the howling suppression model and improving the processing performance of the howling suppression model.
[0116] In an exemplary embodiment, Figure 6 As shown, in step 110, during the audio signal recording process, obtaining a first audio signal and a second audio signal may specifically include the following steps:
[0117] In step S602, during the audio signal collection process, the sound signal and the howling signal of the target object's singing are collected by the audio signal collection component to obtain a first audio signal including the sound signal and the howling signal.
[0118] In practice, during the audio signal collection process, the sound signal of the target object's singing and the howling signal are collected by the audio collection component to obtain a first audio signal containing the sound signal and the howling signal. Specifically, when the audio signal collection component collects the sound signal of the target object, it also collects audio signals in the environment, such as the mixed sound (the mixture of the sound signal and the background audio signal) fed back by the mixing feedback component, environmental noise, etc. These environmental audio signals are collected by the audio collection component, and a howling signal will inevitably be generated in the audio collection component. The audio signal containing the sound signal of the target object and the audio signal containing the howling signal collected by the audio signal collection component are used as the first audio signal.
[0119] Optionally, the audio recording component may be a microphone or other device including an audio signal recording element. The embodiment of the present disclosure does not limit the audio recording component.
[0120] In step S604, a background audio signal corresponding to the sound signal is obtained by the audio mixing feedback component as a second audio signal.
[0121] In implementation, a background audio signal corresponding to the sound signal is obtained, and then the background audio signal is used as the second audio signal to form a reference for comparison with the first audio signal.
[0122] In this embodiment, the first audio signal is recorded and acquired by the audio recording component, and the second audio signal is collected and acquired by the mixing feedback component, and then used as input data of the howling suppression model to implement howling suppression processing on the first audio signal.
[0123] In an exemplary embodiment, when applying the howling suppression model to perform model processing on the first audio signal, the howling suppression model needs to be trained. Specifically, Figure 7 As shown, before step S120, the method further includes:
[0124] In step S702, audio training samples are obtained.
[0125] The audio training sample includes a third audio feature extracted from a third audio signal and a fourth audio feature extracted from a fourth audio signal. The third audio signal includes a sound signal of a target object and a howling signal, and the fourth audio signal is a background audio signal corresponding to the sound signal.
[0126] During implementation, the electronic device constructs an audio training sample. Specifically, during the audio signal collection process, the audio signal collection component collects the sound signal of the target object. Since the howling signal has the characteristics of a single frequency and the frequency values of each frequency point are all above 800 Hz (Hertz), therefore, during the process of the electronic device collecting the sound signal of the target object, a single-frequency noise of 800 Hz is randomly added to simulate the generation of a howling signal, thereby obtaining a third audio signal containing the target object sound signal and the howling signal. Then, based on the sound signal of the target object in the third audio signal, the electronic device obtains a background audio signal corresponding to the sound signal, and uses the background audio signal as the fourth audio signal. Then, the electronic device performs feature extraction on the third audio signal and the fourth audio signal, respectively, to obtain a third audio feature corresponding to the third audio signal and a fourth audio feature corresponding to the fourth audio signal, thereby constructing an audio training sample containing the third audio feature and the fourth audio feature. When it is necessary to train the howling suppression model, the electronic device obtains an audio training sample.
[0127] In step S704, a preset howling suppression model is trained according to the audio training sample, and training audio features are output.
[0128] During implementation, the electronic device performs model training on a preset howling suppression model based on the third audio feature and the fourth audio feature contained in the audio training sample. The howling suppression model performs howling suppression processing on the third audio feature based on the reference of the fourth audio feature through a convolutional layer, a gated recurrent unit, a fully connected layer and an activation layer, and outputs the training audio feature after the howling suppression processing of the third audio feature.
[0129] In step S706, feature restoration processing is performed on the training audio features according to a preset feature restoration algorithm to obtain a training audio signal.
[0130] During implementation, the electronic device performs feature restoration processing on the output training audio features according to a preset feature restoration algorithm to obtain a training audio signal. The feature restoration method of the training audio signal is similar to the audio feature restoration method in the above step 140. The embodiment of the present disclosure does not elaborate on the process of how to restore the training audio features to the training audio signal.
[0131] In step S708, the loss calculation is performed on the training audio signal according to the preset standard audio signal until the loss result corresponding to the training audio signal meets the preset loss condition, and it is determined that the howling suppression model training is completed.
[0132] The standard audio signal is a sound signal of a target object, that is, a sound signal of a pure human voice without any noise such as a howling signal.
[0133] In implementation, for the howling suppression model, a standard audio signal and a loss calculation method are pre-stored in the electronic device. Specifically, the electronic device performs loss calculation on the training audio signal according to the preset standard audio signal, obtains the amplitude spectrum distance between each sampling point of the training audio signal and the standard audio signal, and then continues to iterate the howling suppression model until the loss result corresponding to the training audio signal (i.e., the size of the amplitude spectrum distance) meets the preset loss condition. The electronic device determines that the howling suppression model training is completed. Among them, the loss condition is that the loss result is stable within a preset range during the model training process (i.e., the fluctuation amplitude of the amplitude spectrum distance between the two times does not exceed the preset amplitude threshold). Therefore, based on this loss condition, when the model reaches a stable state after all iterations and the loss result corresponding to the training audio signal is stable, the howling suppression model training is determined to be completed.
[0134] In this embodiment, the howling suppression model is trained using audio training samples. Then, based on a preset loss algorithm and a standard audio signal, a loss calculation is performed on the training audio signal obtained by the model training. Based on a preset loss condition, it is determined whether the howling suppression model is trained. Then, based on the trained howling suppression model, howling suppression is performed on an audio signal containing a howling signal to improve the clarity of the audio signal.
[0135] In an exemplary embodiment, Figure 8 As shown, in step S708, the loss calculation is performed on the training audio signal according to the preset standard audio signal until the loss result corresponding to the training audio signal meets the preset loss condition. The specific processing process of determining that the howling suppression model training is completed includes:
[0136] In step S802, the training audio signal and the preset standard audio signal are compressed by a preset multiple to obtain a compressed training audio signal and a compressed standard audio signal.
[0137] During implementation, the electronic device compresses the training audio signal and the preset standard audio signal by a preset multiple to obtain a compressed training audio signal and a compressed standard audio signal.
[0138] Specifically, when calculating loss between a standard audio signal and a training audio signal, the standard audio signal and the training audio signal are first compressed before the loss calculation is performed. This is primarily to avoid different impacts on loss due to differences in the amplitudes of the audio feature points contained in the training audio signal in the balanced spectrum. For example, the amplitudes corresponding to the audio feature points contained in the output training audio signal are 90 and 1, respectively. The amplitudes of the two audio feature points corresponding to the training audio signal feature points in the standard audio signal are 100 and 5, respectively. From the absolute magnitude of the amplitude spectrum distance, 100-90=10 is greater than the absolute magnitude of the amplitude spectrum distance corresponding to the other audio feature point, 5-1=4. However, the final impact on hearing is indeed greater due to the difference of 4 at the other audio point, resulting in inaccurate loss calculation results. Therefore, the electronic device first compresses the training audio signal and the standard audio signal before performing the loss calculation to improve the accuracy of the loss results.
[0139] In step S804, the amplitude spectrum distance between each sampling point in the compressed training audio signal and the compressed standard audio signal is determined, and the amplitude spectrum distance is used as the loss result corresponding to the training audio signal.
[0140] During implementation, the electronic device determines the amplitude spectrum distance between each sampling point in the compressed training audio signal and the compressed standard audio signal, and uses the amplitude spectrum distance as the loss result corresponding to the training audio signal. Specifically, the electronic device calculates the amplitude spectrum distance between each sampling point of the training audio signal and the standard audio signal after compressing the training audio signal obtained by restoring the training audio features and the preset standard audio signal by a preset multiple (for example, 0.3 times). The electronic device processes the training audio features output by the howling suppression model, that is, restores the features of the training audio features to obtain the training audio signal after feature restoration, and then performs loss calculation based on the training audio signal and the standard audio signal to determine the amplitude spectrum distance corresponding to the training audio signal obtained after each restoration and the standard audio signal.
[0141] In step S806 , when the loss result corresponding to the training audio signal meets the preset loss condition, it is determined that the howling suppression model training is completed.
[0142] In implementation, within a preset number of model iterations in the electronic device, if the loss result corresponding to the training audio signal determined during the model training process meets a preset loss condition, the electronic device determines that the howling suppression model training is completed.
[0143] In this embodiment, the training audio signal and the standard audio signal are compressed using a preset loss calculation method to reduce interference caused by amplitude fluctuations of the training audio signal itself, improve the accuracy of loss calculation, and thus improve the accuracy of the howling suppression model.
[0144] It should be understood that although Figure 1 、 Figure 2 、 Figures 4 to 8 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1 、 Figure 2 、 Figures 4 to 8 At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The order of execution of these steps or stages is not necessarily one by one, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.
[0145] It can be understood that the same / similar parts between the various embodiments of the above method in this specification can be referred to each other, and each embodiment focuses on the differences from other embodiments. For related parts, please refer to the description of other method embodiments.
[0146] Figure 9 FIG. 1 is a block diagram of an audio signal processing apparatus according to an exemplary embodiment. Figure 9 The device 900 includes an acquisition unit 902, a feature extraction unit 904, a processing unit 906 and a restoration unit 906.
[0147] The acquisition unit 902 is configured to acquire a first audio signal and a second audio signal during the audio signal recording process; the first audio signal includes a sound signal of a target object and a howling signal, and the second audio signal is a background audio signal corresponding to the sound signal;
[0148] The feature extraction unit 904 is configured to perform feature extraction on the first audio signal and the second audio signal according to a preset feature extraction algorithm to obtain a first audio feature and a second audio feature;
[0149] The processing unit 906 is configured to input the first audio feature and the second audio feature into a pre-trained howling suppression model for processing to obtain an audio feature after howling suppression processing;
[0150] The restoration unit 908 is configured to execute a preset feature restoration algorithm to restore the audio features after the howling suppression process to obtain a target audio signal.
[0151] In an exemplary embodiment, the feature extraction unit 904 includes:
[0152] The first processing subunit is configured to perform short-time Fourier transform processing on the first audio signal and the second audio signal to obtain a first converted signal and a second converted signal in the frequency domain after processing;
[0153] a sampling subunit, configured to sample the first conversion signal and the second conversion signal in the frequency domain according to a preset sampling strategy, and obtain a plurality of sampling point data corresponding to the first conversion signal and a plurality of sampling point data corresponding to the second conversion signal;
[0154] The second processing sub-unit is configured to perform compression and band division processing on multiple sampling point data corresponding to the first conversion signal and multiple sampling points corresponding to the second conversion signal, respectively, to obtain a first audio feature corresponding to the first audio signal and a second audio feature corresponding to the second audio signal.
[0155] In an exemplary embodiment, the howling suppression model includes a convolutional layer, a gated recurrent unit, a fully connected layer, and an activation layer. The processing unit 906 includes:
[0156] an input subunit, configured to input the first audio feature and the second audio feature into the howling suppression model;
[0157] The output subunit is configured to process the first audio feature and the second audio feature in sequence through the convolution layer, the gated recurrent unit, the fully connected layer and the activation layer in the howling suppression model, and output the audio feature after the howling suppression processing.
[0158] In an exemplary embodiment, the restoring unit 908 includes:
[0159] a conversion subunit configured to restore the audio features after the howling suppression process to obtain a restored audio signal in the frequency domain;
[0160] The restoration subunit is configured to perform an inverse transformation on the restored audio signal in the frequency domain according to an inverse short-time Fourier transform algorithm to obtain a restored target audio signal in the time domain.
[0161] In an exemplary embodiment, the acquiring unit 902 includes:
[0162] The first acquisition subunit is configured to collect the sound signal and the howling signal of the target object singing through the audio signal recording component during the audio signal recording process, and obtain a first audio signal including the sound signal and the howling signal;
[0163] The second acquisition subunit is configured to acquire a background audio signal corresponding to the sound signal as a second audio signal through a mixing feedback component.
[0164] In an exemplary embodiment, the apparatus 900 further includes:
[0165] a training acquisition unit configured to acquire an audio training sample; the audio training sample includes a third audio feature extracted from a third audio signal feature and a fourth audio feature extracted from a fourth audio signal feature; the third audio signal includes a sound signal of a target object and a howling signal, and the fourth audio signal is a background audio signal corresponding to the sound signal;
[0166] A model training unit is configured to perform model training on a preset howling suppression model according to audio training samples and output training audio features;
[0167] The training restoration unit is configured to perform feature restoration processing on the training audio features according to a preset feature restoration algorithm to obtain a training audio signal;
[0168] The training discrimination unit is configured to perform loss calculation on the training audio signal according to a preset standard audio signal until the loss result corresponding to the training audio signal meets the preset loss condition, thereby determining that the howling suppression model training is completed; the standard audio signal is the sound signal of the target object.
[0169] In an exemplary embodiment, training the discriminant unit includes:
[0170] a loss calculation subunit configured to compress the training audio signal and a preset standard audio signal by a preset multiple to obtain a compressed training audio signal and a compressed standard audio signal; determine an amplitude spectrum distance between each sampling point in the compressed training audio signal and the compressed standard audio signal, and use the amplitude spectrum distance as a loss result corresponding to the training audio signal;
[0171] The determining subunit is configured to determine that the howling suppression model training is completed when the loss result corresponding to the training audio signal meets a preset loss condition.
[0172] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0173] Figure 10 FIG1 is a block diagram of an electronic device 1000 for audio signal processing according to an exemplary embodiment. For example, the electronic device 1000 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0174] Reference Figure 10 , the electronic device 1000 may include one or more of the following components: a processing component 1002 , a memory 1004 , a power component 1006 , a multimedia component 1008 , an audio component 1010 , an input / output (I / O) interface 1012 , a sensor component 1014 , and a communication component 1016 .
[0175] The processing component 1002 generally controls the overall operation of the electronic device 1000, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 1002 may include one or more processors 1020 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 1002 may include one or more modules to facilitate interaction between the processing component 1002 and other components. For example, the processing component 1002 may include a multimedia module to facilitate interaction between the multimedia component 1008 and the processing component 1002.
[0176] The memory 1004 is configured to store various types of data to support operations on the electronic device 1000. Examples of such data include instructions for any application or method operating on the electronic device 1000, contact data, phone book data, messages, pictures, videos, etc. The memory 1004 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, optical disk, or graphene memory.
[0177] The power supply assembly 1006 provides power to the various components of the electronic device 1000. The power supply assembly 1006 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 1000.
[0178] The multimedia component 1008 includes a screen that provides an output interface between the electronic device 1000 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 1008 includes a front camera and / or a rear camera. When the electronic device 1000 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0179] The audio component 1010 is configured to output and / or input audio signals. For example, the audio component 1010 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 1000 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 1004 or transmitted via the communication component 1016. In some embodiments, the audio component 1010 also includes a speaker for outputting audio signals.
[0180] I / O interface 1012 provides an interface between processing component 1002 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.
[0181] The sensor assembly 1014 includes one or more sensors for providing various aspects of status assessment for the electronic device 1000. For example, the sensor assembly 1014 can detect the open / closed state of the electronic device 1000, the relative positioning of components, such as the display and keypad of the electronic device 1000. The sensor assembly 1014 can also detect changes in the position of the electronic device 1000 or components of the electronic device 1000, the presence or absence of user contact with the electronic device 1000, the orientation or acceleration / deceleration of the device 1000, and temperature changes of the electronic device 1000. The sensor assembly 1014 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 1014 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 1014 may also include an accelerometer, a gyroscope, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0182] The communication component 1016 is configured to facilitate wired or wireless communication between the electronic device 1000 and other devices. The electronic device 1000 can access a wireless network based on a communication standard, such as WiFi, an operator network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 1016 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1016 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0183] In an exemplary embodiment, the electronic device 1000 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described methods.
[0184] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 1004 including instructions, and the instructions can be executed by the processor 1020 of the electronic device 1000 to perform the above method. For example, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0185] In an exemplary embodiment, a computer program product is further provided. The computer program product includes instructions, and the instructions can be executed by the processor 1020 of the electronic device 1000 to implement the above method.
[0186] It should be noted that the above-mentioned devices, electronic devices, computer-readable storage media, computer program products, etc. can also include other implementation methods according to the description of the method embodiments. The specific implementation methods can refer to the description of the relevant method embodiments and will not be described one by one here.
[0187] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.
[0188] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A method for processing an audio signal, characterized in that: The method comprises: During the audio signal recording process, a first audio signal and a second audio signal are obtained; the first audio signal includes a sound signal of a target object and a howling signal, and the second audio signal is a background audio signal corresponding to the sound signal; performing feature extraction on the first audio signal and the second audio signal according to a preset feature extraction algorithm to obtain a first audio feature and a second audio feature; Inputting the first audio feature and the second audio feature into a pre-trained howling suppression model for processing to obtain audio features after howling suppression processing; The audio features after the howling suppression process are restored according to a preset feature restoration algorithm to obtain a target audio signal.
2. The audio signal processing method according to claim 1, wherein: The extracting features of the first audio signal and the second audio signal according to a preset feature extraction algorithm to obtain first audio features and second audio features includes: Performing short-time Fourier transform processing on the first audio signal and the second audio signal to obtain a first converted signal and a second converted signal in the frequency domain after processing; Sampling the first conversion signal and the second conversion signal in the frequency domain according to a preset sampling strategy to obtain a plurality of sampling point data corresponding to the first conversion signal and a plurality of sampling point data corresponding to the second conversion signal, respectively; Compression and band division processing is performed on multiple sampling point data corresponding to the first conversion signal and multiple sampling points corresponding to the second conversion signal to obtain a first audio feature corresponding to the first audio signal and a second audio feature corresponding to the second audio signal.
3. The audio signal processing method according to claim 1, wherein: The howling suppression model includes a convolutional layer, a gated recurrent unit, a fully connected layer, and an activation layer. The first audio feature and the second audio feature are input into the pre-trained howling suppression model for processing to obtain audio features after howling suppression processing, including: inputting the first audio feature and the second audio feature into the howling suppression model; The first audio feature and the second audio feature are processed sequentially through a convolutional layer, a gated recurrent unit, a fully connected layer, and an activation layer in the howling suppression model, and the audio feature after the howling suppression processing is output.
4. The audio signal processing method according to claim 1, wherein: The restoring process of the audio features after the howling suppression process is performed according to a preset feature restoration algorithm to obtain a target audio signal includes: Performing restoration processing on the audio features after the howling suppression processing to obtain a restored audio signal in the frequency domain; According to an inverse short-time Fourier transform algorithm, an inverse transform is performed on the restored audio signal in the frequency domain to obtain a restored target audio signal in the time domain.
5. The audio signal processing method according to claim 1, wherein: The step of obtaining the first audio signal and the second audio signal during the audio signal recording process includes: During the audio signal collection process, the audio signal collection component collects the sound signal and the howling signal of the target object singing, and obtains a first audio signal including the sound signal and the howling signal; A background audio signal corresponding to the sound signal is obtained through a mixing feedback component as a second audio signal.
6. The audio signal processing method according to claim 1, wherein: In the audio signal recording process, before obtaining the first audio signal and the second audio signal, the method further includes: Obtain an audio training sample; the audio training sample includes a third audio feature extracted from a third audio signal feature and a fourth audio feature extracted from a fourth audio signal feature; the third audio signal includes a sound signal of a target object and a howling signal, and the fourth audio signal is a background audio signal corresponding to the sound signal; Performing model training on a preset howling suppression model according to the audio training sample, and outputting training audio features; Performing feature restoration processing on the training audio features according to a preset feature restoration algorithm to obtain a training audio signal; The loss calculation of the training audio signal is performed according to a preset standard audio signal until the loss result corresponding to the training audio signal meets a preset loss condition, determining that the howling suppression model training is completed; the standard audio signal is the sound signal of the target object.
7. The audio signal processing method according to claim 6, wherein: The step of performing loss calculation on the training audio signal according to a preset standard audio signal until a loss result corresponding to the training audio signal satisfies a preset loss condition, determining that the howling suppression model training is completed, includes: Compressing the training audio signal and the preset standard audio signal by a preset multiple to obtain a compressed training audio signal and a compressed standard audio signal; determining an amplitude spectrum distance between each sampling point in the compressed training audio signal and the compressed standard audio signal, and using the amplitude spectrum distance as a loss result corresponding to the training audio signal; When the loss result corresponding to the training audio signal meets a preset loss condition, it is determined that the howling suppression model training is completed.
8. An audio signal processing device, characterized in that: The device comprises: An acquisition unit is configured to acquire, during the audio signal recording process, a first audio signal and a second audio signal; the first audio signal includes a sound signal of a target object and a howling signal; and the second audio signal is a background audio signal corresponding to the sound signal; a feature extraction unit configured to perform feature extraction on the first audio signal and the second audio signal according to a preset feature extraction algorithm to obtain a first audio feature and a second audio feature; a processing unit configured to input the first audio feature and the second audio feature into a pre-trained howling suppression model for processing to obtain audio features after howling suppression processing; The restoration unit is configured to execute a preset feature restoration algorithm to restore the audio features after the howling suppression process to obtain a target audio signal.
9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the audio signal processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the audio signal processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Audio howling suppression method, device and system and neural network training method
CN111883163A
Voice signal processing method and device, electronic equipment and readable storage medium
CN113823304A