Audio Signal Processing Method, Apparatus, Electronic Device, and Storage Medium
Through the audio signal processing method combined with beamforming algorithm and acoustic model, the problems of inaccurate positioning and echo interference are solved, and higher positioning accuracy and sound quality are achieved, preventing howling and improving user experience.
Patent Information
- Application Number
- CN202210416890.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-20
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-04-20
AI Technical Summary
The sound quality of existing electronic devices decreases due to inaccurate positioning accuracy and echo interference, which can easily cause howling and affect user experience.
The beamforming algorithm is used to locate the sound source, and the noise reduction processing is performed in combination with the acoustic model, and the effective audio signal is obtained through the echo cancellation algorithm.
It improves the accuracy of audio signal direction positioning, enhances signal-to-noise ratio, prevents howling, and improves sound quality and user experience.
Smart Images

Figure CN114664321B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of signal processing, and in particular, to an audio signal processing method, apparatus, electronic device, and storage medium. Background Art
[0002] With the improvement of people's living standards and the rise of intelligent voice, more and more people like to sing impromptu without leaving home, and electronic devices with built-in microphones and synchronous playback are becoming increasingly popular. At present, the positioning accuracy of most electronic devices is inaccurate, and the echo generated by the played sound will interfere with the positioning, thus introducing noise, greatly reducing the sound quality, and even causing a howling phenomenon, seriously affecting the user's use and experience. Summary of the Invention
[0003] The present invention provides an audio signal processing method, apparatus, electronic device, and storage medium. By eliminating echo, the present invention improves the accuracy of audio signal direction positioning, prevents the howling phenomenon, improves the sound quality, and further improves the user's use and experience.
[0004] An audio signal processing method includes:
[0005] Obtaining an audio signal to be processed;
[0006] Using a beamforming algorithm to perform sound source localization on the audio signal to be processed, and determining a sound source direction corresponding to the audio signal to be processed;
[0007] Based on the sound source direction, performing noise reduction processing on the audio signal to be processed through an acoustic model to obtain a first audio signal;
[0008] Performing filtering processing on the first audio signal to obtain a second audio signal;
[0009] Performing echo cancellation processing on the second audio signal to obtain an effective audio signal.
[0010] An audio signal processing apparatus includes:
[0011] An obtaining module, configured to obtain an audio signal to be processed;
[0012] A positioning module, configured to use a beamforming algorithm to perform sound source localization on the audio signal to be processed, and determine a sound source direction corresponding to the audio signal to be processed;
[0013] A noise reduction module, configured to perform noise reduction processing on the audio signal to be processed through an acoustic model based on the sound source direction to obtain a first audio signal;
[0014] A filtering module, configured to perform filtering processing on the first audio signal to obtain a second audio signal;
[0015] An echo cancellation module for performing echo cancellation processing on the second audio signal to obtain an effective audio signal.
[0016] An electronic device includes a microphone, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned audio signal processing method is implemented.
[0017] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned audio signal processing method is implemented.
[0018] The audio signal processing method, device, electronic device, and storage medium provided by the present invention obtain a to-be-processed audio signal; use a beamforming algorithm to perform sound source localization on the to-be-processed audio signal to determine the sound source direction corresponding to the to-be-processed audio signal; based on the sound source direction, perform noise reduction processing on the to-be-processed audio signal through an acoustic model to obtain a first audio signal; perform filtering processing on the first audio signal to obtain a second audio signal; perform echo cancellation processing on the second audio signal to obtain an effective audio signal. In this way, the present invention improves the accuracy of audio signal direction localization through echo cancellation, prevents the howling phenomenon, improves the signal-to-noise ratio of the signal, improves the sound quality, and further improves the user's use and experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 is a flowchart of an audio signal processing method in an embodiment of the present invention;
[0021] Figure 2 is a flowchart of step S10 of the audio signal processing method in an embodiment of the present invention;
[0022] Figure 3 is a flowchart of step S20 of the audio signal processing method in an embodiment of the present invention;
[0023] Figure 4 is a flowchart of step S30 of the audio signal processing method in an embodiment of the present invention;
[0024] Figure 5It is a flowchart of step S50 of the audio signal processing method in an embodiment of the present invention;
[0025] Figure 6 It is a flowchart of step S502 of the audio signal processing method in an embodiment of the present invention;
[0026] Figure 7 It is a schematic block diagram of the audio signal processing device in an embodiment of the present invention. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0028] The present invention provides an audio signal processing method. In one embodiment, as Figure 1 shown, its technical solutions mainly include the following steps S10 - S50:
[0029] S10, acquire an audio signal to be processed.
[0030] Understandably, the audio signal to be processed is acquired by a microphone in a microphone array in an electronic device. The array mode of the microphone array can be set according to requirements. For example, the microphone array can be a circular array, a rectangular array, etc. The number of microphones in the microphone array can be four, six, eight, etc. The audio signal to be processed is a digital signal after acoustic - electrical conversion and analog - to - digital conversion. The audio signal to be processed includes signals such as the acquired sound source signal and noise signal. The sound source signal is the sound generated by a person speaking or singing. The noise signal includes signals such as environmental noise signal, echo signal, and background noise signal. The environmental noise signal is the sound generated by the friction or vibration of objects in the surrounding environment. The echo signal is the sound played by the device and then re - acquired. The background noise signal is the interference signal introduced by components during the conversion of the audio signal to be processed.
[0031] In one embodiment, as Figure 2 shown, in step S10, that is, acquiring the audio signal to be processed, includes:
[0032] S101, collect a sound wave signal through a microphone, and perform acoustic - electrical conversion on the sound wave signal to obtain an analog audio signal.
[0033] Understandably, the microphone array collects acoustic wave signals in the surrounding environment, converts the obtained acoustic wave signals into electrical signals through acoustic-electric conversion to obtain electrical signals, i.e., analog audio signals. The acoustic-electric conversion is the process of converting acoustic wave signals into electrical signals. The analog audio signal is the electrical signal after acoustic-electric conversion. The electrical signal is represented by voltage or current as a function of time, and drawing the function waveform is the electrical signal. The acoustic wave signal is all the sounds collected by the microphone array.
[0034] S102. Perform analog-to-digital conversion on the analog audio signal to obtain the audio signal to be processed.
[0035] Understandably, perform analog-to-digital conversion processing on the analog audio signal to obtain a digital audio signal, that is, determine the obtained digital audio signal as the audio signal to be processed. The analog-to-digital conversion is the process of converting an analog signal into a digital signal. The digital audio signal is the audio signal after the analog audio signal is subjected to analog-to-digital conversion.
[0036] In the present invention, by performing acoustic-electric conversion processing on the collected acoustic wave signals to obtain analog audio signals, and then performing analog-to-digital conversion processing to obtain audio signals to be processed. In this way, through the conversion of the acoustic wave signals, the noise reduction processing and echo cancellation of the audio signals to be processed are made easier and more convenient.
[0037] S20. Apply a beamforming algorithm to perform sound source localization on the audio signal to be processed and determine the sound source direction corresponding to the audio signal to be processed.
[0038] Understandably, the sound source signal in the audio signal to be processed is localized through a beamforming algorithm, that is, the source direction of the sound source signal is determined to obtain the position of the audio signal to be processed and the sound source direction corresponding to the audio signal to be processed. The beamforming algorithm is an algorithm for combining multiple microphone signals, suppressing interference signals in non-target directions, and enhancing sound signals in the target direction. The principle of the beamforming algorithm is to adjust the basic unit parameters of the phase array so that signals at some angles obtain constructive interference, while signals at other angles obtain destructive interference, weight and sum, and filter the signals of each array element, and finally output the audio signal in the target direction. Among them, the beamforming algorithm can be algorithms such as Fourier transform, short-time Fourier transform, fast Fourier transform, and wavelet transform. The sound source localization is the process of determining the position of the sound source signal through the beamforming algorithm. The sound source direction is the direction of receiving the sound source signal.
[0039] In one embodiment, as Figure 3 shown, in step S20, that is, applying the beamforming algorithm to perform sound source localization on the audio signal to be processed and determine the sound source direction corresponding to the audio signal to be processed, includes:
[0040] S201. Perform scanning in the direction of each positioning angle on the audio signal to be processed, and obtain positioning values corresponding to each positioning angle.
[0041] Understandably, divide the area around the microphone array corresponding to the audio signal to be processed. The angles of the divided areas can be divided according to the actual situation. Determine the angles corresponding to the divided areas as positioning angles, perform scanning in the direction of the obtained positioning angles, obtain the audio signals to be processed corresponding to each positioning angle direction, calculate the audio signals to be processed corresponding to each positioning angle direction through the beamforming algorithm, and obtain the positioning values corresponding to each positioning angle, that is, one positioning value is output for one positioning angle. The angle direction is the regional direction of the positioning angle, and the positioning angle is the angle corresponding to the divided area from one boundary to another boundary. For example: the angle direction of the positioning angle of 0 - 5°, the angle direction of the positioning angle of 60° - 65°, the angle direction of the positioning angle of 100° - 105°, the angle direction of the positioning angle of 205° - 210°, the angle direction of the positioning angle range of 335° - 340°. The positioning value is the output value after calculating the audio signal to be processed through the beamforming algorithm.
[0042] Among them, the process of performing scanning in the direction of each positioning angle on the audio signal to be processed is to convert the energy expressions of the audio signal to be processed in the direction of each positioning angle. The energy expressions are as follows:
[0043]
[0044] Among them, θ is the positioning angle, which can take the median of the sound source angle range; A[θ] is the energy expression of the audio signal to be processed at θ; k is a preset wave number; x1 is the position of the microphone numbered 1 in the microphone matrix in the x - axis direction of the microphone matrix; y1 is the position of the microphone numbered 1 in the microphone matrix in the y - axis direction of the microphone matrix; x n is the position of the microphone numbered n in the microphone matrix in the x - axis direction of the microphone matrix; y n is the position of the microphone numbered n in the microphone matrix in the y - axis direction of the microphone matrix.
[0045] Sum all the values in A[θ] to obtain the positioning value corresponding to this positioning angle (θ).
[0046] S202. Obtain the maximum positioning value from all the positioning values, and determine the positioning angle corresponding to the maximum positioning value as the sound source angle.
[0047] Understandably, all the positioning values are compared to obtain the maximum positioning value among all the positioning values, and the positioning angle corresponding to the maximum positioning value is used as the sound source angle. For example, after the calculation of the beamforming algorithm, the positioning value corresponding to the positioning angle direction of 120° - 125° is the largest, then the positioning angle of 120° - 125° is determined as the sound source angle. The maximum positioning value is the maximum output value obtained by calculating the to-be-processed audio signal through the beamforming algorithm, and the sound source angle is the sum of five positioning angles, that is, the angle corresponding to the area where the sound source signal is received.
[0048] S203. Determine the sound source direction according to the sound source angle.
[0049] Understandably, after obtaining the sound source angle, the sound source direction can be determined according to the sound source angle. The direction pointed to by the median value of the sound source angle can be used as the sound source direction, and the sound source direction is the direction where the sound source signal is located.
[0050] In the present invention, by scanning the to-be-processed audio signal in the directions of each positioning angle, positioning values corresponding to each positioning angle are obtained; the maximum positioning value is obtained from all the positioning values, and the positioning angle corresponding to the maximum positioning value is determined as the sound source angle; the sound source direction is determined according to the sound source angle. In this way, through the division of the area around the microphone array, the accurate positioning of the sound source signal is realized, the accuracy of positioning the sound source direction is improved, and the user experience is further enhanced.
[0051] S30. Based on the sound source direction, perform noise reduction processing on the to-be-processed audio signal through an acoustic model to obtain a first audio signal.
[0052] Understandably, according to the sound source direction, the to-be-processed audio signal corresponding to the sound source direction is acquired, and the to-be-processed audio signal is input into the acoustic model. In the acoustic model, noise reduction processing is performed through the delay-and-sum algorithm to obtain a first audio signal. The acoustic model is a mathematical model constructed by mathematical algorithms, and the acoustic model can also be a neural network model. The mathematical model is a scientific or engineering model constructed by using mathematical logic methods and mathematical languages. The neural network model is a network structure constructed based on neurons through network topology. The noise reduction processing is to reduce or eliminate the noise signals other than the sound source direction in the to-be-processed audio signal, that is, to transpose and clear the energy expressions of other sound source directions and retain the signals in the sound source direction, thereby obtaining the first audio signal.
[0053] When the acoustic model is a trained neural network model, features of the sound source direction are extracted from the audio signal to be processed and confirmed based on the extracted features of the sound source direction, so as to determine the first audio signal from the audio signal to be processed. The training process of the acoustic model is as follows: Obtain a historical audio signal sample set, where the historical audio signal samples include a plurality of historical audio signal samples. One historical audio signal sample is associated with a sound source sample label and a corresponding real audio result. Input the historical audio signal sample and the sound source sample label into a neural network model with initial parameters. Through the neural network model, extract features of the historical audio signal sample based on the sound source sample label, identify the features of the audio signal with the sound source sample label in the historical audio signal sample, and output the audio signal as the recognition audio result. According to the recognition audio result and the real audio result, determine the loss value. When the loss value does not reach the preset convergence condition, iteratively update the initial parameters of the neural network model, and re-execute the step of extracting features of the historical audio signal sample based on the sound source sample label through the neural network model until the loss value reaches the convergence condition. Record the neural network model after convergence as the trained acoustic model.
[0054] In one embodiment, as Figure 4 shown, in step S30, that is, based on the sound source direction, perform noise reduction processing on the audio signal to be processed through an acoustic model to obtain a first audio signal, including:
[0055] S301, through the acoustic model, perform matrix conversion on the audio signal to be processed according to the sound source direction to obtain a matrix audio signal.
[0056] Understandably, input the audio signal to be processed corresponding to the sound source direction into the acoustic model, and perform matrix conversion on the audio signal to be processed through a beamforming algorithm in the acoustic model to obtain a matrix audio signal. The matrix conversion is a process of representing multiple audio signals to be processed collected by a microphone array in matrix form, and the matrix audio signal is the audio signal after matrix conversion of the audio signal to be processed.
[0057] S302, perform noise cancellation on the matrix audio signal based on a preset frequency band to obtain the first audio signal.
[0058] Understandably, the matrix audio signal obtained after matrix conversion is processed according to a preset frequency band to obtain the matrix audio signal within the preset frequency band range. Noise cancellation processing is performed on the matrix audio signal within the preset frequency band range to obtain a first audio signal. The preset frequency band is a signal frequency range set in advance, and this range is 100 Hz (bass) to 10,000 Hz (soprano). The noise cancellation is to perform cancellation processing on the noise signal in the matrix audio signal within the preset frequency band range, that is, the process of reducing or eliminating the noise signal.
[0059] Among them, noise reduction processing is performed on the audio signal to be processed through a beamforming algorithm, and the expression is as follows:
[0060]
[0061] y(θ′) = w(θ′) H A[θ′];
[0062] Among them, θ′ is the positioning angle of the sound source direction; A[θ′] is the energy expression of the audio signal to be processed at θ′; w[θ′] is a preset filter matrix for the positioning angle; H is the conjugate transpose; y(θ′) is the first audio signal.
[0063] In the present invention, the audio signal to be processed corresponding to the sound source direction is subjected to matrix conversion through an acoustic model to obtain a matrix audio signal, and noise cancellation based on a preset frequency band is performed on the matrix audio signal to obtain the first audio signal. In this way, the conversion of the audio signal to be processed and the noise cancellation of the matrix audio signal are realized, the interference of the noise signal on the sound source signal is reduced, the signal-to-noise ratio of the signal is improved, the quality of the sound is improved, and the user's use and experience are further improved.
[0064] S40. Filter the first audio signal to obtain a second audio signal.
[0065] Understandably, the first audio signal is filtered through an adaptive filter to obtain a second audio signal. Filtering the first audio signal can effectively suppress residual noise, such as non-coherent noise, scattering noise and other noises. The adaptive filter is based on the estimation of the statistical characteristics of the input and output signals, and adopts a specific algorithm to automatically adjust the filter coefficients to make it reach the best filtering characteristics. The adaptive filter has a small amount of calculation and is especially suitable for real-time processing. The filter can also be a Wiener filter, a Kalman filter or other filters.
[0066] Among them, the principle of the adaptive filter is that the input signal passes through a digitally tunable filter to generate an output signal. The output signal is compared with the desired signal to form an error signal. The filter parameters are adjusted through an adaptive algorithm, ultimately minimizing the mean square value of the error signal. Adaptive filtering can automatically adjust the filter parameters at the current moment based on the results of the filter parameters obtained at the previous moment to adapt to the unknown or time-varying statistical characteristics of the signal and noise, thereby achieving optimal filtering. The automatic adjustment process is a process of adapting the filter parameters at the current moment to the filter parameters at the previous moment to match or enhance the filter parameters at the previous moment based on the filter parameters at the previous moment.
[0067] S50. Perform echo cancellation processing on the second audio signal to obtain an effective audio signal.
[0068] Understandably, the second audio signal is processed by an echo cancellation algorithm to obtain an effective audio signal. The echo cancellation algorithm includes an echo suppression algorithm and an echo cancellation algorithm. The echo suppression algorithm is to compare the level of the sound played by the speaker with the sound collected by the current microphone through a comparator. If the sound played by the speaker is higher than the threshold, it is allowed to be transmitted to the speaker, and the microphone is turned off to prevent it from picking up the sound played by the speaker and causing a far-end echo. If the level of the sound picked up by the microphone is higher than the threshold, the speaker is prohibited to achieve the purpose of echo cancellation. The echo cancellation algorithm is based on the correlation between the speaker signal and the multipath echo generated by it, establishing a voice model of the far-end signal, using it to estimate the echo, and continuously modifying the coefficients of the filter to make the estimated value closer to the real echo, and then subtracting the echo estimated value from the input signal of the microphone to achieve the purpose of echo cancellation. The effective audio signal is the audio signal after noise reduction and echo cancellation of the audio signal to be processed, that is, the sound signal played from the speaker in the electronic device.
[0069] The audio signal processing method provided by the present invention includes obtaining an audio signal to be processed; using a beamforming algorithm to perform sound source localization on the audio signal to be processed to determine the sound source direction corresponding to the audio signal to be processed; based on the sound source direction, performing noise reduction processing on the audio signal to be processed through an acoustic model to obtain a first audio signal; performing filtering processing on the first audio signal to obtain a second audio signal; performing echo cancellation processing on the second audio signal to obtain an effective audio signal. In this way, echo cancellation improves the accuracy of sound source direction localization of the audio signal, improves the signal-to-noise ratio of the signal, prevents the howling phenomenon, improves the quality of the sound, and improves the user's use and experience.
[0070] In an embodiment, as Figure 5As shown, in step S50, that is, performing echo cancellation processing on the second audio signal to obtain an effective audio signal, includes:
[0071] S501, performing filtering processing on the second audio signal to obtain an intermediate audio signal.
[0072] Understandably, filtering the second audio signal through an adaptive filter to obtain an intermediate audio signal. The filtering processing is to determine the range of the second audio signal. For example, the range is 100 - 3600 Hz when singing and 500 - 2000 Hz when speaking. The intermediate audio signal is the audio signal after filtering the second audio signal.
[0073] S502, performing attenuation processing on the intermediate audio signal to obtain the effective audio signal.
[0074] Understandably, performing attenuation processing on the intermediate audio signal through an attenuation algorithm to obtain the effective audio signal. Eliminating the echo signal in the intermediate audio signal through attenuation processing to obtain the sound signal played from the speaker. Among them, the attenuation processing can also be physical attenuation, that is, making the echo signal attenuate by increasing the distance, or adding sound insulation items between the speaker and the microphone to make the echo signal attenuate.
[0075] The present invention performs filtering processing on the second audio signal to obtain an intermediate audio signal; performing attenuation processing on the intermediate audio signal to obtain the effective audio signal. In this way, the acquisition of the effective audio signal is realized, the interference of the echo signal is avoided, the signal-to-noise ratio of the signal is improved, and the sound quality is further improved.
[0076] In an embodiment, as Figure 6 shown, in step S50, that is, performing attenuation processing on the intermediate audio signal to obtain the effective audio signal, includes:
[0077] S5021, performing time-frequency conversion analysis on the intermediate audio signal to obtain a frequency-domain effective audio signal.
[0078] Understandably, performing time-frequency conversion analysis on the intermediate audio signal to obtain a frequency-domain effective audio signal, that is, converting the intermediate audio signal from the time domain to the frequency domain to obtain the frequency-domain effective audio signal.
[0079] S5022, performing frequency division processing on the frequency-domain effective audio signal to obtain the effective audio signal.
[0080] Understandably, the frequency-domain effective audio signal is subjected to frequency division processing through a pre-stack frequency division noise attenuation technique. The pre-stack frequency division noise attenuation technique is a processing process of identifying noise in each frequency band of the effective audio signal in the frequency domain with the weighted median as a parameter, determining the attenuation curve according to the numerical relationship of the noise, and suppressing the noise in the signal in different frequency bands. According to the energy amplitude attribute and frequency characteristics of the noise, the frequency band parameters for suppressing the noise are determined, and the noise is suppressed through threshold setting to obtain the effective audio signal.
[0081] The present invention realizes obtaining a frequency-domain effective audio signal through time-frequency conversion analysis of the relay audio signal; performing frequency division processing on the frequency-domain effective audio signal to obtain the effective audio signal. In this way, the elimination processing of the echo signal is realized, the interference of the echo signal is avoided, the signal-to-noise ratio of the signal is improved, and the quality of the sound is improved.
[0082] In one embodiment, after the step S50, that is, after performing echo cancellation processing on the second audio signal to obtain an effective audio signal, it includes:
[0083] Based on the effective audio signal, the microphone is used to perform secondary echo cancellation on the newly collected sound wave signal to obtain a new audio signal to be processed.
[0084] Understandably, the process of the secondary echo cancellation is, after obtaining the effective audio, using the effective audio signal as a reference signal, collecting a new sound wave signal through a microphone array, performing echo cancellation on the newly collected sound wave signal, that is, eliminating the signals in different frequency bands from the reference signal in the new sound wave signal, and enhancing the signals in the same frequency band as the reference signal in the new sound wave signal to obtain a new audio signal to be processed. The new audio signal to be processed is subjected to noise reduction processing, filtering processing, and echo cancellation to obtain a new effective audio signal, and then the new effective audio signal is used as a reference signal for echo cancellation processing.
[0085] In this way, based on the effective audio signal, rapid noise reduction processing and echo cancellation processing of the newly collected sound wave signal are realized, the interference of the noise signal is avoided, the quality of the sound is improved, and the user's use and experience are further improved.
[0086] In one embodiment, after the step S50, that is, after performing echo cancellation processing on the second audio signal to obtain an effective audio signal, it includes:
[0087] According to the beamforming algorithm and the noise cancellation algorithm, perform background noise processing on the effective audio signal to obtain a target sound source signal.
[0088] Understandably, during the signal processing, multiple signal conversions and signal processing operations are performed. The components through which the signal conversions and signal processing pass will generate background noise signals. After multiple accumulations and gain amplifications, the background noise signals will also increase, interfering with the played sound. The beamforming algorithm is used to perform noise reduction processing and echo cancellation processing on the acoustic wave signal, and the interference subtraction algorithm is used to perform background noise processing on the effective audio signal to obtain the target sound source signal. It is also possible to use the harmonic frequency suppression method, that is, to complete noise reduction by using the method of enhancing the audio signal. Based on the periodic principle of noise, the fundamental frequency tracking is implemented by using the adaptive comb filter of harmonic noise to complete noise reduction. The interference subtraction algorithm is an algorithm that suppresses noise by subtracting the noise spectrum.
[0089] Among them, the principle of the interference subtraction algorithm is based on the periodicity of the acoustic wave signal. The time-domain adaptive noise cancellation method can be utilized by generating a reference signal. The reference signal is formed by delaying the main signal by one period and requires a complex pitch estimation algorithm. Using the fast Fourier transform, subtract the estimated noise amplitude spectrum and inverse-transform the subtracted spectrum amplitude. Then, using the phase of the original noise, calculate the short-time amplitude and phase spectrum of the noisy signal, that is, first decompose the acoustic wave signal into different frequency groups by using a bank of band-pass filters, and then estimate the noise power of each sub-band during the period without the acoustic wave signal. Noise suppression can be obtained by using the attenuation factor, where the attenuation factor corresponds to the ratio of the estimated noise power of each sub-band to the instantaneous signal power.
[0090] Through the beamforming algorithm and the noise cancellation algorithm of the present invention, background noise processing is performed on the effective audio signal, that is, the background noise signal in the acoustic wave signal is removed. In this way, the signal-to-noise ratio of the signal is improved, the sound played by the speaker is clearer, and the sound quality is improved.
[0091] In an embodiment, an audio signal processing device is provided, which corresponds one-to-one to the audio signal processing method in the above embodiment. As Figure 7 shown, the audio signal processing device includes an acquisition module 11, a positioning module 12, a noise reduction module 13, a filtering module 14, and an elimination module 15. The detailed descriptions of each functional module are as follows:
[0092] The acquisition module 11 is used to acquire the audio signal to be processed;
[0093] The positioning module 12 is used to use the beamforming algorithm to perform sound source localization on the audio signal to be processed and determine the sound source direction corresponding to the audio signal to be processed;
[0094] The noise reduction module 13 is used to perform noise reduction processing on the audio signal to be processed based on the sound source direction through an acoustic model to obtain a first audio signal;
[0095] A filtering module 14 is configured to filter the first audio signal to obtain a second audio signal;
[0096] An echo cancellation module 15 is configured to perform echo cancellation processing on the second audio signal to obtain a valid audio signal.
[0097] For the specific limitations of the audio signal processing device, reference may be made to the limitations on the audio signal processing method in the foregoing text, which will not be elaborated herein. Each module in the above audio signal processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the microphone of the electronic device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.
[0098] In one embodiment, an electronic device is provided. The electronic device can be a client or a server. The electronic device includes a processor, a microphone, a network interface, and a database connected through a system bus. Among them, the processor of the electronic device is configured to provide computing and control capabilities. The microphone of the electronic device includes a readable storage medium and an internal memory. The readable storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the readable storage medium. The network interface of the electronic device is configured to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an audio signal processing method.
[0099] In one embodiment, an electronic device is provided, including a microphone, a processor, and a computer program stored on a memory and executable on the processor. When the processor executes the computer program, it implements the audio signal processing method in the above embodiment.
[0100] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the audio signal processing method in the above embodiment.
[0101] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0102] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0103] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.
Claims
1. An audio signal processing method, characterized in that Including: Obtain the audio signal to be processed; Apply the beamforming algorithm to perform sound source localization on the audio signal to be processed, and determine the sound source direction corresponding to the audio signal to be processed; Based on the sound source direction, perform noise reduction processing on the audio signal to be processed through an acoustic model to obtain a first audio signal; Perform filtering processing on the first audio signal to obtain a second audio signal; Perform echo cancellation processing on the second audio signal to obtain a valid audio signal; The performing noise reduction processing on the audio signal to be processed through an acoustic model based on the sound source direction to obtain a first audio signal includes: Perform matrix conversion on the audio signal to be processed through the acoustic model according to the sound source direction to obtain a matrix audio signal; Perform noise cancellation on the matrix audio signal based on a preset frequency band to obtain the first audio signal; Wherein, the beamforming algorithm is used to perform noise reduction processing on the audio signal to be processed, and the expression is as follows: ; ; wherein, is the positioning angle of the sound source direction; is the energy expression of the audio signal to be processed under ; w is a preset filtering matrix for the positioning angle; H is the conjugate transpose; y( ) is the first audio signal.
2. The audio signal processing method according to claim 1, wherein The performing echo cancellation processing on the second audio signal to obtain a valid audio signal includes: Perform filtering processing on the second audio signal to obtain an intermediate audio signal; Perform attenuation processing on the intermediate audio signal to obtain the valid audio signal.
3. The audio signal processing method according to claim 1, wherein The applying the beamforming algorithm to perform sound source localization on the audio signal to be processed and determining the sound source direction corresponding to the audio signal to be processed includes: Scan the audio signal to be processed in the direction of each positioning angle to obtain positioning values corresponding to each positioning angle; Obtain the maximum positioning value from all the positioning values, and determine the positioning angle corresponding to the maximum positioning value as the sound source angle; Determine the sound source direction according to the sound source angle.
4. The audio signal processing method according to claim 2, characterized in that The performing attenuation processing on the intermediate audio signal to obtain the valid audio signal includes: Perform time-frequency conversion analysis on the intermediate audio signal to obtain a frequency-domain valid audio signal; Perform frequency division processing on the frequency-domain valid audio signal to obtain the valid audio signal.
5. The audio signal processing method according to claim 1, characterized in that The obtaining the audio signal to be processed includes: Collect a sound wave signal through a microphone, and perform acoustic-electric conversion on the sound wave signal to obtain an analog audio signal; Perform analog-to-digital conversion on the analog audio signal to obtain the audio signal to be processed.
6. The audio signal processing method according to claim 1, wherein, After performing echo cancellation processing on the second audio signal to obtain a valid audio signal, it includes: Based on the valid audio signal, perform secondary echo cancellation on the newly collected sound wave signal through a microphone to obtain a new audio signal to be processed.
7. An audio signal processing device, characterized in that, Including: An obtaining module, configured to obtain the audio signal to be processed; A positioning module, configured to apply the beamforming algorithm to perform sound source localization on the audio signal to be processed and determine the sound source direction corresponding to the audio signal to be processed; A noise reduction module, configured to perform noise reduction processing on the audio signal to be processed through an acoustic model based on the sound source direction to obtain a first audio signal; A filtering module, configured to perform filtering processing on the first audio signal to obtain a second audio signal; An elimination module, configured to perform echo cancellation processing on the second audio signal to obtain a valid audio signal; The noise reduction module is further configured to: The acoustic model performs matrix conversion on the audio signal to be processed according to the sound source direction, obtaining a matrix audio signal; Performing noise cancellation on the matrix audio signal based on a preset frequency band to obtain the first audio signal; Among them, the beamforming algorithm is used to perform noise reduction processing on the audio signal to be processed, and the expression is as follows: ; ; Among them, is the positioning angle of the sound source direction; is the energy expression of the audio signal to be processed under; w is a preset filtering matrix for the positioning angle; H is the conjugate transpose; y( ) is the first audio signal.
8. An electronic device, comprising a microphone, a processor, and a computer program stored in a memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the audio signal processing method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the audio signal processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Echo cancellation method and device
CN108711433A
Voice signal processing method and device, computer readable medium and electronic equipment
CN111435598A
Sound source positioning system
CN112965033A
Multi-channel speech enhancement method and device
CN113030862A