Method, apparatus, electronic device, and readable storage medium for processing voice signals
By using neural network models to calculate the frequency band gain of voice signals based on frequency domain expression, the problem of interference signal elimination in multi-person voice calls is solved, better voice signal processing effect is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202110784615.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-12
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-07-12
AI Technical Summary
The prior art is difficult to effectively eliminate interfering signals in multi-person voice calls, especially electrical echo, acoustic echo and ambient noise, affecting the experience of voice interaction.
By filtering the received distal voice signal to obtain the echo prediction signal, collect the nearest voice signal, and based on its first frequency domain expression and the second frequency domain expression of the echo prediction signal, the frequency band gain of the nearest voice signal is obtained using a pre-trained neural network model, thereby eliminating the interference signal.
This method can more effectively eliminate interference signals, retain effective parts of the near-end voice signal, and improve user experience, especially in multi-end speech scenarios.
Smart Images

Figure CN113823304B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of signal processing. Specifically, this application relates to a method, apparatus, electronic device, and readable storage medium for processing voice signals. Background Art
[0002] With the increasing popularity of voice call technologies such as VoIP (Voice over Internet Protocol), more and more attention has been paid to the quality of voice calls. When multiple people (two or more) are on a call, in addition to the voice of the proximal speaker entering the microphone, there may also be interfering sounds such as electrical echo, acoustic echo, and proximal ambient noise. If these interfering sounds are transmitted to the distal end and heard by the distal speaker, it will seriously affect the voice interaction experience. Therefore, it is necessary to eliminate the interfering signals proximally.
[0003] In the prior art, the performance of eliminating interfering signals in multi-party speech has always been a difficult point in the industry. Although there are already various methods for eliminating interfering signals in the prior art, the effect of each method needs to be improved. Summary of the Invention
[0004] Embodiments of this application provide a method, apparatus, electronic device, and readable storage medium for processing voice signals, achieving the purpose of better eliminating interfering signals. The technical solutions are as follows:
[0005] According to one aspect of this application, there is provided a method for processing voice signals, the method including:
[0006] Filtering the received distal voice signal to obtain an echo prediction signal;
[0007] Collecting the proximal voice signal;
[0008] Obtaining a first frequency-domain representation of the proximal voice signal and a second frequency-domain representation of the echo prediction signal;
[0009] Based on the first frequency-domain representation and the second frequency-domain representation, obtaining the frequency band gain of the proximal voice signal through a pre-trained neural network model, where the frequency band gain characterizes the weight of the effective voice signal in the proximal voice signal;
[0010] According to the frequency band gain, eliminating the interfering signal from the proximal voice signal to obtain the processed proximal voice signal.
[0011] According to another aspect of this application, there is also provided a device for processing voice signals, the device including:
[0012] A signal filtering module for filtering the received distal voice signal to obtain an echo prediction signal;
[0013] A signal acquisition module for acquiring a proximal speech signal;
[0014] A frequency-domain expression acquisition module for acquiring a first frequency-domain expression of the proximal speech signal and a second frequency-domain expression of the echo prediction signal;
[0015] A frequency-band gain determination module for obtaining the frequency-band gain of the proximal speech signal based on the first frequency-domain expression and the second frequency-domain expression through a pre-trained neural network model, where the frequency-band gain characterizes the weight of the effective speech signal in the proximal speech signal;
[0016] An interference signal cancellation module for canceling the interference signal of the proximal speech signal according to the frequency-band gain to obtain the processed proximal speech signal.
[0017] In an optional implementation manner, when the frequency-band gain determination module is used to obtain the frequency-band gain of the proximal speech signal based on the first frequency-domain expression and the second frequency-domain expression through a pre-trained neural network model, it is specifically used for:
[0018] Determining the frequency-domain information difference between the proximal speech signal and the echo prediction signal based on the first frequency-domain expression and the second frequency-domain expression;
[0019] Obtaining the frequency-band gain of the proximal speech signal through the trained neural network model based on the frequency-domain information difference.
[0020] In an optional implementation manner, when the frequency-band gain determination module is used to obtain the frequency-band gain of the proximal speech signal through the trained neural network model based on the frequency-domain information difference, it is specifically used for:
[0021] Concatenating the first frequency-domain expression and the frequency-domain information difference to obtain the concatenated frequency-domain information;
[0022] Obtaining the frequency-band gain of the proximal speech signal through the trained neural network model based on the concatenated frequency-domain information.
[0023] In an optional implementation manner, when the frequency-domain expression acquisition module is used to acquire the first frequency-domain expression of the proximal speech signal and the second frequency-domain expression of the echo prediction signal, it is specifically used for:
[0024] Obtaining the first spectrum of each frame of the first signal included in the proximal speech signal and the second spectrum of each frame of the second signal included in the echo prediction signal;
[0025] Obtaining the first frequency-domain expression of each frame of the first signal based on the first spectrum of each frame of the first signal;
[0026] Obtaining the second frequency-domain expression of each frame of the second signal based on the second spectrum of each frame of the second signal;
[0027] When the band gain determination module is used to obtain the band gain of each frame of the first signal included in the near-end speech signal through a pre-trained neural network model based on the first frequency-domain expression and the second frequency-domain expression, it is specifically used for:
[0028] For each frame of the first signal, based on the first frequency-domain expression of this frame of the first signal and the second frequency-domain expression of the second signal corresponding to this frame of the first signal in the echo prediction signal, obtain the band gain of this frame of the first signal through the pre-trained neural network model.
[0029] In an optional implementation manner, the spectrum of each frame of the first signal and each frame of the second signal includes the amplitude values of multiple frequency points. When the frequency-domain expression acquisition module is used to obtain the frequency-domain expression of each frame of the signal based on the spectrum of this frame of the signal, it is specifically used for:
[0030] Based on the amplitude value of each frequency point included in the spectrum of this frame of the signal, obtain the frequency-domain expression corresponding to each frequency point of this frame of the signal; wherein, the band gain corresponding to this frame of the first signal includes the band gains corresponding to the frequency points included in the first spectrum of this frame of the first signal.
[0031] When the interference signal cancellation module is used to cancel the interference signal of the near-end speech signal according to the band gain and obtain the processed near-end speech signal, it is specifically used for:
[0032] Determine the residual signal between the near-end speech signal and the echo prediction signal;
[0033] Obtain the third spectrum of each frame of the signal included in the residual signal;
[0034] For each frame of the first signal, based on the band gains of the frequency points corresponding to this frame of the first signal, perform weighted calculation on the amplitude values of the frequency points included in the spectrum of the corresponding frame in the residual signal to obtain the fourth spectrum corresponding to this frame of the first signal;
[0035] Perform frequency-time transformation based on the fourth spectra corresponding to each first signal to obtain the processed near-end speech signal.
[0036] In an optional implementation manner, the spectrum of each frame of the first signal and each frame of the second signal includes the amplitude values of multiple frequency points. When the frequency-domain expression acquisition module is used to obtain the frequency-domain expression of each frame of the signal based on the spectrum of this frame of the signal, it is specifically used for:
[0037] Divide the spectrum of this frame of the signal into M sub-bands, and fuse the amplitude values of the frequency points corresponding to each sub-band to obtain the fused amplitude value, where M≥1;
[0038] Based on the fused amplitude values corresponding to each sub-band, obtain the frequency-domain expression corresponding to each sub-band; wherein, the frequency-band gain corresponding to the first signal of this frame includes the frequency-band gains corresponding to each sub-band included in the first spectrum of the first signal of this frame.
[0039] When the interference signal cancellation module is used to cancel the interference signal of the proximal speech signal according to the frequency-band gain to obtain the processed proximal speech signal, it is specifically used for:
[0040] Determine the residual signal between the proximal speech signal and the echo prediction signal;
[0041] Obtain the third spectrum of each frame signal included in the residual signal, and divide the third spectrum of each frame signal into M sub-bands;
[0042] For each sub-band of each frame of the first signal, perform weighted calculation on the amplitude values of each frequency point included in the corresponding sub-band of the third spectrum of the corresponding frame in the residual signal based on the frequency-band gain corresponding to each sub-band to obtain the fourth spectrum corresponding to each sub-band;
[0043] Based on the fourth spectra corresponding to each first signal, perform frequency-time transformation to obtain the processed proximal speech signal.
[0044] In an alternative implementation, the first frequency-domain expression and the second frequency-domain expression both include at least one of power spectrum, amplitude spectrum, logarithmic power spectrum or logarithmic amplitude spectrum.
[0045] In an alternative implementation, the neural network model is trained in the following manner:
[0046] Obtain a plurality of training samples, each training sample includes a distal sample speech signal, a proximal sample speech signal and annotation information, and the annotation information characterizes the true frequency-band gain of the proximal sample speech signal;
[0047] Perform filtering processing on the distal sample speech signal of each training sample to obtain a sample echo prediction signal;
[0048] Determine the third frequency-domain expression of the proximal sample speech signal and the fourth frequency-domain expression of the sample echo prediction signal of each training sample;
[0049] Based on the third frequency-domain expression and the fourth frequency-domain expression corresponding to each training sample, use a machine learning method to iteratively train the initial neural network model to obtain the predicted frequency-band gain corresponding to each training sample;
[0050] Among them, for each training, if it is determined that the training end condition is met based on the true band gain and the predicted band gain corresponding to each training sample, a trained neural network model is obtained. If the training end condition is not met, the model parameters of the neural network model are adjusted, and the neural network model is continuously trained based on the third frequency domain expression and the fourth frequency domain expression corresponding to each training sample.
[0051] In an alternative implementation, for each training sample, it further includes:
[0052] Based on the third frequency domain expression and the fourth frequency domain expression, determine the sample frequency domain information difference between the proximal sample speech signal and the sample echo prediction signal; wherein, during training, the input of the neural network model includes the sample frequency domain information difference corresponding to each training sample, or includes the sample frequency domain information difference corresponding to each training sample and the third frequency domain expression.
[0053] According to another aspect of the present application, there is also provided an electronic device, which includes:
[0054] A processor and a memory, the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the method for processing speech signals of the present application.
[0055] According to still another aspect of the present application, there is also provided a computer-readable storage medium, which is used to store a computer program. When the computer program runs on a computer, the computer is enabled to execute the method for processing speech signals of the present application.
[0056] According to still another aspect of the present application, there is also provided a computer program product or a computer program. When it runs on a computer device, the computer device is enabled to execute the method for processing speech signals of the present application. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute the method for processing speech signals of the present application.
[0057] The beneficial effects brought by the technical solution provided by the present application are:
[0058] The method, apparatus, electronic device, and readable storage medium for processing voice signals provided in this application use the frequency-domain expressions of the near-end voice signal and the echo prediction signal as the input of the neural network model. The pre-trained neural network model will output the band gain of the near-end voice signal. Thus, the band gain can more reliably represent the weight of the valid voice signal in the near-end voice signal. Subsequently, when the interference signal in the near-end voice signal is eliminated through the band gain, the valid voice signal in the near-end voice signal can be better retained, further improving the interference signal elimination performance and better eliminating the interference signal. Especially when multiple parties are speaking, the near-end human voice can be better retained, enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments of this application.
[0060] Figure 1 Schematic diagram of an application scenario provided for an embodiment of this application;
[0061] Figure 2 Schematic flowchart of a method for processing voice signals provided for an embodiment of this application;
[0062] Figure 3 Schematic diagram of a neural network model provided for an embodiment of this application;
[0063] Figure 4 Schematic diagram of a neural network module provided for an embodiment of this application;
[0064] Figure 5 Schematic diagram of a filtering process provided for an embodiment of this application;
[0065] Figure 6 Schematic diagram of interference signal cancellation provided for an embodiment of this application;
[0066] Figure 7 Schematic diagram of the structure of a device for processing voice signals provided for an embodiment of this application;
[0067] Figure 8 Schematic diagram of the structure of an electronic device provided for an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] The following details the embodiments of this application. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain this application and should not be construed as limiting this application.
[0069] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of this application means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.
[0070] To better illustrate the solutions provided by the embodiments of this application, the technical terms involved in the embodiments of this application will be briefly introduced and explained below.
[0071] Near end: It refers to the local end of the communication link established during a voice call.
[0072] Far end: It refers to the opposite end of the communication link established during a voice call.
[0073] Near-end voice signal: It refers to the voice signal collected by the call device used by the near-end user during a voice call.
[0074] Far-end voice signal: It refers to the voice signal received by the call device used by the near-end user from the call device used by the far-end user through the communication link during a voice call.
[0075] For a two-person or multi-person call process, traditional PSTN (Public Switched Telephone Network) phone calls may cause electrical echo due to the impedance mismatch problem of the hybrid coil. During the call process of current digital devices (such as mobile phones, PCs (Personal Computers), etc.), especially in the hands-free state, after the voice of the far-end speaker is transmitted to the near end, it will be sent to the speaker (loudspeaker) for playback. The sound played by the speaker will form acoustic echo after being transmitted through the air and entering the microphone at the near end. In addition to the voice of the near-end speaker and acoustic echo entering the microphone, there may also be ambient noise at the near end.
[0076] In the embodiments of the present application, in order to prevent the electrical and acoustic echo signals entering the near-end microphone from being transmitted back to the far-end and making the far-end speaker hear their own voice again (i.e., echo), and at the same time to minimize the transmission of near-end noise to the far-end, interference signal cancellation needs to be performed at the near-end, including echo cancellation (Automatic Echo Cancellation, AEC), noise suppression, etc.
[0077] The technical solution provided by the embodiments of the present application relates to the fields of artificial intelligence and speech technology. Optionally, in the embodiments of the present application, the step of determining the frequency band gain of the near-end speech signal can be implemented by a neural network model. Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in theory, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0078] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.
[0079] Among them, the key technologies of Speech Technology include automatic speech recognition technology, speech synthesis technology, and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and speech has become one of the most promising human-computer interaction methods in the future. With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common applications include smart home, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robots, intelligent healthcare, intelligent customer service, vehicle networking, autonomous driving, and intelligent transportation. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0080] The execution entity of the embodiments of this application can be a communication device used by a proximal user, such as a mobile terminal, etc. Among them, the communication device can include an information interaction module, a playback module (such as a speaker, etc., hereinafter introduced by taking the speaker as an example), a collection module (such as a microphone, etc., hereinafter introduced by taking the microphone as an example), etc. In practical applications, the mobile terminal can include, for example, mobile phones, smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart watches, smart TVs, intelligent vehicle-mounted devices, personal digital assistants, portable multimedia players, etc. Those skilled in the art can understand that, except for elements specifically for mobile purposes, the structure according to the embodiments of this application can also be applied to fixed-type terminals, such as digital TVs, desktop computers, etc.
[0081] Optionally, the communication device used by the proximal user and the communication device used by the distal user can be nodes in a distributed system. Among them, the distributed system can be a blockchain system, and the blockchain system can be a distributed system formed by connecting multiple nodes in the form of network communication. Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.
[0082] The underlying blockchain platform may include processing modules such as user management, basic services, smart contracts, and operation detection. Among them, the user management module is responsible for managing the identity information of all blockchain participants, including maintaining the generation of public and private keys (account management), key management, and maintaining the correspondence between the real identity of users and blockchain addresses (permission management). And under authorized circumstances, it supervises and audits the transaction situations of certain real identities, and provides rule configuration for risk control (risk control audit); the basic service module is deployed on all blockchain node devices to verify the validity of business requests, and records the valid requests on the storage after consensus. For a new business request, the basic service first performs interface adaptation parsing and authentication processing (interface adaptation), then encrypts the business information through a consensus algorithm (consensus management), transmits it intact and consistently to the shared ledger after encryption (network communication), and performs record storage; the smart contract module is responsible for the registration and issuance of contracts, contract triggering, and contract execution. Developers can define contract logic through a certain programming language, publish it to the blockchain (contract registration), and trigger the execution by calling keys or other events according to the logic of the contract terms to complete the contract logic. At the same time, it also provides functions for contract upgrade and cancellation; the operation detection module is mainly responsible for the deployment, configuration modification, contract setting, cloud adaptation during the product release process, and visual output of the real-time status during product operation, such as: alarming, detecting network conditions, detecting the health status of node devices, etc. The platform product service layer provides the basic capabilities and implementation frameworks of typical applications. Developers can build on these basic capabilities and overlay the characteristics of the business to complete the blockchain implementation of the business logic. The application service layer provides application services based on the blockchain solution for business participants to use.
[0083] The voice signal processing method provided by the embodiments of the present application can be applied in two-person or multi-person voice interaction scenarios, including but not limited to instant messaging, network calls, live connections, etc.
[0084] As an example, the voice signal processing method provided by the embodiments of the present application can be applied to an application scenario as Figure 1 shown. The user at the proximal end makes a voice call with the user at the distal end. After the voice signal sent by the distal user is transmitted to the proximal end, it is played through the speaker at the proximal end. The voice of the proximal user is collected by the microphone at the proximal end, and at the same time, the microphone may collect some interference signals. In the embodiments of the present application, these interference signals will be eliminated, and then the processed proximal voice signal will be sent to the distal end.
[0085] Next, the embodiments of the present application will be described in conjunction with the accompanying drawings. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0086] In an embodiment of the present application, a method for processing a voice signal is provided. As Figure 2 shown, this method can be executed by any electronic device, specifically, it can be a terminal device of any user in a voice call scenario. The method includes:
[0087] Step S201: Filter the received remote voice signal to obtain an echo prediction signal;
[0088] For example, a linear filter can be used to filter the remote voice signal to predict the acoustic echo that the remote voice signal may form after being collected by the proximal microphone, that is, to obtain the echo prediction signal. In other embodiments, other filtering methods can also be used, such as a filtering method based on a neural network, etc. The embodiments of the present application do not make limitations in this regard.
[0089] In practical applications, in an echo cancellation scenario, the signal before being sent to the speaker for playback can also be referred to as the reference signal for echo cancellation.
[0090] Step S202: Collect the proximal voice signal;
[0091] It should be noted that the step numbers in the above steps S201 and S202 do not constitute a limitation on the order of execution of the two steps, that is, the execution order of steps S201 and S202 can be without a sequence. For example, step S201 can be executed first and then step S202, or step S202 can be executed first and then step S201, or steps S201 and S202 can be executed simultaneously. That is, in the process of implementing the embodiments of the present application, the execution order of filtering the received remote voice signal to obtain the echo prediction signal and collecting the proximal voice signal is not limited.
[0092] Step S203: Obtain the first frequency domain expression of the proximal voice signal and the second frequency domain expression of the echo prediction signal;
[0093] Among them, this step can also be understood as feature extraction, extracting corresponding features from the proximal voice signal and the echo prediction signal as the input of the subsequent neural network model.
[0094] For the first frequency domain expression and the second frequency domain expression in the embodiments of the present application, they can both include but are not limited to at least one of a power spectrum, an amplitude spectrum, a logarithmic power spectrum, or a logarithmic amplitude spectrum.
[0095] In other embodiments, since the logarithmic power spectrum or the logarithmic amplitude spectrum can more closely estimate the response of the human auditory system, the first frequency domain expression and the second frequency domain expression can include at least one of the logarithmic power spectrum or the logarithmic amplitude spectrum.
[0096] Step S204: Based on the first frequency-domain representation and the second frequency-domain representation, obtain the frequency-band gain of the near-end speech signal through a pre-trained neural network model. The frequency-band gain characterizes the weight of the effective speech signal in the near-end speech signal;
[0097] Among them, the pre-trained neural network model contains pre-trained model parameters and can better output the frequency-band gain of the near-end speech signal based on the features obtained from the first frequency-domain representation and the second frequency-domain representation.
[0098] Step S205: According to the frequency-band gain, eliminate the interference signal from the near-end speech signal to obtain the processed near-end speech signal.
[0099] Among them, the frequency-band gain output by the neural network model can characterize the weight of the effective speech signal in the near-end speech signal, that is, it can be effectively used to retain the effective speech signal in the near-end speech signal. That is to say, this frequency-band gain can suppress the interference signal, thereby achieving the purpose of eliminating the interference signal.
[0100] The speech signal processing method provided by the embodiments of the present application uses the frequency-domain representations of the near-end speech signal and the echo prediction signal as the input of the neural network model. The pre-trained neural network model will output the frequency-band gain of the near-end speech signal. Thus, the frequency-band gain can more reliably characterize the weight of the effective speech signal in the near-end speech signal. Then, when eliminating the interference signal from the near-end speech signal through the frequency-band gain, it can better retain the effective speech signal in the near-end speech signal, further improving the performance of eliminating the interference signal and better eliminating the interference signal. Especially when multiple people are speaking, it can better retain the near-end human voice and improve the user experience.
[0101] In a possible implementation provided by the embodiments of the present application, step S204 may include:
[0102] Step S2041: Based on the first frequency-domain representation and the second frequency-domain representation, determine the frequency-domain information difference between the near-end speech signal and the echo prediction signal;
[0103] Step S2042: Based on the frequency-domain information difference, obtain the frequency-band gain of the near-end speech signal through the trained neural network model.
[0104] Among them, the frequency-domain information difference between the proximal speech signal and the echo prediction signal is determined, that is, the difference between the first frequency-domain expression and the second frequency-domain expression is calculated as the input of the neural network model. Since the first frequency-domain expression and the second frequency-domain expression are respectively extracted from the proximal speech signal and the echo prediction signal, the difference between the two can represent the frequency-domain information difference between the proximal speech signal and the echo prediction signal. And this frequency-domain information difference can provide information on the predicted proximal speech signal after echo cancellation. After obtaining the band gain of the proximal speech signal based on the frequency-domain information difference, the proximal speech signal can be echo-cancelled according to the band gain, so as to effectively suppress the echo signal and keep the effective proximal speech signal as much as possible.
[0105] Moreover, taking the difference between the first frequency-domain expression and the second frequency-domain expression as the input of the neural network model can reduce the number of input features. Considering that in the prior art, some neural network-based echo cancellation algorithms have an excessive amount of calculation, resulting in echo leakage caused by system instability. The neural network adopted in the embodiments of the present application requires fewer input features, which can reduce the amount of calculation to a certain extent and improve the stability of the system and the reliability of the echo cancellation performance.
[0106] It can be understood that for the embodiments of the present application, the neural network model also uses the corresponding difference between the first frequency-domain expression and the second frequency-domain expression for training in the training stage.
[0107] In the embodiments of the present application, in order to enable the band gain output by the neural network model to not only suppress the echo signal but also suppress the environmental noise, another possible implementation manner is provided. Specifically, step S2042 may include:
[0108] The first frequency-domain expression and the frequency-domain information difference are spliced to obtain the spliced frequency-domain information;
[0109] Based on the spliced frequency-domain information, the band gain of the proximal speech signal is obtained through the trained neural network model.
[0110] That is to say, in the embodiments of the present application, the difference between the first frequency-domain expression and the second frequency-domain expression is calculated, and this difference is spliced with the first frequency-domain expression as the input of the neural network model. In the training stage of the neural network model, whether the proximal speech signal contains environmental noise is distinguished and labeled. Since the difference between the first frequency-domain expression and the second frequency-domain expression can provide information on the predicted proximal speech signal after echo cancellation, and the first frequency-domain expression extracted from the proximal speech signal can provide information on whether the proximal speech signal contains noise. After obtaining the band gain of the proximal speech signal based on the spliced frequency-domain information, the proximal speech signal can be echo- and environmental-noise cancelled according to the band gain, so as to effectively cancel the echo and environmental noise and keep the proximal effective speech signal as much as possible.
[0111] The neural network adopted in the embodiments of the present application also requires a relatively small number of input features, which can reduce the amount of calculation to a certain extent, improve the stability of the system, and enhance the reliability of echo and environmental noise cancellation performance.
[0112] It can be understood that for the embodiments of the present application, the neural network model is also trained in the training stage by splicing the difference between the corresponding first frequency-domain expression and the second frequency-domain expression with the first frequency-domain expression.
[0113] A possible implementation manner is provided in the embodiments of the present application. The feature extraction process corresponding to step S203 may include:
[0114] Obtain the first spectrum of each frame of the first signal included in the near-end voice signal and the second spectrum of each frame of the second signal included in the echo prediction signal;
[0115] Based on the first spectrum of each frame of the first signal, obtain the first frequency-domain expression of each frame of the first signal;
[0116] Based on the second spectrum of each frame of the second signal, obtain the second frequency-domain expression of each frame of the second signal;
[0117] Among them, the first spectrum of each frame of the first signal may be obtained by performing time-frequency transformation on each frame of the first signal, and the second spectrum of each frame of the second signal may be obtained by performing time-frequency transformation on each frame of the second signal.
[0118] Specifically, the time-frequency transformation process may adopt the short-time Fourier transform, etc., which is not limited in the embodiments of the present application. Among them, the embodiments of the present application also do not limit the length of the short-time Fourier transform. Typical values are, for example, integer powers of 2 such as 128, 256, 512, 1024, etc., to achieve higher calculation efficiency.
[0119] Those skilled in the art can understand that the near-end voice signal is generally composed of multiple frames of signals (the above-mentioned is called the first signal). In the embodiments of the present application, when extracting the first frequency-domain expression of the near-end voice signal, each frame of the first signal included in the near-end voice signal will be processed separately, that is, the first frequency-domain expression of each frame of the first signal is obtained. Specifically, the first spectrum of each frame of the first signal included in the near-end voice signal may be obtained first, and then the first frequency-domain expression of each frame of the first signal may be obtained based on the first spectrum of each frame of the first signal.
[0120] Further, the spectrum of each frame of the first signal and each frame of the second signal includes the amplitude values of multiple frequency points. For each frame of the signal, obtaining the frequency-domain expression of the frame of the signal based on the spectrum of the frame of the signal includes:
[0121] Based on the amplitude value of each frequency point included in the spectrum of the frame signal, the frequency-domain expression corresponding to each frequency point of the frame signal is obtained.
[0122] Since the processing methods of each frame of the first signal included in the proximal speech signal are the same, the following takes the l-th frame as an example for introduction. For example, the l-th frame of the speech signal is a 20-ms speech frame. The short-time Fourier transform is performed on the l-th frame signal in the proximal speech signal to obtain a frame of spectrum signal D(k, l), where d represents the proximal speech signal, k represents the frequency point, and l represents the current frame number. Taking the short-time Fourier transform with a length of 512 points as an example, the spectrum of the l-th frame of the speech signal can be represented by K = 512 / 2 + 1, that is, 257 frequency points, that is, k = 0, 1,..., 256.
[0123] Further, based on the first spectrum of each frame of the first signal, the first frequency-domain expression of each frame of the first signal is obtained. Continuing to take the l-th frame as an example, the first frequency-domain expression of the l-th frame of the first signal is obtained based on the first spectrum D(k, l) of the l-th frame of the first signal. For example, the logarithmic power spectrum Pd(k, l) = log(|D(k, l)| 2 ) of the l-th frame of the first signal is calculated for the first spectrum D(k, l) of the l-th frame of the first signal, or the amplitude power spectrum Ad(k, l) = log(|D(k, l)|) of the l-th frame of the first signal is calculated for the first spectrum D(k, l) of the l-th frame of the first signal, etc.
[0124] Similarly, in the embodiments of the present application, when extracting the first frequency-domain expression of the echo prediction signal, each frame of the second signal included in the echo prediction signal is processed separately, that is, the second frequency-domain expression of each frame of the second signal is obtained. Specifically, the second spectrum of each frame of the second signal included in the echo prediction signal can be obtained first, and then based on the second spectrum of each frame of the second signal, the second frequency-domain expression of each frame of the second signal is obtained. It can be understood that since the echo generated by the distal speech signal is generated when the proximal speech signal is collected, the number of frames of the echo prediction signal and the proximal speech signal are in one-to-one correspondence. Among them, there may be frame signals with empty content in the echo prediction signal. In the embodiments of the present application, the same processing method is adopted for each frame of the second signal included in the echo prediction signal, regardless of the content of the frame.
[0125] Continuing with the l-th frame as an example, perform a short-time Fourier transform on the l-th frame signal in the echo prediction signal to obtain a frame of spectral signal Y(k, l), where Y represents the echo prediction signal, k represents the frequency point, and l represents the current frame number. Further, based on the second spectrum of each frame of the second signal, obtain the second frequency-domain expression of each frame of the second signal. That is, obtain the second frequency-domain expression of the l-th frame of the second signal based on the second spectrum Y(k, l) of the l-th frame of the second signal. For example, calculate the logarithmic power spectrum Py(k, l) = log(|Y(k, l)| 2 ) of the l-th frame of the second signal based on the second spectrum Y(k, l) of the l-th frame of the second signal, or calculate the amplitude power spectrum Ay(k, l) = log(|Y(k, l)|) of the l-th frame of the second signal based on the second spectrum Y(k, l) of the l-th frame of the second signal, etc.
[0126] For this implementation method, based on the first frequency-domain expression and the second frequency-domain expression, the frequency band gains of each frame of the first signal included in the near-end speech signal can be obtained through a pre-trained neural network model, including:
[0127] For each frame of the first signal, based on the first frequency-domain expression of this frame of the first signal and the second frequency-domain expression of the second signal corresponding to this frame of the first signal in the echo prediction signal, obtain the frequency band gain of this frame of the first signal through a pre-trained neural network model.
[0128] That is to say, the frequency band gains of the near-end speech signal include the frequency band gains corresponding to each frame of the first signal included therein.
[0129] Among them, the frequency band gain corresponding to each frame of the first signal includes the frequency band gains corresponding to each frequency point included in the first spectrum of this frame of the first signal.
[0130] Continuing with the l-th frame as an example, the features obtained based on Pd(k, l) and Py(k, l) will be input into the neural network model to obtain the frequency band gain g(k, l) of the l-th frame of the first signal in the near-end speech signal. Or the features that can be obtained based on Ad(k, l) and Ay(k, l) will be input into the neural network model to obtain the frequency band gain g(k, l) of the l-th frame of the first signal in the near-end speech signal.
[0131] It can be understood that if the logarithmic power spectrum is used as the input feature of the neural network model, the logarithmic power spectrum is also correspondingly used for the training of the neural network model during the training stage; if the logarithmic amplitude spectrum is used as the input feature of the neural network model, the logarithmic amplitude spectrum is also correspondingly used for the training of the neural network model during the training stage.
[0132] In an embodiment of the present application, a possible implementation is provided. For step S2041, it may include: for each frame of the first signal, based on the difference between the first frequency-domain expression of the frame of the first signal and the second frequency-domain expression of the second signal corresponding to the frame of the first signal in the echo prediction signal, the frequency-band gain of the near-end speech signal is obtained through a pre-trained neural network model.
[0133] Continuing to take the l-th frame and the frequency-domain expression as the logarithmic power spectrum as an example, the difference Q(k, l) = Pd(k, l) - Py(k, l) of the logarithmic power spectrum of the l-th frame can be calculated as the input of the neural network model. Then, if K = 257, a total of 257 eigenvalue inputs at the moment of the l-th frame are input into the neural network model, and the neural network model will output the spectral gain g(k, l) at the moment of the l-th frame.
[0134] Alternatively, the frequency-domain expression can also be the logarithmic magnitude spectrum. Calculate the difference Q(k, l) = Ad(k, l) - Ay(k, l) of the logarithmic magnitude spectrum of the l-th frame as the feature input to the neural network, and the neural network model can also output the spectral gain g(k, l) at the moment of the l-th frame.
[0135] In an embodiment of the present application, for step S2042, it may include:
[0136] For each frame of the first signal, the difference between the first frequency-domain expression of the frame of the first signal and the second frequency-domain expression of the second signal corresponding to the frame of the first signal in the echo prediction signal is concatenated with the first frequency-domain expression of the frame of the first signal, and then input into the trained neural network model to obtain the frequency-band gain of the near-end speech signal.
[0137] Continuing to take the l-th frame as an example, after calculating the difference Q(k, l) of the logarithmic power spectrum of the l-th frame, it is concatenated with the logarithmic power spectrum of the first signal of the l-th frame as a feature input to the neural network, that is, [Q(k, l) Pd(k, l)], where k = 0, 1,..., K - 1. Therefore, the total number of eigenvalue inputs for the l-th frame is 2 * K, and the neural network model will output the spectral gain g(k, l) at the moment of the l-th frame.
[0138] In an embodiment of the present application, a possible implementation is provided. The feature extraction process corresponding to step S203 can further divide the calculated spectrum into fewer sub-bands to reduce the calculation amount. Specifically, the spectrum of each frame of the first signal and each frame of the second signal includes the amplitude values of multiple frequency points. For each frame of the signal, based on the spectrum of the frame of the signal, the frequency-domain expression of the frame of the signal is obtained, including:
[0139] Divide the spectrum of the frame signal into M sub-bands, fuse the amplitude values of each frequency point corresponding to each sub-band to obtain the fused amplitude value, where M≥1; based on the fused amplitude value corresponding to each sub-band, obtain the frequency-domain expression corresponding to each sub-band.
[0140] Continuing to take the l-th frame as an example, the number of frequency points of the spectrum signal D(k, l) is K, then these K frequency points can be divided into M sub-bands as needed, where M < K. For example, when K = 257, 257 frequency points can be divided into 33 Mel sub-bands according to the human psychoacoustic model. Or other division methods can also be used, and the embodiments of the present application do not limit this here.
[0141] Further, the first spectrum of the first signal of each frame is divided into M sub-bands. For each frame of the first signal, fuse the amplitude values of each frequency point corresponding to each sub-band to obtain the fused amplitude value, and based on the fused amplitude value corresponding to each sub-band, obtain the frequency-domain expression corresponding to each sub-band. Continuing to take the l-th frame as an example, the first spectrum D(k, l) of the first signal of the l-th frame is divided into M sub-bands, and calculate the first frequency-domain expression of each sub-band of the first signal of the l-th frame respectively. For example, calculate the logarithmic power spectrum Pd(m, l) = log(|D(m, l)| 2 ) of each sub-band of the first signal of the l-th frame based on the first spectrum D(k, l) of the first signal of the l-th frame. Or calculate the amplitude power spectrum Ad(m, l) = log(|D(m, l)|) of each sub-band of the first signal of the l-th frame based on the first spectrum D(k, l) of the first signal of the l-th frame, etc. Among them, D(m, l) is obtained by fusing the frequency points corresponding to the sub-band, and the specific fusion method can be triangular filtering, addition, weighted addition, etc., but is not limited thereto.
[0142] The second spectrum of the second signal of each frame is divided into M sub-bands. For each frame of the second signal, fuse the amplitude values of each frequency point corresponding to each sub-band to obtain the fused amplitude value, and based on the fused amplitude value corresponding to each sub-band, obtain the frequency-domain expression corresponding to each sub-band. Continuing to take the l-th frame as an example, the second spectrum Y(k, l) of the second signal of the l-th frame is divided into M sub-bands, and calculate the second frequency-domain expression of each sub-band of the second signal of the l-th frame respectively. For example, calculate the logarithmic power spectrum Py(m, l) = log(|Y(m, l)| 2 ) of each sub-band of the second signal of the l-th frame based on the second spectrum Y(k, l) of the second signal of the l-th frame. Or calculate the amplitude power spectrum Ay(m, l) = log(|Y(m, l)|) of each sub-band of the second signal of the l-th frame based on the second spectrum Y(k, l) of the second signal of the l-th frame, etc. Among them, Y(m, l) is obtained by fusing the frequency points corresponding to the sub-band.
[0143] For this implementation manner, in step S204, the band gain corresponding to each frame of the first signal includes the band gains corresponding to the sub-bands included in the first spectrum of the first signal of this frame.
[0144] Continuing to take the l-th frame as an example, the features obtained based on Pd(m, l) and Py(m, l) can be input into the neural network model to obtain the band gain g(m, l) of each sub-band of the first signal of the l-th frame in the near-end speech signal. Or the features obtained based on Ad(m, l) and Ay(m, l) can be input into the neural network model to obtain the band gain g(m, l) of each sub-band of the first signal of the l-th frame in the near-end speech signal.
[0145] In an embodiment of the present application, a possible implementation manner is provided. For step S2041, it may include: for each frame of the first signal, based on the difference between the first frequency-domain expression of each sub-band of the first signal of this frame and the second frequency-domain expression of the corresponding sub-band of the second signal of the corresponding frame of the first signal in the echo prediction signal, the band gain of the near-end speech signal is obtained through a pre-trained neural network model.
[0146] Continuing to take the l-th frame and the frequency-domain expression as the logarithmic power spectrum as an example, the difference Q(m, l) = Pd(m, l) - Py(m, l) of the logarithmic power spectra of the M sub-bands of the l-th frame can be calculated, where m = 0, 1, … M - 1. As the input of the neural network model, there are M eigenvalue inputs to the neural network model at the moment of the l-th frame, and the output of the neural network model at the moment of the l-th frame is also M nodes, respectively corresponding to the band gains g(m, l) of the M sub-bands.
[0147] Or the frequency-domain expression can also be the logarithmic magnitude spectrum. Calculate the difference Q(m, l) = Ad(m, l) - Ay(m, l) of the logarithmic magnitude spectra of the M sub-bands of the l-th frame as the feature input to the neural network, and the neural network model can also output the spectral gains g(m, l) of the M sub-bands at the moment of the l-th frame.
[0148] In an embodiment of the present application, for step S2042, the number of eigenvalues can be further reduced. An optional implementation manner is:
[0149] For each frame of the first signal, based on the difference between the first frequency-domain expression of each sub-band of the first signal of this frame and the second frequency-domain expression of the corresponding sub-band of the second signal of the corresponding frame of the first signal in the echo prediction signal, they are respectively concatenated with the first frequency-domain expression of each sub-band of the first signal of this frame, and the band gain of the near-end speech signal is obtained through a pre-trained neural network model.
[0150] Continuing with the l-th frame as an example, after calculating the difference Q(m, l) of the logarithmic power spectra of the M subbands of the l-th frame, it is concatenated with the logarithmic power spectra of the M subbands of the first signal of the l-th frame as features to be input into the neural network, that is, [Q(m, l) Pd(m, l)], where m = 0, 1, …, M - 1. Therefore, the total number of feature values input for the l-th frame is 2 * M. The neural network model will output the spectral gain at the l-th frame. Those skilled in the art can set the number of output features according to the actual situation and perform training so that the neural network model will output the corresponding spectral gain. For example, the spectral gain output at the l-th frame can be g(k, l) or g(m, l).
[0151] In the embodiment of the present application, for step S2042, an alternative implementation manner is optionally as follows:
[0152] For each frame of the first signal, based on the difference between the first frequency-domain expression of each subband of the first signal of this frame and the second frequency-domain expression of the corresponding subband of the second signal of the corresponding frame of the first signal in the echo prediction signal, it is respectively concatenated with the first frequency-domain expression of each frequency point of the first signal of this frame, and the frequency-band gain of the near-end speech signal is obtained through a pre-trained neural network model.
[0153] Continuing with the l-th frame as an example, after calculating the difference Q(m, l) of the logarithmic power spectra of the M subbands of the l-th frame, it is concatenated with the logarithmic power spectra of each frequency point of the first signal of the l-th frame as features to be input into the neural network, that is, [Q(m, l) Pd(k, l)]. Therefore, the total number of feature values input for the l-th frame is M + K. The neural network model will output the spectral gain at the l-th frame. Those skilled in the art can set the number of output features according to the actual situation and perform training so that the neural network model will output the corresponding spectral gain. For example, the spectral gain output at the l-th frame can be g(k, l) or g(m, l).
[0154] In the embodiment of the present application, for step S2042, another alternative implementation manner is optionally as follows:
[0155] For each frame of the first signal, based on the difference between the first frequency-domain expression of the first signal of this frame and the second frequency-domain expression of the second signal of the corresponding frame of the first signal in the echo prediction signal, it is respectively concatenated with the first frequency-domain expression of each subband of the first signal of this frame, and the frequency-band gain of the near-end speech signal is obtained through a pre-trained neural network model.
[0156] Continuing with the l-th frame as an example, after calculating the difference Q(k, l) of the logarithmic power spectrum of the l-th frame, it is concatenated with the logarithmic power spectra of the M sub-bands into which the first signal of the l-th frame is divided as the feature input to the neural network, that is, [Q(k, l) Pd(m, l)]. Therefore, the total number of feature values input for the l-th frame is K + M. The neural network model will output the spectral gain at the l-th frame. Those skilled in the art can set the number of output features according to the actual situation and perform training so that the neural network model will output the corresponding spectral gain. For example, the spectral gain output at the l-th frame can be g(k, l) or g(m, l).
[0157] It can be understood that for the above different implementation manners, the neural network model also uses the corresponding input features for training during training.
[0158] In an embodiment of the present application, a possible implementation manner is provided. Step S205 may include:
[0159] Determine the residual signal of the near-end speech signal and the echo prediction signal;
[0160] Based on the residual signal and the frequency band gain of the near-end speech signal, obtain the processed near-end speech signal.
[0161] In practical applications, after obtaining the spectral gain g(k, l) or g(m, l) of each frame, the processed near-end speech signal can be obtained by multiplying the residual signal in the frequency domain by g(k, l) or g(m, l) and then transforming it back to the time domain.
[0162] Specifically, step S205 may include:
[0163] Step S2051: Determine the residual signal of the near-end speech signal and the echo prediction signal;
[0164] Among them, subtracting the near-end speech signal from the echo prediction signal can obtain the corresponding residual signal.
[0165] Step S2052: Obtain the third spectrum of each frame of the third signal included in the residual signal;
[0166] Among them, the third spectrum of each frame of the third signal may be obtained by performing time-frequency transformation on each frame of the third signal, such as using short-time Fourier transform, etc. Since the number of frames of the echo prediction signal and the near-end speech signal are in one-to-one correspondence, the number of frames of the residual signal of the near-end speech signal and the echo prediction signal is also in one-to-one correspondence with the number of frames of the near-end speech signal, and the processing manner of each frame of the third signal included in the residual signal is also the same. Continuing with the l-th frame as an example, for example, performing short-time Fourier transform on the l-th frame signal in the residual signal to obtain a frame of spectral signal E(k, l), where E represents the residual signal, k represents the frequency point, and l represents the current frame number.
[0167] When calculating the band gain of the near-end speech signal, if each frame of the first signal and each frame of the second signal are each divided into M sub-bands, that is, if the band gain corresponding to each frame of the first signal includes the band gains corresponding to the respective sub-bands included in the first spectrum of that frame of the first signal, then the third spectrum of each frame of the signal included in the residual signal is obtained, and the third spectrum of each frame of the signal is divided into M sub-bands;
[0168] Continuing to take the l-th frame as an example, the frequency points of the spectrum E(k, l) of the residual signal are divided into M sub-bands to obtain E(m, l).
[0169] Step S2053: For each frame of the first signal, based on the band gain corresponding to that frame of the first signal and the third spectrum of the third signal corresponding to that frame, the fourth spectrum corresponding to each frame is obtained;
[0170] Specifically, if the band gain corresponding to each frame of the first signal includes the band gains corresponding to the respective frequency points included in the first spectrum of that frame of the first signal, then for each frame of the first signal, based on the band gains of the respective frequency points corresponding to that frame of the first signal, the amplitude values of the respective frequency points included in the spectrum of the corresponding frame in the residual signal are weighted and calculated to obtain the fourth spectrum corresponding to that frame of the first signal;
[0171] Continuing to take the l-th frame as an example, each frequency point of the third spectrum of the l-th frame of the third signal is multiplied by the gain of each frequency point to obtain the fourth spectrum Z(k, l) = E(k, l) * g(k, l) after further echo cancellation.
[0172] If the band gain corresponding to each frame of the first signal includes the band gains corresponding to the respective sub-bands included in the first spectrum of that frame of the first signal, then for each sub-band of each frame of the first signal, based on the band gain corresponding to each sub-band, the amplitude values of the respective frequency points included in the corresponding sub-band of the third spectrum of the corresponding frame in the residual signal are weighted and calculated to obtain the fourth spectrum corresponding to each sub-band.
[0173] Continuing to take the l-th frame as an example, each frequency point of each sub-band of the l-th frame of the third spectrum is multiplied by the gain of the corresponding sub-band to obtain the fourth spectrum after further echo cancellation, that is, Z(k, l) = E(k, l) * g(m, l). As an example, when dividing k frequency points into m sub-bands, assume that the 1st to 3rd frequency points correspond to the first sub-band, the 4th to 8th frequency points correspond to the second sub-band, and so on. Correspondingly, the gain of the first sub-band is multiplied by the amplitudes of the 1st to 3rd frequency points, the gain of the second sub-band is multiplied by the amplitudes of the 4th to 8th frequency points, and so on.
[0174] Step S2054: Based on the respective fourth spectra corresponding to the respective first signals, a frequency-time transformation is performed to obtain the processed near-end speech signal.
[0175] Perform frequency-time transformation on the fourth spectra corresponding to each frame, and transform the fourth spectra in the frequency domain into signals in the time domain, then the processed near-end speech signal can be obtained.
[0176] Specifically, the frequency-time transformation can adopt the inverse transformation of the time-frequency transformation used in the foregoing steps. For example, if the above time-frequency transformation adopts the short-time Fourier transform, then in this step, the inverse short-time Fourier transform can be adopted.
[0177] Continuing to take the l-th frame as an example, perform the inverse short-time Fourier transform on the fourth spectrum Z(k, l) or Z(m, l) corresponding to the l-th frame to obtain the l-th frame z(l) of the time-domain digital audio signal frame. Then, connecting the signals of each frame together can obtain the complete time-domain signal z, and z is the processed near-end speech signal that further eliminates echo and tries to retain the effective speech signal as much as possible.
[0178] In the embodiments of the present application, a feasible implementation manner is provided for the neural network model, as Figure 3 shown, the neural network model can be divided into three parts, including an input layer, a hidden layer, and an output layer. The input layer has multiple (including but not limited to K, M, 2K, 2M, K + M, etc.) input nodes, and each input node corresponds to a feature value. For example, there are K input nodes, and K is equal to the aforementioned 257. The input layer can be a fully connected layer, an LSTM (Long-Short Term Memory), a GRU (Gate Recurrent Unit), or a CNN (Convolutional Neural Networks), etc. types of network layers. In some cases, the hidden layer may not exist, and in some cases, the hidden layer may have more than one layer, and there may be multiple hidden layers. Those skilled in the art can set it according to actual needs, and the embodiments of the present application do not make limitations here. The type of the hidden layer can also be a fully connected layer, an LSTM, a GRU, or a CNN, etc. types of network layers, and the number of nodes can generally be between 5 and 5000, without limitation. The output layer can be a fully connected layer, adopting the sigmoid non-linear function, and the number of nodes is the number of frequency bands (number of frequency points) or the number of sub-bands. For example, it can be the same as the previous K, that is, 257. Each node outputs a gain. Continuing to take the l-th frame with K = 257 as an example, the output of the output layer at the l-th frame is g(k, l), where k = 0, 1,..., 256. In other embodiments, the neural network model can also be other structures.
[0179] In an embodiment of the present application, a possible implementation manner is provided. The feature extraction process corresponding to step S203 can be performed by a feature extraction layer, and the feature extraction layer and the foregoing neural network model can be encapsulated into a neural network module. In this neural network module, the foregoing neural network model can also be referred to as a network layer. As an example, as Figure 4 shown, the neural network module may include a feature extraction layer, a network input layer, a network hidden layer, and a network output layer. The neural network module can directly process the proximal speech signal and the echo prediction signal, and output the band gain of the proximal speech signal.
[0180] In an embodiment of the present application, a possible implementation manner is provided. After step S201, the following steps may further be included:
[0181] Adjust the filter parameters according to the residual signal between the proximal speech signal and the echo prediction signal;
[0182] Filter the subsequent received distal speech signal based on the adjusted filter parameters.
[0183] In an embodiment of the present application, the filter parameters are dynamically adjusted through the residual signal to determine the change of the echo path and improve the accuracy of echo cancellation.
[0184] Figure 5 Taking the example of using a linear filter to perform the filtering process, an example of the filtering process is shown. As Figure 5 shown, the received distal speech signal, that is, the reference signal x, is sent to the speaker for playback. After passing through the echo path, it enters the proximal microphone to form an echo y'. That is to say, when the proximal audio signal d is collected by the microphone, it may simultaneously collect the speech s of the proximal speaker and the acoustic echo y', and may also collect ambient noise n, that is, d = y' + s + n. The echo prediction signal y is obtained by filtering x through a linear filter, and then d is subtracted from y to obtain the residual signal e after linear echo cancellation. At the same time, the linear filter parameters are dynamically adjusted through e to determine the change of the echo path. Based on the adjusted filter parameters, the subsequent received distal speech signal x is filtered.
[0185] Figure 6 A complete example of the speech signal processing method provided in the embodiment of the present application is shown. As Figure 6As shown, the embodiment of the present application adopts a method of a linear filter echo cancellation module plus a neural network module. Specifically, the received remote voice signal x is sent to the speaker for playback, and after passing through the echo path, it enters the proximal microphone to form an echo y'. When the proximal audio signal d is collected by the microphone, it may simultaneously collect the voice s of the proximal speaker and the acoustic echo y', and may also collect ambient noise n, that is, d = y' + s + n. After filtering x through the linear filter, an echo prediction signal y is obtained. The neural network module (corresponding to Figure 6 the neural network in
[0186] An implementation manner is provided in the embodiment of the present application. The foregoing neural network model is trained in the following manner:
[0187] Obtain a plurality of training samples. Each training sample includes a remote sample voice signal, a proximal sample voice signal, and annotation information, and the annotation information represents the true frequency band gain of the proximal sample voice signal;
[0188] For example, annotation information such as "0" and "1" can be used to distinguish the interfering speech part and the valid speech part in the true frequency band gain, or other annotation methods can also be used.
[0189] Perform filtering processing on the remote sample voice signal of each training sample to obtain a sample echo prediction signal;
[0190] The specific filtering processing method can be referred to the description above and will not be elaborated here.
[0191] Determine the third frequency domain expression of the proximal sample voice signal and the fourth frequency domain expression of the sample echo prediction signal of each training sample;
[0192] The specific method for determining the frequency domain expression can be referred to the description above and will not be elaborated here.
[0193] Based on the third frequency domain expression and the fourth frequency domain expression corresponding to each training sample, use a machine learning method to iteratively train the initial neural network model to obtain the predicted frequency band gain corresponding to each training sample;
[0194] Among them, for each training, if it is determined that the training end condition is met based on the true band gain and the predicted band gain corresponding to each training sample, a trained neural network model is obtained. If the training end condition is not met, the model parameters of the neural network model are adjusted, and the neural network model is continuously trained based on the third frequency domain expression and the fourth frequency domain expression corresponding to each training sample.
[0195] Furthermore, for each training sample, it may further include:
[0196] Based on the third frequency domain expression and the fourth frequency domain expression, determine the sample frequency domain information difference between the proximal sample speech signal and the sample echo prediction signal; wherein, during training, the input of the neural network model includes the sample frequency domain information difference corresponding to each training sample, or includes the sample frequency domain information difference corresponding to each training sample and the third frequency domain expression.
[0197] The neural network model trained in this way can, when applied, correspond to inputting the input features of various situations introduced in the foregoing content and output the corresponding band gain. When using this band gain to eliminate the interference signal, it can better retain the effective speech signal in the proximal speech signal, further improve the elimination performance of the interference signal, and better eliminate the interference signal. Especially when multiple parties are speaking, it can better retain the proximal human voice and improve the user experience.
[0198] Those skilled in the art can understand that the foregoing "first frequency domain expression" to "fourth frequency domain expression" only represent the distinction of the frequency domain expressions of different signals, and cannot be understood as the limitation of their magnitudes or quantities. Similarly, the "first spectrum" to "fourth spectrum" only represent the distinction of the spectra of different signals, and the "first signal" to "third signal" only represent the distinction of different types of signals, and cannot be understood as the limitation of them.
[0199] The embodiment of the present application provides a processing device for speech signals, such as Figure 7 shown, the processing device 70 may include: a signal filtering module 701, a signal acquisition module 702, a frequency domain expression acquisition module 703, a band gain determination module 704, and an interference signal elimination module 705, wherein,
[0200] The signal filtering module 701 is used to filter the received distal speech signal to obtain an echo prediction signal;
[0201] The signal acquisition module 702 is used to acquire the proximal speech signal;
[0202] The frequency domain expression acquisition module 703 is used to acquire the first frequency domain expression of the proximal speech signal and the second frequency domain expression of the echo prediction signal;
[0203] The band gain determination module 704 is configured to obtain the band gain of the near-end speech signal based on the first frequency-domain expression and the second frequency-domain expression through a pre-trained neural network model, where the band gain characterizes the weight of the valid speech signal in the near-end speech signal;
[0204] The interference signal cancellation module 705 is configured to cancel the interference signal from the near-end speech signal according to the band gain to obtain the processed near-end speech signal.
[0205] In an optional implementation manner, when the band gain determination module 704 is configured to obtain the band gain of the near-end speech signal based on the first frequency-domain expression and the second frequency-domain expression through a pre-trained neural network model, it is specifically configured to:
[0206] Determine the frequency-domain information difference between the near-end speech signal and the echo prediction signal based on the first frequency-domain expression and the second frequency-domain expression;
[0207] Obtain the band gain of the near-end speech signal through the trained neural network model based on the frequency-domain information difference.
[0208] In an optional implementation manner, when the band gain determination module 704 is configured to obtain the band gain of the near-end speech signal through the trained neural network model based on the frequency-domain information difference, it is specifically configured to:
[0209] Concatenate the first frequency-domain expression and the frequency-domain information difference to obtain the concatenated frequency-domain information;
[0210] Obtain the band gain of the near-end speech signal through the trained neural network model based on the concatenated frequency-domain information.
[0211] In an optional implementation manner, when the frequency-domain expression acquisition module 703 is configured to obtain the first frequency-domain expression of the near-end speech signal and the second frequency-domain expression of the echo prediction signal, it is specifically configured to:
[0212] Obtain the first spectrum of each frame of the first signal included in the near-end speech signal and the second spectrum of each frame of the second signal included in the echo prediction signal;
[0213] Obtain the first frequency-domain expression of each frame of the first signal based on the first spectrum of each frame of the first signal;
[0214] Obtain the second frequency-domain expression of each frame of the second signal based on the second spectrum of each frame of the second signal;
[0215] When the band gain determination module 704 is configured to obtain the band gain of each frame of the first signal included in the near-end speech signal based on the first frequency-domain expression and the second frequency-domain expression through a pre-trained neural network model, it is specifically configured to:
[0216] For each frame of the first signal, based on the first frequency-domain expression of the first signal in this frame and the second frequency-domain expression of the second signal corresponding to this frame of the first signal in the echo prediction signal, the frequency-band gain of this frame of the first signal is obtained through a pre-trained neural network model.
[0217] In an alternative implementation, the spectra of each frame of the first signal and each frame of the second signal each include amplitude values of multiple frequency points. For each frame of the signal, when the frequency-domain expression obtaining module 703 is used to obtain the frequency-domain expression of this frame of the signal based on the spectrum of this frame of the signal, it is specifically used for:
[0218] Based on the amplitude value of each frequency point included in the spectrum of this frame of the signal, the frequency-domain expression corresponding to each frequency point of this frame of the signal is obtained; wherein, the frequency-band gain corresponding to this frame of the first signal includes the frequency-band gains corresponding to the frequency points included in the first spectrum of this frame of the first signal;
[0219] When the interference signal cancellation module 705 is used to cancel the interference signal from the near-end speech signal according to the frequency-band gain to obtain the processed near-end speech signal, it is specifically used for:
[0220] Determine the residual signal between the near-end speech signal and the echo prediction signal;
[0221] Obtain the third spectrum of each frame of the signal included in the residual signal;
[0222] For each frame of the first signal, based on the frequency-band gains corresponding to the frequency points of this frame of the first signal, perform weighted calculation on the amplitude values of the frequency points included in the spectrum of the corresponding frame in the residual signal to obtain the fourth spectrum corresponding to this frame of the first signal;
[0223] Based on the fourth spectra corresponding to each first signal, perform frequency-time transformation to obtain the processed near-end speech signal.
[0224] In an alternative implementation, the spectra of each frame of the first signal and each frame of the second signal each include amplitude values of multiple frequency points. For each frame of the signal, when the frequency-domain expression obtaining module 703 is used to obtain the frequency-domain expression of this frame of the signal based on the spectrum of this frame of the signal, it is specifically used for:
[0225] Divide the spectrum of this frame of the signal into M sub-bands, fuse the amplitude values of the frequency points corresponding to each sub-band to obtain the fused amplitude value, where M≥1;
[0226] Based on the fused amplitude value corresponding to each sub-band, obtain the frequency-domain expression corresponding to each sub-band; wherein, the frequency-band gain corresponding to this frame of the first signal includes the frequency-band gains corresponding to the sub-bands included in the first spectrum of this frame of the first signal;
[0227] When the interference signal cancellation module 705 is used to cancel the interference signal of the near-end speech signal according to the band gain and obtain the processed near-end speech signal, it is specifically used for:
[0228] Determine the residual signal between the near-end speech signal and the echo prediction signal;
[0229] Obtain the third spectrum of each frame signal included in the residual signal, and divide the third spectrum of each frame signal into M sub-bands;
[0230] For each sub-band of each frame of the first signal, based on the band gain corresponding to each sub-band, perform weighted calculation on the amplitude values of each frequency point included in the corresponding sub-band of the third spectrum of the corresponding frame in the residual signal to obtain the fourth spectrum corresponding to each sub-band;
[0231] Perform frequency-time transformation based on the fourth spectra corresponding to the first signals to obtain the processed near-end speech signal.
[0232] In an optional implementation manner, both the first frequency-domain expression and the second frequency-domain expression include at least one of power spectrum, amplitude spectrum, logarithmic power spectrum, or logarithmic amplitude spectrum.
[0233] In an optional implementation manner, the neural network model is trained through the following method:
[0234] Obtain a plurality of training samples, each training sample includes a far-end sample speech signal, a near-end sample speech signal, and annotation information, and the annotation information represents the true band gain of the near-end sample speech signal;
[0235] Perform filtering processing on the far-end sample speech signal of each training sample to obtain a sample echo prediction signal;
[0236] Determine the third frequency-domain expression of the near-end sample speech signal and the fourth frequency-domain expression of the sample echo prediction signal of each training sample;
[0237] Based on the third frequency-domain expression and the fourth frequency-domain expression corresponding to each training sample, use a machine learning method to perform iterative training on the initial neural network model to obtain the predicted band gain corresponding to each training sample;
[0238] Among them, for each training, if it is determined that the training end condition is satisfied based on the true band gain and the predicted band gain corresponding to each training sample, the trained neural network model is obtained. If the training end condition is not satisfied, the model parameters of the neural network model are adjusted, and the neural network model is continuously trained based on the third frequency-domain expression and the fourth frequency-domain expression corresponding to each training sample.
[0239] In an optional implementation manner, for each training sample, it further includes:
[0240] Based on the third frequency-domain representation and the fourth frequency-domain representation, determine the sample frequency-domain information difference between the proximal sample speech signal and the sample echo prediction signal; wherein, during training, the input of the neural network model includes the sample frequency-domain information difference corresponding to each training sample, or includes the sample frequency-domain information difference corresponding to each training sample and the third frequency-domain representation.
[0241] Those skilled in the art can clearly understand that for the speech signal processing device provided in the embodiments of the present application, its implementation principle and the resulting technical effects are the same as those of the foregoing method embodiments. For the sake of convenience and conciseness of description, for the parts not mentioned in the device embodiments, reference can be made to the corresponding content in the foregoing method embodiments, and details will not be repeated here.
[0242] Among them, the speech signal processing device provided in the embodiments of the present application can be a computer program (including program code) running in a computer device. For example, the speech signal processing device is an application software; this device can be used to execute the corresponding content in the foregoing method embodiments.
[0243] In some embodiments, the speech signal processing device provided in the embodiments of the present application can be implemented in a combination of software and hardware. As an example, the speech signal processing device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the speech signal processing method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.
[0244] In other embodiments, the speech signal processing device provided in the embodiments of the present application can be implemented in software. It can be software in the form of programs and plugins, etc., and includes a series of modules for implementing the speech signal processing method provided in the embodiments of the present application.
[0245] The modules described in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the module itself in some cases.
[0246] Based on the same principle as the method shown in the embodiments of the present application, embodiments of the present application also provide an electronic device, which may include but is not limited to: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the processing method of the voice signal shown in any embodiment of the present application by calling the computer program.
[0247] In an alternative embodiment, an electronic device is provided, as Figure 8 shown, Figure 8 the electronic device 800 shown includes: a processor 801 and a memory 803. Among them, the processor 801 and the memory 803 are connected, such as connected through a bus 802. Optionally, the electronic device 800 may further include a transceiver 804, and the transceiver 804 may be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data, etc. It should be noted that in practical applications, the transceiver 804 is not limited to one, and the structure of the electronic device 800 does not constitute a limitation on the embodiments of the present application.
[0248] The processor 801 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in connection with the disclosure of the present application. The processor 801 may also be a combination for implementing computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0249] The bus 802 may include a path for transmitting information between the above components. The bus 802 may be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard architecture) bus, etc. The bus 802 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0250] The memory 803 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0251] The memory 803 is used to store the application program code (computer program) for executing the solution of this application, and is controlled by the processor 801 for execution. The processor 801 is used to execute the application program code stored in the memory 803 to implement the content shown in the foregoing method embodiments.
[0252] Among them, the electronic device can also be a terminal device. Figure 8 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.
[0253] The embodiments of this application provide a computer-readable storage medium, on which a computer program is stored. When it runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiments.
[0254] According to another aspect of this application, there is also provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for processing voice signals provided in the above various implementation manners of the embodiments.
[0255] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0256] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the methods and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and this module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0257] The computer-readable storage medium provided by the embodiments of this application can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0258] The above computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to execute the method shown in the above embodiments.
[0259] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present application that have similar functions.
Claims
1. A method for processing a voice signal, characterized in that, including: filtering the received remote voice signal to obtain an echo prediction signal; acquiring a proximal voice signal; obtaining a first frequency-domain representation of the proximal voice signal and a second frequency-domain representation of the echo prediction signal; obtaining a frequency-band gain of the proximal voice signal based on the first frequency-domain representation and the second frequency-domain representation through a pre-trained neural network model, where the frequency-band gain characterizes the weight of the valid voice signal in the proximal voice signal; eliminating interference signals from the proximal voice signal according to the frequency-band gain to obtain a processed proximal voice signal; The eliminating interference signals from the proximal voice signal according to the frequency-band gain to obtain a processed proximal voice signal includes: determining a residual signal between the proximal voice signal and the echo prediction signal; obtaining a third spectrum of each frame signal included in the residual signal and dividing the third spectrum of each frame signal into M sub-bands; weighting and calculating the amplitude values of each frequency point included in the corresponding sub-band of the third spectrum of the corresponding frame in the residual signal based on the frequency-band gain corresponding to each sub-band to obtain a fourth spectrum corresponding to each sub-band; performing frequency-time transformation based on each fourth spectrum to obtain a processed proximal voice signal.
2. The processing method according to claim 1, characterized in that, The obtaining a frequency-band gain of the proximal voice signal based on the first frequency-domain representation and the second frequency-domain representation through a pre-trained neural network model includes: determining the frequency-domain information difference between the proximal voice signal and the echo prediction signal based on the first frequency-domain representation and the second frequency-domain representation; obtaining the frequency-band gain of the proximal voice signal through a trained neural network model based on the frequency-domain information difference.
3. The processing method according to claim 2, characterized in that, The obtaining a frequency-band gain of the proximal voice signal through a trained neural network model based on the frequency-domain information difference includes: concatenating the first frequency-domain representation and the frequency-domain information difference to obtain concatenated frequency-domain information; obtaining the frequency-band gain of the proximal voice signal through a trained neural network model based on the concatenated frequency-domain information.
4. The processing method according to any one of claims 1 to 3, characterized in that, The obtaining a first frequency-domain representation of the proximal voice signal and a second frequency-domain representation of the echo prediction signal includes: obtaining a first spectrum of each frame of first signal included in the proximal voice signal and a second spectrum of each frame of second signal included in the echo prediction signal; obtaining a first frequency-domain representation of each frame of first signal based on the first spectrum of each frame of first signal; obtaining a second frequency-domain representation of each frame of second signal based on the second spectrum of each frame of second signal; The obtaining a frequency-band gain of each frame of first signal included in the proximal voice signal based on the first frequency-domain representation and the second frequency-domain representation through a pre-trained neural network model includes: for each frame of first signal, obtaining the frequency-band gain of this frame of first signal through a pre-trained neural network model based on the first frequency-domain representation of this frame of first signal and the second frequency-domain representation of the second signal corresponding to this frame of first signal in the echo prediction signal.
5. The processing method according to claim 4, characterized in that, The spectrum of each frame of the first signal and each frame of the second signal includes the amplitude values of multiple frequency points. For each frame of signal, based on the spectrum of this frame of signal, the frequency-domain expression of this frame of signal is obtained, including: Based on the amplitude value of each frequency point included in the spectrum of this frame of signal, the frequency-domain expression corresponding to each frequency point of this frame of signal is obtained; wherein, the frequency-band gain corresponding to this frame of the first signal includes the frequency-band gains corresponding to the frequency points included in the first spectrum of this frame of the first signal; The eliminating the interference signal from the proximal speech signal according to the frequency-band gain to obtain the processed proximal speech signal includes: Determining the residual signal between the proximal speech signal and the echo prediction signal; Obtaining the third spectrum of each frame of signal included in the residual signal; For each frame of the first signal, based on the frequency-band gains of the frequency points corresponding to this frame of the first signal, weighted calculation is performed on the amplitude values of the frequency points included in the spectrum of the corresponding frame in the residual signal to obtain the fourth spectrum corresponding to this frame of the first signal; Based on the fourth spectra corresponding to each of the first signals, frequency-time transformation is performed to obtain the processed proximal speech signal.
6. The processing method according to claim 4, characterized in that, The spectrum of each frame of the first signal and each frame of the second signal includes the amplitude values of multiple frequency points. For each frame of signal, based on the spectrum of this frame of signal, the frequency-domain expression of this frame of signal is obtained, including: Dividing the spectrum of this frame of signal into M sub-bands, fusing the amplitude values of the frequency points corresponding to each sub-band to obtain the fused amplitude value, where M≥1; Based on the fused amplitude value corresponding to each sub-band, the frequency-domain expression corresponding to each sub-band is obtained; wherein, the frequency-band gain corresponding to this frame of the first signal includes the frequency-band gains corresponding to the sub-bands included in the first spectrum of this frame of the first signal; The eliminating the interference signal from the proximal speech signal according to the frequency-band gain to obtain the processed proximal speech signal includes: For each sub-band of each frame of the first signal, based on the frequency-band gain corresponding to each sub-band, weighted calculation is performed on the amplitude values of the frequency points included in the corresponding sub-band in the third spectrum of the corresponding frame in the residual signal to obtain the fourth spectrum corresponding to each sub-band; Based on the fourth spectra corresponding to each of the first signals, frequency-time transformation is performed to obtain the processed proximal speech signal.
7. The processing method according to claim 1, characterized in that, The first frequency-domain expression and the second frequency-domain expression both include at least one of power spectrum, amplitude spectrum, logarithmic power spectrum or logarithmic amplitude spectrum.
8. The processing method according to any one of claims 1-3, characterized in that, The neural network model is trained in the following way: Obtaining a plurality of training samples, each training sample including a distal sample speech signal, a proximal sample speech signal and annotation information, and the annotation information characterizing the true frequency-band gain of the proximal sample speech signal; Performing filtering processing on the distal sample speech signal of each training sample to obtain a sample echo prediction signal; Determining the third frequency-domain expression of the proximal sample speech signal and the fourth frequency-domain expression of the sample echo prediction signal of each training sample; Based on the third frequency-domain expression and the fourth frequency-domain expression corresponding to each of the training samples, use a machine learning method to iteratively train the initial neural network model to obtain the predicted band gain corresponding to each of the training samples; Among them, for each training, if it is determined that the training end condition is met based on the true band gain and the predicted band gain corresponding to each of the training samples, then the trained neural network model is obtained. If the training end condition is not met, then adjust the model parameters of the neural network model, and continue to train the neural network model based on the third frequency-domain expression and the fourth frequency-domain expression corresponding to each of the training samples.
9. The processing method according to claim 8, characterized in that, For each of the training samples, it further includes: Based on the third frequency-domain expression and the fourth frequency-domain expression, determine the sample frequency-domain information difference between the near-end sample voice signal and the sample echo prediction signal; wherein, during training, the input of the neural network model includes the sample frequency-domain information difference corresponding to each of the training samples, or includes the sample frequency-domain information difference corresponding to each of the training samples and the third frequency-domain expression.
10. A processing device for a voice signal, characterized in that, It includes: A signal filtering module for filtering the received far-end voice signal to obtain an echo prediction signal; A signal acquisition module for acquiring a near-end voice signal; A frequency-domain expression acquisition module for acquiring the first frequency-domain expression of the near-end voice signal and the second frequency-domain expression of the echo prediction signal; A band gain determination module for obtaining the band gain of the near-end voice signal through a pre-trained neural network model based on the first frequency-domain expression and the second frequency-domain expression, where the band gain characterizes the weight of the effective voice signal in the near-end voice signal; An interference signal cancellation module for canceling the interference signal of the near-end voice signal according to the band gain to obtain the processed near-end voice signal; When the interference signal cancellation module is used to cancel the interference signal of the near-end voice signal according to the band gain to obtain the processed near-end voice signal, it specifically is used for: Determine the residual signal between the near-end voice signal and the echo prediction signal; Obtain the third spectrum of each frame signal included in the residual signal, and divide the third spectrum of each frame signal into M sub-bands; Based on the band gain corresponding to each sub-band, perform weighted calculation on the amplitude values of the frequency points included in the corresponding sub-band of the third spectrum of the corresponding frame in the residual signal to obtain the fourth spectrum corresponding to each sub-band; Based on each fourth spectrum, perform frequency-time transformation to obtain the processed near-end voice signal.
11. The processing device according to claim 10, characterized in that, When the band gain determination module is used to obtain the band gain of the near-end voice signal through a pre-trained neural network model based on the first frequency-domain expression and the second frequency-domain expression, it specifically is used for: Based on the first frequency-domain expression and the second frequency-domain expression, determine the frequency-domain information difference between the near-end voice signal and the echo prediction signal; Based on the frequency-domain information difference, obtain the band gain of the near-end voice signal through the trained neural network model.
12. The processing device according to claim 11, characterized in that, When the band gain determination module is used to obtain the band gain of the near-end speech signal through a trained neural network model based on the frequency-domain information difference, it is specifically used for: Concatenate the first frequency-domain expression and the frequency-domain information difference to obtain the concatenated frequency-domain information; Based on the concatenated frequency-domain information, obtain the band gain of the near-end speech signal through a trained neural network model.
13. The processing device according to any one of claims 10 to 12, characterized in that, When the frequency-domain expression acquisition module is used to obtain the first frequency-domain expression of the near-end speech signal and the second frequency-domain expression of the echo prediction signal, it is specifically used for: Obtain the first spectrum of each frame of the first signal included in the near-end speech signal and the second spectrum of each frame of the second signal included in the echo prediction signal; Based on the first spectrum of each frame of the first signal, obtain the first frequency-domain expression of each frame of the first signal; Based on the second spectrum of each frame of the second signal, obtain the second frequency-domain expression of each frame of the second signal; When the band gain determination module is used to obtain the band gain of each frame of the first signal included in the near-end speech signal through a pre-trained neural network model based on the first frequency-domain expression and the second frequency-domain expression, it is specifically used for: For each frame of the first signal, based on the first frequency-domain expression of this frame of the first signal and the second frequency-domain expression of the second signal corresponding to this frame of the first signal in the echo prediction signal, obtain the band gain of this frame of the first signal through a pre-trained neural network model.
14. The processing device according to claim 13, wherein, The spectrum of each frame of the first signal and each frame of the second signal includes the amplitude values of multiple frequency points. For each frame of the signal, when the frequency-domain expression acquisition module is used to obtain the frequency-domain expression of this frame of the signal based on the spectrum of this frame of the signal, it is specifically used for: Based on the amplitude value of each frequency point included in the spectrum of this frame of the signal, obtain the frequency-domain expression corresponding to each frequency point of this frame of the signal; wherein, the band gain corresponding to this frame of the first signal includes the band gains corresponding to the frequency points included in the first spectrum of this frame of the first signal; When the interference signal cancellation module is used to cancel the interference signal of the near-end speech signal according to the band gain and obtain the processed near-end speech signal, it is specifically used for: Determine the residual signal between the near-end speech signal and the echo prediction signal; Obtain the third spectrum of each frame of the signal included in the residual signal; For each frame of the first signal, based on the band gains of the corresponding frequency points of this frame of the first signal, perform weighted calculation on the amplitude values of the frequency points included in the spectrum of the corresponding frame in the residual signal to obtain the fourth spectrum corresponding to this frame of the first signal; Based on the fourth spectra corresponding to each of the first signals, perform frequency-time transformation to obtain the processed near-end speech signal.
15. The processing device according to claim 13, wherein, The spectrum of each frame of the first signal and each frame of the second signal includes the amplitude values of multiple frequency points. For each frame of the signal, when the frequency-domain expression acquisition module is used to obtain the frequency-domain expression of this frame of the signal based on the spectrum of this frame of the signal, it is specifically used for: Divide the spectrum of this frame of the signal into M sub-bands, and fuse the amplitude values of the corresponding frequency points of each sub-band to obtain the fused amplitude value, where M≥1; Based on the fused amplitude value corresponding to each sub-band, obtain the frequency-domain expression corresponding to each sub-band; wherein, the frequency-band gain corresponding to the first signal of this frame includes the frequency-band gains corresponding to the sub-bands included in the first spectrum of the first signal of this frame. When the interference signal cancellation module is used to cancel the interference signal of the near-end speech signal according to the frequency-band gain and obtain the processed near-end speech signal, it is specifically used for: For each sub-band of each frame of the first signal, based on the frequency-band gain corresponding to each sub-band, perform weighted calculation on the amplitude values of each frequency point included in the corresponding sub-band of the third spectrum of the corresponding frame in the residual signal to obtain the fourth spectrum corresponding to each sub-band. Based on the fourth spectra corresponding to each of the first signals, perform frequency-time transformation to obtain the processed near-end speech signal.
16. The processing device according to claim 10, wherein, Both the first frequency-domain expression and the second frequency-domain expression include at least one of power spectrum, amplitude spectrum, logarithmic power spectrum, or logarithmic amplitude spectrum.
17. The processing device according to any one of claims 10 - 12, wherein, The neural network model is trained in the following manner: Obtain a plurality of training samples, each training sample includes a far-end sample speech signal, a near-end sample speech signal, and annotation information, and the annotation information characterizes the true frequency-band gain of the near-end sample speech signal. Perform filtering processing on the far-end sample speech signal of each training sample to obtain a sample echo prediction signal. Determine the third frequency-domain expression of the near-end sample speech signal and the fourth frequency-domain expression of the sample echo prediction signal for each training sample. Based on the third frequency-domain expression and the fourth frequency-domain expression corresponding to each training sample, use a machine learning method to perform iterative training on the initial neural network model to obtain the predicted frequency-band gain corresponding to each training sample. Among them, for each training, if it is determined that the training end condition is satisfied based on the true frequency-band gain and the predicted frequency-band gain corresponding to each training sample, then obtain the trained neural network model; if the training end condition is not satisfied, then adjust the model parameters of the neural network model, and continue to train the neural network model based on the third frequency-domain expression and the fourth frequency-domain expression corresponding to each training sample.
18. The processing device according to claim 17, wherein, For each training sample, it further includes: Based on the third frequency-domain expression and the fourth frequency-domain expression, determine the sample frequency-domain information difference between the near-end sample speech signal and the sample echo prediction signal; wherein, during training, the input of the neural network model includes the sample frequency-domain information difference corresponding to each training sample, or includes the sample frequency-domain information difference corresponding to each training sample and the third frequency-domain expression.
19. An electronic device, wherein, It includes: A processor and a memory. The memory stores a computer program, and the computer program is loaded and executed by the processor to implement the method according to any one of claims 1-9.
20. A computer-readable storage medium, wherein, The computer storage medium is used to store a computer program, and when the computer program runs on a computer, it causes the computer to execute the method according to any one of claims 1-9.
21. A computer program product, wherein, The computer program product includes computer instructions, and the processor executes the computer instructions to implement the method according to any one of claims 1-9.
Citation Information
Patent Citations
Voice signal enhancement method and device, electronic equipment and storage medium
CN111968658A