Method, apparatus, electronic device, and computer program for processing audio signals
The method processes audio signals during multi-person voice calls to detect when a user is speaking while muted, prompting them to unmute and improving communication efficiency and user experience.
Patent Information
- Application Number
- JP2023551247
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-16
- Filing Date
- 2022-08-10
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2042-08-10
AI Technical Summary
During multi-person voice calls, users often forget that their microphone is muted, leading to inefficient man-machine interaction as they cannot respond to others until they manually unmute their microphone.
A method and apparatus that process audio signals by collecting audio even when the microphone is muted, analyzing gain parameters across frequency bands, and prompting the user to unmute when target voice is detected.
This solution improves communication efficiency by promptly alerting users to unmute their microphones, reducing the need for repeated speech and enhancing user experience.
Smart Images

Figure 0007694968000001 
Figure 0007694968000002 
Figure 0007694968000003
Abstract
Description
Technical Field
[0001] This application relates to the technical field of audio, and particularly to a method and apparatus for processing an audio signal, an electronic device, and a storage medium.
[0002] This application claims the priority of a Chinese patent application with an application number of 202111087468.5 and an invention title of "Method, Apparatus, Electronic Device, and Storage Medium for Processing Audio Signals", which was filed on September 16, 2021, and the entire content thereof is incorporated herein by reference.
Background Art
[0003] With the development of audio technology and the diversification of terminal functions, it is possible to make voice calls between different terminals based on the VoIP (Voice over Internet Protocol) technology.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Embodiments of this application provide a method and apparatus for processing an audio signal, an electronic device, and a storage medium, which can improve the man-machine interaction efficiency in the state where a user's microphone is off during a multi-person voice call. The technical solution is as follows.
[0005] In one aspect, a method for processing an audio signal is provided, which is executed by a terminal. The method includes: obtaining an audio signal collected by an application program in a target scene, where the application program has an account logged in, and the target scene refers to a state where the account is in a microphone-muted state during a multi-person voice call; The step of obtaining gain parameters on a plurality of frequency bands in a first frequency band range of each of a plurality of audio frames in the audio signal; When it is determined that the target voice is included in the audio signal based on the gain parameter, the step of outputting a prompt message, wherein the prompt message is used to prompt to release the microphone mute state of the account.
[0006] In one aspect, there is provided a processing device for an audio signal, which is arranged in a terminal, and the device includes: A first acquisition module used to acquire an audio signal collected by an application program in a target scene, wherein the application program has an account logged in, and the target scene refers to the account being in a microphone mute state in a multi-person voice call. A second acquisition module used to acquire gain parameters on a plurality of frequency bands in a first frequency band range of each of a plurality of audio frames in the audio signal; An output module used to output a prompt message when it is determined that the target voice is included in the audio signal based on the gain parameter, wherein the prompt message is used to prompt to release the microphone mute state of the account.
[0007] In one aspect, there is provided an electronic device, which includes one or more processors and one or more memories, and at least one computer program is stored in the one or more memories, and the at least one computer program is loaded and executed by the one or more processors to realize the processing method of the audio signal.
[0008] In one aspect, a storage medium is provided, and at least one computer program is stored in the storage medium. The at least one computer program is loaded and executed by a processor to implement the method for processing the audio signal.
[0009] In one aspect, a computer program product or a computer program is provided. The computer program product or the computer program includes one or more program codes, and the one or more program codes are stored in a computer-readable storage medium. One or more processors of an electronic device can read the one or more program codes from the computer-readable storage medium, and the one or more processors execute the one or more program codes so that the electronic device can execute the method for processing the audio signal.
[0010] To more clearly explain the technical solutions in the embodiments of the present application, the drawings that need to be used in the following description of the embodiments are briefly described below.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Embodiments for Carrying Out the Invention
[0012] In a multi-terminal real-time audio and video call scenario, it is quite common that a user corresponding to one terminal speaks while users corresponding to multiple terminals are silent. On the other hand, some users can avoid disturbing the speaking user by selecting to turn off the microphone (or called microphone mute, that is, turning off the microphone of their own terminal) when they are silent.
[0013] In the above scenario, when it is the turn of the user who has turned off the microphone to start speaking, the user often forgets that they are in the microphone-off state. Therefore, if the microphone is not turned on again, the user may directly start speaking, and since the microphone is still off, the audio signal of the user cannot be collected and transmitted to other terminals. At this time, other terminals need to prompt the user to notice that they are in the microphone-off state, and the user needs to repeat the previous speech after turning on the microphone, so the man-machine interaction efficiency is low.
[0014] The terms related to the embodiments of the present application are interpreted below.
[0015] Voice over Internet Protocol (VoIP): VoIP is a voice call technology that enables voice calls and multimedia conferences via the Internet Protocol (IP, also known as the Internetworking Protocol), i.e., communicates via the Internet. Other unofficial names for VoIP include IP phone, Internet phone, broadband phone, broadband phone service, etc. VoIP is used in many Internet access devices including VoIP phones, smartphones, and personal computers, enabling calls and sending short messages via cellular networks and WiFi (Wireless Fidelity).
[0016] In VoIP technology, the transmitting device performs encoding and compression on the audio signal using an audio compression algorithm, then packets the encoded and compressed audio signal according to the IP protocol to obtain voice data packets, and transmits the voice data packets via the IP network to the IP address corresponding to the receiving device. The receiving device analyzes and decompresses the voice data packets, then restores the voice data packets to the original audio signal, thereby achieving the purpose of transporting the audio signal via the Internet.
[0017] Voice Activity Detection (VAD): Also known as voice endpoint detection, voice boundary detection, mute suppression, voice activity measurement, etc. The purpose of VAD is to identify and cancel long mute periods from an audio signal stream to achieve the effect of saving telephone channel resources in situations where business quality is not reduced. VAD is an important component of VoIP telephone applications, which can save precious bandwidth resources and is beneficial for reducing the end-to-end delay felt by users.
[0018] Quadrature Mirror Filter (QMF): QMF is a group of filters that is always used to perform frequency band separation on an input signal. For example, it separates the input signal into a high-frequency band signal (abbreviated as high-frequency signal) and a low-frequency band signal (abbreviated as low-frequency signal). Therefore, the QMF filter group is a common means of sub-band signal decomposition, which can reduce the signal bandwidth and enable each sub-band to be processed smoothly by the channel.
[0019] According to the spectrum division table formulated by the Institute of Electrical and Electronics Engineers (IEEE), the frequency band range of the low-frequency signal is 30 - 300 kHz, the frequency band range of the intermediate-frequency signal is 300 - 3000 kHz, and the frequency band range of the high-frequency signal is 3 - 30 MHz. On the other hand, those with a frequency band range of 30 - 300 MHz are ultra-high-frequency signals, and those with a frequency band range of 300 - 1000 MHz or higher are extremely high-frequency signals. Here, Hz refers to Hertz, which is the physical unit of frequency, kHz means kilohertz, and MHz means megahertz.
[0020] Acoustic Echo Cancellation (AEC): Acoustic echo is caused by the sound of the speaker being fed back to the microphone multiple times in a hands-free or conference application. In some scenarios, the processing method of acoustic echo cancellation is as follows: 1) The multi-party call system of terminal A receives the audio signal of terminal B; 2) The audio signal of terminal B is sampled, and this sampling is called the reference signal for echo cancellation; 3) Then, the audio signal of terminal B is sent to the speaker of terminal A and the acoustic echo canceller; 4) The audio signal of terminal B is picked up by the microphone of terminal A together with the voice of the person speaking by the user corresponding to terminal A; 5) The signal picked up by the microphone is sent to the acoustic echo canceller, compared with the original sampled reference signal, and the reference signal (i.e., the audio signal of terminal B) is removed from the signal picked up by the microphone to achieve the purpose of acoustic echo cancellation.
[0021] Noise Suppression (NS): Noise suppression technology is used to cancel background noise in the audio signal, improve the signal-to-noise ratio and intelligibility of the audio signal, and make it clearly audible to people and machines. Single-channel noise suppression usually includes two parts: noise estimation and gain coefficient estimation.
[0022] Recurrent Neural Network (RNN): An RNN is a type of recurrent neural network that takes sequence data as input, performs recursion in the evolution direction of the sequence, and connects all nodes (recurrent units) in a chain. For example, the audio frame sequence of an audio signal is a typical type of sequence data. RNN has memory, shared parameters, and complete tuning, so it has certain advantages when learning the non-linear characteristics of sequence data. RNN is applied in the fields of natural language processing (NLP), such as noise suppression, speech processing, speech recognition, language modeling, machine translation, etc., and is also used for forecasting various time sequences.
[0023] Automatic Gain Control (AGC): Automatic gain control refers to an automatic control method that automatically adjusts the gain of an amplifier circuit according to the signal strength. The definition of AGC is consistent with Automatic Level Control (ALC), but their operating mechanisms are different. Here, ALC refers to the ability of a repeater to increase the input signal level and improve its ability to control the output signal level when the repeater operates at maximum gain and the output is at maximum power. Comparatively speaking, ALC achieves the purpose of controlling the output signal level by feedback controlling the intensity of the input signal, while AGC achieves this purpose by feedback controlling the gain of the repeater.
[0024] Gain parameter (Gain): Also called the gain value. Generally speaking, the meaning of gain is, simply put, the amplification multiple or amplification rate. In a sound system, generally, the input level of the signal source determines the gain of amplification. The gain parameter relevant in the embodiments of the present application refers to the amplification rate on each individual frequency band within a given first frequency band range predicted when the noise suppression model performs noise suppression on each individual audio frame. The purpose of noise suppression is to amplify the human voice and reduce noise. Therefore, the gain parameter on the frequency band of the human voice of each individual audio frame is greater than the gain parameter on the noise frequency band. Optionally, the gain parameter is a numerical value greater than or equal to 0 and less than or equal to 1.
[0025] Energy parameter: Also called the energy value. The energy parameter of an audio frame is used to characterize the signal amplitude of the audio frame.
[0026] FIG. 1 is a schematic diagram of the implementation environment of the method for processing an audio signal provided by the embodiments of the present application. As shown in reference to FIG. 1, the implementation environment includes a first terminal 120, a server 140, and a second terminal 160.
[0027] An application program that supports multi-person voice calls is installed and operating on the first terminal 120. Here, the multi-person voice calls include multi-person audio calls or multi-person video calls based on VoIP technology. Optionally, the application program includes, but is not limited to, social applications, enterprise applications, IP phone applications, remote conference applications, remote joint diagnosis applications, call applications, etc. The embodiments of the present application do not limit the type of the application program.
[0028] The first terminal 120 and the second terminal 160 are directly or indirectly communicatively connected to the server 140 by a wired or wireless communication method.
[0029] The server 140 includes at least one of a single server, multiple servers, a cloud computing platform, or a virtualization center. The server 140 is used to provide background services for an application program that supports multi-party voice calls. Optionally, the server 140 is responsible for major computing tasks, and the first terminal 120 and the second terminal 160 are responsible for secondary computing tasks, or the server 140 is responsible for secondary computing tasks. The first terminal 120 and the second terminal 160 are responsible for major computing tasks, or a distributed computing architecture is adopted among the three of the server 140, the first terminal 120, and the second terminal 160 to perform collaborative computing.
[0030] Optionally, the server 140 is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0031] An application program that supports multi-party voice calls is installed and operating on the second terminal 160. Here, the multi-party voice calls include multi-party audio calls or multi-party video calls based on VoIP technology. Optionally, the application program includes, but is not limited to, social applications, enterprise applications, IP phone applications, remote conferencing applications, remote joint diagnosis applications, call applications, etc. The embodiments of the present application do not limit the type of the application program.
[0032] Taking the scenario of a two-person voice call as an example, the first terminal 120 is the terminal used by the first user. The first user launches a social application on the first terminal 120, logs in to the social application with the first account, and based on the call option in the chat interface with the second account, triggers the first terminal 120 to send a call request for the second account to the server 140. This call request is used to request that the second account join the two-person voice call. The server 140 forwards this call request to the second terminal 160 logged in with the second account. If the second account agrees to participate in the two-person voice call, the first terminal 120 and the second terminal 160 can conduct online voice communication based on VoIP technology. Here, the scenario of two terminals making a multi-person voice call is used as an example for explanation. The embodiments of the present application can further be applied to voice call scenarios of three or more people, but detailed descriptions are omitted here. In the two-person voice call scenario, if the first user or the second user does not want to talk temporarily, they can turn off the microphone for their corresponding account at any time in the call interface of the social application (or called microphone mute, that is, turn off the microphone of their own terminal), so as to avoid generating noise in the two-person voice call and affecting the call quality.
[0033] Taking the scene of a remote meeting among multiple people as an example, the first terminal 120 is the terminal used by the host of the meeting. The host of the meeting starts a remote meeting application on the first terminal 120 and creates a new network meeting, specifying the start time of the network meeting. The server 140 assigns a meeting number to the network meeting. After the start time of the network meeting arrives, the host of the meeting inputs the meeting number in the remote meeting application, thereby accessing the network meeting. Similarly, the second terminal 160 is the terminal used by any one of the participants in the network meeting. The participant inputs the meeting number in the remote meeting application, thereby accessing the network meeting. In a normal situation, during the process of the network meeting, the host of the meeting needs to give a speech. In such a case, the participants are set to turn off the microphones of their corresponding accounts, which can prevent disturbing the host's lecture.
[0034] Optionally, the application programs installed on the first terminal 120 and the second terminal 160 are the same, or the application programs installed on the two terminals are the same type of application programs on different operating system platforms, or the application programs installed on the two terminals are different versions of the same type of application programs developed for terminals of different model numbers. For example, if the first terminal 120 is a desktop computer, a PC (Personal Computer) side application is installed. If the second terminal 160 is a smartphone, a mobile side application is installed.
[0035] The first terminal 120 may generically refer to one of a plurality of terminals, and the second terminal 160 may generically refer to one of a plurality of terminals. The embodiments of the present application will be described by enumerating only the first terminal 120 and the second terminal 160. The device types of the first terminal 120 and the second terminal 160 may be the same or different, and the device type includes, but is not limited to, at least one of a smartphone, a tablet computer, a smart speaker, a smartwatch, a notebook computer, or a desktop computer. For example, the first terminal 120 may be a desktop computer, the second terminal 160 may be a smartphone, or both the first terminal 120 and the second terminal 160 may be smartphones or other handheld mobile communication devices.
[0036] Those skilled in the art can know that the quantity of the above terminals may be more or less. For example, the above terminal may be only one, or the above terminal may be in a quantity of dozens, or hundreds, or more. The embodiments of the present application do not limit the quantity and device type of the terminals.
[0037] Based on the above implementation environment, in an audio-video communication system, especially in scenarios of voice calls among multiple people (such as real-time audio-video calls among multiple people, remote conferences among multiple people, etc.), there is often a situation where one person is speaking while multiple people are silent. Some users choose to turn off the microphone when silent to avoid disturbing the speaking user. When it's the turn of the user who has turned off the microphone to start speaking, the user may often forget that they are in the microphone-off state and thus directly start speaking without turning on the microphone again (i.e., without releasing the microphone-off). Since the microphone is still off, the audio signal of this user cannot be collected and transmitted to other terminals. At this time, this user believes they are speaking in the voice call among multiple people, but cannot respond to other users. If this user doesn't notice that they are in the microphone-off state themselves, they will only notice it after being prompted by other users, and then this user needs to repeat the previous speech after turning on the microphone. Therefore, the man-machine interaction efficiency is low, which has a serious impact on the user experience.
[0038] In view of the above situation, an embodiment of the present application provides a method for processing an audio signal. If a user sets their account to the microphone mute state in a multi-person voice call, in the microphone mute state, the application program on the terminal can still collect the user's audio signal through the microphone, but the collected audio signal will not be transmitted to other accounts participating in the multi-person voice call. The application program performs signal analysis and processing on the audio signal collected by the microphone, and determines whether the audio signal contains target voice by using the gain parameters on multiple frequency bands in the first frequency band range of each of the multiple audio frames in the audio signal. If the audio signal contains target voice, it means that the user has forgotten to unmute the microphone and started speaking, so a prompt message is output to prompt the user to unmute the microphone. Conversely, if the audio signal does not contain target voice, it indicates that the proportion of noise in the audio signal is very high, which means that the user is not speaking or the user is chatting (not actively wanting to talk in a multi-person voice call), so there is no need to perform any prompt.
[0039] Figure 2 is a flowchart of the method for processing an audio signal provided by an embodiment of the present application. As shown in Figure 2, this embodiment is executed by an electronic device. Taking the example that the electronic device is a terminal, the terminal refers to any one of the terminals participating in a multi-person voice call, for example, the first terminal 120 or the second terminal 160 in the above implementation environment, which will be described in detail below.
[0040] 201: The terminal acquires the audio signal collected by the application program in the target scene. The target scene is that the account logged in to the application program is in the microphone mute state in a multi-person voice call.
[0041] The terminal is an electronic device used by any one of the users participating in the multi-person voice call, and an application program that supports multi-person voice calls is installed and operating on the terminal. An account is logged in to the application program, and the target scene refers to the fact that the account is in the microphone mute state in the multi-person voice call. Optionally, the application program includes, but is not limited to, social applications, enterprise applications, IP phone applications, remote conferencing applications, remote co-diagnosis applications, call applications, etc. The embodiments of the present application do not limit the type of the application program.
[0042] In some embodiments, the application program varies depending on the type of the terminal device. For example, if the terminal is a notebook computer or a desktop computer, the application program is a PC-side application. If the terminal is a smartphone, the application program is a mobile-side application. The embodiments of the present application do not limit this.
[0043] 202: The terminal obtains gain parameters on a plurality of frequency bands in the first frequency band range of each of the plurality of audio frames in the audio signal.
[0044] In some embodiments, the terminal preprocesses the audio signal to obtain a first signal, and then inputs a plurality of audio frames in the first signal into a noise suppression model. The noise suppression model processes each audio frame in the plurality of audio frames and outputs gain parameters on individual frequency bands in the first frequency band range of each audio frame. Here, the gain parameter on the frequency band of the human voice of the audio frame is larger than the gain parameter on the noise frequency band.
[0045] In the above process, for each audio frame in the plurality of audio frames, the gain parameters on each frequency band within the first frequency band range of the audio frame are all determined, so that in the noise suppression process, a higher gain parameter is assigned to the human voice frequency band than to the noise frequency band, so that the effect of effectively enhancing the human voice component in the audio signal and suppressing the noise component in the audio signal can be achieved. Therefore, the gain parameters on each frequency band of each audio frame can contribute to identifying whether each audio frame contains a target voice, and thus it can be determined whether the entire audio signal contains a target voice.
[0046] 203: When the terminal determines that the audio signal contains a target voice based on the gain parameter, outputting a prompt message, the prompt message being used to prompt to release the microphone mute state.
[0047] Wherein, the prompt information is used to prompt the account to unmute the microphone, and the target voice is the speech of the target object in the voice call with the plurality of people, or the target voice is the sound of the target object, wherein the target object refers to a user who participates in the voice call with the plurality of people through the terminal.
[0048] In some embodiments, if the target voice is a statement in a voice call among a plurality of people of the target object, a prompt message will be output externally only when it is detected that the audio signal includes the statements in the voice call among the plurality of people of the target object. If the audio signal only contains the sound of the target object but the sound is not a statement in the voice call among the plurality of people, it means that the user is chatting, but it also means that the content of the chat may not be desired to be transmitted in the voice call among the plurality of people. Or, if the audio signal does not contain the sound of the target object and it means that although the user is not making a sound (voice), there is a possibility that some background noise is being collected. In both of the above cases, a prompt message will not be output externally. Since it is possible to accurately identify when the user wants to speak in a voice call among a plurality of people and whether to output a prompt message at this time, it is possible to avoid disturbing the user by outputting a prompt message when the user is chatting.
[0049] In some embodiments, if the target voice is the sound of the target object, a prompt message will be output externally when it is detected that the audio signal includes the sound of the target object. If the audio signal does not contain the sound of the target object, a prompt message will not be output externally. In this way, the sensitivity of detecting the sound of the target object can be improved, and it is possible to avoid the occurrence of a scene where the user speaks a relatively short sentence but may be determined by the machine as chatting and not prompted. Therefore, the detection sensitivity for the sound of the target object is improved.
[0050] In some embodiments, the terminal can determine whether the audio signal contains a target voice based on the gain parameters on each frequency band of each audio frame. If the target voice is included, it means that the user starts speaking in the microphone mute state, which causes invalid communication, and in this case, a prompt message is output to the outside to prompt the user to unmute the microphone. If the target voice is not included, it means that the user has not started speaking, or the user is chatting (does not want to actively speak in a multi-person voice call), so the microphone is still kept in the mute state and there is no need to perform any prompt.
[0051] In some embodiments, when determining whether the audio signal includes a target voice, the terminal makes a VAD decision based on gain parameters on each frequency band of each audio frame, i.e., based on the gain parameters output by the noise suppression model for each audio frame, thereby determining whether the audio signal includes a target voice, simplifying the VAD decision-making flow, and shortening the time length of the VAD decision-making.
[0052] The above method is usually applied to a scene where the target voice is the voice of the target object, and the only requirement is to determine whether the audio signal contains the voice of the target object. By using the gain parameters in the individual frequency bands of the individual audio frames, it is possible to relatively well determine whether the audio signal contains the voice of the target object. Of course, the above method may also be used in a scene where the target voice is the speech in the voice calls of the plurality of people of the target object. When chatting, usually, there is no continuous voice fluctuation, so it is only necessary to set the conditions for VAD decision-making more strictly. For example, it may be determined that the audio signal contains the target voice as long as the voice activity parameter, that is, the VAD value, of a plurality of consecutive audio frames is 1. The embodiments of the present application are not limited to this.
[0053] In some embodiments, when determining whether the audio signal contains the target voice, the terminal makes a comprehensive determination by combining the gain parameter in the individual frequency band of the individual audio frame and the energy parameter of the individual audio frame, that is, makes a VAD decision based on the gain parameter output by the noise suppression model for the individual audio frame and the energy parameter of the individual audio frame, thereby determining whether the audio signal contains the target voice. Since two-dimensional influencing factors of the gain parameter and the energy parameter are introduced, it is possible to more accurately identify whether the audio signal contains the target voice, thereby improving the accuracy of VAD decision-making.
[0054] The above method is usually applied to a scenario where the target voice is a statement in a voice call among multiple people of the target object. It is necessary not only to identify the voice of the target object in the audio signal, but also to determine whether the voice of the target object is casual conversation or a statement. When the target object makes a statement, the volume is relatively large, that is, not only the VAD value of the signal during the statement is 1, but also it has a relatively large energy parameter. On the other hand, when having casual conversation, the volume is relatively small, that is, the signal during casual conversation has only a VAD value of 1 but a relatively small energy parameter. Therefore, by combining the two dimensions of the gain parameter and the energy parameter to make a comprehensive decision, it is possible to relatively well determine whether the audio signal contains the statements in the voice call among multiple people of the target object. Of course, the above method may also be used in a scenario where the target voice is the voice of the target object, accurately detecting audio signals with several VAD values of 1 but relatively small energy parameters (for example, the distance from the microphone is relatively far), thereby improving the accuracy of VAD decision-making.
[0055] In some embodiments, when determining whether the audio signal contains the target voice, the terminal performs noise suppression on each audio frame based on the gain parameter on the individual frequency band of each audio frame, obtains each target audio frame after noise suppression, then calculates the energy parameter for each target audio frame, and further performs voice activity detection on the energy parameter of each target audio frame by using the VAD algorithm, so as to determine whether the audio signal contains the target voice. Similarly, it can accurately identify whether what is contained in the audio signal is the target voice or noise, thereby improving the accuracy of VAD decision-making.
[0056] The above method is similarly applicable to two scenarios: whether the target voice is a statement in the voice calls of the multiple persons of the target object or the voice of the target object. When iteratively training the VAD algorithm, the training data only needs to be adjusted based on the difference in the target voice that needs to be prompted. Therefore, it has relatively high portability and migration possibility, and has high availability and a wide range of application scenarios.
[0057] In some embodiments, when outputting a prompt message, the terminal adapts based on the difference in the terminal type. If the terminal is a non-mobile device such as a personal computer or a notebook computer, the terminal outputs the prompt message on the desktop side. If the terminal is a mobile device, the terminal outputs the prompt message on the mobile side, thereby enabling compatibility with different types of terminals.
[0058] In some embodiments, the terminal only outputs the prompt message externally, but since the user needs to manually release the microphone mute state, the autonomy to control whether the user releases the microphone mute state can be guaranteed. In some embodiments, when the terminal detects that the target voice is included in the audio signal, the terminal automatically releases the microphone mute state and prompts externally that the microphone mute state has been released. At this time, the user does not need to manually release the microphone mute state, and the complexity of the user operation can be reduced.
[0059] In some embodiments, the output method of the prompt message includes, but is not limited to, text format output, voice format output, animation format output, dynamic format output, etc. The embodiments of the present application do not limit the output method of the prompt message.
[0060] In some embodiments, the terminal displays a text prompt message on the call interface of the multi-person voice call, and the text prompt message is used to prompt the user to unmute the microphone. For example, the text prompt message is "Since the microphone is muted, please unmute the microphone before speaking." Optionally, the text prompt message pops up on the call interface in the form of a pop-up window, or the text prompt message floats on the call interface in the form of a floating layer, or the text prompt message is scrolled and displayed on the call interface in the form of subtitles or prompted by blinking, but the embodiments of the present application do not limit the display method of the text prompt message. Optionally, the text prompt message automatically disappears after being displayed on the call interface for a period of time, or the text prompt message is continuously displayed on the call interface until the user manually turns off the text prompt message, but the embodiments of the present application do not limit the off method of the text prompt message.
[0061] In some embodiments, the terminal plays the voice prompt message externally, and the voice prompt message is used to prompt the user to unmute the microphone. For example, the voice prompt message is "Since the microphone is muted, please unmute the microphone before speaking."
[0062] In some embodiments, the terminal plays an animation prompt message or a dynamic prompt message in the call interface of the multiple-person voice call. The animation prompt message or the dynamic prompt message is used to prompt the user to unmute the microphone. Optionally, the animation prompt message or the dynamic prompt message automatically disappears after being played once in the call interface, or the animation prompt message or the dynamic prompt message is played in a loop in the call interface until the user manually turns off the animation prompt message or the dynamic prompt message. The embodiments of the present application do not limit the off method of the animation prompt message or the dynamic prompt message.
[0063] The above selectable technical solutions can be combined arbitrarily to form selectable embodiments of the present disclosure, which are not described in detail here.
[0064] In the method provided by the embodiment of the present application, when the microphone is muted in a multi-party voice call, the application program still collects the user's audio signal, but does not send the collected audio signal to other accounts participating in the multi-party voice call. The application program performs signal analysis and processing on the audio signal, and uses gain parameters on multiple frequency bands in the first frequency band range of each of the multiple audio frames in the audio signal to determine whether the audio signal contains a target voice, and if the audio signal contains the target voice, it means that the user has forgotten to unmute the microphone and started speaking, so as to output a prompt message to the outside and prompt the user to unmute the microphone in a timely manner, thereby reducing the loss of communication efficiency caused by the user not realizing that the microphone is muted, improving man-machine interaction efficiency, and optimizing the user experience.
[0065] 3 is a flow chart of an audio signal processing method provided by an embodiment of the present application, and as shown in FIG. 3, the embodiment is executed by an electronic device. Take the electronic device as an example of a terminal, the terminal refers to any one of terminals participating in a multi-party voice call, such as the first terminal 120 or the second terminal 160 in the above implementation environment.
[0066] In the embodiment of the present application, the terminal will be described in detail how to determine whether the audio signal contains a target voice based on the gain parameter on each frequency band of each audio frame, i.e., make a VAD decision based on the gain parameter output by the noise suppression model for each audio frame. The embodiment includes the following steps:
[0067] 301: A terminal accesses a multi-party voice call in an application program.
[0068] The multi-party voice call includes multi-party audio-video calls based on VoIP technology. For example, there are multi-party audio calls, multi-party video calls, or some users access in the audio call mode and some users access in the video call mode, etc. However, the embodiments of the present application do not limit the type of the multi-party voice call. Optionally, the multi-party voice call is a two-party real-time audio-video call (for example, a two-party voice call or a two-party video call) started for an account designated based on a social application, or a multi-party real-time audio-video call (for example, a multi-party voice call or a multi-party video call) started within an account group designated based on a social application, or a multi-party remote conference (for example, a multi-party voice conference or a multi-party video conference) started based on a conference application, etc.
[0069] In some embodiments, the user starts an application program that supports the multi-party voice call on the terminal. For example, the start operation is that the user performs a touch operation on the icon of the application program on the desktop of the terminal, or the user inputs a start command for the application program to the smart assistant, and the start command includes a voice command or a text command. However, the embodiments of the present application do not limit the type of the start command. Optionally, when the user sets an automatic start condition for the application program, when the terminal detects an automatic start condition that matches the application program, the operating system automatically starts the application program. For example, the automatic start condition is an opening automatic start or a timing automatic start. For example, the application program is automatically started 5 minutes before a specified meeting starts, etc. However, the embodiments of the present application do not limit the automatic start condition of the application program.
[0070] After the application program is launched, the main interface of the application program is displayed, and an account login option is displayed in the main interface. The user performs a trigger operation on the account login option. If the user's account is logged in in the application program and the login is completed, the user returns to the main interface. In some embodiments, after the account login is completed, the user accesses a multi-person voice call based on the application program. The terminal displays a call interface for the multi-person voice call. In the call interface, each account accessing the multi-person voice call and a microphone setting control member are displayed. The microphone setting control member is used to turn on or cancel the microphone mute state in the multi-person voice call of the present account.
[0071] In some embodiments, the way for the user to access the multi-person voice call in a multi-person real-time audio and video call scene includes displaying a call request interface in the application program in response to receiving a call request from a target account. Optionally, an avatar picture, an accept option, and a reject option of the target account are displayed in the call request interface. The user performs a trigger operation on the accept option to enable access to the multi-person voice call. Optionally, if the target account is the initiator account of the multi-person voice call, the corresponding scene is that the initiator starts a call request to the user, or if the target account is any one of the participant accounts that wants to access the multi-person voice call, the corresponding scene is that the participant invites the user to join the multi-person voice call, but the embodiments of the present application do not limit this.
[0072] In some embodiments, in a multi-person conference scenario, the way for a user to access the multi-person voice call is to query and display the target conference corresponding to the conference number by the user entering the conference number of the target conference in the conference search box of the application program. If the user clicks on the join option of the target conference to enable access to the multi-person voice call, or when the user convenes or marks the target conference and turns on the conference reminder function for the target conference, if the user launches the application program within the target period (for example, 5 minutes before the start) before the target conference starts, the application program will automatically pop up the conference start reminder information and the join option of the target conference, enabling the user to click on the join option of the target conference to access the multi-person voice call.
[0073] In some embodiments, for different types of multi-person voice calls, the display methods for each account accessing the multi-person voice call on the call interface are not the same. For example, for a multi-person audio call, each account's avatar picture is displayed on the call interface; for a multi-person video call, each account's video stream is displayed on the call interface; for a multi-person conference, the conference theme and the presentation (PowerPoint, PPT) introduced by the conference speaker are displayed on the call interface.
[0074] 302: The terminal sets the account logged in to the application program in the multi-person voice call to the microphone mute state.
[0075] The terminal displays a microphone setting control member in the call interface. The enabled state of the microphone setting control member corresponds to the microphone-on state, and the disabled state of the microphone setting control member corresponds to the microphone-mute state. If the account is currently in the microphone-mute state, that is, if the microphone setting control member is currently in the disabled state, when the user clicks on the microphone setting control member, the terminal can switch the microphone setting control member from the disabled state to the enabled state and release the microphone-mute state. If the account is currently in the microphone-on state, that is, if the microphone setting control member is currently in the enabled state, when the user clicks on the microphone setting control member, the terminal can switch the microphone setting control member from the enabled state to the disabled state, enter the microphone-mute state, and enable the execution of the following step 303.
[0076] 303: The terminal obtains the audio signal collected by the application program in the target scene. The target scene is that the account logged in to the application program is in the microphone-mute state in a multi-person voice call.
[0077] In some embodiments, when the user sets the microphone setting control member to the disabled state in the call interface so that the account is in the microphone-mute state in the multi-person voice call, it conforms to the target scene. In the microphone-mute state related to the embodiments of the present application, the terminal does not turn off the microphone, still calls the microphone to collect the audio signal, but does not transmit the audio signal to other accounts participating in the multi-person voice call.
[0078] In some embodiments, the way the terminal collects the audio signal is as follows. The terminal calls a recording interface (Application Programming Interface, API) through the application program, and based on the recording interface, drives a microphone to collect and obtain the audio signal.
[0079] 304: The terminal preprocesses the audio signal to obtain a first signal.
[0080] The way the terminal preprocesses the audio signal includes at least one of framing, windowing, Fourier transform, frequency band separation, or acoustic echo cancellation, but is not limited thereto. Embodiments of the present application do not limit the preprocessing method.
[0081] In some embodiments, the terminal performs natural framing on the audio signal to obtain a plurality of first audio frames, that is, completes the framing process on the audio signal.
[0082] In some embodiments, based on natural framing, the terminal reframes the audio signal to obtain a plurality of second audio frames. Optionally, the reframing method includes the terminal performing windowing processing on the plurality of first audio frames to obtain the plurality of second audio frames. The second audio frame is a first audio frame divided over a finite time, where the finite time is any one time length greater than or equal to 0.
[0083] In some embodiments, the terminal inputs the plurality of first audio frames into a window function and moves the window function in the time domain of the plurality of first audio frames, thereby dividing the plurality of first audio frames into a plurality of second audio frames with equal time lengths, that is, obtaining a plurality of second audio frames by reframing the plurality of first audio frames again. Optionally, the window function includes, but is not limited to, a Hamming window, a Hanning window, or a rectangular window, etc. The embodiments of the present application do not limit the type of the window function.
[0084] In some embodiments, the plurality of second audio frames have an overlap rate of a target ratio, that is, the step size of moving the window function in the time domain is smaller than 1, and the target ratio is any numerical value greater than 0. For example, when the step size is 0.6, the overlap rate of adjacent second audio frames divided by the window function is 40%. By setting a certain overlap rate, it is possible to avoid losing the edge time domain characteristics of each second audio frame cut by the window function due to random errors or system errors in the windowing process.
[0085] In some embodiments, the terminal performs a Fourier transform on the audio signal based on the windowing process to obtain a plurality of third audio frames. Optionally, each of the second audio frames after being divided by the window function can be regarded as a stationary signal. Therefore, the terminal performs a Fourier transform on the plurality of second audio frames to obtain the plurality of third audio frames, that is, the audio signal can be converted from the time domain to the frequency domain, and the time-frequency conversion of the audio signal can be completed.
[0086] Optionally, the method of performing Fourier transform on each second audio frame includes, but is not limited to, Fast Fourier Transform (FFT), Short-Time Fourier Transform (STFT), Discrete Cosine Transform (DCT), etc. Embodiments of the present application do not limit the method of Fourier transform.
[0087] In some embodiments, based on time-frequency conversion, the terminal performs different processing on audio signals with different sampling rates. Optionally, the terminal obtains the sampling rate of the audio signal, and if the sampling rate is greater than the sampling rate threshold, it determines that the audio signal is a super-resolution signal. For the super-resolution signal, the terminal performs frequency band separation to separate the low-frequency signal and the high-frequency signal in the super-resolution signal, and only performs subsequent VAD decision-making on the low-frequency signal, which can reduce the computational complexity of VAD decision-making. For non-super-resolution signals (such as high-resolution signals), the terminal does not need to perform frequency band separation and directly performs subsequent VAD decision-making on the entire audio signal, which can simplify the processing flow of the audio signal.
[0088] In some embodiments, for a super-resolution signal whose sampling rate is greater than the sampling rate threshold, the method by which the terminal performs frequency band separation is to input a plurality of third audio frames after Fourier transform into a QMF analysis filter, filter the plurality of third audio frames based on the QMF analysis filter, and output the high-frequency components and low-frequency components in the plurality of third audio frames respectively. Here, the high-frequency components obtained by filtering are the high-frequency signals in the audio signal, and the low-frequency components obtained by filtering are the low-frequency signals in the audio signal. For example, according to the spectrum division table formulated by IEEE, the frequency band range of the low-frequency signal is 30~300kHz, the frequency band range of the intermediate-frequency signal is 300~3000kHz, and the frequency band range of the high-frequency signal is 3~30MHz.
[0089] In a real-time scenario, assuming that the audio signal collected by the microphone is 16kHz bandwidth data, after frequency band separation is performed by the QMF analysis filter, an 8kHz high-frequency signal and an 8kHz low-frequency signal are output. On the other hand, subsequent noise suppression and VAD decision-making only act on the 8kHz low-frequency signal, and the computational complexity of noise suppression and VAD decision-making can be reduced.
[0090] It should be noted that the above frequency band separation is an optional step in the preprocessing. For example, frequency band separation is only performed on the super-resolution signal, while there is no need to perform frequency band separation on the non-super-resolution signal. The embodiments of the present application do not limit whether to perform frequency band separation on the audio signal.
[0091] In some embodiments, for the low-frequency signal obtained by performing frequency band separation on the super-resolution signal or the non-super-resolution signal, the terminal cancels the acoustic echo in the low-frequency signal or the non-super-resolution signal by performing acoustic echo cancellation, thereby improving the subsequent noise suppression and the accuracy of VAD decision-making. Optionally, the terminal inputs the low-frequency signal or the non-super-resolution signal into an acoustic echo canceller, and the acoustic echo canceller cancels the acoustic echo in the low-frequency signal or the non-super-resolution signal to obtain a first signal after preprocessing.
[0092] As needs to be explained, the above acoustic echo cancellation is an optional step in preprocessing. For example, when the terminal detects that the hands-free state is on in the multi-party voice call, the audio signal emitted by other terminals in the hands-free state is collected by the microphone of the terminal to form an acoustic echo. Therefore, it is necessary to perform acoustic echo cancellation on the audio signal to improve the subsequent noise suppression and the accuracy of VAD decision-making. When the terminal detects that the hands-free state is off in the multi-party voice call, the user receives the multi-party voice call through the earphone. At this time, there is no acoustic echo formed, or the user directly receives the multi-party voice call through the receiver in the non-hands-free state, which means that the influence of the acoustic echo at this time is relatively small. Then, it is not necessary to perform acoustic echo cancellation on the audio signal, thereby saving the computational amount in the processing process of the audio signal. Further, for example, when the terminal detects that there is no acoustic echo canceller arranged, the present application does not limit whether to perform acoustic echo cancellation on the audio signal.
[0093] The first signal refers to the audio signal obtained through preprocessing. The above process is described by taking the example of performing both frequency band separation and acoustic echo cancellation. In some embodiments, if frequency band separation and acoustic echo cancellation are not performed, the frequency domain signal obtained by time-frequency conversion is the first signal, i.e., the first signal is the frequency domain signal obtained by time-frequency conversion. If frequency band separation is performed but acoustic echo cancellation is not, the low-frequency signal obtained by frequency band separation is the first signal. If frequency band separation is not performed but acoustic echo cancellation is, the first signal is obtained after acoustic echo cancellation, but the embodiments of the present application do not limit this.
[0094] 305: The terminal inputs a plurality of audio frames in the first signal into a noise suppression model, and the noise suppression model processes each audio frame in the plurality of audio frames, and outputs gain parameters on individual frequency bands in the respective first frequency band ranges of the individual audio frames. Here, the gain parameter on the frequency band of the human voice in the audio frame is greater than the gain parameter on the noise frequency band.
[0095] In some embodiments, the plurality of audio frames refer to all the audio frames included in the first signal, or the plurality of audio frames refer to some of the audio frames in the first signal. For example, a plurality of key frames in the first signal are extracted as the plurality of audio frames, or one audio frame is sampled for each preset step size for the first signal, and the plurality of audio frames obtained by sampling are used as the plurality of audio frames. Here, the preset step size refers to any integer greater than or equal to 1.
[0096] In some embodiments, the terminal obtains, for each individual audio frame in the plurality of audio frames, a gain parameter on each individual frequency band within the respective first frequency band range of the audio frame. Here, the first frequency band range includes at least the frequency band of human voice. Optionally, in addition to the frequency band of human voice, the first frequency band range further includes a noise frequency band.
[0097] Optionally, the plurality of frequency bands divided within the first frequency band range may be set by a technician, or an equal number of divisions specified for the first frequency band range may be performed. However, the embodiments of the present application do not limit the frequency band division method of the first frequency band range.
[0098] In some embodiments, the first frequency band range is a frequency band range set by a technician or a default frequency band range set by the system. For example, the first frequency band range is 0 - 8000 Hz, or the first frequency band range is 0 - 20000 Hz. However, the embodiments of the present application do not limit the first frequency band range.
[0099] In some embodiments, the noise suppression model is a machine learning model obtained based on sample data training. Optionally, the structure of the noise suppression model includes, but is not limited to, RNN, LSTM (Long Short-Term Memory), GRU (Gate Recurrent Unit), CNN (Convolutional Neural Networks), etc. The embodiments of the present application do not limit the structure of the noise suppression model.
[0100] In one implementation scenario, the noise suppression model is an RNN used for noise suppression. For the RNN, the input is a plurality of audio frames in the audio signal, i.e., the first signal, obtained after preprocessing, and the output is a plurality of gain parameters for individual audio frames. The RNN includes at least one hidden layer, and each hidden layer contains a plurality of neurons. The number of neurons in each hidden layer is the same as the number of input audio frames. The neurons in each hidden layer are all connected, and the adjacent hidden layers are connected in series. For each individual neuron in each hidden layer, the frequency features output by the previous neuron in the current hidden layer and the neuron at the corresponding position in the previous hidden layer are taken as the input of this neuron.
[0101] Based on the above RNN structure, the terminal inputs a plurality of audio frames in the first signal into at least one hidden layer of the RNN, that is, it refers to inputting the plurality of audio frames into a plurality of neurons in the first hidden layer of the RNN respectively. Here, one neuron corresponds to one audio frame. For the i-th (i≥1) neuron in the first hidden layer, the frequency feature output by the (i - 1)-th neuron in the first hidden layer and the i-th audio frame are used as inputs, and weight processing is performed on the frequency feature output by the (i - 1)-th neuron in the first hidden layer and the i-th audio frame. The obtained frequency feature is input into the (i + 1)-th neuron in the first hidden layer and the i-th neuron in the second hidden layer respectively. By analogy, for any one neuron in any one hidden layer of the RNN, any one neuron performs weight processing on the frequency feature output by the previous neuron in the any one hidden layer and the frequency feature output by the neuron at the corresponding position in the previous hidden layer, and the frequency feature obtained by the weight processing is input into the next neuron in the any one hidden layer and the neuron at the corresponding position in the next hidden layer respectively... Finally, the last hidden layer inputs the respective target frequency features for each audio frame, and softmax (exponential normalization) processing is performed on the target frequency features of each audio frame, thereby predicting a plurality of gain parameters for each audio frame respectively. Each gain parameter corresponds to one frequency band in the first frequency band range.
[0102] Since the voice energy in the frequency band of human voice is relatively large, the signal-to-noise ratio is relatively high. Using the noise suppression model of the above RNN architecture, after training, noise and human voice can be accurately identified. Thereby, a relatively large gain parameter is assigned to human voice, and a relatively small gain parameter is assigned to noise, so that the noise suppression model has a very high identification accuracy even for non-stationary noises such as keyboard sounds. Compared with CNN that performs complex convolution calculations, the computational consumption of RNN is relatively low, which can better meet the real-time call scenario and will not excessively occupy computing resources and affect the call quality.
[0103] Figure 4 is a frequency band diagram of Opus provided by the embodiment of the present application. As shown in 400, a frequency band diagram divided based on the Opus encoding method is shown. Here, Opus is a format of irreversible voice encoding. For example, 0 - 8000Hz in the Opus frequency band diagram is used as the first frequency band range, and referring to the frequency band division method in the Opus frequency band diagram, the first frequency band range 0 - 8000Hz is divided into 18 frequency bands. Each point represents one frequency band value, and the 18 frequency band values of 0 - 8000Hz include 0, 200, 400, 600, 800, 1000, 1200, 1400, 1600, 2000, 2400, 2800, 3200, 4000, 4800, 5600, 6800, 8000, where the unit of the frequency band value is Hz. After the terminal inputs a plurality of audio frames in the first signal into the RNN, the RNN outputs 18 gain parameters for each audio frame. Here, each gain parameter corresponds to one frequency band of 0 - 8000Hz in the Opus frequency band diagram.
[0104] In the above steps 304 to 305, the terminal obtains gain parameters on a plurality of frequency bands in the respective first frequency band ranges of the plurality of audio frames in the audio signal. In the noise suppression process, in order to assign a higher gain parameter to the frequency band of the human voice than the noise frequency band, the component of the human voice in the audio signal can be effectively enhanced, and the effect of suppressing the noise component in the audio signal can be achieved. Therefore, the gain parameters on the respective frequency bands of each audio frame can contribute to identifying whether each audio frame contains target voice, thereby determining whether the entire audio signal contains target voice.
[0105] 306: For each individual audio frame, the terminal determines the gain parameters on the individual frequency bands in the second frequency band range of the audio frame based on the gain parameters on the individual frequency bands in the first frequency band range of the audio frame, and the second frequency band range is a subset of the first frequency band range.
[0106] In some embodiments, the frequency band of the human voice is included in the first frequency band range, and the noise frequency band is also included. On the other hand, since the VAD decision only needs to make a detailed determination about the frequency band of the human voice and does not need to be concerned about the noise frequency band, the subset consisting of the terminal obtaining the frequency band of the human voice from the first frequency band range is the second frequency band range, and the terminal obtains the gain parameters on the individual frequency bands in the first frequency band range of each individual audio frame by means of a noise suppression model. On the other hand, since the second frequency band range is also a subset of the first frequency band range, it is obvious that the gain parameters on the individual frequency bands in the second frequency band range of each individual audio frame can be determined.
[0107] As needs to be explained, the second frequency band range enables adaptive changes for users of different genders or different ages. For example, the voice frequency of females is usually higher than that of males. Therefore, the terminal can arrange different second frequency band ranges for different users, but the embodiments of the present application do not limit the second frequency band range.
[0108] In one implementation scenario, the first frequency band range refers to a total of 18 frequency bands from 0 to 8000 Hz in the Opus frequency band diagram. On the other hand, the second frequency band range refers to a total of 9 frequency bands from 200 to 2000 Hz, which are 200, 400, 600, 800, 1000, 1200, 1400, 1600, 2000, or the second frequency band range refers to a total of 5 frequency bands from 300 to 1000 Hz, which are 300, 400, 600, 800, 1000. Here, the unit of the frequency band value is Hz.
[0109] 307: The terminal determines the voice state parameter of the audio frame based on the gain parameter on each frequency band in the second frequency band range of the audio frame.
[0110] In some embodiments, for each individual audio frame, the terminal multiplies the gain parameter on each frequency band in the second frequency band range of the audio frame by the weight coefficient of the corresponding frequency band, obtains the weighted gain parameter on each frequency band in the second frequency band range of the audio frame, adds the weighted gain parameters on each frequency band in the second frequency band range of the audio frame, obtains the comprehensive gain parameter of the audio frame, and determines the voice state parameter of the audio frame based on the comprehensive gain parameter of the audio frame.
[0111] In the above process, the second frequency band range includes the frequency bands of most people's voices in the first frequency band range. That is, since most of the energy of people's voices is within the second frequency band range (for example, 200 - 2000 Hz, or 300 - 1000 Hz, etc.), the gain parameters on the individual frequency bands within the second frequency band range of each audio frame can best represent whether someone is currently speaking (that is, whether the target voice is included in the current audio frame).
[0112] In some embodiments, for the case where the target voice is the sound of the target object, by making it possible to arrange a relatively wide second frequency band range, it is easier to identify the sound of the target object on the frequency bands of more people's voices. For the situation where the target voice is the speech in the multi-person voice call of the target object, by making it possible to arrange a relatively narrow second frequency band range, it is easier to exclude the sounds emitted when chatting on some of the relatively low frequency bands of people's voices. However, the embodiments of the present application are not limited to this.
[0113] Optionally, in the terminal, the correspondence relationship between the individual frequency bands in the second frequency band range and the weight coefficients is pre-stored. For each individual frequency band within the second frequency band range, the weight coefficient corresponding to the frequency band is determined based on the correspondence relationship, and the gain parameter on the frequency band of the audio frame is multiplied by the weight coefficient corresponding to the frequency band to obtain the weighted gain parameter on the frequency band of the audio frame.
[0114] Optionally, for each audio frame, the terminal adds the weighted gain parameters on all frequency bands within the second frequency band range of the audio frame to obtain the comprehensive gain parameter of the audio frame. Based on the magnitude relationship between the comprehensive gain parameter and the activation threshold, it is possible to determine the voice state parameter of the audio frame. Optionally, the voice state parameter includes "including target voice" and "not including target voice". For example, the voice state parameter is a boolean type of data. The value of the boolean type of data is True, which means "including target voice", and the value of the boolean type of data is False, which means "not including target voice". Or, the voice state parameter is a binarized data. The value of the binarized data is 1, which means "including target voice", and the value of the binarized data is 0, which means "not including target voice". Or, the voice state parameter is string data or the like, but the embodiments of the present application do not limit the data type of the voice state parameter.
[0115] In some embodiments, when the comprehensive gain parameter is greater than the activation threshold after being amplified by the target multiple, the terminal determines that the voice state parameter includes the target voice, and when the comprehensive gain parameter is less than or equal to the activation threshold after being amplified by the target multiple, the terminal determines that the voice state parameter does not include the target voice. Here, the target multiple is any numerical value greater than 1. For example, the target multiple is 10000. Here, the activation threshold is any numerical value greater than 0. For example, the activation threshold is 6000.
[0116] In one implementation scenario, taking the example where the second frequency band range is 200 - 2000 Hz, the target multiple is 10000, and the activation threshold is 6000. After the user turns on multiple voice calls, in the microphone muted state, the user speaks one voice into the microphone. After the microphone collects the audio signal, for each frame (assuming the length of each frame is 20 ms), the gain parameter on each frequency band within 200 - 2000 Hz is obtained respectively. Here, the gain parameter is a numerical value greater than or equal to 0 and less than or equal to 1. Weighted integration is performed on the gain parameters on each frequency band within 200 - 2000 Hz of each frame to obtain the comprehensive gain parameter of each frame. The comprehensive gain parameter of each frame is amplified by 10000 times. If the amplified value is greater than 6000, this frame is considered to be activated, the VAD value of this frame is set to 1, which means the voice state parameter of this frame contains the target voice. If the amplified value is less than or equal to 6000, this frame is considered not to be activated, the VAD value of this frame is set to 0, which means the voice state parameter of this frame does not contain the target voice.
[0117] In the above process, for each individual audio frame, by performing weighted integration on the gain parameters on each frequency band within the second frequency band range, the comprehensive gain parameter of the audio frame is obtained. After amplifying the comprehensive gain parameter, it is used to determine the voice state of the current audio frame. That is, the voice state parameter of the audio frame is determined. According to the comprehensive gain parameter of each audio frame, it can accurately determine whether each audio frame contains the target voice, and accurate frame-level human voice identification can be achieved.
[0118] In the above steps 306-307, the terminal determines the voice state parameters of the plurality of audio frames based on the gain parameters on the plurality of frequency bands of the plurality of audio frames. Here, the voice state parameters are used to characterize whether the corresponding audio frame contains target voice. The terminal can determine that the audio signal contains target voice based on the voice state parameters of the plurality of audio frames. In the embodiments of the present application, weighted integration is performed based on the gain parameters on each frequency band within the second frequency band range to obtain the comprehensive gain parameter of each audio frame, and the voice state parameter of each audio frame is determined based on the comprehensive gain parameter. This is taken as an example for explanation because most people's voice frequency bands are included within the second frequency band range. In some other embodiments, the terminal further performs weighted integration based on the gain parameters on each frequency band within the first frequency band range to obtain the comprehensive gain parameter of each audio frame, and determines the voice state parameter of each audio frame based on the comprehensive gain parameter, so that the processing flow of the audio signal can be simplified.
[0119] In some embodiments, the terminal obtains the energy parameter of each audio frame, and combines the comprehensive gain parameter and the energy parameter of each audio frame to determine the voice state parameter of each audio frame. Alternatively, the terminal performs noise suppression on the first signal based on the gain parameter within the first frequency band range of each audio frame, inputs the signal after noise suppression into the VAD algorithm for VAD detection, and obtains the voice state parameter of each audio frame. This will be described in detail in subsequent embodiments, but the embodiments of the present application do not limit the acquisition method of the voice state parameter of each audio frame.
[0120] 308: The terminal determines the activation state of the audio frame group to which the audio frame belongs based on the audio state parameters of the audio frame and the first target number of audio frames before the audio frame. The audio frame group includes the audio frame and the first target number of audio frames before the audio frame.
[0121] Here, the audio frame refers to any one of the plurality of audio frames. In other words, the above step 308 is executed for each individual audio frame in the plurality of audio frames.
[0122] In some embodiments, since the user usually continuously emits an audio signal to the microphone, the audio signal collected by the microphone is one audio stream. In the scene of the audio stream, for any one audio frame in the audio stream, it is necessary to comprehensively determine whether the target audio is included in the audio signal within the period covered by these audio frames by referring to the audio state parameters of the audio frame and the target number of audio frames before the audio frame. Here, the target number is determined based on the first target number and the second target number. For example, the target number is a value obtained by multiplying the value obtained by adding 1 to the first target number and the value obtained by adding 1 to the second target number and then subtracting 1. The first target number is any integer greater than or equal to 1, and the second target number is any integer greater than or equal to 1. For example, if the first target number is 4 and the second target number is 29, the target number is (4 + 1)×(29 + 1)-1 = 149.
[0123] In some embodiments, for any one audio frame, the terminal determines, as an audio frame group to which the audio frame belongs, the audio frame and the first target number of audio frames before the audio frame, and then obtains the voice state parameter of each audio frame in the audio frame group. Optionally, when the number of audio frames including the target voice in the audio frame group exceeds the number threshold, it is determined that the activation state of the audio frame group is activated; when the number of audio frames including the target voice in the audio frame group does not exceed the number threshold, it is determined that the activation state of the audio frame group is deactivated. Here, the value range of the number threshold is greater than or equal to 1 and less than or equal to the value obtained by adding 1 to the first target number. For example, if the first target number is 4, the value range of the number threshold is greater than or equal to 1 and less than or equal to 5.
[0124] In the above process, for each audio frame group, if the target voice is included in the audio frames exceeding a certain number threshold, the entire audio frame group is considered to be activated, and it is possible to relatively well determine whether the current audio frame group includes the target voice. Since non-stationary noise usually does not appear densely within the same audio frame group, the situation of misjudging whether the audio frame group is activated due to individual non-stationary noises (such as keyboard sounds, etc.) is reduced, and the accuracy of identifying whether the audio signal includes the target voice is improved.
[0125] In some embodiments, if the voice state parameters of the audio frames in which there are consecutive preset thresholds in the audio frame group all include the target voice, it is determined that the activation state of the audio frame group is activated. If the voice state parameters of the audio frames in which there are no consecutive preset thresholds in the audio frame group all include the target voice, it is determined that the activation state of the audio frame group is deactivated. Here, the value range of the preset threshold is greater than or equal to 1 and less than or equal to the number obtained by adding 1 to the first target quantity. For example, if the first target quantity is 4, the value range of the preset threshold is greater than or equal to 1 and less than or equal to 5.
[0126] In the above process, for each individual audio frame group, if the target voice is included in the audio frames in which there are consecutive preset thresholds, the entire audio frame group is considered to be activated, and it is possible to relatively accurately determine whether the current audio frame group includes the target voice. Since non-stationary noise or the user's casual conversation usually does not continuously appear in a plurality of consecutive audio frames within the same audio frame group, the situation of misjudging whether the audio frame group is activated due to individual non-stationary noise (such as keyboard sound, etc.) is reduced, and the accuracy of identifying whether the audio signal includes the target voice is improved.
[0127] In one implementation scenario, the discrimination method using the above audio frame group as a unit is called the short filtering algorithm policy. Assuming that the length of each individual audio frame is 20 ms (milliseconds), when the number of the first target is 4, one current audio frame and the four audio frames before the current audio frame are included in each individual audio frame group, that is, five audio frames are included in each individual audio frame group, and the length of each individual audio frame group is 100 ms. Optionally, each individual audio frame group is called one block. The voice state parameter of each individual audio frame, that is, the VAD value, being 1 means that the target voice is included, and the voice state parameter, that is, the VAD value, being 0 means that the target voice is not included.
[0128] In some embodiments, one statistic is performed for each individual block. Assuming that the quantity threshold is 4, if the number of audio frames with a VAD value of 1 in the current block exceeds 4, the current block is considered to be activated; if the number of audio frames with a VAD value of 1 in the current block does not exceed 4, the current block is considered not to be activated.
[0129] In some embodiments, one statistic is performed for each individual block. Assuming that the preset threshold is 4, if there is a situation where the VAD values of four consecutive audio frames in the current block are 1, the current block is considered to be activated; if there is no situation where the VAD values of four consecutive audio frames in the current block are 1, the current block is considered not to be activated.
[0130] When the terminal determines that the activation state of the audio frame group and the audio frame groups of the second target quantity before the audio frame group meets the second condition, the terminal determines that the target voice is included in the audio signal.
[0131] In some embodiments, in the audio frame group and the audio frame groups of the second target quantity before the audio frame group, if the number of audio frame groups in the activation state of activation exceeds the target threshold, it means that the second condition is met, and thereby it is determined that the target voice is included in the audio signal. In the audio frame group and the audio frame groups of the second target quantity before the audio frame group, if the number of audio frame groups in the activation state of activation does not exceed the target threshold, it means that the second condition is not met, and thereby it is determined that the target voice is not included in the audio signal. That is, the second condition is that the number of audio frame groups in the activation state of activation in the audio frame group and the audio frame groups of the second target quantity before the audio frame group exceeds the target threshold. Here, the value range of the target threshold is greater than or equal to 1 and less than or equal to the numerical value obtained by adding 1 to the second target quantity. For example, if the second target quantity is 29, the value range of the target threshold is greater than or equal to 1 and less than or equal to 30.
[0132] In the above process, in the audio frame group and the audio frame groups of the second target quantity before the audio frame group, if the audio frame groups exceeding a certain target threshold are activated, it is considered that the target voice is included in the entire audio signal, which can reduce the interference caused by some random errors and improve the accuracy of identifying whether the target voice is included in the audio signal.
[0133] In some embodiments, if the activation state of the audio frame group in which consecutive specified thresholds exist is activated in the audio frame group and the second target quantity of audio frame groups before the audio frame group, it means that it meets the second condition, and thereby it is determined that the target voice is included in the audio signal. If the activation state of the audio frame group in which no consecutive specified thresholds exist is activated in the audio frame group and the second target quantity of audio frame groups before the audio frame group, it means that it does not meet the second condition, and thereby it is determined that the target voice is not included in the audio signal. That is, the second condition is that the activation state of the audio frame group in which consecutive specified thresholds exist is activated in the audio frame group and the second target quantity of audio frame groups before the audio frame group. Here, the value range of the specified threshold is greater than or equal to 1 and less than or equal to the numerical value obtained by adding 1 to the second target quantity. For example, if the second target quantity is 29, the value range of the specified threshold is greater than or equal to 1 and less than or equal to 30.
[0134] In the above process, if the activation state of the audio frame group in which consecutive specified thresholds exist is activated in the audio frame group and the second target quantity of audio frame groups before the audio frame group, it is regarded that the target voice is included in the entire audio signal, reducing the interference caused by some random errors, and improving the accuracy of identifying whether the target voice is included in the audio signal.
[0135] In one implementation scenario, a discrimination method based on an audio frame group with a numerical value obtained by adding 1 to the above second target quantity is called a long filtering algorithm policy. Assuming that the length of each audio frame is 20 ms, when the first target quantity is 4, the length of each audio frame group (referred to as one block) is 100 ms. When the second target quantity is 29, the current audio frame group and the 29 audio frame groups before this audio frame group are called one duration. Since each duration contains a total of 30 audio frame groups, the length of each duration is 3 s (seconds), that is, one 3-s duration contains 30 100-ms blocks. Optionally, statistical analysis is performed on the audio signal using a sliding window policy. Assuming that the step size of the sliding window is one block and the length of one block is 100 ms, when the size of the sliding window is 30, one sliding window can just cover one duration, thereby enabling one statistical analysis for one duration each time it slides. In other words, statistical analysis is performed by adopting a sliding window with a size of 30 and a step size of 100 ms on the audio signal.
[0136] In some embodiments, assuming that the target threshold is 10, if the number of blocks activated within one sliding window, i.e., one duration, exceeds 10, it means that the second condition is met, and it is determined that the audio signal contains target speech. That is, based on the gain parameters on the multiple frequency bands of the multiple audio frames, it is determined that the audio signal contains target speech, and the following step 310 is executed to output a prompt message externally; otherwise, no prompt processing is performed.
[0137] In some embodiments, assuming that the specified threshold value is 10, if there is a case where 10 consecutive blocks are activated within one sliding window, i.e., duration, it means that the second condition is met, and it is determined that the target voice is included in the audio signal. That is, based on the gain parameters on the multiple frequency bands of the multiple audio frames, it is determined that the target voice is included in the audio signal, and the following step 310 is executed to output a prompt message externally, otherwise, no prompt processing is performed.
[0138] In some embodiments, when the audio signal is an audio stream, if it is detected that the current sliding window meets the second condition, i.e., it is determined that the target voice is included in the audio signal, the sliding window moves and detects on the audio signal according to a step size of 100 ms. Therefore, after the terminal outputs a prompt message externally, it resets the duration of the sliding window and all statistical states for the blocks. In other words, when receiving the audio stream continuously, each time, based on the short filtering algorithm policy and the long filtering algorithm policy, the target number of audio frames within 3 s from the current time is detected. If the second condition is met, a prompt message is output externally, and all statistical states for the duration of the sliding window and the blocks are reset. If the second condition is not met, the sliding window is controlled to continue sliding backward according to a step size of 100 ms. Optionally, if the length of the currently collected audio signal is less than 3 s, or the length of the newly collected audio signal after the sliding window is reset is less than 3 s, at this time, it is in the window filling state, and it is not determined whether the target voice is included in the audio signal in the window filling state, and the corresponding identification result is not determined until the sliding window is filled for the first time.
[0139] A possible embodiment is provided in which, in the above steps 308-309, it is determined that the audio signal contains a target voice if the voice state parameters of any one audio frame and a target number of audio frames preceding the audio frame meet a first condition, where the target quantity is determined based on the first target quantity and the above second target quantity, i.e., the first condition is that the activation states of the audio frame group to which the audio frame belongs and the second target number of audio frame groups preceding the audio frame group meet a second condition.
[0140] 310: The terminal outputs a prompt message, which is used to prompt the terminal to unmute the microphone.
[0141] The above step 310 is similar to the above step 203, so a detailed description thereof will be omitted here.
[0142] In the above process, when the terminal determines that the audio signal contains a target voice based on the gain parameter, it outputs a prompt message, so as to prompt the user to unmute the microphone in a timely manner, thereby avoiding invalid communication and improving the efficiency of man-machine interaction.
[0143] FIG. 5 is a schematic diagram of the principle of the audio signal processing method provided by the embodiment of the present application. As shown in 500, the microphone collects an audio signal, and after framing, windowing, and Fourier transform, it determines the sampling rate of the audio signal. If the sampling rate is greater than the sampling rate threshold, it is a super-resolution signal; if the sampling rate is less than or equal to the sampling rate threshold, it is a high-resolution signal. Frequency band separation is performed on the super-resolution signal to separate the audio signal into a low-frequency signal and a high-frequency signal. The low-frequency signal is directly input into an acoustic echo canceller (AEC module) to cancel the acoustic echo. There is no need to perform frequency band separation on the high-resolution signal, and the high-resolution signal is directly input into the AEC module to cancel the acoustic echo. The audio signal after acoustic echo cancellation is the first signal, and a plurality of audio frames in the first signal are input into the RNN noise suppression model. The RNN noise suppression model outputs the gain parameters on each frequency band from 0 to 8000 Hz for each individual audio frame, inputs the gain parameters on each frequency band from 0 to 8000 Hz for each individual audio frame into the VAD decision-making module, extracts the gain parameters on each individual frequency band from 200 to 2000 Hz for each individual audio frame and performs weighted integration to obtain the comprehensive gain parameter of each individual audio frame, and then determines the voice state parameter VAD value of each individual audio frame.At this time, if the terminal is in the microphone mute state, the VAD value of each audio frame is input in the microphone mute prompt module, and VAD statistical filtering is performed based on the short filtering algorithm policy (that is, the number of audio frames activated at a certain instant time, for example, in the current block is counted), and microphone mute prompt filtering is performed based on the long filtering algorithm policy (that is, the number of blocks activated within a certain long time, for example, within the current duration is counted). If the number of blocks activated within the current duration exceeds the target threshold, it is determined that the target voice is included in the audio signal. If the number of blocks activated within the current duration does not exceed the target threshold, it is determined that the target voice is not included in the audio signal.
[0144] If the target voice is included in the audio signal, it means that the user utters the target voice in the microphone mute state, that is, the reporting condition is met. In this case, a prompt message is output externally. If the target voice is not included in the audio signal, it means that the user does not utter the target voice in the microphone mute state, that is, the reporting condition is not met. In this case, no prompt message is output. Optionally, after the VAD decision module outputs the VAD value of each individual audio frame, if the terminal is in the microphone-on state, at this time, the audio signal collected by the microphone needs to be normally transmitted to other terminals participating in the multi-person voice call by sending it properly, so as to ensure the normal transmission of the audio signal. For the super-resolution signal, frequency band synthesis needs to be performed on the low-frequency signal obtained by frequency band separation and the original high-frequency signal to restore the original super-resolution signal again, and further encoding and transmission need to be performed on the super-resolution signal. Of course, for the high-resolution signal, since there is no need to perform frequency band separation, there is no need to perform frequency band synthesis either, and direct encoding and transmission can be performed. In some embodiments, the terminal sends the encoded audio signal to the server, and the server transfers the encoded audio signal to other terminals participating in the multi-person voice call.
[0145] For example, for a 16 kHz super-resolution signal collected by a microphone, frequency band separation is performed based on a QMF analysis filter to output an 8 kHz high-frequency signal and an 8 kHz low-frequency signal. On the other hand, subsequent noise suppression and VAD decision act only on the 8 kHz low-frequency signal. If the terminal is in the microphone-on state at this time, it is necessary to use a QMF synthesis filter to synthesize the 8 kHz high-frequency signal and the 8 kHz low-frequency signal into a 16 kHz super-resolution signal again, and then perform encoding and transmission on the super-resolution signal.
[0146] In some embodiments, when the terminal is in the microphone-on state, before performing frequency band synthesis and encoding transmission on the audio signal, the gain parameter of the amplification circuit is automatically adjusted according to the signal strength, thereby improving the transmission effect of the audio signal. Even so, it is still supported to perform AGC processing on the audio signal.
[0147] Since the above selectable technical solutions can adopt any combination to form selectable embodiments of the present disclosure, detailed descriptions are omitted here.
[0148] The method provided by the embodiments of the present application is that when in the microphone-mute state during a multi-person voice call, the application program still collects the user's audio signal, but does not transmit the collected audio signal to other accounts participating in the multi-person voice call. The application program performs signal analysis and processing on the audio signal, and determines whether the audio signal contains target voice by using the gain parameters on multiple frequency bands in the first frequency band range of each of the multiple audio frames in the audio signal. If the audio signal contains target voice, it means that the user forgets to release the microphone-mute state and starts speaking. Accordingly, a prompt message is output to the outside, and the user is timely prompted to release the microphone-mute state, reducing the loss of communication efficiency caused by the user not noticing that they are in the microphone-mute state, improving the man-machine interaction efficiency, and optimizing the user experience.
[0149] In the above embodiments, it is shown how the terminal determines whether the audio signal contains the target voice based on the gain parameters on the individual frequency bands of the individual audio frames. On the other hand, in the embodiments of the present application, how the terminal combines the gain parameters on the individual frequency bands of the individual audio frames and the energy parameters of the individual audio frames to comprehensively determine whether the audio signal contains the target voice, that is, based on the energy parameters of the individual audio frames and the gain parameters output by the noise suppression model for the individual audio frames, whether to comprehensively make a VAD decision will be described below.
[0150] FIG. 6 is a flowchart of a method for processing an audio signal provided by an embodiment of the present application. As shown in reference to FIG. 6, the embodiment is executed by an electronic device. Taking the example that the electronic device is a terminal, the terminal refers to any one of the terminals participating in a multi-person voice call, for example, the first terminal 120 or the second terminal 160 in the above implementation environment. The embodiment includes the following steps.
[0151] 601: The terminal accesses a multi-person voice call in the application program.
[0152] The above step 601 is similar to the above step 301, and detailed description is omitted here.
[0153] 602: The terminal sets the account logged in to the application program in the multi-person voice call to the microphone mute state.
[0154] Since the above step 602 is similar to the above step 302, detailed description is omitted here.
[0155] 603: The terminal acquires the audio signal collected by the application program in the target scene, where the target scene is that the account logged in to the application program is in the microphone mute state in a multi-person voice call.
[0156] Since step 603 above is similar to step 303 above, detailed description is omitted here.
[0157] 604: The terminal preprocesses the audio signal to obtain a first signal.
[0158] Since step 604 above is similar to step 304 above, detailed description is omitted here.
[0159] 605: The terminal inputs a plurality of audio frames in the first signal into a noise suppression model, and the noise suppression model processes each audio frame in the plurality of audio frames, and outputs a gain parameter on each frequency band in the respective first frequency band range of each audio frame. Here, the gain parameter on the frequency band of the human voice in the audio frame is larger than the gain parameter on the noise frequency band.
[0160] Step 605 above is similar to step 305 above, and detailed description is omitted here.
[0161] 606: For each audio frame, the terminal determines the gain parameter on each frequency band in the second frequency band range of the audio frame based on the gain parameter on each frequency band in the first frequency band range of the audio frame, where the second frequency band range is a subset of the first frequency band range.
[0162] Since step 606 above is similar to step 306 above, detailed description is omitted here.
[0163] 607: The terminal acquires the energy parameter of the audio frame.
[0164] In some embodiments, the terminal determines the modulus of the amplitude of the audio frame as the energy parameter of the audio frame. Since the terminal executes step 607 for each individual audio frame, the energy parameters of the plurality of audio frames in the audio signal can be acquired.
[0165] 608: The terminal determines the voice state parameter of the audio frame based on the gain parameter on each frequency band in the second frequency band range of the audio frame and the energy parameter of the audio frame.
[0166] In some embodiments, for each individual audio frame, the terminal determines the comprehensive gain parameter of the audio frame based on the gain parameters on the plurality of frequency bands of the audio frame. Since the acquisition method of the comprehensive gain parameter is similar to step 307 above, detailed description is omitted here.
[0167] In some embodiments, when the overall gain parameter of the audio frame is greater than the activation threshold after being amplified by the target multiple and the energy parameter of the audio frame is greater than the energy threshold, the terminal determines that the voice state parameter of the audio frame includes the target voice. When the overall gain parameter of the audio frame is less than or equal to the activation threshold after being amplified by the target multiple, or when the energy parameter of the audio frame is less than or equal to the energy threshold, the terminal determines that the voice state parameter of the audio frame does not include the target voice. Here, the target multiple is any numerical value greater than 1. For example, the target multiple is 10000. Here, the activation threshold is any numerical value greater than 0. For example, the activation threshold is 6000. Here, the energy threshold is any numerical value greater than or equal to 0 and less than or equal to 100. For example, the energy threshold is 30.
[0168] In one implementation scenario, taking the second frequency band range of 200 - 2000 Hz, the target multiple of 10000, the activation threshold of 6000, and the energy threshold of 30 as an example. After the user turns on a multi - person voice call and then speaks one voice into the microphone in the microphone - muted state, after the microphone collects the audio signal, it obtains the gain parameters on each frequency band within 200 - 2000 Hz for each frame (assuming the length of each frame is 20 ms). Here, the gain parameter is a value greater than or equal to 0 and less than or equal to 1. Weighted integration is performed on the gain parameters on each frequency band within 200 - 2000 Hz of each frame to obtain the comprehensive gain parameter of each frame. The comprehensive gain parameter of each frame is amplified by 10000 times. If the amplified value is greater than 6000, the voice state of the current frame is considered to be activated. At the same time, the energy parameter of the current frame is calculated. If the energy parameter is greater than 30, the energy parameter of the current frame is also considered to be activated. In the VAD decision - making, only for the audio frames in which the voice state and the energy parameter are simultaneously activated, the voice state parameter, that is, the VAD value, is set to 1. Otherwise, as long as the voice state is not activated (the amplified gain parameter is less than or equal to 6000) or the energy parameter is not activated (the energy parameter is less than or equal to 30), the voice state parameter, that is, the VAD value, is set to 0.
[0169] In the above process, in the process of making VAD decisions for individual audio frames, it is required to satisfy the conditions corresponding to both the gain parameter and the energy parameter respectively, and set the VAD value of the current frame to 1, that is, calculate the VAD value of the current frame by comprehensively considering both elements of gain and energy. Here, the energy parameter can roughly estimate the distance between the user and the microphone by intuitively reflecting the volume of the user's speech, prevent the sound in the far field from being misjudged as the voice of a person in the near field, and further improve the accuracy of voice identification.
[0170] In steps 605 to 608 above, based on the gain parameters of the plurality of audio frames in the plurality of frequency bands and the energy parameters of the plurality of audio frames, the terminal can determine the voice state parameters of the plurality of audio frames, perform voice activity detection based on the RNN noise suppression model and energy detection, and thereby accurately identify the target voice and noise on the premise of controlling a relatively small computational complexity, especially having a very high identification accuracy for non-stationary noise, reducing the situations of false reports and reporting errors, sensitively capturing the user's speaking state, and reporting and outputting the prompt message in a timely manner.
[0171] 609: Based on the voice state parameters of the audio frame and the first target number of audio frames before the audio frame, the terminal determines the activation state of the audio frame group to which the audio frame belongs. The audio frame group includes the audio frame and the first target number of audio frames before the audio frame.
[0172] Since step 609 is similar to step 308 above, detailed description is omitted here.
[0173] 610: When the activation state of the audio frame group and the audio frame group with the second target quantity before the audio frame group meet the second condition, the terminal determines that the target voice is included in the audio signal.
[0174] Since step 610 above is similar to step 309 above, detailed description is omitted here.
[0175] 611: The terminal outputs a prompt message, and the prompt message is used to prompt to release the microphone mute state.
[0176] Since step 611 above is similar to step 310 above, detailed description is omitted here.
[0177] FIG. 7 is a schematic diagram of the principle of the audio signal processing method provided by the embodiment of the present application. As shown in 700, the microphone collects an audio signal, and after framing, windowing, and Fourier transform, determines the sampling rate of the audio signal. If the sampling rate is greater than the sampling rate threshold, it is a super-resolution signal; if the sampling rate is less than or equal to the sampling rate threshold, it is a high-resolution signal. Frequency band separation is performed on the super-resolution signal to separate the audio signal into a low-frequency signal and a high-frequency signal. The low-frequency signal is directly input into the AEC module to cancel acoustic echo, and there is no need to perform frequency band separation on the high-resolution signal. The high-resolution signal is directly input into the AEC module to cancel acoustic echo. The audio signal after acoustic echo cancellation is the first signal. A plurality of audio frames in the first signal are input into the RNN noise suppression model. The RNN noise suppression model outputs the gain parameters on each frequency band from 0 to 8000 Hz for each individual audio frame, and inputs the gain parameters on each frequency band from 0 to 8000 Hz of each individual audio frame into the VAD decision module. In addition, energy calculation is performed on each individual audio frame, and the energy parameters of each individual audio frame are also input into the VAD decision module. In the VAD decision module, the gain parameters on each individual frequency band from 200 to 2000 Hz for each individual audio frame are extracted for weighted integration to obtain the comprehensive gain parameter of each individual audio frame. Then, the comprehensive gain parameter and the energy parameter are combined to comprehensively determine the voice state parameter VAD value of each individual audio frame. Only when both conditions of gain and energy are simultaneously satisfied and activated, the VAD value of the audio frame is set to 1; otherwise, the VAD value of the audio frame is set to 0 as long as either one of the gain and energy conditions is not activated.
[0178] At this time, if the terminal is in the microphone mute state, the VAD value of each audio frame is input to the microphone mute prompt module, and VAD statistical filtering is performed based on the short filtering algorithm policy (that is, the number of audio frames activated at a certain instant time, for example, in the current block is counted), and microphone mute prompt filtering is performed based on the long filtering algorithm policy (that is, the number of blocks activated within a certain long time, for example, within the current duration is counted). If the number of blocks activated within the current duration exceeds the target threshold, it is determined that the target voice is included in the audio signal, and if the number of blocks activated within the current duration does not exceed the target threshold, it is determined that the target voice is not included in the audio signal.
[0179] If the target voice is included in the audio signal, it means that the user utters the target voice in the microphone muted state, that is, the reporting condition is met. In this case, a prompt message is output externally. If the target voice is not included in the audio signal, it means that the user does not utter the target voice in the microphone muted state, that is, the reporting condition is not met. In this case, no prompt message is output. Optionally, after the VAD decision module outputs the VAD value of each individual audio frame, if the terminal is in the microphone-on state, at this time, the audio signal collected by the microphone needs to be normally transmitted to other terminals participating in the multi-person voice call by sending it normally, so as to ensure the normal transmission of the audio signal. For the super-resolution signal, it is necessary to perform frequency band synthesis on the low-frequency signal obtained by frequency band separation and the original high-frequency signal to restore the original super-resolution signal again, and then perform encoding and transmission on the super-resolution signal. Of course, for the high-resolution signal, since there is no need to perform frequency band separation, there is no need to perform frequency band synthesis either, and encoding and transmission can be directly performed. In some embodiments, the terminal sends the encoded audio signal to the server, and the server transfers the encoded audio signal to other terminals participating in the multi-person voice call.
[0180] For example, for a 16 kHz super-resolution signal collected by a microphone, frequency band separation is performed based on the QMF analysis filter to output an 8 kHz high-frequency signal and an 8 kHz low-frequency signal. On the other hand, subsequent noise suppression and VAD decision only act on the 8 kHz low-frequency signal. If the terminal is in the microphone-on state at this time, it is necessary to use the QMF synthesis filter to synthesize the 8 kHz high-frequency signal and the 8 kHz low-frequency signal into a 16 kHz super-resolution signal again, and then perform encoding and transmission on the super-resolution signal.
[0181] In some embodiments, when the terminal is in the microphone-on state, before performing frequency band synthesis and encoding transmission on the audio signal, the gain parameter of the amplification circuit is automatically adjusted according to the signal strength, thereby improving the transmission effect of the audio signal. Even so, it is still supported to perform AGC processing on the audio signal.
[0182] The above selectable technical solutions can be combined arbitrarily to form selectable embodiments of the present disclosure, and detailed descriptions are omitted here.
[0183] The method provided by the embodiments of the present application is that when in the microphone-mute state during a multi-person voice call, the application program still collects the user's audio signal, but does not transmit the collected audio signal to other accounts participating in the multi-person voice call. The application program performs signal analysis and processing on the audio signal, and determines whether the audio signal contains target voice by using the gain parameters on multiple frequency bands within the first frequency band range of each of the multiple audio frames in the audio signal. If the audio signal contains target voice, it means that the user has forgotten to release the microphone-mute state and started speaking. Accordingly, a prompt message is output externally, and the user is timely prompted to release the microphone-mute state, reducing the loss of communication efficiency caused by the user not noticing that they are in the microphone-mute state, improving the man-machine interaction efficiency, and optimizing the user experience.
[0184] In each of the above embodiments, either directly utilize the gain parameter of each audio frame output by the RNN to make a VAD decision, or combine the gain parameter of each audio frame output by the RNN with the energy parameter of each audio frame to simultaneously make a VAD decision. Each of the above two methods does not require the adoption of a conventional VAD detection algorithm. On the other hand, in the embodiments of the present application, by combining the RNN noise suppression model and the VAD detection algorithm, it relates to a method for identifying whether a target voice is included in an audio signal, which will be described in detail below.
[0185] FIG. 8 is a flowchart of a method for processing an audio signal provided by an embodiment of the present application. As shown in FIG. 8, this embodiment is executed by an electronic device. Taking the example that the electronic device is a terminal, the terminal is any one of the terminals participating in a multi-person voice call, for example, the first terminal 120 or the second terminal 160 in the above implementation environment. This embodiment includes the following steps.
[0186] 801: The terminal accesses a multi-person voice call in an application program.
[0187] The above step 801 is similar to the above step 301, and detailed description is omitted here.
[0188] 802: The terminal sets the account logged in to the application program in the multi-person voice call to the microphone mute state.
[0189] The above step 802 is similar to the above step 302, and detailed description is omitted here.
[0190] 803: The terminal acquires the audio signal collected by the application program in the target scene, where the target scene is that the account logged in to the application program is in the microphone mute state in a multi-person voice call.
[0191] The above step 803 is similar to the above step 303, and the detailed description is omitted here.
[0192] 804: The terminal preprocesses the audio signal to obtain a first signal.
[0193] The above step 804 is similar to the above step 304, and the detailed description is omitted here.
[0194] 805: The terminal inputs a plurality of audio frames in the first signal into a noise suppression model, and the noise suppression model processes each audio frame in the plurality of audio frames, and outputs a gain parameter on each frequency band in the respective first frequency band range of each audio frame. Here, the gain parameter on the frequency band of the human voice of the audio frame is larger than the gain parameter on the noise frequency band.
[0195] The above step 805 is similar to the above step 305, and the detailed description is omitted here.
[0196] 806: The terminal performs noise suppression on the plurality of audio frames based on the gain parameters on the plurality of frequency bands of the plurality of audio frames, and obtains a plurality of target audio frames.
[0197] In some embodiments, for each individual audio frame, the terminal amplifies or attenuates the signal component of the corresponding frequency band in the audio frame based on the gain parameter on each individual frequency band within the first frequency band range of the audio frame, to obtain one target audio frame, and performs the above operation on each individual audio frame among a plurality of audio frames, to obtain a plurality of target audio frames.
[0198] 807: The terminal performs voice activity detection (VAD) based on the energy parameter of the plurality of target audio frames, and obtains the VAD values of the plurality of target audio frames.
[0199] In some embodiments, for each individual target audio frame, the terminal obtains the modulus of the amplitude of the target audio frame as the energy parameter of the target audio frame, and performs the above operation on each individual target audio frame among a plurality of target audio frames, to obtain the energy parameters of the plurality of target audio frames.
[0200] In some embodiments, voice activity detection is performed on the energy parameters of the plurality of target audio frames by using a VAD detection algorithm, and the VAD value of each of the plurality of target audio frames is output. Optionally, the VAD detection algorithm includes, but is not limited to, a VAD detection algorithm based on a Gaussian Mixture Model (GMM), a VAD detection algorithm based on a dual threshold, a VAD detection algorithm based on a statistical model, a VAD detection algorithm based on Empirical Mode Decomposition (EMD), a VAD detection algorithm based on a correlation coefficient method, or a VAD detection algorithm based on a wavelet transform method, etc. Embodiments of the present application do not limit this.
[0201] In one implementation scenario, GMM-VAD is taken as an example for explanation. The GMM-VAD algorithm assumes that both human voices and noise conform to a Gaussian distribution, and that the noise is smoother than human voices and the noise energy is smaller than the energy of human voices, that is, the mean value and variance of the noise signal are smaller than the mean value and variance of the human voice signal. Therefore, two Gaussian models can be used to fit the human voice signal and the noise signal in the input signal (that is, the plurality of target audio frames refers to the first signal with noise suppressed), respectively, and the two can be separated according to the above assumptions. After fitting and separating by the Gaussian model, six parameters of the mean value, variance, and weight of the human voice signal and the mean value, variance, and weight of the noise signal will be output.
[0202] For example, the input signal (that is, the plurality of target audio frames is the first signal with noise suppressed) is divided into six frequency bands: 80Hz - 250Hz, 250Hz - 500Hz, 500Hz - 1KHz, 1KHz - 2KHz, 2KHz - 3KHz, and 3KHz - 4KHz. For each individual frequency band, the signal is fitted using a GMM model.
[0203] When the GMM-VAD algorithm is initialized, the above six parameters will use initial values (for example, pre-trained parameters). Every time a new target audio frame is input into the GMM model, the likelihood probability is calculated based on the existing GMM model, and it is determined whether the current target audio frame is a human voice or noise. Next, according to the judgment result of the GMM model, the above six parameters are updated using maximum likelihood estimation, and then the GMM model is updated. The above process is repeatedly executed to determine whether each target audio frame is a human voice or noise. If the target audio frame is a human voice, the VAD value of the target audio frame is set to 1; if the target audio frame is noise, the VAD value of the target audio frame is set to 0.
[0204] 808: When the VAD values of the plurality of target audio frames meet the third condition, the terminal determines that the audio signal contains target voice.
[0205] In some embodiments, the terminal may also determine whether the audio signal contains target voice by judging the VAD values of the plurality of target audio frames based on a short filtering algorithm policy and a long filtering algorithm policy respectively.
[0206] Optionally, for each target audio frame, the terminal determines the activation state of the target audio frame group to which the target audio frame belongs based on the VAD value of the target audio frame and the VAD values of the first target number of target audio frames before the target audio frame. The target audio frame group includes the target audio frame and the first target number of target audio frames before the target audio frame. When the activation states of the target audio frame group and the target audio frame group of the second target number before the target audio frame meet the second condition, it means that the VAD values of the plurality of target audio frames meet the third condition, and it is determined that the audio signal contains target voice. Since the above judgment method is similar to the above steps 308-309, detailed description is omitted here.
[0207] 809: The terminal outputs a prompt message, which is used to prompt to release the microphone mute state.
[0208] Since the above step 809 is similar to the above step 310, detailed description is omitted here.
[0209] FIG. 9 is a schematic diagram of the principle of the audio signal processing method provided by the embodiment of the present application. As shown in 900, the microphone collects an audio signal, and after framing, windowing, and Fourier transform, determines the sampling rate of the audio signal. If the sampling rate is greater than the sampling rate threshold, it is a super-resolution signal, and if the sampling rate is less than or equal to the sampling rate threshold, it is a high-resolution signal. Frequency band separation is performed on the super-resolution signal to separate the audio signal into a low-frequency signal and a high-frequency signal. The low-frequency signal is directly input into the AEC module to cancel acoustic echo, and there is no need to perform frequency band separation on the high-resolution signal. The high-resolution signal is directly input into the AEC module to cancel acoustic echo. The audio signal after acoustic echo cancellation is the first signal. A plurality of audio frames in the first signal are input into the RNN noise suppression model. The RNN noise suppression model outputs the gain parameters on each frequency band from 0 to 8000 Hz for each individual audio frame, and performs noise suppression on each individual audio frame based on each gain parameter to obtain a plurality of target audio frames. Energy calculation is performed on each individual target audio frame to obtain the energy parameter of each individual target audio frame. The energy parameter of each individual target audio frame is input into the GMM-VAD module, and the GMM model is used to predict whether the target audio frame is target voice or noise for each individual target audio frame. If the target audio frame is target voice, the VAD value of the target audio frame is set to 1, and if the target audio frame is noise, the VAD value of the target audio frame is set to 0.
[0210] At this time, if the terminal is in the microphone mute state, the VAD value of each target audio frame is input in the microphone mute prompt module, and VAD statistical filtering is performed based on the short filtering algorithm policy (that is, the number of target audio frames activated at a certain instantaneous time, for example, in the current block is counted), and microphone mute prompt filtering is performed based on the long filtering algorithm policy (that is, the number of blocks activated within a certain long time, for example, within the current duration is counted). If the number of blocks activated within the current duration exceeds the target threshold, it is determined that the target voice is included in the audio signal, and if the number of blocks activated within the current duration does not exceed the target threshold, it is determined that the target voice is not included in the audio signal.
[0211] If the target voice is included in the audio signal, it means that the user is emitting the target voice in the microphone mute state, that is, the reporting condition is met. In this case, a prompt message is output externally. If the target voice is not included in the audio signal, it means that the user is not emitting the target voice in the microphone mute state, that is, the reporting condition is not met. In this case, no prompt message is output. Optionally, after the GMM-VAD module outputs the VAD value of each individual target audio frame, if the terminal is in the microphone-on state, at this time, the audio signal collected by the microphone needs to be normally transmitted to other terminals participating in the multi-person voice call by sending it normally, so as to ensure the normal transmission of the audio signal. For the super-resolution signal, frequency band synthesis needs to be performed on the low-frequency signal obtained by frequency band separation and the original high-frequency signal to restore the original super-resolution signal again, and then encoding and transmission need to be performed on the super-resolution signal. Of course, for the high-resolution signal, since there is no need to perform frequency band separation, there is no need to perform frequency band synthesis either, and encoding and transmission can be directly performed. In some embodiments, the terminal sends the encoded audio signal to the server, and the server transfers the encoded audio signal to other terminals participating in the multi-person voice call.
[0212] For example, for a 16 kHz super-resolution signal collected by a microphone, frequency band separation is performed based on the QMF analysis filter to output an 8 kHz high-frequency signal and an 8 kHz low-frequency signal. On the other hand, subsequent noise suppression and VAD decision-making only act on the 8 kHz low-frequency signal. If the terminal is in the microphone-on state at this time, it is necessary to use the QMF synthesis filter to synthesize the 8 kHz high-frequency signal and the 8 kHz low-frequency signal into a 16 kHz super-resolution signal again, and then perform encoding and transmission on the super-resolution signal.
[0213] In some embodiments, when the terminal is in a microphone-on state, before performing frequency band synthesis and encoding / transmission on the audio signal, the gain parameter of the amplifier circuit is automatically adjusted according to the signal strength, thereby improving the transmission effect of the audio signal, and further supports performing AGC processing on the audio signal.
[0214] FIG. 10 shows a text prompt message provided by an embodiment of the present application. 10 is a schematic diagram. As shown in Fig. 10, in a multi-party voice call interface 1000, if the audio signal contains a target voice, the terminal displays a text prompt message 1001 "microphone is muted, please unmute the microphone before speaking" and displays a microphone setting control member 1002 in a disabled state. The text prompt message 1001 is used to prompt the user to click the microphone setting control member 1002 in a disabled state to set the microphone setting control member 1002 to an enabled state from a disabled state, thereby unmuting the microphone mute state.
[0215] All the above optional technical solutions can be adopted in any combination to form optional embodiments of the present disclosure, so detailed descriptions are omitted here.
[0216] The method provided by the embodiments of the present application is that when in the microphone mute state during a multi-person voice call, the application program still collects the user's audio signal, but does not transmit the collected audio signal to other accounts participating in the multi-person voice call. The application program performs signal analysis and processing on the audio signal, and determines whether the audio signal contains target voice by using the gain parameters on multiple frequency bands in the first frequency band range of each of the multiple audio frames in the audio signal. If the audio signal contains target voice, it means that the user has forgotten to unmute the microphone and started speaking. Accordingly, a prompt message is output externally, and the user is timely prompted to unmute the microphone, reducing the loss of communication efficiency caused by the user not noticing that they are in the microphone mute state, improving the man-machine interaction efficiency, and optimizing the user experience.
[0217] In the test scenario, some pure noises, pure voices (male voice, female voice, Chinese, English), and voices with noise in multiple scenes are respectively selected to test the stability and sensitivity degree of the audio signal processing method provided by each of the above embodiments. Here, the noises include stationary noises (car noise, wind noise, street, subway, coffee shop, etc.) and non-stationary noises (construction site, keyboard, table, knocking, human voice, etc.) are respectively introduced. Since the method provided by the embodiments of the present application does not rely on the conventional energy-based VAD detection, while improving the accuracy of detecting human voice in the audio signal to a certain extent, and at the same time does not rely on a complex CNN model, the calculation consumption can also be guaranteed. The method provided by the embodiments of the present application may be used in each audio-video call scene or audio-video conference, such as voice call, video call, multi-person voice call, multi-person video call, screen sharing, etc., and may also be used in multiple live or communication products and social software, meeting the calculation needs of the minimum energy consumption on the mobile side.
[0218] FIG. 11 is a structural schematic diagram of an audio signal processing apparatus provided by an embodiment of the present application. As shown in FIG. 11, the apparatus includes: A first acquisition module 1101 used to acquire an audio signal collected by an application program in a target scene, where the target scene is a scene where the accounts logged in to the application program are in a microphone mute state in a multi-person voice call, the first acquisition module 1101; A second acquisition module 1102 used to acquire gain parameters on a plurality of frequency bands in a first frequency band range of each of a plurality of audio frames in the audio signal; An output module 1103 used to output a prompt message when it is determined that the audio signal includes target voice based on the gain parameter, where the prompt message is used to prompt to release the microphone mute state, the output module 1103.
[0219] When the device provided by the embodiment of this application is in the microphone mute state during a multi-person voice call, the application program still collects the user's audio signal, but does not send the collected audio signal to other accounts participating in the multi-person voice call. The application program performs signal analysis and processing on the audio signal, and judges whether the audio signal contains target voice by using the gain parameters on a plurality of frequency bands in the first frequency band range of each of the plurality of audio frames in the audio signal. If the audio signal contains target voice, it means that the user has forgotten to release the microphone mute state and started speaking. Accordingly, a prompt message is output externally, and the user is timely prompted to release the microphone mute state, reducing the loss of communication efficiency caused by the user not noticing that they are in the microphone mute state, and improving the man-machine interaction efficiency.
[0220] In one possible embodiment, the second acquisition module 1102 is a preprocessing unit used to preprocess the audio signal to obtain a first signal, and a processing unit used to input a plurality of audio frames in the first signal into a noise suppression model, process each audio frame in the plurality of audio frames by the noise suppression model, and output the gain parameters on the individual frequency bands in the first frequency band range of each audio frame, where the gain parameter on the frequency band of the human voice of the audio frame is larger than the gain parameter on the noise frequency band, and includes a processing unit.
[0221] In one possible embodiment, the noise suppression model is a regression neural network, which includes at least one hidden layer. A plurality of neurons are included in each hidden layer, and the number of neurons in each hidden layer is the same as the number of input audio frames. The processing unit For any one neuron in any one hidden layer of the regression neural network, the neuron performs a weighting process on the frequency features output by the previous neuron in the same hidden layer and the frequency features output by the neuron at the corresponding position in the previous hidden layer, and the frequency features obtained through the weighting process are used to be input into the next neuron in the same hidden layer and the neuron at the corresponding position in the next hidden layer respectively.
[0222] In one possible embodiment, based on the configuration of the device in FIG. 11, the device A first determination module used to determine the voice state parameters of the plurality of audio frames based on the gain parameters on the plurality of frequency bands of the plurality of audio frames, where the voice state parameters are used to characterize whether the corresponding audio frame contains target voice. The device further includes a second determination module used to determine that the audio signal contains target voice when any one audio frame and the voice state parameters of the target number of audio frames before the audio frame meet the first condition.
[0223] In one possible embodiment, based on the configuration of the device in FIG. 11, the first determination module A first determination unit used to determine gain parameters on individual frequency bands in a second frequency band range of an audio frame based on gain parameters on individual frequency bands in a first frequency band range of the audio frame, where the second frequency band range is a subset of the first frequency band range, and a first determination unit. A second determination unit used to determine an audio state parameter of the audio frame based on gain parameters on individual frequency bands in the second frequency band range of the audio frame.
[0224] In one possible embodiment, based on the configuration of the apparatus in FIG. 11, the second determination unit A multiplication subunit used to multiply gain parameters on individual frequency bands in the second frequency band range of the audio frame by weight coefficients of corresponding frequency bands to obtain weighted gain parameters on individual frequency bands in the second frequency band range of the audio frame. An addition subunit used to add weighted gain parameters on each frequency band in the second frequency band range of the audio frame to obtain an overall gain parameter of the audio frame. A determination subunit used to determine an audio state parameter of the audio frame based on the overall gain parameter of the audio frame.
[0225] In one possible embodiment, the determination subunit When the overall gain parameter amplified by a target multiple is greater than an activation threshold, it is used to determine that the audio state parameter includes a target voice. When the overall gain parameter amplified by the target multiple is less than or equal to the activation threshold, it is used to determine that the audio state parameter does not include a target voice.
[0226] In one possible embodiment, based on the configuration of the apparatus in FIG. 11, the apparatus further includes a third acquisition module used to acquire energy parameters of the plurality of audio frames, The first determination module includes a third determination unit used to determine audio state parameters of the plurality of audio frames based on the gain parameters of the plurality of audio frames in the plurality of frequency bands and the energy parameters of the plurality of audio frames.
[0227] In one possible embodiment, the third determination unit for each individual audio frame, based on the gain parameters of the audio frame in the plurality of frequency bands, determines the overall gain parameter of the audio frame; when the overall gain parameter of the audio frame amplified by the target multiple is greater than the activation threshold and the energy parameter of the audio frame is greater than the energy threshold, determines that the audio state parameter of the audio frame includes the target audio; when the overall gain parameter of the audio frame amplified by the target multiple is less than or equal to the activation threshold, or when the energy parameter of the audio frame is less than or equal to the energy threshold, determines that the audio state parameter of the audio frame does not include the target audio.
[0228] In one possible embodiment, based on the configuration of the apparatus in FIG. 11, the second determination module A fourth determination unit used to determine the activation state of an audio frame group to which an audio frame belongs based on the audio state parameters of the audio frame and the first target number of audio frames before the audio frame for any one of the audio frames, where the audio frame group includes the audio frame and the first target number of audio frames before the audio frame. A fifth determination unit used to determine that the audio signal includes target audio when the activation states of the audio frame group and the second target number of audio frame groups before the audio frame group meet the second condition, where the target number is determined based on the first target number and the second target number.
[0229] In one possible embodiment, the fourth determination unit is used to determine that the activation state of the audio frame group is activated when the number of audio frames in the audio frame group whose audio state parameters include target audio exceeds a quantity threshold, and is used to determine that the activation state of the audio frame group is deactivated when the number of audio frames in the audio frame group whose audio state parameters include target audio does not exceed the quantity threshold.
[0230] In one possible embodiment, based on the configuration of the device in FIG. 11, the device includes a noise suppression module used to perform noise suppression on the plurality of audio frames based on the gain parameters on the plurality of frequency bands of the plurality of audio frames and obtain a plurality of target audio frames. A voice activity detection (VAD) module used to perform voice activity detection (VAD) based on the energy parameters of the plurality of target audio frames and obtain the VAD values of the plurality of target audio frames. A third determination module used to determine that the target audio is included in the audio signal when the VAD values of the plurality of target audio frames meet the third condition.
[0231] In one possible embodiment, the target audio is the speech in the voice call of the plurality of people of the target object, or the target audio is the voice of the target object.
[0232] All of the above selectable technical solutions can be combined arbitrarily to form selectable embodiments of the present disclosure, and detailed descriptions are omitted here.
[0233] It should be noted that when the audio signal processing device provided in the above embodiment processes the audio signal, only the classification of each functional module is listed and described. However, in actual applications, the above functions can be assigned to different functional modules as needed to complete. That is, by dividing the internal structure of the electronic device into different functional modules, all or part of the functions described above can be completed. In addition, the audio signal processing device provided in the above embodiment belongs to the same concept as the embodiment of the audio signal processing method, and its implementation process can be referred to in detail in the embodiment of the audio signal processing method, so detailed descriptions are omitted here.
[0234] FIG. 12 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As shown in FIG. 12, the electronic device being the terminal 1200 will be described as an example. Optionally, the device type of the terminal 1200 includes a smartphone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer, or a desktop computer. The terminal 1200 may further be referred to by other names such as a user device, a mobile terminal, a laptop terminal, a desktop terminal, etc.
[0235] Generally, the terminal 1200 includes a processor 1201 and a memory 1202.
[0236] Optionally, the processor 1201 includes one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Optionally, the processor 1201 is implemented by adopting at least one type of hardware form among DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). In some embodiments, the processor 1201 includes a main processor and a coprocessor. The main processor is a processor used to process data in the wake-up state and is also called a CPU (Central Processing Unit). The coprocessor is a low-power processor used to process data in the standby state. In some embodiments, a GPU (Graphics Processing Unit) is integrated into the processor 1201, and the GPU is used to perform rendering and drawing of content that needs to be displayed on the display screen. In some embodiments, the processor 1201 further includes an AI (Artificial Intelligence) processor, and the AI processor is used to process calculation operations related to machine learning.
[0237] In some embodiments, the memory 1202 includes one or more computer-readable storage media, and optionally, the computer-readable storage media is non-transitory. Optionally, the memory 1202 further includes a high-speed random access memory and non-volatile memory such as, for example, one or more magnetic disk storage devices, flash memory storage devices, etc. In some embodiments, the non-transitory computer-readable storage media in the memory 1202 is used to store at least one program code, and the at least one program code is used to implement the audio signal processing method provided by each embodiment in the present application when executed by the processor 1201.
[0238] In some embodiments, the terminal 1200 further optionally includes an audio circuit 1207.
[0239] In some embodiments, the audio circuit 1207 includes a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, convert the sound waves into electrical signals, input the electrical signals to the processor 1201 for processing, or input the electrical signals to the radio frequency circuit 1204 to realize voice communication. For the purpose of stereo collection or noise reduction, there are multiple microphones, which are respectively installed at different parts of the terminal 1200. Optionally, the microphone is an array microphone or an omnidirectional sound-collecting microphone. The speaker is used to convert the electrical signal from the processor 1201 or the radio frequency circuit 1204 into sound waves. Optionally, the speaker is a conventional thin-film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into audible sound waves by humans, but also perform applications such as ranging by converting the electrical signal into non-audible sound waves by humans. In some embodiments, the audio circuit 1207 further includes a headphone jack.
[0240] FIG. 13 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Since the arrangement or performance of the electronic device 1300 is different, a relatively large difference occurs. The electronic device 1300 includes one or more central processing units (CPUs) 1301 and one or more memories 1302. Here, at least one computer program is stored in the memory 1302. The at least one computer program is loaded and executed by the one or more processors 1301 to implement the audio signal processing method provided by each of the above embodiments. Optionally, the electronic device 1300 further includes components such as a wired or wireless network interface, a keyboard, and an input / output interface to facilitate input / output. The electronic device 1300 further includes other components used to implement the functions of the device, and detailed descriptions thereof are omitted here.
[0241] In an exemplary embodiment, a computer-readable storage medium, such as a memory including at least one computer program, is further provided. The at least one computer program is executed by a processor in a terminal to enable completion of the audio signal processing method in each of the above embodiments. For example, the computer-readable storage medium includes a ROM (Read-Only Memory), a RAM (Random-Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0242] In an exemplary embodiment, a computer program product or a computer program is further provided, which includes one or more program codes, and the one or more program codes are stored in a computer-readable storage medium. One or more processors of an electronic device can read the one or more program codes from the computer-readable storage medium, and the one or more processors execute the one or more program codes to enable the electronic device to complete the method for processing an audio signal in the above embodiment when executed.
[0243] As can be understood by those skilled in the art, all or some of the steps for implementing the above embodiment may be completed by hardware, or may be completed by instructing related hardware by a program. Optionally, the program is stored in a computer-readable storage medium, and optionally, the above-mentioned storage medium is a read-only memory, a magnetic disk, an optical disk, or the like.
[0244] The above are only optional embodiments of the present application and are not used to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included within the protection scope of the present application.
Description of Reference Numerals
[0245] 120 First terminal 140 Server 160 Second terminal 1000 Call interface 1001 Text prompt message 1002 Microphone setting control member 1101 First acquisition module 1102 Second acquisition module 1103 Output module 1200 Terminal 1201 Processor 1202 Memory 1204 Radio frequency circuit 1207 Audio Circuit 1300 Electronic Device 1301 Processor 1302 Memory
Claims
1. A method for processing an audio signal, performed by a terminal, the method comprising: A step of acquiring an audio signal collected by an application program in a target scene, where an account is logged into the application program, and the target scene refers to a situation where the account is in a microphone-mute state in a multi-party voice call; obtaining gain parameters over a plurality of frequency bands in a first frequency band range for each of a plurality of audio frames in the audio signal; outputting a prompt message when it is determined that the audio signal includes a target voice based on the gain parameter, the prompt message being used for prompting the user to unmute a microphone of the account; Including, When it is determined that the audio signal includes a target voice based on the gain parameter, the step of outputting a prompt message includes: determining speech state parameters of the plurality of audio frames based on gain parameters over the plurality of frequency bands of the plurality of audio frames, the speech state parameters being used to characterize whether a corresponding audio frame includes a target speech or not; determining that the audio signal includes a target voice based on speech state parameters of the plurality of audio frames; Including, determining speech state parameters for the plurality of audio frames based on gain parameters over the plurality of frequency bands for the plurality of audio frames, For each individual audio frame, based on the gain parameters on the individual frequency bands within the first frequency band range of the audio frame, determining the gain parameters on the individual frequency bands corresponding to the second frequency band range of the audio frame selected from the individual first frequency bands of the audio frame, wherein the second frequency band range is a subset of the first frequency band range; Determining the voice state parameters of the audio frame based on the gain parameters on the individual frequency bands within the second frequency band range of the audio frame; comprising; The step of determining the voice state parameters of the audio frame based on the gain parameters on the individual frequency bands within the second frequency band range of the audio frame comprises: Multiplying the gain parameters on the individual frequency bands within the second frequency band range of the audio frame by the weight coefficients of the corresponding frequency bands to obtain the weighted gain parameters on the individual frequency bands within the second frequency band range of the audio frame; Adding the weighted gain parameters on each frequency band within the second frequency band range of the audio frame to obtain the comprehensive gain parameter of the audio frame; Determining the voice state parameters of the audio frame based on the comprehensive gain parameter of the audio frame; comprising; The step of determining the voice state parameters of the audio frame based on the comprehensive gain parameter of the audio frame comprises: When the value obtained by amplifying the comprehensive gain parameter by a predetermined multiple is greater than the activation threshold, determining that the voice state parameter includes the target voice; When the value obtained by amplifying the comprehensive gain parameter by the predetermined multiple is smaller than or equal to the activation threshold value, determining that the voice state parameter does not include the target voice; An audio signal processing method including.
2. An audio signal processing method executed by a terminal, the method comprising: Obtaining an audio signal collected by an application program in a target scene, wherein an account is logged in to the application program, and the target scene refers to a state where the account is in a microphone mute state in a multi-person voice call; Obtaining gain parameters on a plurality of frequency bands in a first frequency band range of each of a plurality of audio frames in the audio signal; When it is determined that the audio signal includes a target voice based on the gain parameter, outputting a prompt message, wherein the prompt message is used to prompt to release the microphone mute state of the account; including When it is determined that the audio signal includes a target voice based on the gain parameter, the step of outputting a prompt message includes: Determining a voice state parameter of the plurality of audio frames based on the gain parameters on the plurality of frequency bands of the plurality of audio frames, wherein the voice state parameter is used to characterize whether the corresponding audio frame includes a target voice; Determining that the audio signal includes a target voice based on the voice state parameters of the plurality of audio frames; including The method includes: Further including the step of obtaining energy parameters of the plurality of audio frames, The step of determining audio state parameters of the plurality of audio frames based on the gain parameters of the plurality of audio frames on the plurality of frequency bands is The step of determining audio state parameters of the plurality of audio frames based on the gain parameters of the plurality of audio frames on the plurality of frequency bands and the energy parameters of the plurality of audio frames includes The step of determining audio state parameters of the plurality of audio frames based on the gain parameters of the plurality of audio frames on the plurality of frequency bands and the energy parameters of the plurality of audio frames is For each individual audio frame, determining an overall gain parameter of the audio frame based on the gain parameters of the audio frame on the plurality of frequency bands; When the value obtained by amplifying the overall gain parameter of the audio frame by a predetermined multiple is greater than the activation threshold and the energy parameter of the audio frame is greater than the energy threshold, determining that the audio state parameter of the audio frame includes the target audio; When the value obtained by amplifying the overall gain parameter of the audio frame by the predetermined multiple is less than or equal to the activation threshold, or when the energy parameter of the audio frame is less than or equal to the energy threshold, determining that the audio state parameter of the audio frame does not include the target audio; A method for processing an audio signal, including.
3. A method for processing an audio signal executed by a terminal, the method comprising: A step of obtaining an audio signal collected by an application program in a target scene, wherein an account is logged in to the application program, and the target scene refers to a state where the account is in a microphone mute state in a multi-person voice call, the step; A step of obtaining gain parameters on a plurality of frequency bands in a first frequency band range of each of a plurality of audio frames in the audio signal; A step of outputting a prompt message when it is determined that the audio signal contains target voice based on the gain parameter, wherein the prompt message is used to prompt to release the microphone mute state of the account, the step; Including; When it is determined that the audio signal contains target voice based on the gain parameter, the step of outputting a prompt message A step of determining a voice state parameter of the plurality of audio frames based on the gain parameters on the plurality of frequency bands of the plurality of audio frames, wherein the voice state parameter is used to characterize whether the corresponding audio frame contains target voice, the step; A step of determining that the audio signal contains target voice based on the voice state parameters of the plurality of audio frames; Including; When it is determined that the audio signal contains target voice based on the voice state parameters of the plurality of audio frames, the step When the voice state parameters of any one audio frame and a predetermined number of audio frames before the audio frame meet the first condition, the step of determining that the audio signal contains target voice is included; When the voice state parameters of any one audio frame and a predetermined number of audio frames before the audio frame meet the first condition, the step of determining that the audio signal includes target voice is For any one of the audio frames, a step of determining an activation state of an audio frame group to which the audio frame belongs based on voice state parameters of the audio frame and a first predetermined number of audio frames before the audio frame, wherein the audio frame group includes the audio frame and a first predetermined number of audio frames before the audio frame When the activation states of the audio frame group and a second predetermined number of audio frame groups before the audio frame group meet the second condition, a step of determining that the audio signal includes target voice, wherein the predetermined number is determined based on the first predetermined number and the second predetermined number including When the voice state parameters of any one audio frame and a predetermined number of audio frames before the audio frame meet the first condition, the step of determining that the audio signal includes target voice is For any one of the audio frames, a step of determining an activation state of an audio frame group to which the audio frame belongs based on voice state parameters of the audio frame and a first predetermined number of audio frames before the audio frame, wherein the audio frame group includes the audio frame and a first predetermined number of audio frames before the audio frame When the activation state of the audio frame group and the second predetermined number of audio frame groups before the audio frame group meets the second condition, the step of determining that the target voice is included in the audio signal, wherein the predetermined number is determined based on the first predetermined number and the second predetermined number, the step; An audio signal processing method including.
4. The step of obtaining gain parameters on a plurality of frequency bands in the first frequency band range of each of the plurality of audio frames in the audio signal is The step of preprocessing the audio signal to obtain a first signal; Inputting a plurality of audio frames in the first signal into a noise suppression model, processing each audio frame in the plurality of audio frames by the noise suppression model, and outputting gain parameters on individual frequency bands in the first frequency band range of each audio frame, wherein the gain parameter on the frequency band of the human voice of the audio frame is greater than the gain parameter on the noise frequency band, the step; The method according to any one of claims 1 to 3, including.
5. The noise suppression model is a regression type neural network, the regression type neural network includes at least one hidden layer, a plurality of neurons are included in each hidden layer, and a plurality of audio frames in the first signal are input into the noise suppression model, and the noise suppression model processes each audio frame in the plurality of audio frames, and the step of outputting gain parameters on individual frequency bands in the first frequency band range of each audio frame is For any one neuron in any one hidden layer in the recurrent neural network, the step of, by the any one neuron, performing a weighting process on the frequency features output by the previous neuron in the any one hidden layer and the frequency features output by the neuron at the corresponding position in the previous hidden layer, and inputting the frequency features obtained by the weighting process into the next neuron in the any one hidden layer and the neuron at the corresponding position in the next hidden layer respectively, the method according to claim 4.
6. For any one of the audio frames, the step of determining the activation state of the audio frame group to which the audio frame belongs based on the audio frame and the voice state parameters of the first predetermined number of audio frames before the audio frame is as follows: When the number of audio frames including the target voice in the audio frame group exceeds the quantity threshold, the step of determining that the activation state of the audio frame group is activated; When the number of audio frames including the target voice in the audio frame group does not exceed the quantity threshold, the step of determining that the activation state of the audio frame group is not activated, the method according to claim 3.
7. When it is determined that the target voice is included in the audio signal based on the gain parameter, the step of outputting a prompt message is as follows: The step of performing noise suppression on the plurality of audio frames based on the gain parameters on the plurality of frequency bands of the plurality of audio frames to obtain a plurality of target audio frames; Performing voice activity detection (VAD) based on the energy parameters of the plurality of target audio frames, and obtaining the VAD values of the plurality of target audio frames; Determining that the target audio is included in the audio signal when the VAD values of the plurality of target audio frames match a predetermined value, the method according to any one of claims 1 to 3.
8. The target audio is the speech in the voice call of the plurality of people of the target object, or the target audio is the voice of the target object, the method according to any one of claims 1 to 3.
9. An audio signal processing apparatus including a processor, disposed in a terminal, The audio signal processing apparatus configured to implement the method according to any one of claims 1 to 3 and 6 by the processor.
10. An electronic device including one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, and the at least one computer program is loaded and executed by the one or more processors to implement the audio signal processing method according to any one of claims 1 to 3 and 6.
11. A computer program that, when loaded and executed by a processor, implements the audio signal processing method according to any one of claims 1 to 3 and 6.
Citation Information
Patent Citations
Method and device for audio processing of conference system
CN107276777A
Mute prompting method and device, electronic equipment and storage medium
CN111343410A
Voice noise reduction method and device, equipment and medium
CN111429932A
Feeder
JP2011024456A
Mute detector
US20160182727A1