A method of processing a voice signal and a device therefor

By determining the accurate locations of sound sources and interference sources in multiple speech frames preceding the speech frame, the filtering quality problem caused by inaccurate sound source location is solved, thus improving the filtering effect of the speech frame.

CN114283826BActive Publication Date: 2025-12-05HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011037133.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-27
Publication Date
2025-12-05
Estimated Expiration
2040-09-27

AI Technical Summary

Technical Problem

In existing technologies, the location of the sound source is not accurately determined during the speech enhancement process, resulting in a lot of residual noise in the filtered speech frames and poor filtering quality.

Method used

By including the wake-up-triggered voice frames in multiple voice frames prior to the determined voice frame, the location of the sound source and interference source is determined using these frames, and beam selection is performed using the accurate location of the sound source and interference source for filtering.

Benefits of technology

It improves the filtering quality of voice frames, reduces noise residue, and enhances the accuracy of voice wake-up and recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114283826B_ABST
    Figure CN114283826B_ABST
Patent Text Reader

Abstract

The application discloses a speech signal processing method and a related device, which can accurately indicate the actual position of a sound source, thereby improving the filtering quality of a speech frame. The method comprises the following steps: acquiring a first speech frame; if it is determined that the first speech frame is followed by a speech frame triggering wake-up, determining the direction of a first sound source of a first speech according to N speech frames before the first speech frame, wherein the first speech comprises the first speech frame and the N speech frames, the N speech frames comprise the speech frame triggering wake-up, N is an integer greater than or equal to 1; and filtering the first speech frame according to the direction of the first sound source.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of intelligent voice interaction, and in particular, to a voice signal processing method and a related device thereof. BACKGROUND

[0002] The essence of voice enhancement is voice noise reduction (also referred to as voice filtering). The voice collected by the microphone array of a terminal device usually carries a certain degree of noise, and the noise carried in the voice can be suppressed through voice enhancement, thereby improving the intelligibility and voice quality of the voice, and helping to improve the wake-up rate of voice wake-up and the recognition rate of voice recognition.

[0003] In the process of voice enhancement, each voice frame in the voice can be filtered to obtain an enhanced voice. Specifically, for a voice frame, a plurality of beams can be generated based on the voice frame, and the direction of a sound source generating the voice can be determined based on information of the voice frame. Then, in the plurality of beams, a beam corresponding to the direction of the sound source is selected for filtering to obtain a filtered voice frame. However, in the process of determining the direction of the sound source, the information considered is relatively single, which leads to the direction of the sound source being unable to accurately indicate the actual position of the sound source, and thus the filtered voice frame obtained based on the direction of the sound source is still likely to contain a large amount of noise, i.e., the filtering quality of the voice frame is poor.

[0004] Therefore, how to make the direction of the sound source accurately indicate the actual position of the sound source, thereby improving the filtering quality of the voice frame, has become a problem to be solved. SUMMARY

[0005] Embodiments of the present application provide a voice signal processing method and a related device thereof, which can make the direction of the sound source accurately indicate the actual position of the sound source, thereby improving the filtering quality of the voice frame.

[0006] A first aspect of embodiments of the present application provides a voice signal processing method, which comprises:

[0007] When a user needs to perform voice interaction (e.g., voice wake-up and voice recognition, etc.) with a terminal device, a first voice can be input to the terminal device. Before performing voice wake-up and voice recognition based on the first voice, the terminal device usually performs voice enhancement on the first voice. Specifically, the terminal device can acquire the first voice frame by frame, and perform voice enhancement on each voice frame in the first voice. The process of voice enhancement is as follows:

[0008] Before the first speech frame is acquired, the terminal device has acquired all speech frames before the first speech frame, and has performed wake-up detection on each speech frame. Based on the wake-up detection result of the speech frames, it can be determined whether the speech frames before the first speech frame include a speech frame triggering wake-up (i.e., finally wake up the terminal device), that is, whether the Kth speech frame before the first speech frame is a speech frame triggering wake-up, K being an integer greater than or equal to 1.

[0009] After the first speech frame is acquired, if the terminal device determines that the speech frames before the first speech frame include a speech frame triggering wake-up (i.e., determines that the Kth speech frame before the first speech frame is a speech frame triggering wake-up), it can be determined that the current state is a wake-up state, and then the direction of the first sound source of the first speech is determined according to the N speech frames before the first speech frame. The first speech includes the first speech frame and the N speech frames before the first speech frame. The N speech frames include the speech frame triggering wake-up. Generally, the speech frame triggering wake-up is usually the last speech frame in the N speech frames. N is an integer greater than or equal to 1.

[0010] After the direction of the first sound source is obtained, the terminal device filters the first speech frame according to the direction of the first sound source.

[0011] As can be seen from the above method, after the terminal device determines that the first speech frame includes a speech frame triggering wake-up, the direction of the first sound source is determined using the N speech frames before the first speech frame. Since the N speech frames include the speech frame triggering wake-up, when determining the direction of the first sound source, multiple speech frame information related to waking up the terminal device is considered, so that the direction can more accurately indicate the actual position of the first sound source. Therefore, filtering the first speech frame based on the direction can improve the filtering quality of the speech frame.

[0012] In a possible implementation, the terminal device determining the direction of the first sound source of the first speech according to the N speech frames before the first speech frame specifically includes: after the terminal device determines that the first speech frame includes a speech frame triggering wake-up, the terminal device can determine the N speech frames before the first speech frame, and the N speech frames include the speech frame triggering wake-up. Then, the terminal device obtains N estimated directions of the first sound source corresponding to the N speech frames, wherein each speech frame corresponds to an estimated direction of the first sound source. Finally, the terminal device takes the mode in the N estimated directions of the first sound source to obtain the direction of the first sound source of the first speech.

[0013] In the implementation manner, the N speech frames before the first speech frame contain the speech frame triggering the wake-up, and thus the N speech frames can be considered as the speech frames related to the wake-up terminal device, i.e., the speech frames of the wake-up segment in the first speech frame. The acoustic characteristics of the sound source in the wake-up segment are more obvious, and thus the estimated directions of the N first sound sources corresponding to the N speech frames can highlight the first sound source from the complex environment containing the first sound source and the first interference source. Therefore, the direction of the first sound source obtained based on the part of information can more accurately indicate the actual position of the first sound source.

[0014] In a possible implementation manner, the first speech further includes M speech frames before the N speech frames, M is an integer greater than or equal to 1, and the method further includes: after determining the N speech frames before the first speech frame, the terminal device can determine the M speech frames before the N speech frames. Then, the terminal device obtains the estimated directions of the M first interference sources corresponding to the M speech frames, wherein each speech frame corresponds to an estimated direction of a first interference source. Finally, the terminal device obtains the direction of the first interference source of the first speech by taking the mode of the estimated directions of the M first interference sources.

[0015] In the implementation manner, after determining the N speech frames before the first speech frame, the terminal device can determine the M speech frames before the N speech frames, and the M speech frames can be considered as the speech frames unrelated to the wake-up terminal device, i.e., the speech frames of the non-wake-up segment in the first speech frame. The acoustic characteristics of the interference source in the non-wake-up segment are more obvious, and thus the estimated directions of the M first interference sources corresponding to the M speech frames can highlight the first interference source from the complex environment containing the first sound source and the first interference source. Therefore, the direction of the first interference source obtained based on the part of information can more accurately indicate the actual position of the first interference source.

[0016] In a possible implementation manner, the terminal device filters the first speech frame according to the direction of the first sound source specifically includes: the terminal device first obtains a plurality of first beams corresponding to the first speech frame, and different first beams have different directions. Then, the terminal device determines a first target beam in the plurality of first beams according to the direction of the first sound source, and determines a first interference beam in the plurality of first beams according to the direction of the first interference source. Finally, the terminal device filters the first target beam according to the first interference beam. In the implementation manner, the direction of the first sound source can accurately indicate the actual position of the first sound source, and the direction of the first interference source can accurately indicate the actual position of the first interference source. Therefore, the filtering of the first speech frame based on the first target beam and the first interference beam determined respectively based on the two directions can improve the filtering quality of the first speech frame.

[0017] In a possible implementation, after the terminal device filters the first speech frame according to the direction of the first sound source, the method further includes: the terminal device first acquires a second speech frame, the second speech frame being any speech frame after the first speech frame in the first speech. Then, the terminal device acquires a plurality of second beams corresponding to the second speech frame, different second beams having different directions. Next, if the terminal device determines that the second speech frame does not include a speech frame triggering wake-up before, the terminal device determines a second target beam from the plurality of second beams according to the direction of the first sound source, and determines a second interference beam from the plurality of second beams according to the direction of the first interference source. Finally, the terminal device filters the second target beam according to the second interference beam. In the implementation, since the direction of the first sound source can accurately indicate the actual position of the first sound source, and the direction of the first interference source can accurately indicate the actual position of the first interference source, filtering of the second speech frame by the second target beam and the second interference beam determined based on the two directions can improve the filtering quality of the second speech frame.

[0018] In a possible implementation, for any speech frame in the N speech frames, the estimated direction of the first sound source corresponding to the speech frame is determined according to a difference between a fast-updated cross-correlation energy spectrum and a slow-updated cross-correlation energy spectrum, the fast-updated cross-correlation energy spectrum being determined according to a cross-correlation energy spectrum of the speech frame, a cross-correlation energy spectrum of a speech frame before the speech frame, and a preset first weight, and the slow-updated cross-correlation energy spectrum being determined according to the cross-correlation energy spectrum of the speech frame, the cross-correlation energy spectrum of the speech frame before the speech frame, and a preset second weight, the cross-correlation energy spectrum of the speech frame being determined according to the speech frame. In the implementation, for any speech frame, the estimated direction of the first sound source of the speech frame is determined according to a difference between a fast-updated cross-correlation energy spectrum and a slow-updated cross-correlation energy spectrum, which can make the sound source position information reflected by the estimated direction more accurate.

[0019] In a possible implementation, for any speech frame in the M speech frames, the estimated direction of the first interference source corresponding to the speech frame is determined according to a slow-updated cross-correlation energy spectrum. In the implementation, for any speech frame, the estimated direction of the first interference source of the speech frame is determined according to a slow-updated cross-correlation energy spectrum, which can make the interference source position information reflected by the estimated direction more accurate.

[0020] In a possible implementation, if the first voice is a voice that wakes up the terminal device for the first time, before the terminal device acquires the first voice frame, the method further includes: the terminal device first acquires a third voice frame, the third voice frame being any one voice frame before the first voice frame in the first voice. Then, the terminal device acquires a plurality of third beams corresponding to the third voice frame, different third beams having different orientations. Next, if the terminal device determines that no voice frame triggering wake-up is included before the third voice frame, the terminal device determines a third target beam in the plurality of third beams according to the energy of each third beam, and determines a third interference beam in the plurality of third beams according to the orientation relationship between the third beams and the third target beam. Finally, the terminal device filters the third target beam according to the third interference beam, thereby completing filtering processing of the third voice frame, and making the scheme more comprehensive.

[0021] In a possible implementation, if the first voice is a voice that does not wake up the terminal device for the first time, before the terminal device acquires the first voice frame, the method further includes: the terminal device first acquires a third voice frame, the third voice frame being any one voice frame before the first voice frame in the first voice. Then, the terminal device acquires a plurality of third beams corresponding to the third voice frame, different third beams having different orientations. Next, if the terminal device determines that no voice frame triggering wake-up is included before the third voice frame, the terminal device determines a third target beam in the plurality of third beams according to the energy of each third beam and / or the orientation of a second sound source of a second voice, and determines a third interference beam in the plurality of third beams according to the orientation of a second interference source of the second voice, the second voice being a voice that wakes up the terminal device for the last time. Finally, the terminal device filters the third target beam according to the third interference beam, thereby completing filtering processing of the third voice frame, and making the scheme more comprehensive.

[0022] In a possible implementation, the method further includes: the terminal device performs echo cancellation on the first voice frame, the second voice frame, or the third voice frame.

[0023] In a possible implementation, the filtering includes linear filtering and / or nonlinear filtering, for any one voice frame of the first voice, a gain of linear filtering of the voice frame is determined according to a target beam corresponding to the voice frame, an interference beam corresponding to the voice frame, and a third weight, a gain of nonlinear filtering of the voice frame is determined according to the target beam corresponding to the voice frame, the interference beam corresponding to the voice frame, and a fourth weight, and the third weight or the fourth weight is determined according to a signal-to-noise ratio of the voice frame. In the implementation, for any one voice frame, the gain of linear filtering of the voice frame and the gain of nonlinear filtering of the voice frame are set based on the signal-to-noise ratio of the voice frame, thereby controlling the filtering intensity of the voice frame, and the filtering quality of the voice frame can be further improved.

[0024] In a possible implementation, for any one speech frame of the first speech, a wake-up threshold corresponding to the speech frame is determined according to a preset threshold, a signal-to-noise ratio of the speech frame, and a reference signal average energy of the speech frame, the wake-up threshold corresponding to the speech frame is used to determine whether the speech frame is a speech frame triggering wake-up, the reference signal average energy of the speech frame is determined based on an energy of a reference speech frame corresponding to the speech frame and energies of reference speech frames corresponding to speech frames before the speech frame, and the reference speech frame corresponding to the speech frame is a speech frame output by the terminal device when the speech frame is received. In the implementation, for any one speech frame, the wake-up threshold corresponding to the speech frame is adjusted based on the signal-to-noise ratio of the speech frame and the reference signal average energy of the speech frame, and when wake-up detection is performed on the speech frame based on the adjusted wake-up threshold, the accuracy of the wake-up detection can be improved.

[0025] A second aspect of the embodiment of the application provides a device for processing a speech signal, and the device comprises: an acquisition module configured to acquire a first speech frame. A determination module configured to, if it is determined that a speech frame triggering wake-up is included before the first speech frame, determine a direction of a first sound source of a first speech according to N speech frames before the first speech frame, the first speech comprising the first speech frame and the N speech frames, the N speech frames including the speech frame triggering wake-up, and N being an integer greater than or equal to 1. A filtering module configured to filter the first speech frame according to the direction of the first sound source.

[0026] In a possible implementation, the determination module is specifically configured to: acquire N estimated directions of N first sound sources corresponding to the N speech frames before the first speech frame, wherein each speech frame corresponds to an estimated direction of a first sound source. The direction of the first sound source of the first speech is obtained by taking a mode in the N estimated directions of the N first sound sources.

[0027] In a possible implementation, the determination module is further configured to: acquire M estimated directions of M first interference sources corresponding to the M speech frames, wherein each speech frame corresponds to an estimated direction of a first interference source. The direction of the first interference source of the first speech is obtained by taking a mode in the M estimated directions of the M first interference sources.

[0028] In a possible implementation, the filtering module is specifically configured to: acquire a plurality of first beams corresponding to the first speech frame, different first beams having different directions. The first target beam is determined in the plurality of first beams according to the direction of the first sound source. The first interference beam is determined in the plurality of first beams according to the direction of the first interference source. The first target beam is filtered according to the first interference beam.

[0029] In a possible implementation, the acquisition module is further configured to acquire a second speech frame, the second speech frame being any one of the speech frames after the first speech frame in the first speech. The filtering module is further configured to: acquire a plurality of second beams corresponding to the second speech frame, different second beams having different directions. If it is determined that no speech frame triggering the wake-up is included before the second speech frame, determine a second target beam from the plurality of second beams according to the direction of the first sound source. Determine a second interference beam from the plurality of second beams according to the direction of the first interference source. Filter the second target beam according to the second interference beam.

[0030] In a possible implementation, for any one of the N speech frames, the estimated direction of the first sound source corresponding to the speech frame is determined according to a difference between a fast-updated cross-correlation energy spectrum and a slow-updated cross-correlation energy spectrum, the fast-updated cross-correlation energy spectrum being determined according to a cross-correlation energy spectrum of the speech frame, a cross-correlation energy spectrum of a speech frame before the speech frame, and a preset first weight, and the slow-updated cross-correlation energy spectrum being determined according to the cross-correlation energy spectrum of the speech frame, the cross-correlation energy spectrum of the speech frame before the speech frame, and a preset second weight, the cross-correlation energy spectrum of the speech frame being determined according to the speech frame.

[0031] In a possible implementation, for any one of the M speech frames, the estimated direction of the first interference source corresponding to the speech frame is determined according to a slow-updated cross-correlation energy spectrum.

[0032] In a possible implementation, if the first speech is the speech for the first time to wake up the terminal device, the acquisition module is further configured to acquire a third speech frame, the third speech frame being any one of the speech frames before the first speech frame in the first speech. The filtering module is further configured to: acquire a plurality of third beams corresponding to the third speech frame, different third beams having different directions. If it is determined that no speech frame triggering the wake-up is included before the third speech frame, determine a third target beam from the plurality of third beams according to energy of each third beam. Determine a third interference beam from the plurality of third beams according to a direction relationship between the third beams and the third target beam. Filter the third target beam according to the third interference beam.

[0033] In a possible implementation, if the first voice is a voice that does not wake up the terminal device for the first time, the obtaining module is further configured to obtain a third voice frame, the third voice frame being any one of the voice frames before the first voice frame in the first voice. The filtering module is further configured to: obtain a plurality of third beams corresponding to the third voice frame, different third beams having different orientations; if it is determined that the voice frame that triggers the wake-up is not included before the third voice frame, determine a third target beam from the plurality of third beams according to the energy of each third beam and / or the orientation of a second sound source of a second voice, the second voice being a voice that wakes up the terminal device for the last time; determine a third interference beam from the plurality of third beams according to the orientation of a second interference source of the second voice; and filter the third target beam according to the third interference beam.

[0034] In a possible implementation, the filtering includes linear filtering and / or nonlinear filtering, for any one of the voice frames of the first voice, a gain of linear filtering of the voice frame is determined according to a target beam corresponding to the voice frame, an interference beam corresponding to the voice frame, and a third weight, a gain of nonlinear filtering of the voice frame is determined according to the target beam corresponding to the voice frame, the interference beam corresponding to the voice frame, and a fourth weight, and the third weight or the fourth weight is determined according to a signal-to-noise ratio of the voice frame.

[0035] In a possible implementation, for any one of the voice frames of the first voice, a wake-up threshold corresponding to the voice frame is determined according to a preset threshold, a signal-to-noise ratio of the voice frame, and a reference signal average energy of the voice frame, the wake-up threshold corresponding to the voice frame is used to determine whether the voice frame is a voice frame that triggers the wake-up, and the reference signal average energy of the voice frame is determined based on an energy of a reference voice frame corresponding to the voice frame and energies of reference voice frames corresponding to voice frames before the voice frame, the reference voice frame corresponding to the voice frame being a voice frame output by the terminal device when the voice frame is received.

[0036] A third aspect of the embodiment of the present application provides an electronic device, including a processor and a memory, the processor being configured to invoke program instructions stored in the memory to perform the method described in the first aspect or any possible implementation of the first aspect.

[0037] A fourth aspect of the embodiment of the present application provides a computer-readable storage medium, including instructions, when the instructions run on a computer or a processor, causing the computer or the processor to perform the method described in the first aspect or any possible implementation of the first aspect.

[0038] The fifth aspect of the embodiments of the present application provides a computer program product containing instructions, the computer program product comprising program instructions which, when executed on a computer or processor, cause the computer or processor to perform the method as described in the first aspect or any possible implementation manner of the first aspect.

[0039] From the above technical solutions, the embodiments of the present application have the following advantages:

[0040] In the embodiments of the present application, after the terminal device determines the voice frame triggering the wake-up, the terminal device determines the direction of the first sound source by using the N voice frames before the first voice frame. Since the N voice frames include the voice frame triggering the wake-up, when determining the direction of the first sound source, the terminal device considers the information of multiple voice frames related to the wake-up of the terminal device, so that the direction can more accurately indicate the actual position of the first sound source. Therefore, filtering the first voice frame based on the direction can improve the filtering quality of the voice frame. BRIEF DESCRIPTION OF DRAWINGS

[0041] Figure 1 An example diagram of a far-field voice interaction scene is provided for the embodiments of the present application;

[0042] Figure 2 An example diagram of a terminal device processing a voice signal in a traditional scheme is provided;

[0043] Figure 3 An example diagram of a terminal device is provided for the embodiments of the present application;

[0044] Figure 4 An example diagram of a hardware architecture of a voice signal processing device is provided for the embodiments of the present application;

[0045] Figure 5 Another example diagram of a far-field voice interaction scene is provided for the embodiments of the present application;

[0046] Figure 6 An example diagram of a voice is provided for the embodiments of the present application;

[0047] Figure 7 An example flow diagram of a voice signal processing method is provided for the embodiments of the present application;

[0048] Figure 8 Another example flow diagram of a voice signal processing method is provided for the embodiments of the present application;

[0049] Figure 9 An example diagram of an angle of incidence of a beam is provided for the embodiments of the present application;

[0050] Figure 10 An example diagram of an application example of a voice signal processing method is provided for the embodiments of the present application;

[0051] Figure 11 A schematic diagram of a terminal device processing a voice signal is provided for an embodiment of the present application.

[0052] Figure 12 A structural schematic diagram of an apparatus for processing a voice signal is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0053] Embodiments of the present application provide a voice signal processing method and related device, which can accurately indicate the actual position of a sound source, thereby improving the filtering quality of a voice frame.

[0054] The terms "first", "second", and the like in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the terms used in this way can be interchanged, which is only a distinguishing way used in the description of the embodiments of the present application to describe the objects with the same attributes. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that the processes, methods, systems, products or devices containing a series of units do not necessarily limit to those units, but can include other units not clearly listed or inherent to these processes, methods, products or devices.

[0055] Artificial intelligence (AI) technology is a technology discipline that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence. AI technology obtains the best results by perceiving the environment, acquiring knowledge and using knowledge. In other words, artificial intelligence technology is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Voice interaction using artificial intelligence is a common application of artificial intelligence.

[0056] With the popularity of intelligent terminal devices, voice interaction between man and machine, especially in the far-field voice interaction scenario, has gradually become an important human-computer interaction interface scenario and is considered to be the most important user traffic entrance in the future. Among them, the intelligent terminal device provided with a sound collection component can collect voice signals in the surrounding space and process the voice signals in a predetermined manner to realize voice-based human-computer interaction and other applications, and has a wide application prospect in intelligent voice interaction and other artificial intelligence scenarios.

[0057] According to different specific application scenarios, the intelligent terminal device can have different product forms. For example, the intelligent terminal device can include at least one of a smart sound box, a smart television, a smart television set-top box, a smart robot, and a smart vehicle-mounted device. For example, Figure 1 An example diagram of a far-field voice interaction scenario is provided for an embodiment of the present application. As shown in the diagram, Figure 1 A smart sound box is placed in a room, and a user speaks a voice, such as "Xiao X student, lower the volume", at any position in the room. The voice spoken by the user is transmitted through the air to the smart sound box and can be received by a sound collection component (for example, a microphone array) provided in the smart sound box. After processing the received voice, the smart sound box can be woken up by the voice and recognize the control command contained in the voice, so as to lower the volume based on the control command.

[0058] However, various noises usually exist in the environment where the user is located, so the voice collected by the intelligent terminal device usually contains noise and target voice (i.e., the voice spoken by the user). As shown in the diagram, Figure 2 Figure 2 To suppress the interference of noise, before voice wake-up and voice recognition, the intelligent terminal device can perform voice enhancement on the collected voice to filter noise and extract sufficiently pure target voice, i.e., obtain enhanced voice, so as to improve the wake-up rate of voice wake-up and the recognition rate of voice recognition.

[0059] Specifically, the intelligent terminal device can perform a filtering operation on each voice frame in the collected voice, so as to complete voice enhancement. For any voice frame, the intelligent terminal device can determine the direction of the sound source based on the information of the voice frame, and perform filtering on the voice frame based on the direction of the sound source. However, in the foregoing process of determining the direction of the sound source, the considered information is relatively single, so that the obtained direction of the sound source cannot accurately indicate the actual position of the sound source, and the voice frame filtered based on the direction of the sound source obtained in this way still contains a lot of noise, i.e., the filtering quality of the voice frame is poor. To improve the filtering quality of the voice frame, a method for processing a voice signal is provided in an embodiment of the present application.

[0060] The method for processing a voice signal provided in an embodiment of the present application can be applied to various intelligent terminal devices (hereinafter referred to as terminal devices), and correspondingly, the voice signal processing apparatus provided in an embodiment of the present application can be various forms of terminal products, such as a smart phone, a smart sound box, a smart television, a smart television set-top box, a smart robot, a smart vehicle-mounted device, a tablet computer, smart glasses, a wearable device, a camera, and a video camera, etc. Figure 3 An example structural diagram of a terminal is provided for an embodiment of the present application. As shown in the diagram, Figure 3 ​As shown, the terminal 300 can include an antenna system 310, a radio frequency (RF) circuit 320, a processor 330, a memory 340, a camera 350, an audio circuit 360, a display screen 370, one or more sensors 380, and a wireless transceiver 390, etc.

[0061] The antenna system 310 can be one or more antennas, and can also be an antenna array composed of multiple antennas. The radio frequency circuit 320 can include one or more analog radio frequency transceivers, and can also include one or more digital radio frequency transceivers, which are coupled to the antenna system 310. It should be understood that in various embodiments of the present application, coupling means mutual connection in a specific way, including direct connection or indirect connection through other devices, for example, connection through various interfaces, transmission lines, buses, etc. The radio frequency circuit 320 can be used for various types of cellular wireless communication.

[0062] The processor 330 can include a communication processor, which can be used to control the RF circuit 320 to realize the reception and transmission of signals through the antenna system 310, which can be voice signals, media signals or control signals. The processor 330 can include various general-purpose processing devices, for example, can be a general-purpose central processing unit (CPU), a system on chip (SOC), a processor integrated on an SOC, a separate processor chip or a controller, etc.; the processor 330 can also include special-purpose processing devices, for example, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or digital signal processors (DSPs), special-purpose video or graphics processors, graphics processing units (GPUs), and neural network processing units (NPUs), etc. The processor 330 can be a processor group composed of multiple processors, which are coupled to each other through one or more buses. The processor can include analog-to-digital converters (ADCs) and digital-to-analog converters (DACs) to realize the connection of signals between different components of the device. The processor 330 is used to realize the processing of image, audio and video media signals.

[0063] The memory 340 is coupled to the processor 330. Specifically, the memory 340 can be coupled to the processor 330 via one or more memory controllers. The memory 340 can be used for storing computer program instructions, including a computer operating system (OS) and various user application programs, and can also be used for storing user data, such as calendar information, contact information, acquired image information, audio information, or other media files, etc. The processor 330 can read computer program instructions or user data from the memory 340, or store computer program instructions or user data into the memory 340, to implement relevant processing functions. The memory 340 can be a non-volatile memory, such as an EMMC (Embedded Multi Media Card), a UFS (Universal Flash Storage), or a read-only memory (ROM), or other types of static storage devices that can store static information and instructions, and can also be a volatile memory, such as a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, and can also be an EEPROM (Electrically Erasable Programmable Read-Only Memory), a CD-ROM (Compact Disc Read-Only Memory), or other optical disc storage, optical disc storage (including a compact disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium, or other magnetic storage devices, or any other computer-readable storage medium capable of carrying or storing program code in the form of instructions or data structures and capable of being accessed by a computer, but not limited to the above. The memory 340 can exist independently, or the memory 340 can be integrated with the processor 330.

[0064] The camera 350 is configured to capture images or videos, and can be triggered to start by application instructions to realize the functions of taking pictures or videos, such as capturing pictures or videos of any scene. The camera can include components such as an imaging lens, a filter, an image sensor, and the like. Light emitted or reflected by an object enters the imaging lens, passes through the filter, and finally converges on the image sensor. The imaging lens is mainly used to converge the light emitted or reflected by all objects (which can also be referred to as a scene to be photographed, a target scene, or a scene image that a user expects to capture) in the photographing angle of view into an image; the filter is mainly used to filter out unnecessary light waves (such as infrared light waves in addition to visible light) in the light; and the image sensor is mainly used to perform photoelectric conversion on the received light signals, convert them into electrical signals, and input them to the processor 330 for subsequent processing. The camera can be located on the front of the terminal device, or on the back of the terminal device, and the specific number and arrangement of the camera can be flexibly determined according to the needs of designer or vendor strategies, which are not limited in the present application.

[0065] The audio circuit 360 is coupled with the processor 330. The audio circuit 360 can include a microphone 361 and a speaker 362. The microphone 361 can receive sound input from the outside world, and the speaker 362 can realize the playing of audio data. It should be understood that the terminal 300 can have one or more microphones and one or more earphones, and the number of microphones and earphones is not limited in the embodiments of the present application. It should be noted that the microphone 361 can be used to collect voice signals emitted by a user, and send the collected voice signals to the processor 330, so that the processor 330 processes the collected voice signals, for example, voice enhancement, voice wake-up, voice recognition, and the like. Furthermore, the speaker 362 can receive audio signals from the processor 330, and output the audio signals to the outside for use by the user.

[0066] The display screen 370 is configured to display information input by a user, provide various menus of information to a user, which are associated with specific modules or functions within the terminal 300, and accept user input, such as enabling or disabling control information. In particular, the display screen 370 can include a display panel 371 and a touch panel 372. The display panel 371 can be configured using a liquid crystal display (LCD), an organic light-emitting diode (OLED), a light emitting diode (LED) display device, a cathode ray tube (CRT), or the like. The touch panel 372, also referred to as a touch screen, a touch-sensitive screen, or the like, can collect a contact or non-contact operation (such as an operation by a user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 372, which can also include a somatosensory operation; the operation includes a single-point control operation, a multi-point control operation, or the like) of a user on or near the touch panel 372, and drive a corresponding connection device according to a pre-set program. Optionally, the touch panel 372 can include a touch detection device and a touch controller. The touch detection device detects a signal resulting from a touch operation of a user, and transmits the signal to the touch controller. The touch controller receives touch information from the touch detection device, converts the touch information into information that can be processed by the processor 330, and transmits the information to the processor 330. The touch controller can also receive a command from the processor 330 and execute the command. Further, the touch panel 372 can cover the display panel 371. A user can perform an operation on or near the touch panel 372 that covers the display panel 371 based on content displayed on the display panel 371, such as a soft keyboard, a virtual mouse, a virtual key, an icon, or the like. After the touch panel 372 detects an operation on or near the touch panel 372, the touch panel 372 transmits the operation to the processor 330 via the I / O subsystem 30 to determine a user input. Subsequently, the processor 330 provides a corresponding visual output on the display panel 371 via the I / O subsystem 30 based on the user input. Although in the above description, the touch panel 372 and the display panel 371 are implemented as two separate components to perform input and output functions of the terminal 300, in some embodiments, the touch panel 372 and the display panel 371 can be integrated to perform input and output functions of the terminal 300. Figure 3

[0067] ​The sensor 380 can include an image sensor, a motion sensor, a proximity sensor, an ambient noise sensor, a sound sensor, an accelerometer, a temperature sensor, a gyroscope, or other types of sensors, and combinations of various forms thereof. The processor 330 drives the sensor 380 to receive various information such as audio signals, image signals, motion information, etc. through the sensor controller 32 in the I / O subsystem 30, and the sensor 380 transmits the received information to the processor 330 for processing.

[0068] The wireless transceiver 390 can provide wireless connectivity to other devices, which can be peripheral devices such as wireless headsets, Bluetooth headsets, wireless mice, wireless keyboards, etc., or wireless networks such as Wireless Fidelity (WiFi) networks, Wireless Personal Area Networks (WPANs), or other Wireless Local Area Networks (WLANs), etc. The wireless transceiver 390 can be a Bluetooth-compatible transceiver for wirelessly coupling the processor 330 to a Bluetooth headset, a wireless mouse, or other peripheral devices, and can also be a WiFi-compatible transceiver for wirelessly coupling the processor 330 to a wireless network or other devices.

[0069] The terminal 300 can further include other input devices 34 coupled to the processor 330 to receive various user inputs such as inputted numbers, names, addresses, and media selections, etc. The other input devices 34 can include a keyboard, physical buttons (push buttons, rocker buttons, etc.), a dial, a slide switch, a joystick, a click wheel, an optical mouse (an optical mouse is a touch-sensitive surface that does not display visual output, or is an extension of a touch-sensitive surface formed by a touch screen), etc.

[0070] The terminal 300 can further include the I / O subsystem 30 described above, which can include other input device controllers 31 for receiving signals from or sending control or driving information of the processor 330 to the other input devices 34, and can further include the sensor controller 32 and the display controller 33 described above for enabling exchange of data and control information between the sensor 380 and the display 370 and the processor 330, respectively.

[0071] The terminal 300 can also include a power supply 301 for powering the other components of the terminal 300 including 310-390, which can be a rechargeable lithium ion battery or a nickel-metal hydride battery. Further, when the power supply 301 is a rechargeable battery, it can be coupled with the processor 330 through a power management system, such that the power management system manages charge, discharge, and power consumption adjustments, etc.

[0072] It should be appreciated that Figure 3 The terminal 300 in the above is only an example, and the specific form of the terminal 300 is not limited, and the terminal 300 can also include other components that are not shown in the above Figure 3 The terminal 300 in the above is only an example, and the specific form of the terminal 300 is not limited, and the terminal 300 can also include other components that are not shown in the above

[0073] In an optional solution, the RF circuit 320, the processor 330 and the memory 340 can be partially or entirely integrated on one chip, or can be three independent chips. The RF circuit 320, the processor 330 and the memory 340 can include one or more integrated circuits arranged on a printed circuit board (PCB).

[0074] Figure 4 A hardware architecture diagram of the voice signal processing device provided by the embodiment of the present application is shown in FIG. 4, which can be a processor chip, for example, and an exemplary hardware architecture diagram of the processor 330 in the above is shown in FIG. 5. Figure 4 The hardware architecture diagram shown in FIG. 4 can be Figure 3 The voice signal processing method provided by the embodiment of the present application can be applied to the processor chip.

[0075] Referring to Figure 4 The device 400 includes at least one CPU, a memory, a microcontroller unit (MCU), a GPU, an NPU, a memory bus, a receiving interface and a sending interface, etc. Although Figure 4 The device 400 can also include an application processor (AP), a decoder and a dedicated video or image processor, which are not shown in the above.

[0076] The above various parts of the device 400 are coupled through connectors, which include various interfaces, transmission lines or buses, etc. These interfaces are usually electrical communication interfaces, but can also be mechanical interfaces or other forms of interfaces, which are not limited in the embodiment.

[0077] Optionally, the CPU can be a single-CPU processor or a multi-CPU processor; optionally, the CPU can be a processor group composed of multiple processors, and the multiple processors are coupled with each other through one or more buses. The receiving interface can be an interface for data input of the processor chip, and in an optional case, the receiving interface and the sending interface can be a High Definition Multimedia Interface (HDMI), a V-By-One interface, an Embedded Display Port (eDP), a Mobile Industry Processor Interface (MIPI), or a Display Port (DP), etc. The memory can refer to the foregoing description of the memory 340.

[0078] In an optional case, the above-mentioned parts are integrated on the same chip; in another optional case, the CPU, the GPU, the decoder, the receiving interface, and the sending interface are integrated on a chip, and the parts inside the chip access the external memory through a bus. The dedicated video / graphics processor can be integrated on the same chip as the CPU, or can exist as a separate processor chip, for example, the dedicated video / graphics processor can be a dedicated ISP. In an optional case, the NPU can also exist as an independent processor chip. The NPU is used to implement various neural network or deep learning related operations. Optionally, the image processing method and the image processing framework provided in the embodiments of the present application can be implemented by the GPU or the NPU, or can be implemented by a dedicated graphics processor.

[0079] In the embodiments of the present application, the chip is a system manufactured on the same semiconductor substrate by integrated circuit technology, also called a semiconductor chip, which can be a collection of integrated circuits formed on a substrate (usually a semiconductor material such as silicon) by integrated circuit technology, and the outer layer is usually packaged by a semiconductor packaging material. The integrated circuit can include various functional devices, each of which includes a logic gate circuit, a Metal-Oxide-Semiconductor (MOS) transistor, a bipolar transistor, or a diode transistor, and can also include a capacitor, a resistor, or an inductor and other components. Each functional device can work independently or under the action of necessary driving software, and can realize various functions such as communication, calculation, or storage.

[0080] The above is a specific description of the terminal device provided by the embodiments of the present application, and the following describes the image processing method provided by the embodiments of the present application in combination with Figure 5 and Figure 6The speech signal processing method provided in the embodiments of this application will be briefly introduced. Figure 5 This is another schematic diagram of a far-field voice interaction scenario provided in an embodiment of this application. Figure 6 A schematic diagram of the voice provided in an embodiment of this application.

[0081] like Figure 5 As shown, in a far-field voice interaction scenario, the terminal device can intermittently receive multiple voice messages from the user. For example, after the terminal device starts up, at a certain time point T1, it receives the user's voice message "Hi, Little X, play a song by Y," and at the next time point T2, it receives the user's voice message "Hi, Little X, turn down the volume," and so on. These voice messages all contain wake-up words and command words. Therefore, after the terminal device performs wake-up detection and recognition on these voice messages, it can be woken up by these voice messages and recognize the control commands in the voice messages, thereby executing the control commands. Before performing wake-up detection and recognition on these voice messages, the terminal device can first perform voice enhancement on each voice message, and then perform wake-up detection and recognition based on the enhanced voice messages. For ease of explanation, in this part of the voice messages, the voice message that wakes up the terminal device the current time can be called the first voice message, and the voice message that woke up the terminal device the previous time can be called the second voice message. That is, the first voice message and the second voice message are two adjacent voice messages that wake up the terminal device.

[0082] In this part of the speech, each speech can be divided into three consecutive speech segments: the speech segment corresponding to the non-wake word (hereinafter referred to as the non-wake word segment), the speech segment corresponding to the wake word (hereinafter referred to as the wake-up segment), and the speech segment corresponding to the command word (hereinafter referred to as the command segment). Each speech segment contains at least one speech frame. For example... Figure 6 As shown, suppose the terminal device receives a second voice message and a first voice message sequentially. The second voice message is "Hi, Little X, play a song by Y." In this case, "Hi" is the non-wake-up segment, "Little X" is the wake-up segment, and "play a song by Y" is the command segment. The first voice message is "Hi, Little X, turn down the volume." In this case, "Hi" is the non-wake-up segment, "Little X" is the wake-up segment, and "turn down the volume" is the command segment. For ease of diagramming, Figure 6 Only the first voice is shown; the second voice is not shown.

[0083] It should be noted that the terminal device can collect the first voice through the microphone array, and divide the voice data of the first voice into multiple voice frames for processing in the process of collecting the first voice. Each voice frame contains voice data of a certain time length, which can be set according to actual needs, and is not limited here. For example, after the user starts to utter the first voice, the terminal device can also start to collect the first voice. When the terminal device collects 10 ms of voice data, it can take the 10 ms of voice data as the first voice frame of the first voice for subsequent processing. When the terminal device continues to collect 10 ms of voice data, it can take the 10 ms of voice data as the second voice frame of the first voice for subsequent processing, and so on, until the collection of the first voice is completed. It should be noted that since the terminal device acquires the first voice frame by frame, the multiple voice frames of the first voice have a certain chronological order. For example, it is assumed that the terminal device acquires twenty voice frames contained in the first voice in turn. Among them, for the tenth voice frame, the first voice frame to the ninth voice frame is the voice frame before the tenth voice frame, and the eleventh voice frame to the twentieth voice frame is the voice frame after the tenth voice frame, and the subsequent will not be described in detail.

[0084] After receiving the first voice, the terminal device can perform voice enhancement (i.e., filtering operation) on each voice frame of the first voice, then perform wake-up detection on each filtered voice frame to obtain a wake-up detection score of each voice frame, and then compare the wake-up detection score with a wake-up threshold to determine which voice frame triggers wake-up (i.e., finally wakes up the terminal device). Specifically, in the first voice, the wake-up detection scores of the non-wake-up segments and the command segments are usually lower than the wake-up threshold, and the wake-up detection scores of one or more voice frames located at the end of the wake-up segment are usually greater than or equal to the wake-up threshold. Generally, if the terminal device determines that only the wake-up detection score of one voice frame in the wake-up segment is greater than the wake-up threshold, the terminal device determines that the voice frame triggers wake-up. If the terminal device determines that the wake-up detection scores of multiple voice frames in the wake-up segment are all greater than the wake-up threshold, the terminal device can select one of the voice frames as the voice frame that triggers wake-up. For example, the terminal device can select the first voice frame whose wake-up detection score is greater than or equal to the wake-up threshold as the voice frame that triggers wake-up, and for another example, the terminal device can select the second voice frame whose wake-up detection score is greater than or equal to the wake-up threshold as the voice frame that triggers wake-up, and so on, which is not limited here. For ease of description, the first voice frame whose wake-up detection score is greater than or equal to the wake-up threshold is taken as an example of the voice frame that triggers wake-up in the following description.

[0085] After determining the speech frame triggering the wake-up in the first speech, the terminal device can determine the direction of the first sound source of the first speech. It should be noted that the terminal device can determine the direction of the first sound source in the filtering process of a certain specific speech frame, so as to use the direction of the first sound source to perform more effective filtering on the speech frame and subsequent speech frames. For ease of introduction, the speech frame will be referred to as the first speech frame hereinafter. Generally, the first speech frame is usually located after the speech frame triggering the wake-up, that is, the speech frame triggering the wake-up can be set as the Kth speech frame before the first speech frame, K being an integer greater than or equal to 1. For example, when K = 1, the first speech frame and the speech frame triggering the wake-up are two adjacent speech frames, that is, the speech frame triggering the wake-up is the 1st speech frame before the first speech frame. For example, when K = 2, the speech frame triggering the wake-up is the 2nd speech frame before the first speech frame, and so on, which is not specifically limited here. For ease of illustration, the following will be described by way of example with K = 1 (as shown in FIG. 8). Figure 6

[0086] For ease of introduction, in the first speech, any speech frame after the first speech frame is referred to as a second speech frame, and any speech frame before the first speech frame is referred to as a third speech frame. In order to complete the filtering of the first speech, the first speech frame, the second speech frame and the third speech frame all need to be filtered, and the filtering processes of the first speech frame, the second speech frame and the third speech frame are all different. The foregoing filtering process will be introduced in detail below. Figure 7 Figure 7 A flowchart of a method for processing a speech signal provided by an embodiment of the present application is shown in FIG. 8. As shown in FIG. 8, the method comprises the following steps. Figure 7

[0087] 701, obtaining a first speech frame.

[0088] In this embodiment, the terminal device can obtain the first speech frame by frame. Before obtaining the first speech frame, the terminal device has obtained all speech frames before the first speech frame, and has performed wake-up detection on each speech frame. Based on the wake-up detection result of the speech frame, it can be determined whether the speech frame before the first speech frame includes a speech frame triggering the wake-up (i.e., finally wakes up the terminal device), that is, whether the speech frame before the first speech frame is the speech frame triggering the wake-up.

[0089] 702, if it is determined that the speech frame triggering the wake-up is included before the first speech frame, determining the direction of the first sound source of the first speech according to N speech frames before the first speech frame, the first speech comprising the first speech frame and the N speech frames, and the N speech frames including the speech frame triggering the wake-up.

[0090] ​​​If the terminal device determines that there is a voice frame triggering wake-up before the first voice frame, i.e., determines that the voice frame before the first voice frame is the voice frame triggering wake-up, it is determined that the terminal device is currently in a woken-up state, and the direction of the first sound source of the first voice is determined according to the N voice frames before the first voice frame. The first voice includes the first voice frame and the N voice frames before the first voice frame. The N voice frames include the voice frame triggering wake-up, and generally, the voice frame triggering wake-up is usually the last voice frame in the N voice frames. N is an integer greater than or equal to 1.

[0091] 703. Filter the first voice frame according to the direction of the first sound source.

[0092] After obtaining the direction of the first sound source, the terminal device filters the first voice frame according to the direction of the first sound source.

[0093] In this embodiment, after the terminal device determines that the first voice frame includes the voice frame triggering wake-up, the terminal device determines the direction of the first sound source by using the N voice frames before the first voice frame. Since the N voice frames include the voice frame triggering wake-up, when determining the direction of the first sound source, multiple voice frame information related to waking up the terminal device is considered, so that the direction can more accurately indicate the actual position of the first sound source. Therefore, filtering the first voice frame based on the direction can improve the filtering quality of the voice frame.

[0094] In order to further understand the filtering process in this application, the following will be introduced in combination with Figure 8 the foregoing filtering process, Figure 8 another flowchart of the method of processing a voice signal provided by an embodiment of the present application. As shown in the figure, Figure 8 the method includes:

[0095] 801. Obtain a third voice frame, which is any one of the voice frames before the first voice frame in the first voice.

[0096] In this embodiment, the terminal device can obtain the first voice issued by the user frame by frame. It is worth noting that the terminal device can calculate the direction of the first sound source of the first voice in the filtering process of the next voice frame (i.e., the first voice frame) of the voice frame triggering wake-up only after determining the voice frame triggering wake-up. Since there are usually multiple third voice frames before the first voice frame, the terminal device can perform filtering processing on each third voice frame as in steps 802 to 804 until the voice frame triggering wake-up is determined.

[0097] 802. Obtain a plurality of third beams corresponding to the third voice frame.

[0098] After acquiring a third voice frame, the terminal device can generate multiple third beams corresponding to that third voice frame using beamforming algorithms (e.g., super-directional beamforming, conventional beamforming, minimum variance distortionless response (MVDR) beamforming, etc.). Each of these third beams has a specific angle of incidence, and different third beams have different angles of incidence (azimuth). The following section combines... Figure 9 The incident angle of the beam will be introduced. Figure 9 This is a schematic diagram of the incident angle of the beam provided in an embodiment of this application. Figure 9 As shown, suppose four third beams (third beam A, third beam B, third beam C and third beam D) are generated based on a certain third voice frame, and the coverage of each beam may be the same or different. Therefore, relative to the microphone array of the terminal device, the incident angle of third beam A is 0°, the incident angle of third beam B is 90°, the incident angle of third beam C is 180° and the incident angle of third beam D is 270°.

[0099] It should be understood that Figure 8 The example of four beams is provided for illustration only and does not constitute a limitation on the number of beams in this application.

[0100] Furthermore, after acquiring a third speech frame, the terminal device can also determine the average energy of the reference signal corresponding to that third speech frame. Similarly, after executing steps 805 and 811, the terminal device can also determine the average energy of the reference signals corresponding to the first and second speech frames, which will not be elaborated further. Specifically, for any speech frame in the first speech, the average energy of the reference signal of that speech frame can be determined by the following formula:

[0101]

[0102] In the above formula, refeng is the average energy of the reference signal corresponding to the speech frame, np is the frame number of the speech frame, n1 is the frame number of the speech frame at the starting point of the statistics (e.g., the frame number of the first speech frame of the first speech), and ref is the energy of the reference speech frame corresponding to each speech frame during the statistical process. The reference speech frame corresponding to this speech frame is the speech frame output by the terminal device when it receives this speech frame. It should be noted that while the terminal device collects the first speech through the microphone array, it also emits sound (i.e., the reference signal) to the outside world through the speaker. Therefore, there is a one-to-one correspondence between the multiple speech frames of the first speech and the multiple reference speech frames of the reference signal.

[0103] Further, after obtaining the third voice frame, the terminal device can further determine the estimated position of the first sound source corresponding to the third voice frame and the estimated position of the first interference source. It should be noted that the first sound source and the first interference source jointly generate the first voice, for example, the first sound source can be a user, and the first interference source can be a noise source around the user, etc. Similarly, after performing steps 805 and 811, the terminal device can also determine the estimated position of the first sound source corresponding to the first voice frame, the estimated position of the first interference source, and the estimated position of the first sound source corresponding to the second voice frame, the estimated position of the first interference source, and the subsequent description. Specifically, for any voice frame in the first voice, the estimated position of the first sound source corresponding to the voice frame and the estimated position of the first interference source can be determined by the following process:

[0104] For the voice frame, it can be represented as z1, z2,..., z K , K is the number of microphones of the microphone array (i.e., the number of channels of the microphone array). For any two channel data z m (t) and z l (t), the corresponding cross-correlation function can be determined by the following formula:

[0105]

[0106] In the above formula, Z m (ω) and Z l (ω) are the frequency domain representations of z m (t) and z l (t), respectively. τ represents the time difference between z m (t) and z l (t).

[0107] Further, the corresponding relationship between the incident angle θ of the voice frame and τ is:

[0108]

[0109] In the above formula, θ is the angle between the voice frame and the line connecting the microphone pair, d is the distance between the microphone pair, and c is the sound speed.

[0110] Therefore, which can be recorded as a function of θ By forming multiple microphone pairs from each pair of channels in the microphone array and further fusing, the position of the sound source can be more accurately and stably estimated. Therefore, the cross-correlation energy spectrum of the voice frame can be determined by the following formula:

[0111]

[0112] Further, the cross-correlation energy spectrum is a function of time t, and a fast update is performed on the cross-correlation energy spectrum of the speech frame to obtain a fast-updated cross-correlation energy spectrum of the speech frame:

[0113]

[0114] ......

[0116]

[0117]

[0118] Similarly, a slow update can also be performed on the cross-correlation energy spectrum of the speech frame to obtain a slow-updated cross-correlation energy spectrum of the speech frame:

[0119]

[0120] ......

[0122]

[0123]

[0124] In the above formulae, is the fast-updated cross-correlation energy spectrum of the speech frame, is the slow-updated cross-correlation energy spectrum of the speech frame, and α and β are preset first weights, and λ and μ are preset second weights. Generally, β is greater than μ.

[0125] Therefore, the estimated position of the first sound source and the estimated position of the first interference source corresponding to the speech frame can be determined by the following formulae:

[0126]

[0127]

[0128]

[0129] In the above formulae, is the estimated position of the first sound source corresponding to the speech frame, is the estimated position of the first interference source corresponding to the speech frame.

[0130] 803、If it is determined that the Kth speech frame before the third speech frame is a non-wakeup-triggering speech frame, a third target beam and a third interference beam are determined based on the target parameter in multiple third beams.

[0131] After obtaining the third voice frame, the terminal device can determine whether the previous voice frame of the third voice frame is the voice frame triggering the wake-up. If it is determined that the previous voice frame of the third voice frame is the voice frame not triggering the wake-up (that is, it is determined that the voice frame triggering the wake-up is not included before the third voice frame), it is indicated that the direction of the first sound source of the first voice and the direction of the first interference source do not need to be calculated in the filtering process of the third voice frame. For example, it is assumed that the 10th voice frame of the first voice is the voice frame triggering the wake-up. When the terminal device obtains the 5th voice frame, it is determined whether the 4th voice frame is the voice frame triggering the wake-up. Since the 4th voice frame is the voice frame not triggering the wake-up, the direction of the first sound source of the first voice does not need to be calculated in the filtering process of the 5th voice frame. Similarly, the direction of the first sound source of the first voice does not need to be calculated in the filtering process of the 1st voice frame to the 10th voice frame.

[0132] After determining that the previous voice frame of the third voice frame is the voice frame not triggering the wake-up, the terminal device can determine the third target beam and the third interference beam in the third beams corresponding to the third voice frame based on the target parameter. The terminal device can determine the third target beam and the third interference beam in various ways, which will be introduced as follows:

[0133] (1) If it is determined that the first voice is the voice triggering the wake-up of the terminal device for the first time (that is, the voice triggering the wake-up of the terminal device for the first time after the terminal device is started), the terminal device determines the third target beam in the third beams according to the energy of each third beam, and determines the third interference beam in the third beams according to the direction relationship (incident angle relationship) between the third beams and the third target beam. For example, after the terminal device obtains the third beams corresponding to a certain third voice frame, the terminal device can determine the energy of each third beam, and determine the third beam with the maximum energy as the third target beam. Then, the beam farthest from the third target beam is determined as the third interference beam (for example, the incident angle of the third target beam is 0°, and the incident angle of the third interference beam is 180°).

[0134] (2) If it is determined that the first voice is a voice that is not the first to wake up the terminal device, the terminal device first determines a third target beam from the plurality of third beams according to the energy of each third beam and / or the direction of the second sound source of the second voice (i.e., the angle of incidence of the second sound source relative to the microphone array), where the second voice is the voice that woke up the terminal device the last time the first voice woke up the terminal device. Then, a third interference beam is determined from the plurality of third beams according to the direction of the second interference source of the second voice (i.e., the angle of incidence of the second interference source relative to the microphone array). For example, after obtaining the plurality of third beams corresponding to a certain third voice frame, the terminal device can determine the energy of each third beam, and determine the third target beam as the third beam with the largest energy. Then, the direction of the second interference source of the second voice is obtained, and the third beam closest to the second interference source is determined as the third interference beam (for example, the direction of the second interference source is 85°, and the angle of incidence of the third interference beam is 90°). For another example, after obtaining the plurality of third beams corresponding to a certain third voice frame, the terminal device can obtain the direction of the second sound source of the second voice, and determine the third target beam as the third beam closest to the second sound source. Then, the direction of the second interference source of the second voice is obtained, and the third beam closest to the second interference source is determined as the third interference beam. For another example, after obtaining the plurality of third beams corresponding to a certain third voice frame, the terminal device can determine the energy of each third beam, and if the energy of a certain third beam is much larger than the energy of the other third beams, the third beam is determined as the third target beam, and if the energy of all third beams is not much different, the direction of the second sound source of the second voice is obtained, and the third beam closest to the second sound source is determined as the third target beam. Then, the direction of the second interference source of the second voice is obtained, and the third beam closest to the second interference source is determined as the third interference beam.

[0135] Further, for any voice frame in the first voice, while performing beam selection (i.e., determining the target beam and the interference beam corresponding to the voice frame), the signal-to-noise ratio of the voice frame can also be determined. Specifically, in the first voice, the signal-to-noise ratios of the first voice frame and the second voice frame can be determined by the following formula:

[0136]

[0137] In the above formula, N2 is the frame number of the last voice frame in the N voice in step 806, N1 is the frame number of the first voice frame in the N voice, z is the energy of the voice frame, M2 is the frame number of the last voice frame in the M voice in step 807, and M1 is the frame number of the first voice frame in the M voice.

[0138] If the first voice is a voice that is not the first voice to wake up the terminal device, the signal-to-noise ratio of the third voice frame is the signal-to-noise ratio determined based on the above formula in the second voice. If the first voice is the first voice to wake up the terminal device, the signal-to-noise ratio of the third voice frame is a preset value.

[0139] Further, for any voice frame in the first voice, after obtaining the signal-to-noise ratio of the voice frame, the terminal device can adjust the wake-up threshold based on the signal-to-noise ratio of the voice frame and the average energy of the reference signal, so that the adjusted wake-up threshold can more accurately wake up detection of the voice frame, so as to determine whether the voice frame is a voice frame that triggers wake-up. Specifically, the wake-up threshold corresponding to the voice frame can be determined by the following formula:

[0140] Thrnew = Thr + ζ + ε

[0141]

[0142]

[0143] In the above formula, Thr is a preset threshold, and Thrnew is a wake-up threshold.

[0144] 804, filtering the third target beam according to the third interference beam.

[0145] After obtaining the third target beam and the third interference beam corresponding to a certain third voice frame, the third target beam is filtered according to the third interference beam, so as to complete the filtering processing of the third voice frame.

[0146] In addition, the terminal device can also complete the filtering processing of the first voice frame and the second voice frame after executing step 810 and step 814. In this embodiment, the filtering processing includes linear filtering processing and / or nonlinear filtering processing. For any voice frame in the first voice, the gain of linear filtering of the voice frame can be determined by the following formula:

[0147] G kalman ``= ηG kalman

[0148] In the above formula, G kalman is the original gain of linear filtering (determined based on the target beam and the interference beam), η is a third weight, and G kalman ``is the finally determined gain of linear filtering.

[0149] Similarly, the gain of nonlinear filtering of the voice frame can be determined by the following formula:

[0150]

[0151] In the above formula, G nlp`` is the original gain of the non-linear filtering (determined based on the target beam and the interference beam), G is the fourth weight, nlp `` is the final gain of the non-linear filtering.

[0152] In addition, the third weight and the fourth weight corresponding to the speech frame are determined based on a signal-to-noise ratio corresponding to the speech frame:

[0153]

[0154]

[0155] Further, the terminal device can perform wake-up detection on each filtered third speech frame, thereby obtaining a wake-up detection score of each third speech frame. Once it is determined that the wake-up detection score of a certain third speech frame is greater than or equal to a wake-up threshold, the terminal device determines the third speech frame as the speech frame triggering wake-up (i.e., the last third speech frame).

[0156] Still as in the above example, after the filtering processing on the 1st speech frame, wake-up detection can be performed thereon. After it is determined that the wake-up detection score of the 1st speech frame is less than the wake-up threshold, the filtering processing is continued on the 2nd speech frame, and wake-up detection is performed on the 2nd speech frame. Until it is determined that the wake-up detection score of the 10th speech frame is greater than or equal to the wake-up threshold (i.e., the wake-up detection scores of the 1st speech frame to the 9th speech frame are all less than the wake-up threshold), the 10th speech frame can be determined as the speech frame triggering wake-up.

[0157] 805, obtain a first speech frame.

[0158] After the speech frame triggering wake-up is determined, the terminal device can obtain a next speech frame of the speech frame triggering wake-up, i.e., a first speech frame.

[0159] 806, if it is determined that the Kth speech frame before the first speech frame is the speech frame triggering wake-up, obtain estimated directions of N first sound sources corresponding to N speech frames before the first speech frame, and take a mode in the estimated directions of the N first sound sources to obtain a direction of a first sound source of the first speech, wherein one speech frame corresponds to one estimated direction of a first sound source.

[0160] After the terminal device acquires the first voice frame, it determines whether the voice frame before the first voice frame is the voice frame triggering the wake-up. After determining that the voice frame before the first voice frame is the voice frame triggering the wake-up, it is indicated that the direction of the first sound source of the first voice and the direction of the first interference source of the first voice need to be calculated in the filtering process of the first voice frame. Specifically, before acquiring the first voice frame, the estimated direction of the first sound source corresponding to each third voice frame has been calculated and stored in the terminal device (for details, refer to the related description of step 802). Therefore, the terminal device takes the voice frame triggering the wake-up as the reference point, and takes N voice frames forward, thereby obtaining N voice frames before the first voice frame. Then, the terminal device directly acquires the estimated direction of the N first sound sources corresponding to the N voice frames, and takes the mode in the estimated direction of the N first sound sources, thereby obtaining the direction of the first sound source of the first voice.

[0161] It is worth noting that the N voice frames can represent the wake-up information of the terminal device. Specifically, the last voice frame of the N voice frames is the voice frame triggering the wake-up. Since the voice frame triggering the wake-up corresponds to the time when the wake-up ends, the terminal device can estimate the wake-up duration based on the time when the wake-up ends by using the traditional wake-up duration calculation method, and then calculate the time when the wake-up starts by using the wake-up duration and the time when the wake-up ends. Since the time when the wake-up starts corresponds to the first voice frame in the N voice frames, the value of N can be determined.

[0162] Still as in the above example, it is assumed that N=5. After the terminal device acquires the 11th voice frame, it can determine whether the 10th voice frame is the voice frame triggering the wake-up. After determining that the 10th voice frame is the voice frame triggering the wake-up, the estimated direction of the first sound source corresponding to the voice frames from the 6th voice frame to the 10th voice frame is acquired, and then the mode in the estimated direction of the five first sound sources is taken, thereby finally determining the direction of the first sound source of the first voice.

[0163] 807、Acquire the estimated direction of the M first interference sources corresponding to the M voice frames before the N voice frames, and take the mode in the estimated direction of the M first interference sources, thereby obtaining the direction of the first interference source of the first voice, wherein one voice frame corresponds to one estimated direction of the first interference source.

[0164] After determining the direction of the first sound source of the first voice, the terminal device can also determine the direction of the first interference source of the first voice. Specifically, the terminal device can first determine the M voice frames before the N voice frames. Then, the terminal device can directly acquire the estimated direction of the M first interference sources corresponding to the M voice frames, and take the mode in the estimated direction of the M first interference sources, thereby obtaining the direction of the first interference source of the first voice.

[0165] Still as the above example, assuming M=4. After the terminal device determines the direction of the first sound source based on the 6th voice frame to the 10th voice frame, it can obtain the estimated direction of the first interference source corresponding to the 2nd voice frame to the 5th voice frame, and then take the mode of the 5 estimated directions of the first interference source to finally determine the direction of the first interference source of the first voice.

[0166] It should be understood that the determination process of the direction of the second sound source and the direction of the second interference source in step 803 is also the same as the determination process of the direction of the first sound source and the direction of the first interference source, which will not be described here.

[0167] 808, obtain a plurality of first beams corresponding to the first voice frame.

[0168] After the terminal device obtains the first voice frame, it can also generate a plurality of first beams corresponding to the first voice frame.

[0169] It should be noted that the generation process of the plurality of first beams can refer to the generation process of the plurality of third beams described above, which will not be described here.

[0170] It should be understood that steps 808 and 806 can be executed synchronously or asynchronously. For example, steps 808 and 806 can be executed simultaneously. For another example, step 808 can be executed before step 806. For another example, step 808 can be executed after step 806, which is not limited here.

[0171] 809, determine a first target beam from the plurality of first beams according to the direction of the first sound source, and determine a first interference beam from the plurality of first beams according to the direction of the first interference source.

[0172] After the terminal device determines the direction of the first sound source and the direction of the first interference source, it determines the first beam closest to the first sound source as the first target beam and the first beam closest to the first interference source as the first interference beam from the plurality of first beams.

[0173] 810, filter the first target beam according to the first interference beam.

[0174] After the terminal device determines the first target beam and the first interference beam, it filters the first target beam according to the first interference beam, thereby completing the filtering process of the first voice frame.

[0175] It should be noted that the filtering process of the first target beam can refer to the filtering process of the third target beam described above, which will not be described here.

[0176] 811, obtain a second voice frame, which is any voice frame after the first voice frame in the first voice.

[0177] After the filtering operation of the first speech frame is completed, the terminal device can continue to acquire a speech frame after the first speech frame, i.e., a second speech frame. Since there are usually multiple second speech frames after the first speech frame, the terminal device can perform filtering processing such as steps 812 to 815 on each second speech frame as it is acquired, until the first speech acquisition is completed.

[0178] 812. Acquire multiple second beams corresponding to the second speech frame.

[0179] After acquiring a certain second speech frame, the terminal device can also generate multiple second beams corresponding to the second speech frame.

[0180] It should be noted that the generation process of the multiple second beams can refer to the generation process of the multiple third beams described above, which will not be described again here.

[0181] 813. If it is determined that the Kth speech frame before the second speech frame is a non-wakeup-triggering speech frame, a second target beam is determined from the multiple second beams according to the direction of the first sound source, and a second interference beam is determined from the multiple second beams according to the direction of the first interference source.

[0182] After acquiring a certain second speech frame, the terminal device can determine whether the previous speech frame of the second speech frame is a wakeup-triggering speech frame. If it is determined that the previous speech frame of the second speech frame is a non-wakeup-triggering speech frame (i.e., it is determined that there is no wakeup-triggering speech frame before the second speech frame), it means that the direction of the first sound source and the direction of the first interference source do not need to be calculated in the filtering process of the second speech frame. Since the direction of the first sound source and the direction of the first interference source have been calculated in the filtering process of the first speech frame and stored in the terminal device. Therefore, the terminal device can directly acquire the direction of the first sound source and the direction of the first interference source, determine the second beam closest to the first sound source as the second target beam from the multiple second beams, and determine the second beam closest to the first interference source as the second interference beam.

[0183] 814. Filter the second target beam according to the second interference beam.

[0184] After the terminal device determines the second target beam and the second interference beam corresponding to a certain second speech frame, it filters the second target beam according to the second interference beam, thereby completing the filtering processing of the second speech frame.

[0185] It should be noted that the filtering process of the second target beam can refer to the filtering process of the third target beam described above, which will not be described again here.

[0186] When the filtering processing of all second speech frames is completed, the filtering operation of the first speech is completed.

[0187] In this embodiment, after the terminal device determines that the Kth speech frame before the first speech frame is the speech frame triggering the wake-up, the terminal device determines the direction of the first sound source by using N speech frames before the first speech frame. Since the last speech frame in the N speech frames is the Kth speech frame before the first speech frame, when the direction of the first sound source is determined, the information of multiple speech frames related to the wake-up of the terminal device is considered, so that the direction can more accurately indicate the actual position of the first sound source. Therefore, filtering the first speech frame based on the direction can improve the filtering quality of the speech frame.

[0188] For further understanding, a specific application example will be provided below to further introduce the method for processing a speech signal provided by the embodiments of the present application. Figure 10 An application example of the method for processing a speech signal provided by the embodiments of the present application is intended to, as shown in Figure 10 It should be noted that the 10th speech frame is the speech frame triggering the wake-up, K = 1, and the application example includes:

[0189] S1: Obtain the ith speech frame in the current speech.

[0190] S2: Perform echo cancellation on the ith speech frame, and calculate the reference signal average energy refeng of the ith speech frame.

[0191] S3: Form multiple beams of the ith speech frame.

[0192] S4: Estimate the direction of the ith speech frame.

[0193] If i < 10, obtain a target parameter, which can be the energy of the beam, the direction of the sound source determined in the last speech, and the direction of the interference source, etc.

[0194] If i = 10, determine the direction of the sound source and the direction of the interference source of the current speech based on the 1st speech frame to the 9th speech frame.

[0195] If i > 10, obtain the direction of the sound source and the direction of the interference source of the current speech.

[0196] S5: Determine the target beam and the interference beam corresponding to the ith speech frame, and determine the signal-to-noise ratio SNR of the ith speech frame.

[0197] If i < 10, determine the target beam and the interference beam based on the target parameter.

[0198] If i = 10, determine the target beam and the interference beam based on the direction of the sound source and the direction of the interference source of the current speech.

[0199] If i>10, the target beam and the interference beam are determined based on the orientation of the sound source of the current speech and the orientation of the interference source.

[0200] S6: The target beam is filtered according to the interference beam, and a filtered ith speech frame is obtained.

[0201] S7: The wake-up threshold is adjusted according to refeng and SNR of the ith speech frame, and the filtered ith speech frame is subjected to wake-up detection based on the adjusted wake-up threshold.

[0202] S8: Speech recognition is performed on the filtered ith speech frame.

[0203] By Figure 10 and Figure 11 ( Figure 11 As can be seen from a schematic diagram of a terminal device processing a speech signal provided by an embodiment of the present application, in the process of speech enhancement, the present application utilizes the information of speech wake-up (i.e. the information of the speech frame triggering wake-up and the like) to optimize parameters, thereby improving the filtering quality of the speech frame. Furthermore, in the process of speech wake-up, the information of speech enhancement (i.e. refeng and SNR) is also utilized to optimize parameters, thereby improving the accuracy of wake-up detection.

[0204] The above is a specific description of the method for processing a speech signal provided by an embodiment of the present application, and the device for processing a speech signal provided by an embodiment of the present application will be introduced below. Figure 12 A structural schematic diagram of the device for processing a speech signal provided by an embodiment of the present application is shown in Figure 12 As shown in the figure, the device comprises:

[0205] The acquisition module 1201 is configured to acquire a first speech frame.

[0206] The determination module 1202 is configured to, if it is determined that the first speech frame includes a speech frame triggering wake-up, determine the orientation of a first sound source of a first speech according to N speech frames before the first speech frame, the first speech including the first speech frame and the N speech frames, the N speech frames including the speech frame triggering wake-up, and N being an integer greater than or equal to 1.

[0207] The filtering module 1203 is configured to filter the first speech frame according to the orientation of the first sound source.

[0208] In a possible implementation manner, the determination module 1202 is specifically configured to: acquire N estimated orientations of the first sound source corresponding to the N speech frames before the first speech frame, wherein each speech frame corresponds to an estimated orientation of the first sound source. The mode of the N estimated orientations of the first sound source is taken to obtain the orientation of the first sound source of the first speech.

[0209] In a possible implementation, the determining module 1202 is further configured to: obtain the estimated positions of the M first interference sources corresponding to the M speech frames, where each speech frame corresponds to an estimated position of a first interference source. The position of the first interference source of the first speech is obtained by taking the mode of the estimated positions of the M first interference sources.

[0210] In a possible implementation, the filtering module 1203 is specifically configured to: obtain a plurality of first beams corresponding to the first speech frame, different first beams having different positions. The first target beam is determined from the plurality of first beams according to the position of the first sound source. The first interference beam is determined from the plurality of first beams according to the position of the first interference source. The first target beam is filtered according to the first interference beam.

[0211] In a possible implementation, the obtaining module 1201 is further configured to obtain a second speech frame, the second speech frame being any speech frame after the first speech frame in the first speech. The filtering module 1203 is further configured to: obtain a plurality of second beams corresponding to the second speech frame, different second beams having different positions. If it is determined that the second speech frame does not include a speech frame triggering wake-up before, the second target beam is determined from the plurality of second beams according to the position of the first sound source. The second interference beam is determined from the plurality of second beams according to the position of the first interference source. The second target beam is filtered according to the second interference beam.

[0212] In a possible implementation, for any speech frame in the N speech frames, the estimated position of the first sound source corresponding to the speech frame is determined according to a difference between a fast-updated cross-correlation energy spectrum and a slow-updated cross-correlation energy spectrum, the fast-updated cross-correlation energy spectrum being determined according to a cross-correlation energy spectrum of the speech frame, a cross-correlation energy spectrum of a speech frame before the speech frame, and a preset first weight, the slow-updated cross-correlation energy spectrum being determined according to the cross-correlation energy spectrum of the speech frame, the cross-correlation energy spectrum of the speech frame before the speech frame, and a preset second weight, the cross-correlation energy spectrum of the speech frame being determined according to the speech frame.

[0213] In a possible implementation, for any speech frame in the M speech frames, the estimated position of the first interference source corresponding to the speech frame is determined according to the slow-updated cross-correlation energy spectrum.

[0214] In a possible implementation, if the first voice is a voice that wakes up the terminal device for the first time, the obtaining module 1201 is further configured to obtain a third voice frame, the third voice frame being any one of the voice frames before the first voice frame in the first voice. The filtering module 1203 is further configured to: obtain a plurality of third beams corresponding to the third voice frame, different third beams having different orientations. If it is determined that there is no voice frame that triggers wake-up before the third voice frame, a third target beam is determined from the plurality of third beams according to the energy of each third beam. A third interference beam is determined from the plurality of third beams according to the orientation relationship between the third beams and the third target beam. The third target beam is filtered according to the third interference beam.

[0215] In a possible implementation, if the first voice is a voice that wakes up the terminal device for the first time, the obtaining module 1201 is further configured to obtain a third voice frame, the third voice frame being any one of the voice frames before the first voice frame in the first voice. The filtering module 1203 is further configured to: obtain a plurality of third beams corresponding to the third voice frame, different third beams having different orientations. If it is determined that there is no voice frame that triggers wake-up before the third voice frame, a third target beam is determined from the plurality of third beams according to the energy of each third beam and / or the orientation of the second sound source of the second voice, the second voice being a voice that wakes up the terminal device for the last time. A third interference beam is determined from the plurality of third beams according to the orientation of the second interference source of the second voice. The third target beam is filtered according to the third interference beam.

[0216] In a possible implementation, the filtering includes linear filtering and / or nonlinear filtering. For any one of the voice frames of the first voice, a gain of linear filtering of the voice frame is determined according to a target beam corresponding to the voice frame, an interference beam corresponding to the voice frame, and a third weight, a gain of nonlinear filtering of the voice frame is determined according to the target beam corresponding to the voice frame, the interference beam corresponding to the voice frame, and a fourth weight, and the third weight or the fourth weight is determined according to a signal-to-noise ratio of the voice frame.

[0217] In a possible implementation, for any one of the voice frames of the first voice, a wake-up threshold corresponding to the voice frame is determined according to a preset threshold, a signal-to-noise ratio of the voice frame, and a reference signal average energy of the voice frame, the wake-up threshold corresponding to the voice frame being used to determine whether the voice frame is a voice frame that triggers wake-up, and the reference signal average energy of the voice frame being determined based on an energy of a reference voice frame corresponding to the voice frame and energies of reference voice frames corresponding to voice frames before the voice frame, the reference voice frame corresponding to the voice frame being a voice frame output by the terminal device when the voice frame is received.

[0218] It should be noted that the information interaction, execution process and the like between the modules / units of the above apparatus are based on the same concept as the method embodiments of the present application, and the technical effects brought by the same are the same as the method embodiments of the present application. For details, refer to the description in the foregoing method embodiments of the present application, which will not be repeated here.

[0219] The embodiments of the present application also relate to a computer-readable storage medium comprising instructions that, when executed on a computer or processor, cause the computer or processor to perform the method as shown in Figure 7 or Figure 8

[0220] The embodiments of the present application also relate to a computer program product comprising instructions that, when executed on a computer or processor, cause the computer or processor to perform the method as shown in Figure 7 or Figure 8

[0221] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, apparatus and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.

[0222] In several embodiments provided in the present application, it should be understood that the disclosed system, apparatus and method can be implemented in other ways. For example, the above-described apparatus embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0223] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments of the present application.

[0224] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0225] ​​The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

Claims

1. A method of processing a speech signal, characterized by, The method comprises: acquiring a first speech frame; if it is determined that the first speech frame before includes a speech frame triggering wake-up, determining the direction of a first sound source of a first speech according to N speech frames before the first speech frame, the first speech comprising the first speech frame, the N speech frames and M speech frames before the N speech frames, the N speech frames comprising the speech frame triggering wake-up, N being an integer greater than or equal to 1, and M being an integer greater than or equal to 1; acquiring the estimated direction of M first interference sources corresponding to the M speech frames, wherein each speech frame corresponds to the estimated direction of one first interference source; taking the mode in the estimated direction of the M first interference sources to obtain the direction of the first interference source of the first speech; acquiring a plurality of first beams corresponding to the first speech frame, different first beams having different directions; determining a first target beam in the plurality of first beams according to the direction of the first sound source; determining a first interference beam in the plurality of first beams according to the direction of the first interference source; filtering the first target beam according to the first interference beam.

2. The method of claim 1, wherein, The method further comprises: acquiring a second speech frame, the second speech frame being any one speech frame after the first speech frame in the first speech; acquiring a plurality of second beams corresponding to the second speech frame, different second beams having different directions; 3. The method of claim 2, wherein, if it is determined that the second speech frame before does not include a speech frame triggering wake-up, determining a second target beam in the plurality of second beams according to the direction of the first sound source; determining a second interference beam in the plurality of second beams according to the direction of the first interference source; filtering the second target beam according to the second interference beam. For any one speech frame in the N speech frames, the estimated direction of the first sound source corresponding to the speech frame is determined according to the difference between the fast-updated cross-correlation energy spectrum and the slow-updated cross-correlation energy spectrum, the fast-updated cross-correlation energy spectrum being determined according to the cross-correlation energy spectrum of the speech frame, the cross-correlation energy spectrum of the speech frame before, and a preset first weight, the slow-updated cross-correlation energy spectrum being determined according to the cross-correlation energy spectrum of the speech frame, the cross-correlation energy spectrum of the speech frame before, and a preset second weight, and the cross-correlation energy spectrum of the speech frame being determined according to the speech frame. For any one speech frame in the M speech frames, the estimated direction of the first interference source corresponding to the speech frame is determined according to the slow-updated cross-correlation energy spectrum. ​ 4. The method of claim 3, wherein, ​ 5. The method of claim 4, wherein, ​ 6. The method according to any one of claims 1 to 5, characterized in that, If the first voice is a voice for first time waking up the terminal device, before the first voice frame is acquired, the method further includes: acquiring a third voice frame, the third voice frame being any one voice frame before the first voice frame in the first voice; acquiring a plurality of third beams corresponding to the third voice frame, different third beams having different orientations; if it is determined that the third voice frame does not include a voice frame triggering wake-up, determining a third target beam in the plurality of third beams according to energy of each third beam; determining a third interference beam in the plurality of third beams according to an orientation relationship between the third beams and the third target beam; filtering the third target beam according to the third interference beam.

7. The method according to any one of claims 1 to 5, characterized in that, If the first voice is a voice for first time waking up the terminal device, before the first voice frame is acquired, the method further includes: acquiring a third voice frame, the third voice frame being any one voice frame before the first voice frame in the first voice; acquiring a plurality of third beams corresponding to the third voice frame, different third beams having different orientations; if it is determined that the third voice frame does not include a voice frame triggering wake-up, determining a third target beam in the plurality of third beams according to energy of each third beam and / or an orientation of a second sound source of a second voice, the second voice being a voice for last time waking up the terminal device by the first voice; determining a third interference beam in the plurality of third beams according to an orientation of a second interference source of the second voice; filtering the third target beam according to the third interference beam.

8. The method according to any one of claims 1 to 7, characterized in that, The filtering includes linear filtering and / or nonlinear filtering, for any one voice frame of the first voice, a gain of linear filtering of the voice frame is determined according to a target beam corresponding to the voice frame, an interference beam corresponding to the voice frame and a third weight value, a gain of nonlinear filtering of the voice frame is determined according to the target beam corresponding to the voice frame, the interference beam corresponding to the voice frame and a fourth weight value, the third weight value or the fourth weight value is determined according to a signal-to-noise ratio of the voice frame.

9. The method of claim 8, wherein, For any one voice frame of the first voice, a wake-up threshold corresponding to the voice frame is determined according to a preset threshold, a signal-to-noise ratio of the voice frame and a reference signal average energy of the voice frame, the wake-up threshold corresponding to the voice frame is used to determine whether the voice frame is a voice frame triggering wake-up, the reference signal average energy of the voice frame is determined based on energy of a reference voice frame corresponding to the voice frame and energy of a reference voice frame corresponding to a voice frame before the voice frame, the reference voice frame corresponding to the voice frame being a voice frame output by the terminal device when the voice frame is received.

10. An apparatus for processing a speech signal, characterized by The apparatus includes: an acquisition module, configured to acquire a first voice frame; a determination module, configured to: determine a direction of a first sound source of the first voice according to N voice frames before the first voice frame, the first voice including the first voice frame, the N voice frames and M voice frames before the N voice frames, the N voice frames including the voice frame triggering the wake-up, N being an integer greater than or equal to 1, and M being an integer greater than or equal to 1; obtain estimated directions of M first interference sources corresponding to the M voice frames, wherein each voice frame corresponds to an estimated direction of a first interference source; obtain a direction of a first interference source of the first voice by taking a mode value in the estimated directions of the M first interference sources; the filtering module is configured to: obtain a plurality of first beams corresponding to the first voice frame, different first beams having different directions; determine a first target beam in the plurality of first beams according to the direction of the first sound source; determine a first interference beam in the plurality of first beams according to the direction of the first interference source; and filter the first target beam according to the first interference beam.

11. The apparatus of claim 10, wherein, The determining module is specifically configured to: obtain estimated directions of N first sound sources corresponding to N voice frames before the first voice frame, wherein each voice frame corresponds to an estimated direction of a first sound source; obtain the direction of the first sound source of the first voice by taking a mode value in the estimated directions of the N first sound sources.

12. The apparatus of claim 11, wherein, The obtaining module is further configured to obtain a second voice frame, the second voice frame being any one of voice frames after the first voice frame in the first voice; The filtering module is further configured to: obtain a plurality of second beams corresponding to the second voice frame, different second beams having different directions; if it is determined that no voice frame triggering the wake-up is included before the second voice frame, determine a second target beam in the plurality of second beams according to the direction of the first sound source; determine a second interference beam in the plurality of second beams according to the direction of the first interference source; and filter the second target beam according to the second interference beam.

13. The apparatus of claim 12, wherein, For any one of the N voice frames, an estimated direction of a first sound source corresponding to the voice frame is determined according to a difference between a fast-updated cross-correlation energy spectrum and a slow-updated cross-correlation energy spectrum, the fast-updated cross-correlation energy spectrum being determined according to a cross-correlation energy spectrum of the voice frame, a cross-correlation energy spectrum of a voice frame before the voice frame and a preset first weight, the slow-updated cross-correlation energy spectrum being determined according to the cross-correlation energy spectrum of the voice frame, the cross-correlation energy spectrum of the voice frame before the voice frame and a preset second weight, and the cross-correlation energy spectrum of the voice frame being determined according to the voice frame.

14. The apparatus of claim 13, wherein, For any one of the M voice frames, an estimated direction of a first interference source corresponding to the voice frame is determined according to a slow-updated cross-correlation energy spectrum.

15. The apparatus of any one of claims 10 to 14, wherein, If the first voice is a voice for first waking up the terminal device, the obtaining module is further configured to obtain a third voice frame, the third voice frame being any one of voice frames before the first voice frame in the first voice. The filtering module is further configured to: obtain a plurality of third beams corresponding to a third voice frame, different third beams having different orientations; if it is determined that no voice frame triggering wake-up is included before the third voice frame, determine a third target beam from the plurality of third beams according to energy of each third beam; determine a third interference beam from the plurality of third beams according to an orientation relationship between the third beams and the third target beam; filter the third target beam according to the third interference beam.

16. The apparatus of any one of claims 10 to 14, wherein, If the first voice is a voice that does not wake up the terminal device for the first time, the obtaining module is further configured to obtain a third voice frame, the third voice frame being any one of the voice frames before the first voice frame in the first voice; The filtering module is further configured to: obtain a plurality of third beams corresponding to a third voice frame, different third beams having different orientations; if it is determined that no voice frame triggering wake-up is included before the third voice frame, determine a third target beam from the plurality of third beams according to energy of each third beam and / or an orientation of a second sound source of a second voice, the second voice being a voice that wakes up the terminal device last time for the first voice; determine a third interference beam from the plurality of third beams according to an orientation of a second interference source of the second voice; filter the third target beam according to the third interference beam.

17. The apparatus of any one of claims 10 to 16, wherein, The filtering includes linear filtering and / or nonlinear filtering, for any one of the voice frames of the first voice, a gain of linear filtering of the voice frame is determined according to a target beam corresponding to the voice frame, an interference beam corresponding to the voice frame, and a third weight, a gain of nonlinear filtering of the voice frame is determined according to the target beam corresponding to the voice frame, the interference beam corresponding to the voice frame, and a fourth weight, the third weight or the fourth weight is determined according to a signal-to-noise ratio of the voice frame.

18. The apparatus of claim 17, wherein, For any one of the voice frames of the first voice, a wake-up threshold corresponding to the voice frame is determined according to a preset threshold, a signal-to-noise ratio of the voice frame, and a reference signal average energy of the voice frame, the wake-up threshold corresponding to the voice frame is used to determine whether the voice frame is a voice frame triggering wake-up, the reference signal average energy of the voice frame is determined based on energy of a reference voice frame corresponding to the voice frame and energy of a reference voice frame corresponding to a voice frame before the voice frame, the reference voice frame corresponding to the voice frame being a voice frame output by the terminal device when the voice frame is received.

19. An electronic device, comprising: comprising: a processor and a memory, the processor being configured to invoke program instructions stored in the memory to perform the method of any one of claims 1 to 9.

20. A computer-readable storage medium comprising instructions that, when executed on a computer or a processor, cause the computer or the processor to perform the method of any one of claims 1 to 9.

21. A computer program product comprising instructions, the computer program product comprising program instructions that, when executed on a computer or a processor, cause the computer or the processor to perform the method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Microphone array speech enhancement device capable of suppressing mobile noise

    CN102969002A

  • Microphone array-based human voice acquisition method and electronic device

    CN106935246A