Signal processing method, electronic equipment and computer readable storage medium
By using multiple sets of beam parameters on electronic devices to enhance the signal collected by the microphone array, the problem of difficulty in collecting voice signals in noisy environments is solved, directional enhancement and interference suppression are achieved, and the signal-to-noise ratio and user experience of voice signals are improved.
Patent Information
- Application Number
- CN202411394297.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-09-29
AI Technical Summary
In noisy environments, it is difficult for electronic devices to effectively collect and recognize voice signals, resulting in low voice intelligibility and affecting user experience.
Multiple beam parameters are used to enhance the multiple input signals, obtain more airspace information through the microphone array, enhance the voice signal in a direction, and suppress interference signals to improve the signal-to-noise ratio.
Improve the signal-to-noise ratio of voice signals in noisy environments, improve user experience, ensure voice clarity and wake-up rate, and is suitable for voice calls, hearing aids and voice conference scenarios.
Smart Images

Figure CN120469284A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of signal processing technology, and in particular to a signal processing method, an electronic device, and a computer-readable storage medium. Background Art
[0002] Electronic devices are often equipped with microphones to collect sound signals. These signals can be transmitted to enable voice calls or recognized to perform voice wake-up operations.
[0003] However, in some scenarios, such as during a voice conference, if the surrounding environment is noisy, the signal-to-noise ratio of the participant's voice signal will be reduced. As a result, the participating devices may not be able to properly capture the speaker's voice, and thus cannot accurately recognize or transmit the voice signal, resulting in poor voice intelligibility, which affects the user experience. Summary of the Invention
[0004] The present application provides a signal processing method, device, chip, electronic device, computer-readable storage medium and computer program product, which can enhance speech directionally.
[0005] In a first aspect, a signal processing method is provided for use in an electronic device comprising multiple microphones. The method comprises: obtaining multiple input signals, where the multiple input signals are signals collected by the multiple microphones, and the multiple input signals correspond one-to-one to the multiple microphones; enhancing the multiple input signals using multiple sets of beam parameters to obtain multiple enhanced signals, where the multiple sets of beam parameters and the multiple enhanced signals correspond one-to-one to multiple sectors, and the beam parameters are associated with the frequency and arrival direction of the input signals, where the multiple sectors are spatial regions divided around the multiple microphones; and determining and outputting output signals based on the multiple enhanced signals, where the output signals are one or more of the multiple enhanced signals. Each microphone can output the collected sound signal as one output signal, and a microphone array comprising multiple microphones can output multiple output signals. For a detailed description of beam parameters and sectors, please refer to the relevant description herein. The processor can enhance the multiple input signals using any set of beam parameters, for example, by multiplying them by corresponding filter weight vectors, to obtain multiple enhanced signals. It should be noted that each sector corresponds to a set of beam parameters. Each set of beam parameters can output an enhanced signal corresponding to a corresponding sector. The processor determines the signal with the best signal quality from the multiple enhanced signals and outputs it as the output signal; the processor can also directly output the multiple enhanced signals.
[0006] When using a microphone array to collect multiple input signals, the sound signals received from the same sound source are different due to the different positions of different microphones in the microphone array. Therefore, the microphone array can obtain more spatial information when collecting sound signals, which is convenient for directional enhancement of speech.
[0007] This method can perform directional enhancement on the signals collected by multiple microphones. In this way, even in a noisy environment, it can receive sound signals from the direction of the sound with a higher gain and suppress interference signals from other directions, thereby improving the signal-to-noise ratio of useful signals and thus enhancing the user experience.
[0008] In some possible implementations, multiple sets of beam parameters are used to enhance multiple input signals to obtain multiple enhanced signals, including: using a first beam parameter to determine a first fluctuation value of the multiple input signals, the first fluctuation value characterizing the degree of fluctuation of the multiple input signals in multiple incoming wave directions, the first beam parameter is any one of the multiple sets of beam parameters, and the first beam parameter corresponds to the first sector of the multiple sectors; if the first fluctuation value is greater than or equal to a preset threshold, then the updated filter weight vector at the next moment is adjusted according to the current filter weight vector, and the updated filter weight vector is different from the current filter weight vector; if the first fluctuation value is less than the preset threshold, then the updated filter weight vector at the next moment is determined to be the current filter weight vector; according to the updated filter weight vector, the multiple input signals are enhanced to obtain a first enhanced signal to obtain the first enhanced signal, which is one of the multiple enhanced signals.
[0009] The above-mentioned first fluctuation value can characterize the fluctuation degree of the input signal in multiple incoming wave directions. Therefore, the filter weight vector can be adjusted when the fluctuation degree is large, and when the fluctuation degree is small, the filter weight vector can be kept unchanged, thereby realizing real-time adaptive update of the filter weight vector and improving the accuracy of voice directionality.
[0010] In some possible implementations, the first fluctuation value is the difference between the maximum intensity ratio and the minimum intensity ratio, the maximum intensity ratio is the largest of multiple intensity ratios, and the minimum intensity ratio is the smallest of multiple intensity ratios; the multiple signal strengths are the ratios of the signal strength of each directional enhanced signal in the multiple directional enhanced signals to the signal strength of each input signal in the multiple input signals; the multiple directional enhanced signals are enhanced signals obtained by performing beamforming processing on the multiple input signals in multiple incoming directions using the first beam parameters, and the number of multiple intensity ratios is the product of the number of multiple directional enhanced signals and the number of multiple input signals.
[0011] In some possible implementations, the first beam parameters include multiple groups of preset gains, which correspond one-to-one to multiple wave directions. The first beam parameters are used to determine the first fluctuation value of the multi-channel input signal, including: using the multiple groups of preset gains to perform beamforming processing on the multi-channel input signal in multiple wave directions to obtain multiple direction-enhanced signals, and the multiple direction-enhanced signals correspond one-to-one to the multiple wave directions; determining a first intensity ratio of the signal strength of the first direction-enhanced signal to the signal strength of the first input signal, the first direction-enhanced signal is any one of the multiple direction-enhanced signals, the first input signal is any one of the multiple input signals, and the first intensity ratio is one of the multiple intensity ratios; determining the difference between the maximum intensity ratio and the minimum intensity ratio among the multiple intensity ratios as the first fluctuation value.
[0012] In some possible implementations, the output signal is the one with the strongest signal strength among the multiple enhanced signals.
[0013] For each sector, the processor can calculate the ratio of the strength of each initial enhanced signal to the strength of each input signal, denoted as the intensity ratio. The processor can obtain multiple intensity ratios, then select the largest and smallest values and calculate the difference between them. This difference can quantitatively represent the degree of fluctuation in the multiple intensity ratios.
[0014] In some possible implementations, the output signal is the one with the strongest signal strength among the multiple enhanced signals.
[0015] The processor may directly compare the signal strengths of the enhanced signals and select the one with the strongest signal strength from the multiple enhanced signals as the output signal, thereby ensuring that the signal quality of the output signal is the output signal of the optionally best quality.
[0016] In some possible implementations, the output signal is the one with the strongest signal strength among the multiple enhanced signals and the multiple input signals.
[0017] The processor may select the one with the strongest signal strength from the multiple enhanced signals and the multiple input signals as the output signal, ensuring that the signal quality of the output signal is the output signal with the best quality.
[0018] In some possible implementations, determining and outputting an output signal based on the multiple enhanced signals includes: determining the signal strength of each enhanced signal in the multiple enhanced signals, and multiple second intensity ratios of the comprehensive intensity, where the comprehensive intensity is the sum of the intensities of multiple input signals, and the multiple second intensity ratios correspond one-to-one to the multiple enhanced signals; screening out the largest second intensity ratio from the multiple second intensity ratios; determining the enhanced signal corresponding to the largest second intensity ratio as the output signal and outputting it.
[0019] The processor may also use the second intensity ratio to quantify the signal quality of the enhanced signal. The processor selects the enhanced signal corresponding to the one with the largest second intensity ratio as the output signal to ensure the signal quality of the output signal.
[0020] In some possible implementations, the electronic device is in a voice scenario, where the voice scenario includes one of a call scenario, a hearing assistance scenario, and a voice conference scenario.
[0021] When the processor outputs one output signal, that output signal can be used in voice scenarios, such as hearing-assisted scenarios, voice calls, and voice conferencing scenarios. In hearing-assisted scenarios, this method can enhance the voice from the direction of the sound source, allowing users to hear the sound more clearly and facilitate conversations. In voice conferencing scenarios, this method can enhance the voice from the direction of the sound source, improve voice clarity, and reduce interference in noisy environments, thereby improving the user's conference experience.
[0022] In some possible implementations, the method further includes: outputting a plurality of signal-to-noise ratio estimation values, wherein the plurality of signal-to-noise ratio estimation values correspond one-to-one to the plurality of sectors.
[0023] The signal-to-noise ratio estimation value can quantify the signal quality of the enhanced signal, making it easier to selectively output the enhanced signal with good quality, thereby improving the effect of speech processing.
[0024] In some possible implementations, determining and outputting an output signal based on multiple enhanced signals includes: determining at least one signal to be output from the multiple enhanced signals, where a signal-to-noise ratio estimate corresponding to each signal in the at least one signal to be output is greater than or equal to a preset signal-to-noise ratio threshold; and if the number of the at least one signal to be output is less than the number of the multiple enhanced signals, outputting the at least one signal to be output and part or all of the multiple input signals as output signals.
[0025] If there is an enhanced signal among the multiple enhanced signals whose estimated signal-to-noise ratio is less than a preset signal-to-noise ratio threshold, it indicates that such an enhanced signal may not effectively enhance speech. Therefore, the processor can select as output the enhanced signal whose estimated signal-to-noise ratio is greater than or equal to the preset signal-to-noise ratio threshold, while also selecting one or more channels from the original input signals for output.
[0026] This ensures that voice enhancement can be achieved for different sectors in scenarios with more interference. In scenarios with less interference, when the signal quality of some enhanced signals may not be as good as the original input signals, directly outputting the input signals will enable subsequent processing modules to obtain higher quality output signals, making subsequent processing more accurate.
[0027] In some possible implementations, the number of output signal paths is the same as the number of sectors.
[0028] In some possible implementations, the signal-to-noise ratio estimate corresponding to each signal in the at least one signal to be output is greater than or equal to a preset signal-to-noise ratio threshold, including the situation where multiple instantaneous signal-to-noise ratio estimates corresponding to the signal-to-noise ratio estimate corresponding to each signal in the at least one signal to be output at multiple consecutive moments are all greater than the preset signal-to-noise ratio threshold.
[0029] To prevent interference from causing poor signal quality in the output enhanced signal, the processing module can continuously collect multiple instantaneous SNR estimates when determining whether the SNR estimate of an enhanced signal is greater than or equal to a preset SNR threshold. This approach prevents interference from causing poor signal quality in the output enhanced signal, improves the output signal quality, and enhances speech processing effectiveness.
[0030] In some possible implementations, the electronic device is in a non-voice scenario, where the non-voice scenario includes one or more of: voice wake-up and voice recognition.
[0031] When the processor outputs multiple output signals, it can be used in some non-speech scenarios, such as speech recognition and voice wake-up. In the voice wake-up scenario, this method can enhance speech in a targeted manner and improve the wake-up rate in noisy environments.
[0032] In a second aspect, a signal processing device is provided, comprising a unit composed of software and / or hardware, wherein the unit is configured to execute any one of the methods in the technical solution of the first aspect.
[0033] In a third aspect, an embodiment of the present application provides a chip comprising a processor; the processor is used to read and execute a computer program stored in a memory to execute any one of the methods in the technical solution described in the first aspect.
[0034] Optionally, the chip further includes a memory, and the memory is connected to the processor via a circuit or wire.
[0035] Further optionally, the chip also includes a communication interface.
[0036] In a fourth aspect, an embodiment of the present application provides a chip, which is a processor or an advanced processor.
[0037] In a fifth aspect, an electronic device is provided, which includes: a processor, a memory, and an interface; the processor, the memory, and the interface cooperate with each other so that the electronic device executes any one of the methods in the technical solution described in the first aspect.
[0038] In a sixth aspect, an electronic device is provided, which includes any one chip of the technical solutions described in the third aspect and the fourth aspect.
[0039] In the seventh aspect, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, the processor executes any one of the methods in the technical solution described in the first aspect.
[0040] In an eighth aspect, a computer program product is provided, comprising: a computer program code, which, when executed on an electronic device, enables the electronic device to execute any one of the methods in the technical solution described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 1 is a schematic structural diagram of a terminal device 100 provided in an embodiment of the present application;
[0042] Figure 2 is a software structure block diagram of the terminal device 100 provided in an embodiment of the present application;
[0043] Figure 3 This is an architecture diagram of a signal processing method provided in an embodiment of the present application;
[0044] Figure 4 This is a schematic diagram of sector division provided in an embodiment of the present application;
[0045] Figure 5 This is an example of a beam pattern provided in an embodiment of the present application;
[0046] Figure 6 This is a schematic diagram of a signal enhancement process provided in an embodiment of the present application;
[0047] Figure 7 This is a logic diagram of an example of a channel selection module determining an output signal provided in an embodiment of the present application;
[0048] Figure 8 This is a schematic diagram of a beamforming effect in a conference scenario provided by an embodiment of the present application;
[0049] Figure 9 This is another example of a logic diagram of a channel selection module determining an output signal provided by an embodiment of the present application;
[0050] Figure 10 This is another example of a logic diagram of a channel selection module determining an output signal provided by an embodiment of the present application;
[0051] Figure 11 This is a flow chart of a signal processing method provided in an embodiment of the present application;
[0052] Figure 12 This is a structural diagram of a signal processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships can exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.
[0054] In the following, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the technical features indicated. Therefore, a feature specified as "first," "second," or "third" may explicitly or implicitly include one or more of the features.
[0055] The signal processing method provided in the embodiments of the present application can be applied to terminal devices such as mobile phones, tablet computers, wearable devices, vehicle-mounted devices, hearing aids, augmented reality (AR) / virtual reality (VR) devices, extended-range (XR) devices, large-screen devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). The embodiments of the present application do not impose any restrictions on the specific type of terminal device.
[0056] For example, Figure 11 is a schematic diagram of the structure of an example terminal device 100 provided in an embodiment of the present application. The terminal device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0057] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0058] The controller may be the nerve center and command center of the terminal device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0059] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0060] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0061] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.
[0062] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the terminal device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the terminal device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0063] The terminal device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0064] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.
[0065] The speaker 170A, also called a "speaker", is used to convert audio electrical signals into sound signals. The terminal device 100 can listen to music or listen to hands-free calls through the speaker 170A.
[0066] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the terminal device 100 receives a call or voice message, the user can hear the voice by placing the receiver 170B close to the ear.
[0067] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The terminal device 100 can be provided with at least one microphone 170C. In other embodiments, the terminal device 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the terminal device 100 can also be provided with three, four or more microphones 170C to realize sound signal collection, noise reduction, and can also identify the source of sound, realize directional recording function, etc.
[0068] The headphone jack 170D is used to connect a wired headphone and can be the USB interface 130 or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0069] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the terminal device 100. In other embodiments of the present application, the terminal device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0070] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the terminal device 100. In other embodiments of the present application, the terminal device 100 may also adopt a different interface connection method from the above embodiments, or a combination of multiple interface connection methods.
[0071] The software system of the terminal device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. In the embodiment of the present application, the Android system with a layered architecture is used as an example to illustrate the software structure of the terminal device 100.
[0072] Figure 2 This is a software structure diagram of the terminal device 100 in an embodiment of the present application. The layered architecture divides the software into several layers, each with clear roles and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer. The application layer can include a series of application packages.
[0073] like Figure 2 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.
[0074] The application framework layer provides an application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0075] like Figure 2 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.
[0076] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.
[0077] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.
[0078] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.
[0079] The phone manager is used to provide communication functions of the terminal device 100, such as management of call status (including answering, hanging up, etc.).
[0080] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0081] The notification manager enables applications to display notification information in the status bar, which can be used to convey informational messages and disappear automatically after a short stay without user interaction.
[0082] The Android runtime includes the core library and the virtual machine. The Android runtime is responsible for scheduling and management of the Android system.
[0083] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.
[0084] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.
[0085] The system library can include multiple functional modules, such as a surface manager, media libraries, a 3D graphics processing library (such as OpenGL ES), and a 2D graphics engine (such as SGL).
[0086] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.
[0087] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0088] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0089] A 2D graphics engine is a drawing engine for 2D drawings.
[0090] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.
[0091] For ease of understanding, the following examples of this application will be described with Figure 1 and Figure 2 Taking the terminal device with the shown structure as an example, the signal processing method provided in the embodiment of the present application is specifically explained in combination with the accompanying drawings and application scenarios.
[0092] Electronic devices are often equipped with microphones to collect sound signals. These devices can process and transmit the sound signals collected by the microphones to facilitate voice calls, or recognize sound signals to perform voice wake-up operations. To better collect sound signals, some electronic devices can be equipped with multiple microphones, forming a microphone array, to capture multiple sound channels. Microphone arrays can capture sound signals over a larger area, providing more spatial information and enabling a wider range of applications. However, electronic devices may experience poor sound reception in the presence of ambient sound interference. For example, when a user participates in an online meeting using a large-screen device, with multiple users seated around a conference table facing the large-screen device, if the environment is noisy, the large-screen device may not be able to properly capture the voices of speakers from different locations, impacting the user's experience. Furthermore, with some other electronic devices, the direction of the sound source varies, making it impossible to ensure optimal sound signal collection.
[0093] The present invention provides a signal processing method that can enhance the incoming sound signal in the direction of origin. This allows spatial filtering of the signal received by the microphone array even in noisy environments. This method receives the incoming sound signal with a higher gain and suppresses interference signals from other directions, thereby improving the signal-to-noise ratio of the useful signal and speech intelligibility, thereby enhancing the user experience. Furthermore, this method does not require additional hardware and does not increase hardware costs.
[0094] The signal processing method provided in the embodiments of the present application can be applied to a processor (signal processor, DSP) or (advanced signal processor, ADSP). Optionally, the processor can be applied to an electronic device. A microphone array can also be provided on the electronic device. The signal processing method provided in the embodiments of the present application is described below using the processor as an example of the execution subject and combining the functions of the software function modules in the processor.
[0095] Optionally, the software architecture in the processor can be found in Figure 3 Specifically, the processor can be equipped with two software functional modules: a beamforming module and a channel selection module. The beamforming module is used to enhance the multi-channel sound signals collected by the microphone array and input the enhanced signals into the channel selection module. The channel selection module selects one or more enhanced signals for output for subsequent scene processing.
[0096] First, the multiple input signals input to the beamforming module are described.
[0097] In a microphone array in an electronic device, each microphone can collect and output a single sound signal. Therefore, multiple microphones in the microphone array can collect and output multiple sound signals. Because the microphones in the array are positioned differently, the sound signals collected by each microphone may differ. The multiple sound signals collected by the microphone array can serve as multiple input signals and be fed into a beamforming module for beam enhancement. Figure 3 , when there are two microphones in the microphone array, the input signals include two input signals, input signal 1 and input signal 2. Optionally, when the number of microphones in the microphone array is three, four, or more, the multi-channel input signals may further include input signal 3, input signal 4, or more input signals.
[0098] When using a microphone array to collect multiple input signals, the sound signals received from the same sound source are different due to the different positions of different microphones in the microphone array. Therefore, the microphone array can obtain more spatial information when collecting sound signals, which is convenient for directional enhancement of speech.
[0099] Now that we understand the sources of the multi-channel input signals, we will describe in detail how the beamforming module processes the multi-channel input signals to obtain the multi-channel enhanced signals.
[0100] Specifically, the beamforming module includes a plurality of sector beamformers, each of which corresponds to a sector and is used to enhance the sound signal transmitted from the direction corresponding to the corresponding sector.
[0101] For ease of understanding, here we first introduce the sector division method, for details, please refer to Figure 4 As shown in . Figure 4 Figure a in FIG is a schematic diagram of sector division in the vertical direction with the position of the microphone array of the electronic device as the center. Figure 4 The diagram a in FIG shows a schematic diagram of the three sectors divided into 0 degrees to 60 degrees, 60 degrees to 120 degrees, and 120 degrees to 180 degrees. It should be noted that, Figure 4 The angle range of the sectors shown in the figure is a schematic diagram of the division within a plane. The actual sectors are sectors of three-dimensional space. That is, sector 1 corresponding to 0 degrees to 60 degrees is the spatial area of the cone centered on the microphone array, and sector 2 corresponding to 120 degrees to 180 degrees is the spatial area of the cone centered on the microphone array. The vertices of sectors 1 and 2 are opposite, and sector 3 corresponding to 60 degrees to 120 degrees is the spatial area in the space except sectors 1 and 2. For details, please refer to Figure 4 As shown in Figure b.
[0102] It should be noted that the finer the sector division and the greater the number of sectors, the finer the granularity of the division of signal enhancement directions, the more targeted and effective the signal enhancement effect; the coarser the sector division and the fewer the number of sectors, the simpler the signal processing process, saving resources and improving processing efficiency. In some embodiments, dividing the space into three sectors can balance the signal enhancement effect and the complexity of signal processing, and is therefore more reasonable.
[0103] Figure 4 Figure c in FIG. 1 shows another schematic diagram of sector division, which includes three sectors: 0 degrees to 45 degrees, 45 degrees to 90 degrees, and 90 degrees to 180 degrees.
[0104] Optionally, the number of divided sectors may also be two, for example, two sectors from 0 degrees to 90 degrees and from 90 degrees to 180 degrees. The number of sectors may also be four, five or more.
[0105] Optionally, the multiple sectors can be divided evenly or unevenly. For example, the sectors from which the sound is likely to come in most cases can be divided with fine granularity, while the sectors in other directions can be divided with coarse granularity. Figure 4 As shown in Figure c, when using a mobile phone, the user often moves closer to the direction of the microphone array, and rarely speaks from the top of the phone (the opposite end of the microphone array). Therefore, the area from 0 to 90 degrees can be divided into two sectors, and 90 to 180 degrees can be divided into one sector, that is, the fine-grained division of the sound direction is ensured without increasing the number of sectors too much, which is more reasonable. Optionally, the space can also be divided into three sectors: 0 to 40 degrees, 40 to 80 degrees, and 80 to 180 degrees. The embodiment of the present application does not limit the way of dividing the sectors.
[0106] It should be noted that each sector corresponds to a set of debugged beam parameters, which are used to enhance the multiple input signals transmitted from the corresponding sector. The multiple sets of beam parameters for multiple sectors can also be referred to as direction of arrival information. The beam parameters are preset and can be downloaded from a server or pre-set in the processor.
[0107] Each set of beam parameters can perform beamforming in N directions for multiple input signals, that is, perform signal enhancement in N directions. Optionally, N is a natural number, which can be a natural number greater than or equal to 3, such as 5, 8, 10, etc. The N directions here may be different from the previous sector division. Although it is also a division for space, the number of divisions can be more. There is no dependency between these N directions and the divided N sectors, and the two are independent division methods. The direction here refers to the arrival direction corresponding to the signal, also called the arrival angle (ArrivalAngle), and can also be called the direction of arrival, and the unit can be angle (degree).
[0108] Since the processor has different gains for signals arriving from different directions, the gains corresponding to the signals output after N beamforming operations are all different. Figure 5 The following is an example of a beam pattern. The lighter the color, the greater the gain, and the darker the color, the smaller the gain. It can be seen that the gain of the signal is related to the direction of the wave and the frequency of the signal. The gain of the same signal in different directions of the wave also varies. Figure 5 As shown in the figure, at very low frequencies, close to zero frequency, the gain hardly changes with the incoming wave direction. In other words, for low-frequency signals, the gain has no obvious directional difference (i.e., it is non-directional). For signals of the same frequency, the gain varies with the incoming wave direction, and multiple lobes can appear within the range of the incoming wave direction. Figure 5The figure shows the gain points of the main lobe and side lobe. The gain of the signal here is large, showing an enhancement state. The gain of the signal corresponding to the other dark gully locations is very small, showing a suppression state. Figure 5 In the example beam pattern, the gain is the beam parameter mentioned above. The gain corresponding to each frequency (or frequency range) and incoming wave direction (or range of incoming wave directions) is the beam parameter corresponding to that incoming wave direction. This shows that the beam parameter is not a single number, but a set of values corresponding to multiple frequencies and directions. Alternatively, the beam parameter can be obtained through experimental measurement or simulation, which will not be detailed here.
[0109] Next, we will take one of the first sector beamformers as an example to explain the signal enhancement process in detail. Figure 6 As shown, Figure 6 In the example, the multiple input signals are input signal 1 and input signal 2.
[0110] After input signals 1 and 2 are input to the first sector beamformer, the first sector beamformer first enhances input signals 1 and 2 using each of N sets of beam parameters, generating N initial enhanced signals corresponding to each of the N sets of beam parameters. These N initial enhanced signals correspond to N incoming beam directions.
[0111] When N is 3, the beam parameters may include three groups: beam parameter 1, beam parameter 2, and beam parameter 3. The processor uses the first sector beamformer to enhance input signal 1 and input signal 2, that is, multiplying input signal 1 and input signal 2 by beam parameter 1 to obtain initial enhanced signal 1; multiplying input signal 1 and input signal 2 by beam parameter 2 to obtain initial enhanced signal 2; and multiplying input signal 1 and input signal 2 by beam parameter 3 to obtain initial enhanced signal 3.
[0112] Next, the processor calculates respectively: the ratio of the energy of the initial enhanced signal 1 to the energy of the input signal 1, the ratio of the energy of the initial enhanced signal 2 to the energy of the input signal 1, the ratio of the energy of the initial enhanced signal 3 to the energy of the input signal 1, the ratio of the energy of the initial enhanced signal 1 to the energy of the input signal 2, the ratio of the energy of the initial enhanced signal 2 to the energy of the input signal 2, and the ratio of the energy of the initial enhanced signal 3 to the energy of the input signal 2, to obtain six intensity ratios. The signal energy mentioned in the embodiments of the present application can be expressed by signal intensity or signal power.
[0113] The processor selects the largest maximum intensity ratio from the six intensity ratios, then selects the smallest minimum intensity ratio, and then subtracts the minimum intensity ratio from the maximum intensity ratio to obtain a first fluctuation value. The first fluctuation value can represent the degree of fluctuation of the multiple input signals in multiple incoming directions.
[0114] It should be noted that a larger N indicates a finer granularity for directional divisions in the beam parameters, resulting in a higher precision for the discriminant factor generated based on beam parameters corresponding to more directions. A smaller N indicates a coarser granularity for directional divisions, but this simplifies the signal processing process, conserves resources, and improves processing efficiency. In some embodiments, an N of 8 is more reasonable because it balances signal enhancement with signal processing complexity.
[0115] Next, the processor determines whether the first fluctuation value is greater than or equal to a preset threshold. If the first fluctuation value is greater than or equal to the preset threshold, it means that the fluctuation of the input signal in multiple incoming wave directions is relatively large, which means that there is incoming speech at this time, and the weight vector of the subsequent filter needs to be adjusted to enhance the incoming speech. If the first fluctuation value is less than the preset threshold, it means that the fluctuation of the input signal is relatively small, then there is a high probability that there is no incoming speech at this time, and there is no need to adjust the weight vector of the filter to enhance the incoming speech. At this time, the existing filter weight vector can be kept unchanged. It should be noted that the filter weight vector can be used to perform directional enhancement and other-directional suppression on the input signal.
[0116] The processor may also generate a discriminant factor based on the first fluctuation value. Optionally, when the first fluctuation value is greater than or equal to a preset threshold, the discriminant factor is 1; when the first fluctuation value is greater than or equal to the preset threshold, the discriminant factor is 0. This discriminant factor can be used in the subsequent update process of the filter weight vector.
[0117] In order to illustrate the role of the above-mentioned discriminant factors in determining the filter weight vector, the principle of obtaining the filter weight is first introduced here:
[0118] Taking the case of time t as an example, the vector of the signal output by the microphone array is recorded as: X(t,e jω ), the filter weight vector at time t is recorded as W(t,e jω ), the sound source vector is S(t,e jω ), the noise vector in the environment is N(t,e jω ).
[0119] Among them, t represents the time and jω represents the frequency.
[0120] If the signal is filtered using the filter weight vector at time t, the filter output vector is recorded as Y(t,ejω ).
[0121] Specifically, the filter output vector is the product of the filter weight vector and the vector of the signal output by the microphone array. It can be expressed as:
[0122] Y(t,e jω )=W H (t,e jω )X(t,e jω ). Among them, W H (t,e jω ) is W(t,e jω ) is the transposed matrix of .
[0123] Since X(t,e jω )=A(t,e jω )S(t,e jω )+N(t,e jω ), and substituting this relationship into the formula of the above filter output vector, we can obtain the following relationship:
[0124] Y(t,e jω )=W H (t,e jω )X(t,e jω )=W H (t,e jω )A(t,e jω )S(t,e jω )+W H (t,e jω )N(t,e jω ).
[0125] Where A(t,e jω ) is the propagation function of air, which is used to characterize the attenuation of the signal in the air.
[0126] Then the power of the signal output by the filter can be expressed as the mathematical expectation of the filter output vector and the conjugate transpose of the filter output vector, expressed as the formula:
[0127] E{Y(t,e jω )Y * (t,e jω )}=E{W H (t,e jω )X(t,e jω )X H (t,e jω )W(t,e jω )}.
[0128] The expression can be rewritten as: E{Y(t,e jω )Y* (t,e jω )}=W H (t,e jω )Φ XX W(t,e jω ).
[0129] Among them, Φ XX (t,e jω )=E{X(t,e jω )X H (t,e jω )}, it can be seen that Φ XX (t,e jω ) is related to the signal output by the microphone.
[0130] From the above, it can be seen that the power of the signal output by the filter and the power of the signal output by the microphone (that is, the signal input to the filter) are related to the filter weight vector.
[0131] Next, to minimize the power of the output signal, we need to constrain the vector W H (t,e jω )A(t,e jω ), resulting in a convex optimization problem. When solving this convex optimization problem, the processor can use the Lagrange multiplier method to solve it and obtain a closed-form solution. However, the closed-form solution cannot track changes in the environment. In other words, when the sound direction changes, it cannot be updated in real time. Therefore, the filter weight vector can be updated, and the following weight vector update relationship is obtained:
[0132]
[0133] Among them, μ is the discriminant factor, L(e jω ) is the Lagrangian function, To find the gradient of the function, it means to find the gradient of the Lagrangian function.
[0134] Then, when the first fluctuation value is greater than or equal to the preset threshold, the processor outputs a discrimination factor of 1, and the weight of the filter at the next moment is reduced by Δ W *L(e jω ), needs to be updated. When the first fluctuation value is less than the preset threshold, the processor outputs the discrimination factor as 0, that is, W(t+1,e jω )=W(t,e jω ).
[0135] Optionally, the preset threshold may be a value between 0 and 1, for example, 0.35, 0.4, etc., which is not limited in the embodiment of the present application.
[0136] Next, at the next moment (time t+1), the updated filter weight vector can be used to enhance the input signal. Specifically, the processor performs a Fourier transform (FFT) on input signal 1 to obtain transformed signal 1, and the processor also performs a Fourier transform on input signal 2 to obtain transformed signal 2.
[0137] The processor then enhances the input signal using the updated filter weight vector, specifically by multiplying the updated filter weight vector with the transformed signal 1 and the transformed signal 2, thereby obtaining a frequency domain enhanced signal. The processor may perform an inverse Fourier transform (IFFT) on the frequency domain enhanced signal to obtain an enhanced signal (i.e., the first enhanced signal corresponding to the first sector), and output the signal to the channel selection module.
[0138] During signal enhancement, the processor uses the gain differences of the microphone array for signals from different directions to indicate the update of the filter weight vector. This not only enhances the incoming sound signal, but also suppresses interference signals from other directions. This allows for tracking and enhancing incoming sounds, achieving adaptive directional speech enhancement. Even if the sound source location changes, the signal processing method of this application can still adaptively track and enhance the incoming sound, improving the user experience.
[0139] Optionally, when the first sector beamformer outputs the first enhanced signal, a first signal-to-noise ratio estimation value corresponding to the first enhanced signal may also be output. The first signal-to-noise ratio estimation value is used to characterize the signal quality of the first enhanced signal. A larger first signal-to-noise ratio estimation value indicates better signal quality of the first enhanced signal; a smaller first signal-to-noise ratio estimation value indicates worse signal quality of the first enhanced signal.
[0140] Similarly, other sector beamformers may also output corresponding enhanced signals to the channel selection module. Optionally, each sector beamformer may also output a corresponding signal-to-noise ratio estimation value.
[0141] Next, the process of the processor using the channel selection module to switch the output signal is described.
[0142] When the beamforming module outputs multiple enhanced signals to the channel selection module, the channel selection module can select one or more enhanced signals as output signals and output them to the back end for processing.
[0143] Optionally, in a scenario where only one output signal is needed, the processor may use a channel selection module to select an enhanced signal with the highest estimated signal-to-noise ratio as the output signal.
[0144] Optionally, the processor may also calculate the ratio of the energy of each enhanced signal to the comprehensive energy, and then select the enhanced signal corresponding to the signal with the largest ratio as the output signal. It should be noted that the comprehensive energy is the sum of the energies of all input signals.
[0145] The energy of the signal described here can be the intensity of the signal. For example, the processor can also calculate the intensity ratio of the intensity of each enhanced signal to the comprehensive intensity, and then select the enhanced signal corresponding to the one with the largest intensity ratio as the output signal for output. It should be noted that the comprehensive intensity is the sum of the intensities of all input signals. Optionally, the processor can also calculate the intensity ratio of the intensity of each enhanced signal to the comprehensive intensity, and also calculate the intensity ratio of each input signal to the comprehensive intensity, and then select the enhanced signal corresponding to the one with the largest intensity ratio as the output signal for output. For example, see Figure 7 shown.
[0146] When the processor outputs one output signal, that output signal can be used in voice scenarios, such as hearing-assisted scenarios, voice calls, and voice conferencing scenarios. In hearing-assisted scenarios, this method can enhance the voice from the direction of the sound source, allowing users to hear the sound more clearly and facilitate conversations. In voice conferencing scenarios, this method can enhance the voice from the direction of the sound source, improve voice clarity, and reduce interference in noisy environments, thereby improving the user's conference experience.
[0147] Figure 8 The figure shows a schematic diagram of the effect of voice adaptive directional enhancement in a conference scenario. Figure 8 As shown, a microphone array is provided on the large-screen device. Figure 8 The figure shows the directional diagram of speech enhancement in different sectors for different directions. When participants from different directions speak, the large-screen device can use the signal processing method provided in the embodiment of this application to perform directional speech enhancement in the direction of the speaking participant. If the speaker's position changes, the large-screen device can also adaptively track the incoming sound, thereby improving the overall meeting experience.
[0148] If the signal-to-noise ratio of the input signal is relatively low and the estimated signal-to-noise ratio of the enhanced signal is relatively high, the processor can directly output the enhanced signal, for example Figure 9 In figure a, the processor outputs enhanced signal 1, enhanced signal 2 and enhanced signal 3.
[0149] If the interference signal is very weak, and the estimated SNR of the enhanced signal is low, it is likely to be lower than the SNR of the original input signal. In this case, the processor can select as the output signal only those enhanced signals whose SNR estimates exceed a preset SNR threshold, while also selecting the original input signal. This selection ensures that the output signal is high-quality, with a high SNR, thereby improving the quality of the speech signal.
[0150] Specifically, the processor may select, through the channel selection module, from the multiple enhanced signals, an enhanced signal having an estimated signal-to-noise ratio greater than or equal to a preset signal-to-noise ratio threshold. If the estimated signal-to-noise ratios of the multiple enhanced signals are all greater than or equal to the preset signal-to-noise ratio threshold, the processor may directly output the multiple enhanced signals.
[0151] Optionally, the preset signal-to-noise ratio threshold may be a value set based on experience. Optionally, when the signal-to-noise ratio estimate is a normalized value, the preset signal-to-noise ratio threshold may be a value such as 0.2 or 0.3, which is not limited in the present embodiment.
[0152] If there is an enhanced signal among the multiple enhanced signals whose estimated signal-to-noise ratio is less than a preset signal-to-noise ratio threshold, it indicates that such an enhanced signal may not effectively enhance speech. Therefore, the processor can select as output the enhanced signal whose estimated signal-to-noise ratio is greater than or equal to the preset signal-to-noise ratio threshold, while also selecting one or more channels from the original input signals for output.
[0153] This ensures that voice enhancement can be achieved for different sectors in scenarios with more interference. In scenarios with less interference, when the signal quality of some enhanced signals may not be as good as the original input signals, directly outputting the input signals will enable subsequent processing modules to obtain higher quality output signals, making subsequent processing more accurate.
[0154] For example, when a user inputs voice close to the microphone array of a mobile phone, the signal quality of the original input signal will be relatively high, that is, the signal-to-noise ratio will be high; and in the enhanced signal, the signal-to-noise ratio estimate of the enhanced signal corresponding to the sector at the top of the mobile phone (the opposite side of the microphone array) is lower than the signal-to-noise ratio of the original input signal. In this case, there is no need to output the enhanced signal corresponding to the sector at the top of the mobile phone, but one or more channels of the original input signal can be output as the output signal.
[0155] For example, see Figure 9As shown in Figure b, the input signals are input signal 1 and input signal 2. Among enhanced signal 1, enhanced signal 2, and enhanced signal 3, the estimated signal-to-noise ratio of enhanced signal 1 is greater than the preset signal-to-noise ratio threshold, while the estimated signal-to-noise ratios of enhanced signal 2 and enhanced signal 3 are both less than the preset signal-to-noise ratio threshold. The processor can then output enhanced signal 1, input signal 1, and input signal 2.
[0156] If, among enhanced signals 1, 2, and 3, the estimated signal-to-noise ratio values of enhanced signal 1 and enhanced signal 2 are greater than a preset signal-to-noise ratio threshold, and the estimated signal-to-noise ratio value of enhanced signal 3 is less than the preset signal-to-noise ratio threshold, the processor may output enhanced signal 1 and enhanced signal 2 while also selecting one of input signals 1 and 2 for output.
[0157] Optionally, when the processor selects an output from the original input signal, it can arbitrarily select one or more paths according to the number of paths allowed for output, or it can select in sequence from input signal 1 by default, or it can select the one with a high signal-to-noise ratio as the output signal. This embodiment of the present application does not limit this.
[0158] Optionally, in order to avoid the situation where the signal quality of the output enhanced signal is poor due to interference, the processing module may also continuously collect the instantaneous signal-to-noise ratio estimation values at multiple moments for judgment when determining whether the signal-to-noise ratio estimation value of an enhanced signal is greater than or equal to a preset signal-to-noise ratio threshold.
[0159] For example, the processor records M instantaneous signal-to-noise ratio (SNR) estimates for time t and the preceding M consecutive time points. At time t, the processor determines whether the instantaneous SNR estimates for these M time points are all greater than or equal to a preset SNR threshold. If the instantaneous SNR estimates for these M time points are all greater than or equal to the preset SNR threshold, then the SNR estimate for the enhanced signal is determined to be greater than or equal to the preset SNR threshold. If any of the instantaneous SNR estimates for these M time points is less than the preset SNR threshold, then the SNR estimate for the enhanced signal is determined to be neither greater than nor equal to the preset SNR threshold.
[0160] Optionally, when the processor outputs multiple output signals, the number of output signals is equal to the number of sectors. For example, when there are three sectors, three output signals are output.
[0161] For example Figure 10 As shown, the input signals are input signal 1 and input signal 2. Among enhanced signal 1, enhanced signal 2, and enhanced signal 3, the M instantaneous signal-to-noise ratio estimates of enhanced signal 1 at M consecutive moments are all greater than the preset signal-to-noise ratio threshold, while the M instantaneous signal-to-noise ratio estimates of enhanced signal 2 and enhanced signal 3 at M consecutive moments are all less than the preset signal-to-noise ratio threshold. The processor can then output enhanced signal 1, input signal 1, and input signal 2.
[0162] Optionally, when the estimated signal-to-noise ratio value of the enhanced signal is greater than a preset signal-to-noise ratio threshold, the quality of the enhanced signal may be higher than the quality of the original input signal; when the estimated signal-to-noise ratio value of the enhanced signal is less than the preset signal-to-noise ratio threshold, the quality of the enhanced signal may be worse than the quality of the original input signal; and when the estimated signal-to-noise ratio value of the enhanced signal is equal to or approximately equal to the preset signal-to-noise ratio threshold, the quality of the enhanced signal may be equivalent to the quality of the original input signal.
[0163] When the processor outputs multiple output signals, it can be used in some non-speech scenarios, such as speech recognition and voice wake-up. In the voice wake-up scenario, this method can enhance speech in a targeted manner and improve the wake-up rate in noisy environments.
[0164] Figure 11 This is a flowchart of a signal processing method provided in an embodiment of the present application. The method is applied to an electronic device including multiple microphones. The method includes:
[0165] S1101 . Acquire multiple input signals, where the multiple input signals are signals collected by multiple microphones, and the multiple input signals correspond one-to-one to the multiple microphones.
[0166] Each microphone can output the collected sound signal as one output signal, and a microphone array containing multiple microphones can output multiple output signals.
[0167] S1102. Use multiple sets of beam parameters to enhance the multiple input signals to obtain multiple enhanced signals. The multiple sets of beam parameters and the multiple enhanced signals correspond one-to-one to the multiple sectors. The beam parameters are associated with the frequency of the input signal and the direction of arrival of the input signal. The multiple sectors are multiple spatial areas obtained by dividing the space with multiple microphones as the center.
[0168] For a detailed introduction to beam parameters and sectors, please refer to the previous section. The processor can use any set of beam parameters to enhance multiple input signals, for example, by multiplying them by corresponding filter weight vectors to produce multiple enhanced signals. It should be noted that each sector corresponds to a set of beam parameters. Each set of beam parameters can output an enhanced signal corresponding to the corresponding sector.
[0169] S1103 : Determine and output an output signal based on the multiple enhanced signals, where the output signal is one or more of the multiple enhanced signals.
[0170] The processor determines the signal with the best signal quality from the multiple enhanced signals and outputs it as the output signal; the processor can also directly output the multiple enhanced signals.
[0171] When using a microphone array to collect multiple input signals, the sound signals received from the same sound source are different due to the different positions of different microphones in the microphone array. Therefore, the microphone array can obtain more spatial information when collecting sound signals, which is convenient for directional enhancement of speech.
[0172] This method can perform directional enhancement on the signals collected by multiple microphones. In this way, even in a noisy environment, it can receive sound signals from the direction of the sound with a higher gain and suppress interference signals from other directions, thereby improving the signal-to-noise ratio of useful signals and thus enhancing the user experience.
[0173] In some embodiments, S1102 uses multiple sets of beam parameters to perform enhancement processing on multiple input signals to obtain multiple enhanced signals, including: using a first beam parameter to determine a first fluctuation value of the multiple input signals, the first fluctuation value characterizing the degree of fluctuation of the multiple input signals in multiple incoming wave directions, the first beam parameter is any one of the multiple sets of beam parameters, and the first beam parameter corresponds to the first sector of the multiple sectors; if the first fluctuation value is greater than or equal to a preset threshold, then the updated filter weight vector at the next moment is adjusted according to the current filter weight vector, and the updated filter weight vector is different from the current filter weight vector; if the first fluctuation value is less than the preset threshold, then the updated filter weight vector at the next moment is determined to be the current filter weight vector; according to the updated filter weight vector, the multiple input signals are enhanced to obtain a first enhanced signal to obtain the first enhanced signal, which is one of the multiple enhanced signals.
[0174] The above-mentioned first fluctuation value can characterize the fluctuation degree of the input signal in multiple incoming wave directions. Therefore, the filter weight vector can be adjusted when the fluctuation degree is large, and when the fluctuation degree is small, the filter weight vector can be kept unchanged, thereby realizing real-time adaptive update of the filter weight vector and improving the accuracy of voice directionality.
[0175] In some embodiments, the first fluctuation value is the difference between the maximum intensity ratio and the minimum intensity ratio, the maximum intensity ratio is the largest of multiple intensity ratios, and the minimum intensity ratio is the smallest of multiple intensity ratios; the multiple signal strengths are the ratios of the signal strength of each directional enhanced signal in the multiple directional enhanced signals to the signal strength of each input signal in the multiple input signals; the multiple directional enhanced signals are enhanced signals obtained by performing beamforming processing on the multiple input signals in multiple incoming directions using the first beam parameters, and the number of multiple intensity ratios is the product of the number of multiple directional enhanced signals and the number of multiple input signals.
[0176] In some embodiments, the first beam parameters include multiple groups of preset gains, which correspond one-to-one to multiple wave directions. The first beam parameters are used to determine the first fluctuation value of the multi-channel input signal, including: using the multiple groups of preset gains to perform beamforming processing on the multi-channel input signal in multiple wave directions to obtain multiple direction-enhanced signals, and the multiple direction-enhanced signals correspond one-to-one to the multiple wave directions; determining a first intensity ratio of the signal strength of the first direction-enhanced signal to the signal strength of the first input signal, the first direction-enhanced signal is any one of the multiple direction-enhanced signals, the first input signal is any one of the multiple input signals, and the first intensity ratio is one of the multiple intensity ratios; determining the difference between the maximum intensity ratio and the minimum intensity ratio among the multiple intensity ratios as the first fluctuation value.
[0177] Beamforming involves multiplying multiple input signals by their corresponding preset gains. For each sector, the processor calculates the ratio of the strength of each initial boosted signal to the strength of each input signal, recording this as the intensity ratio. The processor obtains multiple intensity ratios, selects the largest and smallest values, and calculates the difference between them. This difference quantifies the fluctuation in the intensity ratios.
[0178] In some embodiments, the output signal is the one with the strongest signal strength among the multiple enhanced signals.
[0179] The processor may directly compare the signal strengths of the enhanced signals and select the one with the strongest signal strength from the multiple enhanced signals as the output signal, thereby ensuring that the signal quality of the output signal is the output signal of the optionally best quality.
[0180] In some embodiments, the output signal is the one with the strongest signal strength among the multiple enhanced signals and the multiple input signals.
[0181] The processor may select the one with the strongest signal strength from the multiple enhanced signals and the multiple input signals as the output signal, ensuring that the signal quality of the output signal is the output signal of the optionally best quality.
[0182] In some embodiments, based on the multiple enhanced signals, an output signal is determined and outputted, including: determining the signal strength of each enhanced signal in the multiple enhanced signals, and multiple second intensity ratios of the comprehensive intensity, where the comprehensive intensity is the sum of the intensities of multiple input signals, and the multiple second intensity ratios correspond one-to-one to the multiple enhanced signals; screening out the largest second intensity ratio from the multiple second intensity ratios; determining the enhanced signal corresponding to the largest second intensity ratio as the output signal and outputting it.
[0183] The processor may also use the second intensity ratio to quantify the signal quality of the enhanced signal. The processor selects the enhanced signal corresponding to the one with the largest second intensity ratio as the output signal to ensure the signal quality of the output signal.
[0184] In some embodiments, the electronic device is in a voice scenario, which includes one of a call scenario, a hearing assistance scenario, and a voice conference scenario.
[0185] When the processor outputs one output signal, that output signal can be used in voice scenarios, such as hearing-assisted scenarios, voice calls, and voice conferencing scenarios. In hearing-assisted scenarios, this method can enhance the voice from the direction of the sound source, allowing users to hear the sound more clearly and facilitate conversations. In voice conferencing scenarios, this method can enhance the voice from the direction of the sound source, improve voice clarity, and reduce interference in noisy environments, thereby improving the user's conference experience.
[0186] In some embodiments, the method further comprises: outputting a plurality of signal-to-noise ratio estimation values, the plurality of signal-to-noise ratio estimation values corresponding one-to-one to the plurality of sectors.
[0187] The signal-to-noise ratio estimation value can quantify the signal quality of the enhanced signal, making it easier to selectively output the enhanced signal with good quality, thereby improving the effect of speech processing.
[0188] In some embodiments, determining and outputting an output signal based on multiple enhanced signals includes: determining at least one signal to be output from the multiple enhanced signals, wherein a signal-to-noise ratio estimate corresponding to each signal in the at least one signal to be output is greater than or equal to a preset signal-to-noise ratio threshold; if the number of the at least one signal to be output is less than the number of the multiple enhanced signals, then outputting the at least one signal to be output and part or all of the multiple input signals as output signals.
[0189] If there is an enhanced signal among the multiple enhanced signals whose estimated signal-to-noise ratio is less than a preset signal-to-noise ratio threshold, it indicates that such an enhanced signal may not effectively enhance speech. Therefore, the processor can select as output the enhanced signal whose estimated signal-to-noise ratio is greater than or equal to the preset signal-to-noise ratio threshold, while also selecting one or more channels from the original input signals for output.
[0190] This ensures that voice enhancement can be achieved for different sectors in scenarios with more interference. In scenarios with less interference, when the signal quality of some enhanced signals may not be as good as the original input signals, directly outputting the input signals will enable subsequent processing modules to obtain higher quality output signals, making subsequent processing more accurate.
[0191] In some embodiments, the number of output signal paths is the same as the number of the plurality of sectors.
[0192] In some embodiments, the signal-to-noise ratio estimate corresponding to each signal in at least one signal to be output is greater than or equal to a preset signal-to-noise ratio threshold, including that multiple instantaneous signal-to-noise ratio estimation values corresponding to the signal-to-noise ratio estimate corresponding to each signal in at least one signal to be output at multiple consecutive moments are all greater than the preset signal-to-noise ratio threshold.
[0193] To prevent interference from causing poor signal quality in the output enhanced signal, the processing module can continuously collect multiple instantaneous SNR estimates when determining whether the SNR estimate of an enhanced signal is greater than or equal to a preset SNR threshold. This approach prevents interference from causing poor signal quality in the output enhanced signal, improves the output signal quality, and enhances speech processing effectiveness.
[0194] In some embodiments, the electronic device is in a non-voice scenario, where the non-voice scenario includes one or more of: voice wake-up and voice recognition.
[0195] When the processor outputs multiple output signals, it can be used in some non-speech scenarios, such as speech recognition and voice wake-up. In the voice wake-up scenario, this method can enhance speech in a targeted manner and improve the wake-up rate in noisy environments.
[0196] The above describes in detail an example of the method provided by the present application. It is understandable that, in order to implement the above functions, the corresponding device includes a hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0197] The present application can divide the signal processing device into functional modules according to the above method examples. For example, each function can be divided into various functional modules, or two or more functions can be integrated into one module. The above integrated modules can be implemented in the form of hardware or software functional modules. It should be noted that the division of modules in this application is schematic and is only a logical functional division. In actual implementation, there may be other division methods.
[0198] Figure 12 FIG2 is a schematic diagram showing the structure of a signal processing device provided by the present application. The device 1200 includes a beamforming module 1201 and a channel selection module 1202 .
[0199] The beamforming module 1201 is used to obtain multiple input signals and enhance the multiple input signals using multiple sets of beam parameters to obtain multiple enhanced signals. The multiple input signals are signals collected by multiple microphones, and the multiple input signals correspond one-to-one to the multiple microphones. The multiple sets of beam parameters and the multiple enhanced signals correspond one-to-one to the multiple sectors. The beam parameters are associated with the frequency and arrival direction of the input signals. The multiple sectors are multiple spatial regions obtained by dividing the space with the multiple microphones as the center.
[0200] The channel selection module 1202 is configured to determine and output an output signal based on the multiple enhanced signals, where the output signal is one or more of the multiple enhanced signals.
[0201] Optionally, the beamforming module 1201 is specifically used to determine a first fluctuation value of multiple input signals using a first beam parameter, where the first fluctuation value represents the degree of fluctuation of the multiple input signals in multiple incoming wave directions, and the first beam parameter is any group of multiple groups of beam parameters, and the first beam parameter corresponds to a first sector of multiple sectors; if the first fluctuation value is greater than or equal to a preset threshold, the updated filter weight vector at the next moment is adjusted according to the current filter weight vector, and the updated filter weight vector is different from the current filter weight vector; if the first fluctuation value is less than the preset threshold, the updated filter weight vector at the next moment is determined to be the current filter weight vector; according to the updated filter weight vector, the multiple input signals are enhanced to obtain a first enhanced signal, which is one of the multiple enhanced signals.
[0202] Optionally, the first fluctuation value is the difference between the maximum intensity ratio and the minimum intensity ratio, the maximum intensity ratio is the largest one among multiple intensity ratios, and the minimum intensity ratio is the smallest one among multiple intensity ratios; the multiple signal strengths are the ratios of the signal strength of each directional enhanced signal in the multiple directional enhanced signals to the signal strength of each input signal in the multiple input signals; the multiple directional enhanced signals are enhanced signals obtained by performing beamforming processing on the multiple input signals in multiple incoming directions using the first beam parameters, and the number of multiple intensity ratios is the product of the number of multiple directional enhanced signals and the number of multiple input signals.
[0203] Optionally, the first beam parameters include multiple groups of preset gains, which correspond one-to-one to multiple wave directions. The beamforming module 1201 is specifically used to use multiple groups of preset gains to perform beamforming processing on multiple input signals in multiple wave directions to obtain multiple direction-enhanced signals, and the multiple direction-enhanced signals correspond one-to-one to multiple wave directions; determine a first intensity ratio of the signal strength of the first direction-enhanced signal to the signal strength of the first input signal, the first direction-enhanced signal is any one of the multiple direction-enhanced signals, the first input signal is any one of the multiple input signals, and the first intensity ratio is one of the multiple intensity ratios; determine the difference between the maximum intensity ratio and the minimum intensity ratio among the multiple intensity ratios as the first fluctuation value.
[0204] Optionally, the output signal is the one with the strongest signal strength among the multiple enhanced signals.
[0205] Optionally, the output signal is the one with the strongest signal strength among the multiple enhanced signals and the multiple input signals.
[0206] Optionally, the beamforming module 1201 is specifically used to determine the signal strength of each enhanced signal among multiple enhanced signals, multiple second intensity ratios of the comprehensive intensity, the comprehensive intensity is the sum of the intensities of multiple input signals, and the multiple second intensity ratios correspond one-to-one to the multiple enhanced signals; screen out the largest second intensity ratio from the multiple second intensity ratios; determine the enhanced signal corresponding to the largest second intensity ratio as the output signal and output it.
[0207] Optionally, the electronic device is in a voice scenario, where the voice scenario includes one of a call scenario, a hearing assistance scenario, and a voice conference scenario.
[0208] Optionally, the beamforming module 1201 is further configured to output multiple signal-to-noise ratio estimation values, where the multiple signal-to-noise ratio estimation values correspond one-to-one to the multiple sectors.
[0209] Optionally, the channel selection module 1202 is specifically configured to determine at least one signal to be output from a plurality of enhanced signals, where a signal-to-noise ratio estimate corresponding to each signal in the at least one signal to be output is greater than or equal to a preset signal-to-noise ratio threshold; and if the number of the at least one signal to be output is less than the number of the plurality of enhanced signals, the at least one signal to be output and part or all of the multiple input signals are output as output signals.
[0210] Optionally, the number of output signal paths is the same as the number of the multiple sectors.
[0211] Optionally, the signal-to-noise ratio estimate corresponding to each signal in the at least one signal to be output is greater than or equal to a preset signal-to-noise ratio threshold, including that multiple instantaneous signal-to-noise ratio estimation values corresponding to the signal-to-noise ratio estimate corresponding to each signal in the at least one signal to be output at multiple consecutive moments are all greater than the preset signal-to-noise ratio threshold.
[0212] Optionally, the electronic device is in a non-voice scenario, and the non-voice scenario includes: one or more of voice wake-up and voice recognition.
[0213] The specific manner in which the device 1200 executes the signal processing method and the beneficial effects produced can be found in the relevant description of the method embodiment, which will not be repeated here.
[0214] The embodiment of the present application also provides an electronic device, including the above-mentioned processor. The electronic device provided by this embodiment can be Figure 1 The terminal device 100 shown is configured to execute the aforementioned signal processing method. If integrated, the terminal device may include a processor, a storage module, and a communication module. The processor may be used to control and manage the terminal device's operations. For example, it may be used to support the terminal device in executing steps performed by the display unit, detection unit, and processing unit. The storage module may be used to support the terminal device in executing and storing program code and data. The communication module may be used to support communication between the terminal device and other devices.
[0215] The processor may be a processor or a controller. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor (DSP) and a microprocessor, and so on. The storage module may be a memory. The communication module may specifically be a device that interacts with other terminal devices, such as a radio frequency circuit, a Bluetooth chip, or a Wi-Fi chip.
[0216] In one embodiment, when the processor is a processor and the storage module is a memory, the terminal device involved in this embodiment may be a terminal device having Figure 1 Device with the structure shown.
[0217] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the processor executes the signal processing method described in any of the above embodiments.
[0218] The embodiment of the present application further provides a computer program product. When the computer program product is run on a computer, the computer is caused to execute the above-mentioned related steps to implement the signal processing method in the above-mentioned embodiment.
[0219] Among them, the electronic device, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0220] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, the replaced units may or may not be physically separated, and the components displayed as units may be one physical unit or multiple physical units, that is, they may be located in one place, or they may be distributed in multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0221] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0222] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0223] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A signal processing method, characterized in that: Applicable to electronic devices, wherein the electronic devices include multiple microphones, including: Acquire multiple input signals, where the multiple input signals are signals collected by multiple microphones, and the multiple input signals correspond one-to-one to the multiple microphones; performing enhancement processing on the multiple input signals using multiple sets of beam parameters to obtain multiple enhanced signals, wherein the multiple sets of beam parameters and the multiple enhanced signals correspond one-to-one to multiple sectors, the beam parameters are associated with the frequency of the input signal and the direction of arrival of the input signal, and the multiple sectors are multiple spatial regions obtained by dividing a space centered on the multiple microphones; An output signal is determined based on the multiple enhanced signals and outputted, where the output signal is one or more of the multiple enhanced signals.
2. The method according to claim 1, characterized in that The enhancing process of the multiple input signals by using multiple sets of beam parameters to obtain multiple enhanced signals includes: Determining a first fluctuation value of the multi-path input signal using a first beam parameter, where the first fluctuation value represents a degree of fluctuation of the multi-path input signal in multiple incoming directions, the first beam parameter being any one of the multiple groups of beam parameters, and the first beam parameter corresponding to a first sector of the multiple sectors; If the first fluctuation value is greater than or equal to a preset threshold, adjusting the current filter weight vector to an updated filter weight vector at the next moment, where the updated filter weight vector is different from the current filter weight vector; If the first fluctuation value is less than the preset threshold, determining the updated filter weight vector at the next moment to be the current filter weight vector; The multi-channel input signals are enhanced according to the updated filter weight vector to obtain a first enhanced signal, where the first enhanced signal is one of the multiple enhanced signals.
3. The method according to claim 2, characterized in that The first fluctuation value is a difference between a maximum intensity ratio and a minimum intensity ratio, the maximum intensity ratio being the largest of a plurality of intensity ratios, and the minimum intensity ratio being the smallest of the plurality of intensity ratios; The multiple signal strengths are ratios of the signal strength of each directional enhanced signal in the multiple directional enhanced signals to the signal strength of each input signal in the multiple input signals; The multiple directional enhancement signals are enhanced signals obtained by performing beamforming processing on the multiple input signals in multiple directions using the first beam parameters, and the number of the multiple intensity ratios is the product of the number of the multiple directional enhancement signals and the number of the multiple input signals.
4. The method according to claim 3, characterized in that The first beam parameters include multiple groups of preset gains, the multiple groups of preset gains correspond one-to-one to multiple incoming wave directions, and the first fluctuation value of the multiple input signals is determined by using the first beam parameters, including: Performing beamforming processing on the multiple input signals in multiple wave directions using the multiple sets of preset gains to obtain the multiple directional enhancement signals, where the multiple directional enhancement signals correspond one-to-one to the multiple wave directions; determining a first strength ratio of a signal strength of a first direction-enhanced signal to a signal strength of a first input signal, where the first direction-enhanced signal is any one of the multiple direction-enhanced signals, the first input signal is any one of the multiple input signals, and the first strength ratio is one of the multiple strength ratios; A difference between the maximum intensity ratio and the minimum intensity ratio among the plurality of intensity ratios is determined as the first fluctuation value.
5. The method according to any one of claims 1 to 4, characterized in that The output signal is the one with the strongest signal strength among the multiple enhanced signals.
6. The method according to any one of claims 1 to 4, characterized in that The output signal is the one with the strongest signal strength among the multiple enhanced signals and the multiple input signals.
7. The method according to claim 6, characterized in that The step of determining and outputting an output signal based on the plurality of enhanced signals comprises: Determine a signal strength of each enhanced signal in the plurality of enhanced signals and a plurality of second strength ratios of a comprehensive strength, wherein the comprehensive strength is a sum of strengths of the multiple input signals, and the plurality of second strength ratios correspond one-to-one to the plurality of enhanced signals; Filtering out the largest second intensity ratio from the plurality of second intensity ratios; An enhanced signal corresponding to the maximum second intensity ratio is determined as the output signal and outputted.
8. The method according to any one of claims 1 to 7, characterized in that The electronic device is in a voice scenario, which includes one of a call scenario, a hearing assistance scenario, and a voice conference scenario.
9. The method according to any one of claims 1 to 4, characterized in that The method further comprises: A plurality of signal-to-noise ratio estimation values are output, wherein the plurality of signal-to-noise ratio estimation values correspond one-to-one to the plurality of sectors.
10. The method according to claim 9, characterized in that The step of determining and outputting an output signal based on the plurality of enhanced signals comprises: Determine at least one to-be-output signal from the multiple enhanced signals, wherein a signal-to-noise ratio estimate corresponding to each signal in the at least one to-be-output signal is greater than or equal to a preset signal-to-noise ratio threshold; If the number of the at least one signal to be output is smaller than the number of the multiple enhanced signals, the at least one signal to be output and part or all of the multiple input signals are output as the output signal.
11. The method according to claim 10, characterized in that The number of the output signal paths is the same as the number of the multiple sectors.
12. The method according to claim 10 or 11, characterized in that The signal-to-noise ratio estimate corresponding to each signal in the at least one signal to be output is greater than or equal to a preset signal-to-noise ratio threshold, including that multiple instantaneous signal-to-noise ratio estimates corresponding to the signal-to-noise ratio estimate corresponding to each signal in the at least one signal to be output at multiple consecutive moments are all greater than the preset signal-to-noise ratio threshold.
13. The method according to any one of claims 1 to 4 and 9 to 12, characterized in that The electronic device is in a non-voice scenario, where the non-voice scenario includes one or more of voice wake-up and voice recognition.
14. An electronic device, characterized in that: The device comprises a microphone array, and further comprises: a processor, a memory and an interface; The microphone array includes a plurality of microphones, and the plurality of microphones are used to collect multi-channel input signals; The processor, the memory, and the interface cooperate with each other so that the electronic device executes the method according to any one of claims 1 to 13.
15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Directional speech enhancement method based on small microphone array
CN101587712A
Method for enhancing microphone array voice based on combined inhibition
CN102509552A
Voice awakening method and device, equipment and medium
CN109949810A
Directional voice enhancement method and system
CN112017681A
Voice signal processing method and device
CN113038318A