Signal processing method, electronic device and computer readable storage medium
By using multi-microphone arrays and beamforming technology, the voice signal is enhanced in a directional manner, solving the problem of voice acquisition and recognition in noisy environments and improving the signal-to-noise ratio and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2024-09-29
- Publication Date
- 2026-07-14
AI Technical Summary
In noisy environments, electronic devices struggle to effectively capture and recognize voice signals, resulting in low voice intelligibility and negatively impacting user experience.
By employing a multi-microphone array and beamforming technology, multiple input signals are acquired, and beam parameters are used for directional enhancement processing to select the output signal with the best signal quality, suppress interference signals, and improve the signal-to-noise ratio.
In noisy environments, this technology improves the reception gain of voice signals, suppresses interference signals, enhances the signal-to-noise ratio and voice intelligibility, and improves the user experience.
Smart Images

Figure CN120469284B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of signal processing technology, specifically to a signal processing method, electronic device, and computer-readable storage medium. Background Technology
[0002] Electronic devices are typically equipped with microphones to collect sound signals. These microphones can then be used to transmit the sound signals for voice calls or to recognize the sound signals for actions such as voice wake-up.
[0003] However, in some scenarios, such as during voice conferences, a noisy environment can reduce the signal-to-noise ratio of the participants' audio signals. In such cases, the devices used in the voice conference may be unable to properly capture the speaker's voice, resulting in inaccurate audio signal recognition or transmission, leading to low speech intelligibility and negatively impacting the user experience. Summary of the Invention
[0004] This application provides a signal processing method, apparatus, chip, electronic device, computer-readable storage medium, and computer program product capable of targeted speech enhancement.
[0005] Firstly, a signal processing method is provided, applied to an electronic device including multiple microphones. The method includes: acquiring multiple input signals, where each input signal is a signal collected by multiple microphones, and each input signal corresponds one-to-one with a microphone; enhancing the multiple input signals using multiple sets of beam parameters to obtain multiple enhanced signals, where each set of beam parameters, multiple enhanced signals, and multiple sectors correspond one-to-one; the beam parameters are related to the frequency and direction of arrival of the input signals, and the multiple sectors are spatial regions divided around the multiple microphones; and determining and outputting an output signal based on the multiple enhanced signals, where the output signal is one or more of the multiple enhanced signals. Each microphone can output one output signal from the acquired sound signal, and a microphone array containing multiple microphones can output multiple output signals. Detailed descriptions of beam parameters and sectors can be found in the relevant descriptions within the text. The processor can use any set of beam parameters to enhance the multiple input signals, for example, by multiplying by the corresponding filter weight vector to obtain multiple enhanced signals. It should be noted that each sector corresponds to a set of beam parameters. Each set of beam parameters can output a corresponding enhanced signal for a specific sector. The processor selects the signal with the best quality from multiple enhanced signals and outputs it; the processor can also directly output multiple enhanced signals.
[0006] In the method of acquiring multiple input signals using a microphone array, the sound signals received from the same sound source are different due to the different positions of the microphones in the microphone array. Therefore, the microphone array can acquire more spatial information when acquiring sound signals, which is convenient for directional enhancement of speech.
[0007] This method can perform directional enhancement on signals collected by multiple microphones, so that even in noisy environments, it can receive sound signals from the direction of sound with high gain and suppress interference signals from other directions, thereby improving the signal-to-noise ratio of useful signals and thus enhancing the user experience.
[0008] In some possible implementations, multiple sets of beam parameters are used to enhance multiple input signals to obtain multiple enhanced signals. This includes: determining a first fluctuation value of the multiple input signals using a first beam parameter, where the first fluctuation value characterizes the fluctuation degree of the multiple input signals in multiple incoming wave directions; the first beam parameter is any one of multiple sets of beam parameters, and the first beam parameter corresponds to the first sector among multiple sectors; if the first fluctuation value is greater than or equal to a preset threshold, the current filter weight vector is adjusted to the updated filter weight vector for the next time step, and the updated filter weight vector is different from the current filter weight vector; if the first fluctuation value is less than the preset threshold, the updated filter weight vector for the next time step is determined to be the current filter weight vector; and enhancing the multiple input signals according to the updated filter weight vector to obtain a first enhanced signal, which is one of the multiple enhanced signals.
[0009] The aforementioned first fluctuation value can characterize the degree of fluctuation of the input signal in multiple incoming wave directions. Therefore, the filter weight vector can be adjusted when the fluctuation degree is large, and the filter weight vector can be kept unchanged when the fluctuation degree is small, thereby realizing real-time adaptive updating of the filter weight vector and improving the accuracy of speech direction finding.
[0010] In some possible implementations, the first fluctuation value is the difference between the maximum intensity ratio and the minimum intensity ratio, the maximum intensity ratio is the largest of the multiple intensity ratios, and the minimum intensity ratio is the smallest of the multiple intensity ratios; the multiple signal intensities are the ratios of the signal intensity of each directional enhancement signal among the multiple directional enhancement signals to the signal intensity of each input signal among the multiple input signals; the multiple directional enhancement signals are enhancement signals obtained by beamforming multiple input signals in multiple directions of arrival using the first beam parameters, and the number of multiple intensity ratios is the product of the number of multiple directional enhancement signals and the number of multiple input signals.
[0011] In some possible implementations, the first beam parameters include multiple sets of preset gains, each set of preset gains corresponding to multiple incoming wave directions. Determining the first fluctuation value of the multiple input signals using the first beam parameters includes: performing beamforming processing on the multiple input signals using multiple sets of preset gains to obtain multiple directional enhanced signals, each directional enhanced signal corresponding to a multiple incoming wave direction; determining a first intensity ratio between the signal intensity of the first directional enhanced signal and the signal intensity of the first input signal, wherein the first directional enhanced signal is any one of the multiple directional enhanced signals, the first input signal is any one of the multiple input signals, and the first intensity ratio is one of the multiple intensity ratios; and determining the difference between the maximum and minimum intensity ratios among the multiple intensity ratios as the first fluctuation value.
[0012] In some possible implementations, the output signal is the one with the strongest signal strength among multiple enhanced signals.
[0013] For each sector, the processor can calculate the ratio of the intensity of each initial boost signal to that of each input signal, denoted as the intensity ratio. The processor can obtain multiple intensity ratios, then select the largest and smallest one and calculate the difference between them. This difference quantifies the degree of fluctuation among the multiple intensity ratios.
[0014] In some possible implementations, the output signal is the one with the strongest signal strength among multiple enhanced signals.
[0015] The processor can directly compare the signal strength of the enhanced signals and select the one with the strongest signal strength from multiple enhanced signals as the output signal, ensuring that the output signal has the best possible signal quality.
[0016] In some possible implementations, the output signal is the one with the strongest signal strength from multiple enhancement signals and multiple input signals.
[0017] The processor can select the strongest signal from multiple enhancement signals and multiple input signals as the output signal, ensuring that the output signal has the best possible quality.
[0018] In some possible implementations, determining and outputting an output signal based on the plurality of enhancement signals includes: determining the signal strength of each enhancement signal among the plurality of enhancement signals, comparing it with a plurality of second intensity ratios of the combined strength, wherein the combined strength is the sum of the strengths of the multiple input signals, and the plurality of second intensity ratios correspond one-to-one with the plurality of enhancement signals; selecting the largest second intensity ratio from the plurality of second intensity ratios; and determining and outputting the enhancement signal corresponding to the largest second intensity ratio.
[0019] The processor can also use a second intensity ratio to quantify the signal quality of the enhanced signal. The processor selects the enhanced signal with the largest second intensity ratio as the output signal to ensure the signal quality of the output signal.
[0020] In some possible implementations, the electronic device is in a voice scenario, which includes one of the following: a call scenario, a hearing aid scenario, and a voice conference scenario.
[0021] When the processor outputs one signal, this signal can be used in voice scenarios, such as hearing aid scenarios, voice calls, and voice conferences. In hearing aid scenarios, this method enhances the speech from the direction of the sound source, allowing users to hear the sound more clearly and facilitating conversations. In voice conference scenarios, this method enhances the speech from the direction of the sound source, improving speech clarity and reducing interference in noisy environments, thereby enhancing the user's conference experience.
[0022] In some possible implementations, the method also includes: outputting multiple signal-to-noise ratio (SNR) estimates, with each SNR estimate corresponding to a specific sector.
[0023] The signal-to-noise ratio (SNR) estimate can quantify the signal quality of the enhanced signal, making it easier to select the best enhanced signal and improve the performance of speech processing.
[0024] In some possible implementations, determining and outputting an output signal based on multiple enhancement signals includes: determining at least one signal to be output from the multiple enhancement signals, wherein the signal-to-noise ratio estimate corresponding to each of the at least one signal to be output is greater than or equal to a preset signal-to-noise ratio threshold; if the number of at least one signal to be output is less than the number of multiple enhancement signals, then the at least one signal to be output and some or all of the multiple input signals are used as the output signal and output.
[0025] If any of the enhanced signals has a signal-to-noise ratio (SNR) estimate lower than a preset SNR threshold, it indicates that such an enhanced signal may not effectively enhance the speech. Therefore, the processor can select enhanced signals with an SNR estimate greater than or equal to the preset SNR threshold as output, while also selecting one or more of the original input signals for output.
[0026] This ensures that speech enhancement can be achieved for different sectors in scenarios with high interference, while in scenarios with low interference, where the signal quality of some enhanced signals may not be as good as the original input signal, directly outputting the input signal allows subsequent processing modules to obtain higher quality output signals, resulting in higher accuracy in subsequent processing.
[0027] In some possible implementations, the number of output signal paths is the same as the number of sectors.
[0028] In some possible implementations, the signal-to-noise ratio (SNR) estimate corresponding to each signal in at least one of the output signals is greater than or equal to a preset SNR threshold, including multiple instantaneous SNR estimates corresponding to the SNR estimate of each signal in at least one of the output signals at multiple consecutive times that are all greater than the preset SNR threshold.
[0029] To avoid poor signal quality in the enhanced output signal due to interference, the processing module can continuously collect real-time signal-to-noise ratio (SNR) estimates at multiple moments when determining whether the estimated SNR of an enhanced signal is greater than or equal to a preset SNR threshold. This method avoids poor-quality enhanced output signals due to interference, improving the quality of the output signal and enhancing the effectiveness of speech processing.
[0030] In some possible implementations, the electronic device is in a non-voice scenario, which includes one or more of the following: voice wake-up and voice recognition.
[0031] When the processor outputs multiple signals, it can be used in non-voice scenarios, such as speech recognition and voice wake-up. In voice wake-up scenarios, this method can directionally enhance speech and improve the wake-up rate in noisy environments.
[0032] In a second aspect, a signal processing apparatus is provided, comprising a unit consisting of software and / or hardware, the unit being used to perform any one of the methods in the technical solution of the first aspect.
[0033] Thirdly, embodiments of this application provide a chip including a processor; the processor is used to read and execute a computer program stored in a memory to perform any one of the methods described in the first aspect.
[0034] Optionally, the chip further includes a memory, which is connected to the processor via a circuit or wire.
[0035] Alternatively, the chip may further include a communication interface.
[0036] Fourthly, embodiments of this application provide a chip, which is a processor or an advanced processor.
[0037] Fifthly, an electronic device is provided, comprising: a processor, a memory, and an interface; the processor, memory, and interface cooperate with each other to enable the electronic device to perform any one of the methods described in the first aspect.
[0038] Sixthly, an electronic device is provided, which includes any one of the chips described in the third and fourth aspects.
[0039] In a seventh aspect, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, the processor performs any one of the methods described in the first aspect.
[0040] Eighthly, a computer program product is provided, the computer program product comprising: computer program code, which, when run on an electronic device, causes the electronic device to perform any one of the methods described in the first aspect. Attached Figure Description
[0041] Figure 1 This is a schematic diagram of the structure of a terminal device 100 provided in an embodiment of this application;
[0042] Figure 2 This is a software structure block diagram of the terminal device 100 provided in the embodiments of this application;
[0043] Figure 3 This is an architectural diagram of an example signal processing method provided in an embodiment of this application;
[0044] Figure 4 This is a schematic diagram of sector division provided in an embodiment of this application;
[0045] Figure 5 This is an example of a beam pattern provided in an embodiment of this application;
[0046] Figure 6 This is a schematic flowchart of a signal enhancement method provided in an embodiment of this application;
[0047] Figure 7 This is a schematic diagram illustrating the logic of a channel selection module determining an output signal, as provided in an embodiment of this application.
[0048] Figure 8 This is a schematic diagram of beamforming effect in a conference scenario provided in an embodiment of this application;
[0049] Figure 9 This is another example of a logic diagram illustrating how a channel selection module determines an output signal, provided in an embodiment of this application.
[0050] Figure 10 This is another example of a logic diagram illustrating how a channel selection module determines an output signal, provided in an embodiment of this application.
[0051] Figure 11 This is a schematic flowchart of an example signal processing method provided in an embodiment of this application;
[0052] Figure 12 This is a schematic diagram of a signal processing device provided in an embodiment of this application. Detailed Implementation
[0053] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.
[0054] Hereinafter, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first," "second," or "third" may explicitly or implicitly include one or more of that feature.
[0055] The signal processing method provided in this application can be applied to terminal devices such as mobile phones, tablets, wearable devices, in-vehicle devices, hearing aids, augmented reality (AR) / virtual reality (VR) devices, extended-range (XR) devices, large-screen devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). This application does not impose any restrictions on the specific type of terminal device.
[0056] For example, Figure 1This is a schematic diagram of the structure of a terminal device 100 provided in an embodiment of this application. The terminal device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0057] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0058] The controller can serve as the central nervous system and command center of the terminal device 100. The controller can generate operation control signals based on the instruction opcode and timing signals to control the fetching and execution of instructions.
[0059] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0060] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0061] The external storage interface 120 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the terminal device 100. The external storage card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.
[0062] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of terminal device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of terminal device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0063] Terminal device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0064] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0065] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The terminal device 100 can listen to music or make hands-free calls through the speaker 170A.
[0066] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the terminal device 100 answers a phone call or voice message, the receiver 170B can be brought close to the listener's ear to hear the voice.
[0067] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Terminal device 100 may be equipped with at least one microphone 170C. In some embodiments, terminal device 100 may be equipped with two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, terminal device 100 may be equipped with three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.
[0068] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.
[0069] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the terminal device 100. In other embodiments of this application, the terminal device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0070] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the terminal device 100. In other embodiments of this application, the terminal device 100 may also adopt different interface connection methods or a combination of multiple interface connection methods as described in the above embodiments.
[0071] The software system of terminal device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of terminal device 100.
[0072] Figure 2 This is a software structure block diagram of the terminal device 100 according to an embodiment of this application. The layered architecture divides the software into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer. The application layer may include a series of application packages.
[0073] like Figure 2 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.
[0074] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0075] like Figure 2 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0076] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0077] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0078] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0079] The phone manager is used to provide communication functions for terminal device 100. For example, it manages call status (including connection, hang-up, etc.).
[0080] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0081] The notification manager allows applications to display notification information in the status bar. It can be used to convey informational messages and can disappear automatically after a short time without user interaction.
[0082] The Android runtime consists of core libraries and a virtual machine. The Android runtime is responsible for scheduling and managing the Android system.
[0083] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0084] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0085] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0086] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0087] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0088] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0089] A 2D graphics engine is a graphics engine for 2D drawing.
[0090] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0091] For ease of understanding, the following embodiments of this application will be described using the following methods: Figure 1 and Figure 2 Taking the terminal device with the structure shown as an example, and in conjunction with the accompanying drawings and application scenarios, the signal processing method provided in this application embodiment will be specifically described.
[0092] Electronic devices are typically equipped with microphones to collect sound signals. These signals are then processed and transmitted to enable voice calls or to perform actions like voice wake-up. Some devices utilize multiple microphones in a microphone array to capture multiple audio signals. Microphone arrays can cover a larger area and provide more spatial information, enabling a wider range of applications. However, electronic devices may experience poor sound pickup in the presence of ambient noise. For example, when using a large-screen device for an online meeting, multiple users may sit around a table facing the device. If the environment is noisy, the device may not be able to accurately capture the voices of speakers from different locations, negatively impacting the user experience. Furthermore, in other electronic devices, the location of the sound source is not fixed, making it difficult to guarantee effective sound signal capture.
[0093] This application provides a signal processing method that can directionally enhance the input sound signal in the direction of its arrival. This allows for spatial filtering of the signal received by the microphone array even in noisy environments, enabling the reception of the sound signal from the direction of arrival with higher gain and suppressing interference signals from other directions. This improves the signal-to-noise ratio of the useful signal, enhances speech intelligibility, and thus improves the user experience. Furthermore, this method requires no additional hardware and does not increase hardware costs.
[0094] The signal processing method provided in this application can be applied to a signal processor (DSP) or an advanced signal processor (ADSP). Optionally, the processor can be applied to an electronic device. A microphone array can also be installed on the electronic device. The following description uses a processor as the executing entity and, in conjunction with the functions of the software modules within the processor, describes the signal processing method provided in this application.
[0095] Alternatively, the software architecture in the processor can be found in [reference needed]. Figure 3 As shown. Specifically, the processor can contain two software modules: a beamforming module and a channel selection module. The beamforming module enhances the multiple audio signals acquired by the microphone array and inputs the enhanced signals into the channel selection module, which then selects one or more enhanced signals for output for subsequent scene processing.
[0096] First, the multiple input signals of the input beamforming module are described.
[0097] In a microphone array within an electronic device, each microphone can acquire and output one audio signal. Therefore, multiple microphones in the array can acquire and output multiple audio signals. Because the microphones are positioned differently, the audio signals acquired by each microphone will differ. These multiple audio signals acquired by the microphone array can be used as multiple input signals and fed into a beamforming module for beam amplification. Figure 3 The diagram illustrates the case where the microphone array has two microphones, and the input signals include two input signals: Input Signal 1 and Input Signal 2. Optionally, when the microphone array has three, four, or more microphones, the multiple input signals may also include Input Signal 3, Input Signal 4, or more input signals.
[0098] In the method of acquiring multiple input signals using a microphone array, the sound signals received from the same sound source are different due to the different positions of the microphones in the microphone array. Therefore, the microphone array can acquire more spatial information when acquiring sound signals, which is convenient for directional enhancement of speech.
[0099] Having understood the sources of multiple input signals, the next step is to describe in detail the process by which the beamforming module processes multiple input signals to obtain multiple enhanced signals.
[0100] Specifically, the beamforming module includes multiple sector beamformers. Each sector beamformer corresponds to one sector and is used to amplify the sound signal coming from the direction corresponding to that sector.
[0101] To facilitate understanding, we will first introduce the sector division method; for details, please refer to [link to relevant documentation]. Figure 4 As shown in the image. Figure 4 Figure a shows a schematic diagram of the vertical sector division centered on the location of the microphone array of the electronic device. Figure 4 Figure a shows a diagram dividing the area into three sectors: 0° to 60°, 60° to 120°, and 120° to 180°. It should be noted that... Figure 4 The angular range of the sectors shown is a schematic diagram of a planar division. The actual sectors are sectors in three-dimensional space. That is, sector 1, corresponding to 0 degrees to 60 degrees, is the spatial region of a cone centered on the microphone array, and sector 2, corresponding to 120 degrees to 180 degrees, is also the spatial region of a cone centered on the microphone array. Since the vertices of sectors 1 and 2 are opposite each other, sector 3, corresponding to 60 degrees to 120 degrees, represents the spatial region excluding sectors 1 and 2. For details, please refer to [link to relevant documentation]. Figure 4 As shown in Figure b.
[0102] It should be noted that the more finely divided and numerous the sectors, the finer the granularity of the signal enhancement direction, resulting in greater targeting and improved signal enhancement effectiveness. Conversely, the coarser the sector division and the fewer the sectors, the simpler the signal processing process, saving resources and improving processing efficiency. In some embodiments, dividing the space into three sectors is more reasonable as it balances the signal enhancement effect with the complexity of signal processing.
[0103] Figure 4 Figure c in the diagram illustrates another type of sector division. This includes three sectors: 0 degrees to 45 degrees, 45 degrees to 90 degrees, and 90 degrees to 180 degrees.
[0104] Optionally, the number of sectors can also be two, such as two sectors from 0 degrees to 90 degrees and from 90 degrees to 180 degrees. The number of sectors can also be four, five or more.
[0105] Optionally, the multiple sectors can be divided uniformly or non-uniformly. For example, sectors representing the most likely sound directions can be divided with fine-grained subdivision, while sectors representing other directions can be divided with coarse-grained subdivision. Figure 4 As shown in Figure c, when using a mobile phone, users often move closer to the microphone array and rarely emit sound from the top of the phone (the opposite end of the microphone array). Therefore, instead of dividing the 0-90 degree range into two sectors, the 90-180 degree range can be divided into one sector. This ensures fine-grained division of the sound direction without increasing the number of sectors excessively, making it more reasonable. Optionally, the space can also be divided into three sectors: 0-40 degrees, 40-80 degrees, and 80-180 degrees. This application does not limit the sector division method.
[0106] It should be noted that each sector corresponds to a set of pre-calibrated beam parameters, used to enhance the signal from multiple input signals within that sector. These multiple sets of beam parameters for multiple sectors can also be referred to as direction-of-arrival (DOA) information. These beam parameters are preset parameters, which can be downloaded from the server or pre-configured in the processor.
[0107] Each set of beam parameters can perform beamforming in N directions for multiple input signals, that is, enhance the signal in N directions. Optionally, N is a natural number, which can be a natural number greater than or equal to 3, such as 5, 8, 10, etc. The N directions here may differ from the sector divisions mentioned earlier. Although it is also a spatial division, the number of divisions can be greater. These N directions and the N sectors are independent of each other. The direction here refers to the direction of arrival of the signal, also known as the angle of arrival, or the direction of wave arrival, and the unit can be degrees.
[0108] Since the processor has different gains for signals arriving from different directions, the output signals after N beamformings all have different gains. Figure 5 This is a schematic diagram of a beam pattern. Lighter colors indicate higher gain, and darker colors indicate lower gain. It can be seen that the signal gain is related to both the direction of arrival and the signal frequency; the gain of the same signal will differ depending on the direction of arrival. For example... Figure 5 As shown, at very low frequencies, close to zero, the gain hardly changes with the direction of arrival; that is, for low-frequency signals, the gain has no significant directional difference (i.e., it is non-directional). For signals of the same frequency, the gain changes with the direction of arrival, and multiple lobes can be presented within the range of the direction of arrival. Figure 5The diagram illustrates the gain points of the main lobe and side lobes, where the signal gain is high, indicating enhancement. The locations of the other darker grooves correspond to areas with very low signal gain, indicating suppression. Figure 5 In the example beam pattern, the gain is the beam parameter mentioned earlier. The gain corresponding to each frequency (or frequency range) and a direction of arrival (or range of directions of arrival) is the beam parameter corresponding to that direction of arrival. Therefore, the beam parameter is not a single number, but a set of values corresponding to multiple frequencies and directions. Optionally, the beam parameter can be data obtained through experimental measurement or simulation, which will not be elaborated here.
[0109] The following section will use one of the first sector beamformers as an example to explain the specific process of signal enhancement. See details below. Figure 6 As shown, Figure 6 The example uses two input signals, input signal 1 and input signal 2, as multiple input signals.
[0110] When input signal 1 and input signal 2 are input to the first sector beamformer, the first sector beamformer first uses each of the N sets of beam parameters to enhance input signal 1 and input signal 2 respectively, resulting in N initial enhanced signals corresponding to the N sets of beam parameters. These N initial enhanced signals correspond to the N directions of arrival.
[0111] When N is 3, the beam parameters can include three sets: beam parameter 1, beam parameter 2, and beam parameter 3. The processor uses a first sector beamformer to enhance input signal 1 and input signal 2, that is, multiplying input signal 1 and input signal 2 by beam parameter 1 to obtain initial enhanced signal 1; multiplying input signal 1 and input signal 2 by beam parameter 2 by the first sector beamformer to obtain initial enhanced signal 2; and multiplying input signal 1 and input signal 2 by beam parameter 3 by the first sector beamformer to obtain initial enhanced signal 3.
[0112] Next, the processor calculates the ratios of the energy of the initial boosted signal 1 to the energy of the input signal 1, the initial boosted signal 2 to the energy of the input signal 1, the initial boosted signal 3 to the energy of the input signal 1, the initial boosted signal 1 to the energy of the input signal 2, the initial boosted signal 2 to the energy of the input signal 2, and the initial boosted signal 3 to the energy of the input signal 2, thus obtaining six intensity ratios. The signal energy mentioned in this embodiment can be represented by signal intensity or signal power.
[0113] The processor selects the largest maximum intensity ratio from these six intensity ratios, then selects the smallest minimum intensity ratio, and finally obtains the first fluctuation value by subtracting the minimum intensity ratio from the maximum intensity ratio. This first fluctuation value can characterize the magnitude of fluctuation of the multi-channel input signal in multiple incoming wave directions.
[0114] It should be noted that the larger N is, the finer the granularity of the direction division in the beam parameters, and the higher the accuracy of the discriminant factor generated based on the beam parameters corresponding to more directions; the smaller N is, the coarser the granularity of the direction division, but the simpler the signal processing process, saving resources and improving processing efficiency. In some embodiments, N of 8 can balance the signal enhancement effect and the complexity of signal processing, and is therefore more reasonable.
[0115] Next, the processor determines whether the first fluctuation value is greater than or equal to a preset threshold. If the first fluctuation value is greater than or equal to the preset threshold, it indicates that the input signal fluctuates significantly in multiple directions of arrival, suggesting the presence of incoming speech. In this case, the weight vector of subsequent filters needs to be adjusted to enhance the incoming speech. If the first fluctuation value is less than the preset threshold, it indicates that the fluctuation of the input signal is relatively small, suggesting that there is likely no incoming speech. Therefore, there is no need to adjust the filter weight vector to enhance the incoming speech; the existing filter weight vector can remain unchanged. It should be noted that the filter weight vector can be used for directional enhancement and other-direction suppression of the input signal.
[0116] The processor can also generate a discriminant factor based on the first fluctuation value. Optionally, the discriminant factor is 1 when the first fluctuation value is greater than or equal to a preset threshold; the discriminant factor is 0 when the first fluctuation value is less than the preset threshold. This discriminant factor can be used in the subsequent update process of the filter weight vector.
[0117] To explain how the aforementioned discriminant factor plays a role in determining the filter weight vector, the principle of obtaining the filter weights will be introduced first:
[0118] Taking the case at time t as an example, the vector of the signal output by the microphone array is denoted as: The filter weight vector at time t is denoted as The sound source vector is The noise vector present in the environment is .
[0119] Where t represents time. Represents frequency.
[0120] If the filter weight vector at time t is used to filter the signal, the resulting filter output vector is denoted as... .
[0121] Specifically, the filter output vector is the product of the filter weight vector and the vector of the signal output from the microphone array. This can be expressed as the formula:
[0122] .in, for The transpose of .
[0123] because Substituting this relationship into the formula for the filter output vector above, we obtain the following relationship:
[0124] .
[0125] in is the air propagation function, used to characterize the degree of signal attenuation in air.
[0126] The power of the signal output by the filter can then be expressed as the mathematical expectation of the filter output vector and its conjugate transpose, as shown in the formula:
[0127] .
[0128] The expression can be rewritten as: .
[0129] in, It can be seen that It is related to the signal output by the microphone.
[0130] In summary, it can be seen that the power of the signal output by the filter and the power of the signal output by the microphone (that is, the signal input to the filter) are related to the filter weight vector.
[0131] Next, to minimize the power of the output signal, we need to constrain the vector. This leads to a convex optimization problem. To solve this problem, the processor can use the Lagrange multiplier method to obtain a closed-form solution. However, the closed-form solution cannot track environmental changes; that is, it cannot update in real time when the direction of the sound changes. Therefore, the filter weight vector can be updated, resulting in the following update relationship for the weight vector:
[0132] .
[0133] in, As a discriminant factor, For Lagrange functions, The function for solving the gradient is denoted as , which represents finding the gradient of the Lagrange function.
[0134] Next, when the first fluctuation value is greater than or equal to the preset threshold, the processor outputs a discrimination factor of 1, and the weight of the filter in the next moment decreases. This needs to be updated. When the first fluctuation value is less than the preset threshold, the processor outputs a discrimination factor of 0, i.e. .
[0135] Optionally, the preset threshold can be a value between 0 and 1, such as 0.35, 0.4, etc., and this application embodiment does not limit it.
[0136] Next, at the next time step (t+1), the updated filter weight vector can be used to enhance the input signal. Specifically, the processor performs a Fourier transform (fft) on input signal 1 to obtain transformed signal 1, and the processor also performs a Fourier transform on input signal 2 to obtain transformed signal 2.
[0137] The processor then enhances the input signal using the updated filter weight vector. Specifically, it multiplies the updated filter weight vector with both the transformed signal 1 and the transformed signal 2 to obtain the frequency-domain enhanced signal. The processor can then perform an inverse Fourier transform (IFT) on the frequency-domain enhanced signal to obtain the enhanced signal (i.e., the first enhanced signal corresponding to the first sector), and output it to the channel selection module.
[0138] During signal enhancement, the processor uses the gain differences of signals from different directions via the microphone array to indicate the update of the filter weight vector. This not only enhances the incoming sound signal but also suppresses interference signals from other directions, enabling adaptive directional speech enhancement by tracking and enhancing the incoming sound. Even if the sound source location changes, the signal processing method of this application can adaptively track and enhance the incoming sound, improving the user experience.
[0139] Optionally, when the first sector beamformer outputs a first enhanced signal, it can also output a first signal-to-noise ratio (SNR) estimate corresponding to the first enhanced signal. The first SNR estimate is used to characterize the signal quality of the first enhanced signal. The larger the first SNR estimate, the better the signal quality of the first enhanced signal; the smaller the first SNR estimate, the worse the signal quality of the first enhanced signal.
[0140] Similarly, other sector beamformers can also output corresponding enhanced signals to the channel selection module. Optionally, each sector beamformer can also output a corresponding signal-to-noise ratio estimate.
[0141] Next, the process of the processor switching output signals using the channel selection module will be described.
[0142] When the beamforming module outputs multiple enhancement signals to the channel selection module, the channel selection module can select one or more enhancement signals as output signals and output them to the backend for processing.
[0143] Optionally, for scenarios where only one output signal is required, the processor can use a channel selection module to select the enhanced signal with the highest signal-to-noise ratio estimate as the output signal.
[0144] Optionally, the processor can also calculate the ratio of the energy of each enhanced signal to the total energy, and then select the enhanced signal with the largest ratio as the output signal. It should be noted that the total energy is the sum of the energies of all input signals.
[0145] The energy of the signal mentioned here can be the signal strength. For example, the processor can also calculate the ratio of the strength of each enhanced signal to the overall strength, and then select the enhanced signal with the largest ratio as the output signal. It should be noted that the overall strength is the sum of the strengths of all input signals. Optionally, the processor can also calculate the ratio of the strength of each enhanced signal to the overall strength separately, and also calculate the ratio of each input signal to the overall strength, and then select the enhanced signal with the largest ratio as the output signal. For example, see... Figure 7 As shown.
[0146] When the processor outputs one signal, this signal can be used in voice scenarios, such as hearing aid scenarios, voice calls, and voice conferences. In hearing aid scenarios, this method enhances the speech from the direction of the sound source, allowing users to hear the sound more clearly and facilitating conversations. In voice conference scenarios, this method enhances the speech from the direction of the sound source, improving speech clarity and reducing interference in noisy environments, thereby enhancing the user's conference experience.
[0147] Figure 8 This diagram illustrates the effect of adaptive speech orientation enhancement in a conference scenario. Figure 8 As shown, a microphone array is installed on the large-screen device. Figure 8 The diagram illustrates the directional maps of voice enhancement for different sectors in different directions. When participants speak from different directions, the large-screen device can use the signal processing method provided in this application embodiment to perform directional voice enhancement in the direction of the speaking participant. If the speaker's position changes, the large-screen device can also adaptively track the direction of the sound, thereby improving the overall meeting experience.
[0148] If the signal-to-noise ratio (SNR) of the input signal is relatively low, but the estimated SNR of the enhanced signal is relatively high, then the processor can directly output the enhanced signal, for example... Figure 9 In Figure a, the processor outputs enhancement signal 1, enhancement signal 2, and enhancement signal 3.
[0149] When the interference signal is very weak, if the estimated signal-to-noise ratio (SNR) of the enhanced signal is low, it is highly likely that the estimated SNR of the enhanced signal will be lower than that of the original input signal. In this case, the processor can select only the enhanced signal with an estimated SNR greater than a preset SNR threshold from among multiple enhanced signals as the output signal, and can also select the original input signal as the output signal. Based on this, by selecting the output signal, the processor can ensure a high-quality output signal with a high SNR, thereby improving the quality of the speech signal.
[0150] Specifically, the processor can select the enhanced signal whose signal-to-noise ratio (SNR) estimate is greater than or equal to a preset SNR threshold from multiple enhanced signals through the channel selection module. If the SNR estimates of multiple enhanced signals are all greater than or equal to the preset SNR threshold, the processor can directly output multiple enhanced signals.
[0151] Optionally, the preset signal-to-noise ratio threshold can be a value set based on experience. Optionally, when the signal-to-noise ratio estimate is a normalized value, the preset signal-to-noise ratio threshold can be a value such as 0.2 or 0.3, and this embodiment of the application does not limit this.
[0152] If any of the enhanced signals has a signal-to-noise ratio (SNR) estimate lower than a preset SNR threshold, it indicates that such an enhanced signal may not effectively enhance the speech. Therefore, the processor can select enhanced signals with an SNR estimate greater than or equal to the preset SNR threshold as output, while also selecting one or more of the original input signals for output.
[0153] This ensures that speech enhancement can be achieved for different sectors in scenarios with high interference, while in scenarios with low interference, where the signal quality of some enhanced signals may not be as good as the original input signal, directly outputting the input signal allows subsequent processing modules to obtain higher quality output signals, resulting in higher accuracy in subsequent processing.
[0154] For example, when a user inputs voice near the microphone array of a mobile phone, the signal quality of the original input signal is relatively high, that is, the signal-to-noise ratio is high. However, in the enhanced signal, the estimated signal-to-noise ratio of the enhanced signal corresponding to the sector at the top of the mobile phone (the side opposite the microphone array) is lower than that of the original input signal. In this case, it is not necessary to output the enhanced signal corresponding to the sector at the top of the mobile phone, but to output one or more of the original input signals as output signals.
[0155] For example, see [link / reference] Figure 9As shown in Figure b, the input signals are input signal 1 and input signal 2. Among enhanced signals 1, 2, and 3, the estimated signal-to-noise ratio (SNR) of enhanced signal 1 is greater than a preset SNR threshold, while the estimated SNRs of enhanced signals 2 and 3 are both less than the preset SNR threshold. The processor can then output enhanced signal 1, input signal 1, and input signal 2.
[0156] If, among enhanced signals 1, 2, and 3, the estimated signal-to-noise ratio (SNR) values of enhanced signals 1 and 2 are greater than a preset SNR threshold, while the estimated SNR value of enhanced signal 3 is less than the preset SNR threshold, the processor can output enhanced signals 1 and 2 simultaneously, and also select one output from input signals 1 and 2.
[0157] Optionally, when the processor selects an output from the original input signal, it can arbitrarily select one or more outputs according to the number of allowed outputs, or it can select from input signal 1 in sequence by default, or it can select the output signal with a high signal-to-noise ratio. This application embodiment does not limit this.
[0158] Optionally, to avoid poor signal quality of the output enhanced signal due to interference, the processing module can also continuously collect real-time signal-to-noise ratio estimates at multiple times to determine whether the signal-to-noise ratio estimate of an enhanced signal is greater than or equal to a preset signal-to-noise ratio threshold.
[0159] For example, the processor records M instantaneous signal-to-noise ratio (SNR) estimates for time t and the preceding M consecutive time points. At time t, the processor determines whether all M instantaneous SNR estimates are greater than or equal to a preset SNR threshold. If all M instantaneous SNR estimates are greater than or equal to the preset SNR threshold, then the processor determines that the SNR estimate of the enhanced signal is greater than or equal to the preset SNR threshold; if any of the M instantaneous SNR estimates is less than the preset SNR threshold, then the processor determines that the SNR estimate of the enhanced signal is neither greater than nor equal to the preset SNR threshold.
[0160] Optionally, when the processor outputs multiple output signals, the number of output signals is equal to the number of sectors. For example, if there are three sectors, then three output signals are output.
[0161] For example Figure 10 As shown, the input signals are input signal 1 and input signal 2. Among enhanced signals 1, 2, and 3, the M instantaneous signal-to-noise ratio (SNR) estimates of enhanced signal 1 at M consecutive time points are all greater than a preset SNR threshold, while the M instantaneous SNR estimates of enhanced signals 2 and 3 at some point are less than the preset SNR threshold. The processor can then output enhanced signal 1, input signal 1, and input signal 2.
[0162] Optionally, if the estimated signal-to-noise ratio (SNR) of the enhanced signal is greater than the preset SNR threshold, the quality of the enhanced signal may be higher than that of the original input signal; if the estimated SNR of the enhanced signal is less than the preset SNR threshold, the quality of the enhanced signal may be lower than that of the original input signal; if the estimated SNR of the enhanced signal is equal to or approximately equal to the preset SNR threshold, the quality of the enhanced signal may be comparable to that of the original input signal.
[0163] When the processor outputs multiple signals, it can be used in non-voice scenarios, such as speech recognition and voice wake-up. In voice wake-up scenarios, this method can directionally enhance speech and improve the wake-up rate in noisy environments.
[0164] Figure 11 A flowchart illustrating an example of a signal processing method provided in this application embodiment. This method is applied to an electronic device, which includes multiple microphones. The method includes:
[0165] S1101. Acquire multiple input signals. The multiple input signals are signals collected by multiple microphones, and each input signal corresponds to one microphone.
[0166] Each microphone can output one audio signal, while a microphone array containing multiple microphones can output multiple audio signals.
[0167] S1102. Multiple sets of beam parameters are used to enhance multiple input signals to obtain multiple enhanced signals. The multiple sets of beam parameters, multiple enhanced signals and multiple sectors correspond one-to-one. The beam parameters are related to the frequency of the input signal and the direction of arrival of the input signal. The multiple sectors are multiple spatial regions obtained by dividing the space with multiple microphones as the center.
[0168] For a detailed introduction to beam parameters and sectors, please refer to the preceding text. The processor can use any set of beam parameters to enhance multiple input signals, for example, by multiplying them by the corresponding filter weight vector to obtain multiple enhanced signals. It should be noted that each sector corresponds to a set of beam parameters. Each set of beam parameters can output a corresponding enhanced signal for that sector.
[0169] S1103. Determine and output the output signal based on the multiple enhancement signals. The output signal is one or more of the multiple enhancement signals.
[0170] The processor selects the signal with the best quality from multiple enhanced signals and outputs it; the processor can also directly output multiple enhanced signals.
[0171] In the method of acquiring multiple input signals using a microphone array, the sound signals received from the same sound source are different due to the different positions of the microphones in the microphone array. Therefore, the microphone array can acquire more spatial information when acquiring sound signals, which is convenient for directional enhancement of speech.
[0172] This method can perform directional enhancement on signals collected by multiple microphones, so that even in noisy environments, it can receive sound signals from the direction of sound with high gain and suppress interference signals from other directions, thereby improving the signal-to-noise ratio of useful signals and thus enhancing the user experience.
[0173] In some embodiments, S1102 uses multiple sets of beam parameters to enhance multiple input signals to obtain multiple enhanced signals, including: determining a first fluctuation value of the multiple input signals using a first beam parameter, the first fluctuation value representing the fluctuation degree of the multiple input signals in multiple incoming wave directions, the first beam parameter being any one of the multiple sets of beam parameters, the first beam parameter corresponding to a first sector among multiple sectors; if the first fluctuation value is greater than or equal to a preset threshold, then adjusting the current filter weight vector to the updated filter weight vector for the next time step, the updated filter weight vector being different from the current filter weight vector; if the first fluctuation value is less than the preset threshold, then determining the updated filter weight vector for the next time step as the current filter weight vector; and enhancing the multiple input signals according to the updated filter weight vector to obtain a first enhanced signal, the first enhanced signal being one of the multiple enhanced signals.
[0174] The aforementioned first fluctuation value can characterize the degree of fluctuation of the input signal in multiple incoming wave directions. Therefore, the filter weight vector can be adjusted when the fluctuation degree is large, and the filter weight vector can be kept unchanged when the fluctuation degree is small, thereby realizing real-time adaptive updating of the filter weight vector and improving the accuracy of speech direction finding.
[0175] In some embodiments, the first fluctuation value is the difference between the maximum intensity ratio and the minimum intensity ratio, the maximum intensity ratio is the largest of the plurality of intensity ratios, and the minimum intensity ratio is the smallest of the plurality of intensity ratios; the plurality of signal intensities are the ratios of the signal intensity of each of the plurality of directional enhancement signals to the signal intensity of each of the plurality of input signals; the plurality of directional enhancement signals are enhancement signals obtained by beamforming the plurality of input signals in multiple directions of arrival using the first beam parameters, and the number of the plurality of intensity ratios is the product of the number of the plurality of directional enhancement signals and the number of the plurality of input signals.
[0176] In some embodiments, the first beam parameters include multiple sets of preset gains, each set of preset gains corresponding to a multiple incoming wave direction. Determining the first fluctuation value of the multiple input signals using the first beam parameters includes: performing beamforming processing on the multiple input signals using multiple sets of preset gains for multiple incoming wave directions to obtain multiple directional enhancement signals, each directional enhancement signal corresponding to a multiple incoming wave direction; determining a first intensity ratio between the signal intensity of the first directional enhancement signal and the signal intensity of the first input signal, wherein the first directional enhancement signal is any one of the multiple directional enhancement signals, the first input signal is any one of the multiple input signals, and the first intensity ratio is one of the multiple intensity ratios; and determining the difference between the maximum intensity ratio and the minimum intensity ratio among the multiple intensity ratios as the first fluctuation value.
[0177] Beamforming involves multiplying multiple input signals by their respective preset gains. For each sector, the processor calculates the ratio of the intensity of each initial boost signal to that of each input signal, denoted as the intensity ratio. The processor obtains multiple intensity ratios, selects the largest and smallest, and calculates their difference. This difference quantifies the degree of fluctuation among the multiple intensity ratios.
[0178] In some embodiments, the output signal is the one with the strongest signal strength among a plurality of enhanced signals.
[0179] The processor can directly compare the signal strength of the enhanced signals and select the one with the strongest signal strength from multiple enhanced signals as the output signal, ensuring that the output signal has the best possible signal quality.
[0180] In some embodiments, the output signal is the one with the strongest signal strength among multiple enhancement signals and multiple input signals.
[0181] The processor can select the strongest signal from multiple enhancement signals and multiple input signals as the output signal, ensuring that the output signal has the best possible quality.
[0182] In some embodiments, determining and outputting an output signal based on the plurality of enhanced signals includes: determining the signal strength of each of the plurality of enhanced signals, a plurality of second intensity ratios of the combined strength, wherein the combined strength is the sum of the strengths of the multiple input signals, and the plurality of second intensity ratios correspond one-to-one with the plurality of enhanced signals; selecting the largest second intensity ratio from the plurality of second intensity ratios; and determining and outputting the enhanced signal corresponding to the largest second intensity ratio.
[0183] The processor can also use a second intensity ratio to quantify the signal quality of the enhanced signal. The processor selects the enhanced signal with the largest second intensity ratio as the output signal to ensure the signal quality of the output signal.
[0184] In some embodiments, the electronic device is in a voice scenario, which includes one of a call scenario, a hearing aid scenario, and a voice conference scenario.
[0185] When the processor outputs one signal, this signal can be used in voice scenarios, such as hearing aid scenarios, voice calls, and voice conferences. In hearing aid scenarios, this method enhances the speech from the direction of the sound source, allowing users to hear the sound more clearly and facilitating conversations. In voice conference scenarios, this method enhances the speech from the direction of the sound source, improving speech clarity and reducing interference in noisy environments, thereby enhancing the user's conference experience.
[0186] In some embodiments, the method further includes: outputting multiple signal-to-noise ratio (SNR) estimates, wherein the multiple SNR estimates correspond one-to-one with multiple sectors.
[0187] The signal-to-noise ratio (SNR) estimate can quantify the signal quality of the enhanced signal, making it easier to select the best enhanced signal and improve the performance of speech processing.
[0188] In some embodiments, determining and outputting an output signal based on a plurality of enhanced signals includes: determining at least one signal to be output from the plurality of enhanced signals, wherein the signal-to-noise ratio estimate corresponding to each of the at least one signal to be output is greater than or equal to a preset signal-to-noise ratio threshold; if the number of at least one signal to be output is less than the number of a plurality of enhanced signals, then at least one signal to be output and a portion or all of the multiple input signals are used as the output signal and output.
[0189] If any of the enhanced signals has a signal-to-noise ratio (SNR) estimate lower than a preset SNR threshold, it indicates that such an enhanced signal may not effectively enhance the speech. Therefore, the processor can select enhanced signals with an SNR estimate greater than or equal to the preset SNR threshold as output, while also selecting one or more of the original input signals for output.
[0190] This ensures that speech enhancement can be achieved for different sectors in scenarios with high interference, while in scenarios with low interference, where the signal quality of some enhanced signals may not be as good as the original input signal, directly outputting the input signal allows subsequent processing modules to obtain higher quality output signals, resulting in higher accuracy in subsequent processing.
[0191] In some embodiments, the number of output signal channels is the same as the number of multiple sectors.
[0192] In some embodiments, the signal-to-noise ratio (SNR) estimate corresponding to each of the at least one output signal is greater than or equal to a preset SNR threshold, including multiple instantaneous SNR estimates corresponding to the SNR estimate of each of the at least one output signal at multiple consecutive times that are all greater than the preset SNR threshold.
[0193] To avoid poor signal quality in the enhanced output signal due to interference, the processing module can continuously collect real-time signal-to-noise ratio (SNR) estimates at multiple moments when determining whether the estimated SNR of an enhanced signal is greater than or equal to a preset SNR threshold. This method avoids poor-quality enhanced output signals due to interference, improving the quality of the output signal and enhancing the effectiveness of speech processing.
[0194] In some embodiments, the electronic device is in a non-voice scenario, which includes one or more of voice wake-up and voice recognition.
[0195] When the processor outputs multiple signals, it can be used in non-voice scenarios, such as speech recognition and voice wake-up. In voice wake-up scenarios, this method can directionally enhance speech and improve the wake-up rate in noisy environments.
[0196] The foregoing has detailed examples of the methods provided in this application. It is understood that the corresponding apparatus, in order to achieve the above functions, includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0197] This application can divide the signal processing device into functional modules based on the above method example. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0198] Figure 12 A schematic diagram of a signal processing apparatus provided in this application is shown. The apparatus 1200 includes a beamforming module 1201 and a channel selection module 1202.
[0199] The beamforming module 1201 is used to acquire multiple input signals and enhance them using multiple sets of beam parameters to obtain multiple enhanced signals. The multiple input signals are signals collected by multiple microphones, and there is a one-to-one correspondence between the multiple input signals and the multiple microphones. The multiple sets of beam parameters and the multiple enhanced signals correspond one-to-one with multiple sectors. The beam parameters are related to the frequency and direction of arrival of the input signals. The multiple sectors are multiple spatial regions obtained by dividing the space with the multiple microphones as the center.
[0200] The channel selection module 1202 is used to determine and output an output signal based on multiple enhancement signals, wherein the output signal is one or more of the multiple enhancement signals.
[0201] Optionally, the beamforming module 1201 is specifically used to determine a first fluctuation value of the multiple input signals using first beam parameters. The first fluctuation value characterizes the fluctuation degree of the multiple input signals in multiple incoming wave directions. The first beam parameters are any one of multiple sets of beam parameters, and the first beam parameters correspond to the first sector among multiple sectors. If the first fluctuation value is greater than or equal to a preset threshold, it is adjusted to the updated filter weight vector at the next moment according to the current filter weight vector. The updated filter weight vector is different from the current filter weight vector. If the first fluctuation value is less than the preset threshold, the updated filter weight vector at the next moment is determined to be the current filter weight vector. Based on the updated filter weight vector, the multiple input signals are enhanced to obtain a first enhanced signal, which is one of multiple enhanced signals.
[0202] Optionally, the first fluctuation value is the difference between the maximum intensity ratio and the minimum intensity ratio, the maximum intensity ratio is the largest among the multiple intensity ratios, and the minimum intensity ratio is the smallest among the multiple intensity ratios; the multiple signal intensities are the ratios of the signal intensity of each directional enhancement signal among the multiple directional enhancement signals to the signal intensity of each input signal among the multiple input signals; the multiple directional enhancement signals are enhancement signals obtained by beamforming multiple incoming wave directions of the multiple input signals using the first beam parameters, and the number of multiple intensity ratios is the product of the number of multiple directional enhancement signals and the number of multiple input signals.
[0203] Optionally, the first beam parameters include multiple sets of preset gains, each set of preset gains corresponding to multiple incoming wave directions. The beamforming module 1201 is specifically used to perform beamforming processing on multiple input signals from multiple incoming wave directions using multiple sets of preset gains to obtain multiple directional enhanced signals, each directional enhanced signal corresponding to multiple incoming wave directions. A first intensity ratio is determined between the signal strength of the first directional enhanced signal and the signal strength of the first input signal. The first directional enhanced signal is any one of the multiple directional enhanced signals, the first input signal is any one of the multiple input signals, and the first intensity ratio is one of the multiple intensity ratios. The difference between the maximum and minimum intensity ratios among the multiple intensity ratios is determined as a first fluctuation value.
[0204] Optionally, the output signal is the one with the strongest signal strength among multiple enhanced signals.
[0205] Optionally, the output signal is the one with the strongest signal strength among multiple enhancement signals and multiple input signals.
[0206] Optionally, the beamforming module 1201 is specifically used to determine the signal strength of each of the multiple enhanced signals, compare it with multiple second intensity ratios of the comprehensive strength, the comprehensive strength being the sum of the strengths of the multiple input signals, and the multiple second intensity ratios corresponding one-to-one with the multiple enhanced signals; select the largest second intensity ratio from the multiple second intensity ratios; determine the enhanced signal corresponding to the largest second intensity ratio as the output signal and output it.
[0207] Optionally, the electronic device is in a voice scenario, which includes one of a call scenario, a hearing aid scenario, and a voice conference scenario.
[0208] Optionally, the beamforming module 1201 is also used to output multiple signal-to-noise ratio estimates, with each of the multiple signal-to-noise ratio estimates corresponding to a specific sector.
[0209] Optionally, the channel selection module 1202 is specifically used to determine at least one signal to be output from multiple enhanced signals, wherein the signal-to-noise ratio estimate corresponding to each of the at least one signal to be output is greater than or equal to a preset signal-to-noise ratio threshold; if the number of at least one signal to be output is less than the number of multiple enhanced signals, then at least one signal to be output and part or all of the multiple input signals are used as output signals and output.
[0210] Optionally, the number of output signal channels is the same as the number of multiple sectors.
[0211] Optionally, the signal-to-noise ratio (SNR) estimate corresponding to each signal in at least one of the output signals is greater than or equal to a preset SNR threshold, including multiple instantaneous SNR estimates corresponding to the SNR estimate of each signal in at least one of the output signals at multiple consecutive times that are all greater than the preset SNR threshold.
[0212] Optionally, the electronic device is in a non-voice scenario, which includes one or more of the following: voice wake-up and voice recognition.
[0213] The specific manner in which the device 1200 performs the signal processing method and the beneficial effects thereof can be found in the relevant descriptions in the method embodiments, and will not be repeated here.
[0214] This application also provides an electronic device, including the processor described above. The electronic device provided in this embodiment may be... Figure 1 The terminal device 100 shown is used to execute the above-described signal processing method. When using integrated units, the terminal device may include a processor, a storage module, and a communication module. The processor can be used to control and manage the operations of the terminal device; for example, it can support the terminal device in executing the steps performed by the display unit, detection unit, and processing unit. The storage module can be used to support the terminal device in executing stored program code and data. The communication module can be used to support communication between the terminal device and other devices.
[0215] The processor can be a processor or a controller. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination of functions that implement computing capabilities, such as a combination of one or more microprocessors, a digital signal processor (DSP), and a microprocessor, etc. The storage module can be a memory. The communication module can specifically be a radio frequency circuit, a Bluetooth chip, a Wi-Fi chip, or other devices that interact with other terminal devices.
[0216] In one embodiment, when the processor is a processor and the storage module is a memory, the terminal device involved in this embodiment can be a device having... Figure 1 The device with the structure shown.
[0217] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the signal processing method described in any of the above embodiments.
[0218] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement the signal processing method described above.
[0219] In this embodiment, the electronic device, computer-readable storage medium, computer program product or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.
[0220] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between devices or units. The replaced units may or may not be physically separate. The component shown as a unit may be one physical unit or multiple physical units, that is, it may be located in one place or distributed in multiple different places. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0221] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0222] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0223] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A signal processing method, characterized in that, Applied to an electronic device, the electronic device including multiple microphones, including: Acquire multiple input signals, wherein the multiple input signals are signals collected by multiple microphones, the multiple input signals correspond one-to-one with the multiple microphones, the multiple microphones form a microphone array and the positions of the multiple microphones are different; Multiple sets of beam parameters are used to enhance the multiple input signals to obtain multiple enhanced signals. The multiple sets of beam parameters and the multiple enhanced signals correspond one-to-one with multiple sectors. The beam parameters are associated with the frequency of the input signal and the direction of arrival of the input signal. The multiple sectors are multiple spatial regions obtained by dividing the space with the multiple microphones as the center. Any set of beam parameters is used to enhance the multiple input signals from the corresponding sector in N directions. The N directions are the directions obtained by dividing the space. Based on the plurality of enhancement signals, an output signal is determined and output, wherein the output signal is one or more of the plurality of enhancement signals; The process of enhancing the multiple input signals using multiple sets of beam parameters yields multiple enhanced signals, including: The first fluctuation value of the multiple input signals is determined by the first beam parameter. The first fluctuation value characterizes the fluctuation degree of the multiple input signals in multiple incoming wave directions. The first beam parameter is any one of the multiple sets of beam parameters. The first beam parameter corresponds to the first sector among the multiple sectors. If the first fluctuation value is greater than or equal to a preset threshold, the filter weight vector is adjusted to the updated filter weight vector for the next moment based on the current filter weight vector, and the updated filter weight vector is different from the current filter weight vector. If the first fluctuation value is less than the preset threshold, then the updated filter weight vector for the next moment is determined to be the current filter weight vector. Based on the updated filter weight vector, the multiple input signals are enhanced to obtain a first enhanced signal, which is one of the multiple enhanced signals.
2. The method according to claim 1, characterized in that, The first fluctuation value is the difference between the maximum intensity ratio and the minimum intensity ratio, wherein the maximum intensity ratio is the largest among a plurality of intensity ratios, and the minimum intensity ratio is the smallest among the plurality of intensity ratios; The multiple intensity ratios are the ratio of the signal strength of each of the multiple directional enhancement signals to the signal strength of each of the multiple input signals. The multiple directional enhancement signals are enhancement signals obtained by performing beamforming processing on the multiple input signals in multiple directions of arrival using the first beam parameters, and the number of the multiple intensity ratios is the product of the number of the multiple directional enhancement signals and the number of the multiple input signals.
3. The method according to claim 2, characterized in that, The first beam parameters include multiple sets of preset gains, each set corresponding to a specific direction of arrival. Determining the first fluctuation value of the multiple input signals using the first beam parameters includes: The multiple sets of preset gains are used to perform beamforming processing on the multiple input signals in multiple directions of arrival to obtain the multiple directional enhancement signals, and the multiple directional enhancement signals correspond one-to-one with the multiple directions of arrival; A first intensity ratio is determined between the signal strength of the first directional enhancement signal and the signal strength of the first input signal, wherein the first directional enhancement signal is any one of the plurality of directional enhancement signals, the first input signal is any one of the plurality of input signals, and the first intensity ratio is one of the plurality of intensity ratios; The difference between the maximum intensity ratio and the minimum intensity ratio among the plurality of intensity ratios is determined as the first fluctuation value.
4. The method according to any one of claims 1 to 3, characterized in that, The output signal is the one with the strongest signal strength among the plurality of enhanced signals.
5. The method according to any one of claims 1 to 3, characterized in that, The output signal is the one with the strongest signal strength among the plurality of enhanced signals and the plurality of input signals.
6. The method according to claim 5, characterized in that, The step of determining and outputting an output signal based on the plurality of enhanced signals includes: The signal strength of each of the plurality of enhanced signals is determined and compared with a plurality of second intensity ratios of the combined strength, wherein the combined strength is the sum of the strengths of the multiple input signals, and the plurality of second intensity ratios correspond one-to-one with the plurality of enhanced signals; The largest second intensity ratio is selected from the plurality of second intensity ratios; The enhanced signal corresponding to the largest second intensity ratio is determined as the output signal and output.
7. The method according to any one of claims 1 to 3 and 6, characterized in that, The electronic device is in a voice scenario, which includes one of a call scenario, a hearing aid scenario, and a voice conference scenario.
8. The method according to any one of claims 1 to 3 and 6, characterized in that, The method further includes: Output multiple signal-to-noise ratio (SNR) estimates, each of which corresponds to one of the multiple sectors.
9. The method according to claim 8, characterized in that, The step of determining and outputting an output signal based on the plurality of enhanced signals includes: At least one signal to be output is determined from the plurality of enhanced signals, wherein the signal-to-noise ratio estimate corresponding to each of the at least one signal to be output is greater than or equal to a preset signal-to-noise ratio threshold. If the number of the at least one signal to be output is less than the number of the plurality of enhancement signals, then at least one signal to be output and part or all of the multiple input signals will be used as the output signal and output.
10. The method according to claim 9, characterized in that, The number of output signal channels is the same as the number of the plurality of sectors.
11. The method according to claim 9 or 10, characterized in that, The signal-to-noise ratio (SNR) estimate corresponding to each of the at least one output signal is greater than or equal to a preset SNR threshold, including multiple instantaneous SNR estimates corresponding to the SNR estimate of each of the at least one output signal at multiple consecutive times that are all greater than the preset SNR threshold.
12. The method according to any one of claims 1 to 3, 9 to 10, characterized in that, The electronic device is in a non-voice scenario, which includes one or more of the following: voice wake-up and voice recognition.
13. An electronic device, characterized in that, Including the microphone array, it also includes: processor, memory, and interface; The microphone array includes multiple microphones, which are used to collect multiple input signals. The processor, the memory, and the interface cooperate with each other to enable the electronic device to perform the method as described in any one of claims 1 to 12.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the method of any one of claims 1 to 12.