Sound Processing Method, Wearable Device, and Readable Storage Medium

By detecting the user's head rotation in a smart wearable device, obtaining head posture and microphone array position information, and performing directional voice enhancement, the problem of insufficient sound enhancement of existing devices in noisy environments is solved, improving user's auditory experience and providing listening assistance.

CN119255135BActive Publication Date: 2025-06-24SHARGE TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411766827.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-06-24
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

Existing smart wearable devices lack effective sound enhancement capabilities, especially in scenes such as noisy environments, meetings and multi-person conversations, resulting in a reduced user's hearing experience.

Method used

When the user is detected to turn the head, the user's head posture information and the position information of the microphone array are obtained, and the original audio signal is directionally enhanced based on this information, thereby achieving enhancement of the target sound source.

Benefits of technology

It realizes directional voice enhancement based on the user's head rotation behavior, suppresses other sound sources and ambient noise, improves the user's hearing experience, and provides hearing assistance to users with hearing defects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119255135B_ABST
    Figure CN119255135B_ABST
Patent Text Reader

Abstract

The present application discloses a sound processing method, a wearable device, and a readable storage medium. The sound processing method includes: obtaining an original audio signal received by the microphone array, where the original audio signal includes audio signals corresponding to at least one sound source; if it is detected that the user turns their head, obtaining the head pose information of the user and the position information of the microphone array, and performing directional voice enhancement on the original audio signal according to the head pose information and the position information to obtain a target audio signal corresponding to a target sound source, where the target sound source is the sound source closest to the front of the user's head; and playing the target audio signal. The above sound processing method can achieve directional voice enhancement according to the user's head turn, suppress other sound sources and environmental noise, and improve the user's auditory experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of voice processing, and in particular, to a sound processing method, a wearable device, and a computer-readable storage medium. Background Art

[0002] Currently, most existing intelligent wearable devices (such as smart glasses) lack effective sound enhancement functions, while traditional wearable devices (such as Bluetooth headsets) are limited by space and hardware processing capabilities and cannot implement advanced audio separation and enhancement algorithms, resulting in the inability to achieve sound enhancement in wearable devices in noisy environments, meetings, multi-person conversations, etc., reducing the user's auditory experience. Summary of the Invention

[0003] The present application provides a sound processing method, a wearable device, and a computer-readable storage medium, which can solve the problem that existing wearable devices cannot achieve sound enhancement.

[0004] In a first aspect, the present application further provides a sound processing method, where the sound processing method includes: obtaining an original audio signal received by the microphone array, where the original audio signal includes audio signals corresponding to at least one sound source; if it is detected that the user turns the head, obtaining the head pose information of the user and the position information of the microphone array, and performing directional voice enhancement on the original audio signal according to the head pose information and the position information to obtain a target audio signal corresponding to a target sound source, where the target sound source is the sound source closest to the front of the user's head; playing the target audio signal.

[0005] In a second aspect, the present application provides a sound processing method applied to a wearable device, where the wearable device includes a microphone array; the method includes:

[0006] Obtaining the head pose information of the user, the original audio signal received by the microphone array, and the position information of the microphone array;

[0007] Performing sound source localization according to the original audio signal, the head pose information, and the position information to obtain an estimated value of the sound source position corresponding to at least one sound source;

[0008] Performing beamforming according to the original audio signal, the head pose information, the position information, and the estimated value of the sound source position corresponding to each sound source to obtain a beamformed signal;

[0009] Performing blind source separation on the original audio signal to obtain an independent audio signal corresponding to each sound source;

[0010] Perform voice enhancement based on the beamformed signal and the independent audio signal corresponding to each sound source to obtain a target audio signal.

[0011] In a third aspect, the present application further provides a wearable device, which includes a memory, a processor, a microphone array, and an inertial measurement unit;

[0012] The memory is used to store a computer program;

[0013] The microphone array is used to receive an audio signal;

[0014] The inertial measurement unit is used to collect the head pose information of the user;

[0015] The processor is used to execute the computer program and implement the voice processing method as described above when executing the computer program.

[0016] In a fourth aspect, the present application further provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the voice processing method as described above.

[0017] The present application discloses a voice processing method, a wearable device, and a computer-readable storage medium. By obtaining the head pose information of the user and the position information of the microphone array when it is detected that the user turns the head, and performing directional voice enhancement on the original audio signal according to the head pose information and the position information, it is possible to achieve directional voice enhancement based on the user's head turning behavior and suppress other sound sources and environmental noise. It can be applied to various application scenarios, can effectively filter out noise, improve the user's auditory experience, and at the same time can play a hearing assistance role to help users with hearing defects improve their hearing. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 is a schematic block diagram of the structure of a wearable device provided by an embodiment of the present application;

[0020] Figure 2 is a schematic flowchart of a voice processing method provided by an embodiment of the present application;

[0021] Figure 3 is a schematic scenario diagram of directional voice enhancement provided by an embodiment of the present application;

[0022] Figure 4 is a software flowchart of a sound processing method provided by an embodiment of the present application;

[0023] Figure 5 is a schematic block diagram of a hardware system of a sound processing method provided by an embodiment of the present application;

[0024] Figure 6 is a schematic flowchart of preprocessing an audio signal provided by an embodiment of the present application;

[0025] Figure 7 is a schematic flowchart of a second sound processing method provided by an embodiment of the present application;

[0026] Figure 8 is a schematic flowchart of directional speech enhancement provided by an embodiment of the present application;

[0027] Figure 9 is a schematic flowchart of sound source localization provided by an embodiment of the present application;

[0028] Figure 10 is a schematic flowchart of beamforming provided by an embodiment of the present application;

[0029] Figure 11 is a schematic flowchart of blind source separation provided by an embodiment of the present application;

[0030] Figure 12 is a schematic diagram of the structure of a convolutional neural network model provided by an embodiment of the present application;

[0031] Figure 13 is a schematic flowchart of sound enhancement provided by an embodiment of the present application;

[0032] Figure 14 is a schematic flowchart of a third sound processing method provided by an embodiment of the present application. Detailed implementation manners

[0033] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0034] The flowcharts shown in the accompanying drawings are merely illustrative examples and do not necessarily include all content and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, combined, or partially merged, so the actual execution order may change according to the actual situation.

[0035] It should be understood that the terms used in the specification of this application are merely for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0036] It should also be understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0037] Currently, most existing smart wearable devices (such as smart glasses) lack effective sound enhancement functions, while traditional wearable devices (such as Bluetooth headsets) are limited by space and hardware processing capabilities and cannot implement advanced audio separation and enhancement algorithms, resulting in the inability to achieve sound enhancement in scenarios such as noisy environments, meetings, and multi-person conversations, reducing the user's auditory experience.

[0038] To this end, embodiments of this application provide a sound processing method, a wearable device, and a computer-readable storage medium. Among them, the sound processing method can be applied to a wearable device. By obtaining the user's head pose information and the position information of the microphone array when detecting that the user turns the head, and performing directional voice enhancement on the original audio signal according to the head pose information and the position information to obtain the target audio signal corresponding to the target sound source, it is possible to achieve directional voice enhancement based on the user's head turning behavior and suppress other sound sources and environmental noise, which can be applicable to various application scenarios, effectively filter out noise, improve the user's auditory experience, and at the same time can play a role in hearing assistance to help users with hearing defects improve their hearing.

[0039] Exemplarily, the wearable device may include, but is not limited to, smart glasses such as AR (Augmented Reality) glasses and VR (Virtual Reality) glasses, and may also include smart electronic devices such as smart helmets.

[0040] Please refer to Figure 1 , Figure 1It is a schematic structural diagram of a wearable device provided by an embodiment of the present application. The wearable device may include a processor 1001, a memory 1002, a microphone array 1003, and an inertial measurement unit 1004. Among them, the processor 1001, the memory 1002, the microphone array 1003, and the inertial measurement unit 1004 may be connected through a bus, and this bus may be any applicable bus such as an Inter-integrated Circuit (I2C) bus.

[0041] Among them, the memory 1002 may include a storage medium and an internal memory. Among them, the storage medium may be a non-volatile storage medium or a volatile storage medium. The storage medium can store an operating system and a computer program, and the internal memory provides an environment for the operation of the computer program in the storage medium. This computer program includes program instructions, and when the program instructions are executed, the processor can be made to execute the sound processing method described in any embodiment.

[0042] The microphone array 1003 is used to receive audio signals. Among them, the microphone array may include microphones. Among them, (in this embodiment of the present application, 3 microphones are taken as an example for description), the original audio signals collected by each microphone in the microphone array 1003 can be expressed as: , where represents the signal collected by the th microphone at time , and the sampling frequency is .

[0043] The geometric information of the microphone array 1003 in the world coordinate system (assuming that the center of the head is the origin of the world coordinate system) can be expressed as: , where represents the microphone number; the microphone spacing can be expressed as: , represents the distance between microphone and microphone .

[0044] The inertial measurement unit 1004 (Inertial Measurement Unit, IMU) is used to collect the head pose information of the user. It should be noted that the inertial measurement unit 1004 is a device for measuring the three-axis attitude angle (or angular rate) and acceleration of an object. In this embodiment of the present application, the head pose information can be expressed as: , where, represents the yaw angle, represents the pitch angle, represents the roll angle.

[0045] The processor 1001 is used to provide computing and control capabilities to support the operation of the entire wearable device 10.

[0046] Among them, the processor 1001 can be a Central Processing Unit (CPU), and this processor can also be a general-purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the general-purpose processor can also be any conventional processor, etc.

[0047] Among them, in one embodiment, the processor 1001 is used to run a computer program stored in the memory to implement the following steps:

[0048] Obtain the original audio signal received by the microphone array, where the original audio signal includes the audio signals corresponding to at least one sound source; if it is detected that the user turns their head, obtain the head pose information of the user and the position information of the microphone array, and perform directional voice enhancement on the original audio signal according to the head pose information and the position information to obtain the target audio signal corresponding to the target sound source, where the target sound source is the sound source closest to the front of the user's head; play the target audio signal.

[0049] In some embodiments, the original audio signal includes the audio signals corresponding to at least one sound source received by at least one microphone channel; before the processor 1001 implements directional voice enhancement on the original audio signal according to the head pose information and the position information to obtain the target audio signal corresponding to the target sound source, it is also used to implement:

[0050] Perform a frequency-domain transformation on the audio signal received by each microphone channel to obtain the time-frequency signal corresponding to each microphone channel; combine the time-frequency signals corresponding to each microphone channel to obtain a multi-channel time-frequency signal; perform covariance matrix estimation on the multi-channel time-frequency signal to obtain the covariance matrix estimation result corresponding to the multi-channel time-frequency signal.

[0051] In some embodiments, after the processor 1001 implements covariance matrix estimation on the multi-channel time-frequency signal to obtain the covariance matrix estimation result corresponding to the multi-channel time-frequency signal, it is also used to implement:

[0052] Estimate the number of sound sources according to the covariance matrix estimation result to obtain the estimated value of the number of sound sources.

[0053] In some embodiments, when the processor 1001 implements directional voice enhancement of the original audio signal according to the head pose information and the position information to obtain the target audio signal corresponding to the target sound source, it is used to implement:

[0054] Perform sound source localization based on the multi-channel time-frequency signal, head pose information, and position information to obtain the estimated sound source position values corresponding to at least one sound source; perform beamforming on the multi-channel time-frequency signal according to the head pose information, position information, and the estimated sound source position values corresponding to each sound source to obtain a beamformed signal; perform blind source separation on the multi-channel time-frequency signal to obtain the independent audio signal corresponding to each sound source; perform sound enhancement based on the beamformed signal and the independent audio signal corresponding to each sound source to obtain the target audio signal.

[0055] In some embodiments, when the processor 1001 implements sound source localization based on the multi-channel time-frequency signal, head pose information, and position information to obtain the estimated sound source position values corresponding to at least one sound source, it is used to implement:

[0056] Determine the estimated time delay value corresponding to the time difference of arrival of the multi-channel time-frequency signal; based on the Kalman filter, perform sound source localization according to the estimated time delay value, head pose information, and position information to obtain the estimated sound source position values corresponding to each sound source.

[0057] In some embodiments, when the processor 1001 implements determining the estimated time delay value corresponding to the time difference of arrival of the multi-channel time-frequency signal based on the multi-channel time-frequency signal, it is used to implement:

[0058] Construct the cross-correlation function corresponding to the multi-channel time-frequency signal; construct the generalized cross-correlation function according to the Fourier transform of the cross-correlation function; determine the estimated time delay value according to the peak position of the generalized cross-correlation function.

[0059] In some embodiments, the position information includes the first position of the microphone array; when the processor 1001 implements sound source localization based on the Kalman filter according to the estimated time delay value, head pose information, and position information to obtain the estimated sound source position values corresponding to each sound source, it is used to implement:

[0060] Determine the rotation matrix of the head coordinate system used for the head pose information relative to the world coordinate system; according to the rotation matrix, transform the first position from the world coordinate system to the head coordinate system to obtain the second position; input the estimated time delay value and the second position into the Kalman filter for cyclic filtering processing to obtain the state estimation values corresponding to each sound source; perform coordinate transformation on the state estimation values corresponding to each sound source to obtain the estimated sound source position values corresponding to each sound source.

[0061] In some embodiments, when the processor 1001 implements coordinate transformation on the state estimation values corresponding to each sound source to obtain the sound source position estimation values corresponding to each sound source, it is used to implement:

[0062] Determine the first position estimation value of each sound source in the head coordinate system according to the state estimation value corresponding to each sound source; according to the rotation matrix, transform the first position estimation value corresponding to each sound source from the head coordinate system to the world coordinate system to obtain the second position estimation value corresponding to each sound source; determine the sound source position estimation value corresponding to each sound source according to the second position estimation value corresponding to each sound source.

[0063] In some embodiments, when the processor 1001 implements beamforming on the multi-channel time-frequency signal according to the head pose information, position information, and the sound source position estimation value corresponding to each sound source to obtain the beamformed signal, it is used to implement:

[0064] Calculate the target sound source direction corresponding to each sound source according to the head pose information and the sound source position estimation value corresponding to each sound source; perform beam calculation on the multi-channel time-frequency signal according to the target sound source direction, position information, and covariance matrix estimation result corresponding to each sound source to obtain the beamformed signal.

[0065] In some embodiments, the sound source position estimation value includes the second position estimation value corresponding to each sound source; when the processor 1001 implements sound source direction calculation according to the head pose information and the sound source position estimation value to obtain the target sound source direction, it is used to implement:

[0066] Perform head front direction calculation on the head pose information to obtain the head front direction; determine the target sound source according to the second position estimation value corresponding to each sound source and the head front direction, where the target sound source is the sound source whose included angle between the second position estimation value and the head front direction is less than the preset angle; determine the target sound source direction according to the direction of the target sound source relative to the microphone array.

[0067] In some embodiments, when the processor 1001 implements beam calculation on the multi-channel time-frequency signal according to the target sound source direction, position information, and covariance matrix estimation result to obtain the beamformed signal, it is used to implement:

[0068] Calculate the steering vector according to the target sound source direction and position information; perform weight calculation according to the steering vector and the covariance matrix estimation result to obtain the beamforming weight value; perform matrix calculation on the beamforming weight value and the multi-channel time-frequency signal to obtain the output signal matrix; determine the beamformed signal according to the output signal matrix.

[0069] In some embodiments, when the processor 1001 implements blind source separation of multi-channel time-frequency signals to obtain independent audio signals corresponding to each sound source, it is used to implement:

[0070] Extract features from the multi-channel time-frequency signals to obtain Mel-frequency cepstral coefficients; input the Mel-frequency cepstral coefficients into a trained deep learning model for mask estimation to obtain first mask estimation values corresponding to each sound source; based on the estimated value of the number of sound sources, perform signal separation on the multi-channel time-frequency signals according to the first mask estimation values corresponding to each sound source to obtain independent audio signals corresponding to each sound source.

[0071] In some embodiments, when the processor 1001 implements signal separation of the multi-channel time-frequency signals according to the estimated value of the number of sound sources and the first mask estimation values corresponding to each sound source to obtain independent audio signals corresponding to each sound source, it is used to implement:

[0072] Perform mask expansion on the first mask estimation values corresponding to each sound source to obtain second mask estimation values after mask expansion for each sound source; superimpose the second mask estimation values corresponding to each sound source onto the multi-channel time-frequency signals to obtain frequency-domain estimation signals of each sound source on each microphone channel; perform time-domain transformation on the frequency-domain estimation signals of each sound source on each microphone channel to obtain multi-channel time-domain signals corresponding to each sound source; determine independent audio signals corresponding to each sound source according to the multi-channel time-domain signals corresponding to each sound source.

[0073] In some embodiments, when the processor 1001 implements sound enhancement based on the beamformed signal and the independent audio signals corresponding to each sound source to obtain a target audio signal, it is used to implement:

[0074] Determine a target sound source from at least one sound source according to the head pose information and the estimated sound source positions corresponding to at least one sound source; perform frequency-domain transformation on the independent audio signals corresponding to each sound source to obtain time-frequency signals of each sound source; perform gain calculation according to the beamformed signal and the time-frequency signals of each sound source to obtain soft mask values of each sound source on each microphone channel; perform signal synthesis on the soft mask values of each sound source on each microphone channel to obtain a single-channel frequency-domain total signal; perform time-domain transformation on the single-channel frequency-domain total signal to obtain a target audio signal.

[0075] In some embodiments, when the processor 1001 implements gain calculation according to the beamformed signal and the time-frequency signals of each sound source to obtain soft mask values of each sound source on each microphone channel, it is used to implement:

[0076] Input the beamformed signal and the time-frequency signals of each sound source into the trained deep learning model for mask estimation to obtain the time-frequency mask values of each sound source on each microphone channel; determine the soft mask values of each sound source on each microphone channel according to the multi-channel time-frequency signals and the time-frequency mask values corresponding to each sound source signal.

[0077] In some embodiments, when the processor 1001 is implemented to synthesize signals of the soft mask values of each sound source on each microphone channel to obtain a single-channel frequency-domain total signal, it is used to implement:

[0078] Sum the soft mask values of each sound source on each microphone channel to obtain a single-channel frequency-domain signal corresponding to each sound source; add the single-channel frequency-domain signals corresponding to all sound sources to obtain a single-channel frequency-domain total signal.

[0079] In some embodiments, when the processor 1001 is implemented to play the target audio signal, it is used to implement:

[0080] According to the head pose information and the estimated sound source position corresponding to each sound source, allocate the target audio signal to the left and right channels of the wearable device; output the audio signal on the left channel and the audio signal on the right channel to the sound generating unit of the wearable device respectively.

[0081] In some embodiments, the target audio signal includes audio signals of at least one sound source; when the processor 1001 is implemented to allocate the target audio signal to the left and right channels of the wearable device according to the head pose information and the estimated sound source position corresponding to each sound source, it is used to implement:

[0082] Calculate the azimuth angle of the estimated sound source position corresponding to each sound source in the head coordinate system; determine the first gain of each sound source on the left channel and the second gain of each sound source on the right channel according to the azimuth angle corresponding to each sound source; multiply the audio signal of each sound source by the corresponding first gain of each sound source and then superimpose it on the left channel, and multiply the audio signal of each sound source by the corresponding second gain of each sound source and then superimpose it on the right channel respectively.

[0083] In some embodiments, the processor 1001 is further used to implement:

[0084] Obtain the head pose information of the user, the original audio signal received by the microphone array, and the position information of the microphone array; perform sound source localization based on the original audio signal, head pose information, and position information to obtain the estimated sound source position corresponding to at least one sound source; perform beamforming based on the original audio signal, head pose information, position information, and the estimated sound source position corresponding to each sound source to obtain a beamformed signal; perform blind source separation on the original audio signal to obtain the independent audio signal corresponding to each sound source; perform sound enhancement based on the beamformed signal and the independent audio signal corresponding to each sound source to obtain the target audio signal.

[0085] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other. Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a sound processing method provided by an embodiment of the present application. As Figure 2 shown, the sound processing method includes steps S101 to S103.

[0086] Step S101: Obtain the original audio signal received by the microphone array, where the original audio signal includes the audio signals corresponding to at least one sound source.

[0087] Exemplarily, the external audio signal can be received by the microphone array as the original audio signal. Among them, the microphone array includes multiple microphone channels, and one microphone corresponds to one microphone channel. The original audio signal can include the audio signals corresponding to at least one sound source received by at least one microphone channel. A sound source refers to the source of sound, such as an object or a person that emits sound.

[0088] Step S102: If it is detected that the user turns the head, obtain the head pose information of the user and the position information of the microphone array, and perform directional speech enhancement on the original audio signal according to the head pose information and the position information to obtain the target audio signal corresponding to the target sound source, where the target sound source is the sound source closest to the front of the user's head.

[0089] For example, it can be detected whether the user turns the head through an inertial measurement unit. When the inertial measurement unit detects a change in the rotation angle of the user's head, it can be confirmed that the user turns the head. For another example, it can be judged whether the user turns the head by detecting the position information of the microphone array.

[0090] Exemplarily, if it is detected that the user turns the head, obtain the head pose information of the user and the position information of the microphone array, and perform directional speech enhancement on the original audio signal according to the head pose information and the position information to obtain the target audio signal corresponding to the target sound source. The target audio signal can be expressed as Among them, the head pose information of the user can be collected by an inertial measurement unit; since the position of the microphone array in the wearable device (such as a glasses frame) is fixed, if the center of the head is assumed as the origin, the X, Y, and Z coordinates of each microphone in the microphone array are fixed. Therefore, the position information of the microphone array can be determined according to the X, Y, and Z coordinates of each microphone in the microphone array.

[0091] Please refer to Figure 3 , Figure 3 which is a schematic scenario diagram of directional voice enhancement provided by an embodiment of the present application. As Figure 3 shown, when the user is watching a concert, due to the noisy environment, the original audio signal received by the wearable device contains audio signals of multiple sound sources. At this time, the user can turn the head to listen to the target sound source that he wants to hear. For example, the user turns the head to face the actor who is singing or speaking. When the wearable device detects that the user turns the head, it can perform directional voice enhancement on the original audio signal according to the head pose information and the position information, obtain the target audio signal corresponding to the target sound source, and play the target audio signal. Thus, it can be realized that the wearable device performs directional voice enhancement on the original audio signal according to the head turn of the user, suppresses other sound sources and environmental noise, and improves the user's auditory experience.

[0092] Step S103, play the target audio signal.

[0093] Exemplarily, after obtaining the target audio signal corresponding to the target sound source, the target audio signal can be played. For example, the target audio signal can be played through a speaker in the wearable device. In some embodiments, the target audio signal can also be power-amplified by a power amplifier first, and then the power-amplified target audio signal can be played through the speaker.

[0094] In the above embodiment, when it is detected that the user turns the head, the head pose information of the user and the position information of the microphone array are obtained, and the original audio signal is directionally voice-enhanced according to the head pose information and the position information to obtain the target audio signal corresponding to the target sound source. It can be realized to perform directional voice enhancement based on the user's head rotation behavior and suppress other sound sources and environmental noise. It can be applied to various application scenarios, can effectively filter out noise, improve the user's auditory experience, and at the same time can play a role in hearing assistance to help users with hearing defects improve their hearing.

[0095] Please refer to Figure 4 , Figure 4 which is a software flowchart of a sound processing method provided by an embodiment of the present application. As Figure 4As shown, the sound processing method mainly includes four parts: sound source localization, beamforming, blind source separation, and sound enhancement and suppression. The following will elaborate on these four parts of sound source localization, beamforming, blind source separation, and sound enhancement and suppression respectively. During the design of the system software, the low-power requirements in the actual application scenario are fully considered. At the same time, the software itself has strong scalability, and some links can be trimmed according to the actual hardware performance and application scenario, making the application scenario more flexible and diverse.

[0096] Please refer to Figure 5 , Figure 5 which is a schematic block diagram of the hardware system of a sound processing method provided by an embodiment of the present application. As Figure 5 shown, the hardware system of the sound processing method may include a microphone array, an inertial measurement unit, a system-on-chip, a digital signal processor, a neural processing unit, and a power amplifier.

[0097] Among them, the microphone array may be a PDM (Pulse Density Modulation) micro digital microphone, including two or more microphones, which are used to capture audio signals. The inertial measurement unit contains sensor devices such as gyroscopes and accelerometers to capture and estimate the head movement of the user. The system-on-chip (SoC) is responsible for collecting sensor signals and performing algorithm processing. The digital signal processor (Digital Signal Processor, DSP) is mainly responsible for processing various digital signals and has a relatively fast processing speed. In the embodiment of the present application, by using the digital signal processor to process various digital signals of sound, the speed of processing sound signals can be effectively improved. The digital signal processor can be an external device or a sub-unit integrated in the SoC. The neural processing unit (Neural Processing Unit, NPU) is mainly used to accelerate neural network calculations; the power amplifier (Power Amplifier, PA) is used to drive the open speaker to emit sound.

[0098] In the embodiment of the present application, as Figure 4 shown, since the original audio signal is an audio signal in the time domain, in order to facilitate subsequent algorithms to process the original audio signal, before performing directional voice enhancement on the original audio signal, preprocessing can also be performed on the original audio signal. The preprocessing mainly includes frequency domain transformation and covariance matrix estimation. Among them, the frequency domain transformation refers to performing a short-time Fourier transform on the original audio signal to obtain a time-frequency signal, taking into account both time domain and frequency domain characteristics. The covariance matrix is used to describe the correlation between the signals received by the microphone array in different channels and reflects the spatial distribution of the sound source and noise. Exemplarily, the frequency domain transformation mainly uses the short-time Fourier transform. The following will elaborate on the frequency domain transformation of the original audio signal in detail.

[0099] Please refer to Figure 6 , Figure 6 which is a schematic flowchart for preprocessing an audio signal provided by an embodiment of the present application. As Figure 6 shown, it may include steps S201 to S203.

[0100] Step S201: Perform a frequency-domain transformation on the audio signal received by each microphone channel to obtain a time-frequency signal corresponding to each microphone channel.

[0101] Exemplarily, when performing a frequency-domain transformation on the audio signal received by each microphone channel, the audio signal corresponding to each microphone channel can be framed, windowed, and Fourier-transformed in sequence to obtain a time-frequency signal corresponding to each microphone channel. The specific steps are as follows:

[0102] (1) Initialize a preset window function. The preset window function can be a Hanning window function or other types of window functions, and the present application does not limit this. In the embodiment of the present application, the Hanning window function is taken as an example for illustration. For example, the Hanning window function is expressed as , where the window length is , the frame length is , and the frame shift is .

[0103] (2) Frame the audio signal. The audio signal of each microphone channel can be divided into multiple overlapping short-time frame signals and frame shift according to the frame length , where represents the time frame index and represents the sample index within the frame, then .

[0104] (3) Window each short-time frame signal. Multiply each frame by the window function to obtain the windowed signal .

[0105] (4) Perform a short-time Fourier transform (STFT) on the windowed signal of each frame to obtain a time-frequency signal corresponding to each microphone channel, which can be expressed as , where .

[0106] Step S202: Combine the time-frequency signals corresponding to each microphone channel to obtain a multi-channel time-frequency signal.

[0107] Exemplarily, the time-frequency signals corresponding to each microphone channel can be combined to obtain a multi-channel time-frequency signal. The multi-channel time-frequency signal can be expressed as:

[0108] .

[0109] In the formula, represents the total number of microphone channels. It can be understood that the multi-channel time-frequency signal includes the time-frequency signals corresponding to multiple microphone channels.

[0110] Step S203: Estimate the covariance matrix of the multi-channel time-frequency signal to obtain the covariance matrix estimation result corresponding to the multi-channel time-frequency signal.

[0111] Exemplarily, estimating the covariance matrix of the multi-channel time-frequency signal may include the following steps:

[0112] (1) Average the multi-channel time-frequency signal in the frequency domain to obtain the multi-channel time-frequency average signal , wherein, is the number of frequency points after STFT transformation.

[0113] (2) Calculate the instantaneous covariance matrix according to the multi-channel time-frequency average signal , wherein the instantaneous covariance matrix , wherein, represents the conjugate transpose, is the matrix of .

[0114] (3) Estimate the instantaneous covariance matrix based on the recursive averaging method to obtain the covariance matrix estimation result. The covariance matrix estimation result can be expressed as , where is the matrix of , is the forgetting factor.

[0115] Please refer to Figure 7 , Figure 7 which is a schematic flowchart of the second sound processing method provided by the embodiments of the present application. As Figure 7 shown, it may include steps S301 to S304.

[0116] Step S301: Perform a frequency-domain transformation on the audio signal received by each microphone channel to obtain the time-frequency signal corresponding to each microphone channel.

[0117] Step S302: Combine the time-frequency signals corresponding to each microphone channel to obtain a multi-channel time-frequency signal.

[0118] Step S303: Estimate the covariance matrix of the multi-channel time-frequency signal to obtain the covariance matrix estimation result corresponding to the multi-channel time-frequency signal.

[0119] It can be understood that steps S301 to S303 are the same as steps S201 to S203 above and will not be elaborated here.

[0120] Step S304: Estimate the number of sound sources based on the covariance matrix estimation result to obtain the estimated value of the number of sound sources.

[0121] In the embodiments of the present application, since there may be multiple sound sources and subsequent signal separation of multiple sound sources is required, it is necessary to determine the number of sound sources.

[0122] As Figure 4 shown, estimating the number of sound sources based on the covariance matrix estimation result may include the following steps:

[0123] (1) Perform eigenvalue decomposition on the covariance matrix estimation result to obtain eigenvalues . Among them, the obtained eigenvector matrix is discarded.

[0124] (2) Calculate the MDL (Minimum description length) criterion value for each possible number of sound sources k. Among them, the MDL criterion value is:

[0125] ,

[0126] In the formula, is the likelihood function, indicating the likelihood of observing the multi-channel time-frequency signal assuming there are sound sources. is the number of samples, that is, the number of frames of STFT.

[0127] Among them,

[0128] is the degree of freedom of the model, indicating the number of parameters to be estimated in the model assuming there are sound sources. Among them, .

[0129] (3) Estimate the number of sound sources. For each frame , select the that minimizes as the estimated value of the number of sound sources for this frame , perform smoothing processing on the of multiple frames to obtain the final estimated value of the number of sound sources , and use subsequently.Represents the estimated value of the number of sound sources.

[0130] In the above embodiments, by estimating the number of sound sources based on the covariance matrix estimation result, the number of sound sources included in the audio signal received by the microphone array can be estimated. By using the MDL criterion based on eigenvalue decomposition to estimate the number of sound sources, the power consumption is lower and the real-time performance is higher.

[0131] Please refer to Figure 8 , Figure 8 which is a schematic flowchart of a directional speech enhancement provided by an embodiment of the present application. As Figure 8 shown, in step S102, performing directional speech enhancement on the original audio signal according to the head pose information and position information to obtain the target audio signal corresponding to the target sound source may include steps S401 to S404.

[0132] Step S401: Perform sound source localization according to the multi-channel time-frequency signal, head pose information, and position information to obtain the estimated value of the sound source position corresponding to at least one sound source.

[0133] In the embodiment of the present application, as Figure 4 shown, by performing sound source localization according to the multi-channel time-frequency signal, head pose information, and position information, the positions of each sound source in space can be estimated, and the Kalman filtering algorithm is combined to provide more accurate direction information for the subsequent beamforming step.

[0134] Please refer to Figure 9 , Figure 9 which is a schematic flowchart of a sound source localization provided by an embodiment of the present application. As Figure 9 shown, in step S401, performing sound source localization according to the multi-channel time-frequency signal, head pose information, and position information may include steps S4011 and S4012.

[0135] Step S4011: Determine the estimated value of the time delay corresponding to the time difference of arrival of the multi-channel time-frequency signal.

[0136] Exemplarily, based on the Generalized Cross Correlation Phase Transformation (GCC-PHAT), according to the multi-channel time-frequency signal, calculate the estimated value of the time delay corresponding to the Time Difference of Arrival (TDOA) of the multi-channel time-frequency signal.

[0137] In some embodiments, determining the delay estimation value corresponding to the time difference of arrival of a multi-channel time-frequency signal may include: constructing a cross-correlation function corresponding to the multi-channel time-frequency signal; constructing a generalized cross-correlation function according to the Fourier transform of the cross-correlation function; and determining the delay estimation value according to the peak position of the generalized cross-correlation function.

[0138] Exemplarily, the time-frequency signals corresponding to each microphone channel in the multi-channel time-frequency signal may be constructed The cross-correlation function can be expressed as:

[0139] , where denotes the complex conjugate of

[0140] Calculate the Fourier transform of the cross-correlation function : , and construct a generalized cross-correlation function, which can be expressed as:

[0141] .

[0142] According to the peak position (i.e., the maximum value) of the generalized cross-correlation function , determine the delay estimation value, where the delay estimation value can be expressed as . For example, the value corresponding to the peak position of the generalized cross-correlation function can be determined as the delay estimation value .

[0143] Step S4012: Based on the Kalman filter, perform sound source localization according to the delay estimation value, the head pose information, and the position information, and obtain the sound source position estimation value corresponding to each sound source.

[0144] Exemplarily, the position information may include the first position of the microphone array. It should be noted that coordinate transformation needs to be performed on the position information of the microphone array. To distinguish the position information before and after coordinate transformation, the position of the microphone array before coordinate transformation can be denoted as the first position .

[0145] In some embodiments, based on a Kalman filter, sound source localization is performed according to the time delay estimation value, head pose information, and position information to obtain a sound source position estimation value corresponding to each sound source, which may include: determining a rotation matrix of the head coordinate system used for the head pose information relative to the world coordinate system; according to the rotation matrix, converting the first position from the world coordinate system to the head coordinate system to obtain a second position; inputting the time delay estimation value and the second position into the Kalman filter for iterative filtering processing to obtain a state estimation value corresponding to each sound source; performing coordinate transformation on the state estimation value corresponding to each sound source to obtain a sound source position estimation value corresponding to each sound source.

[0146] Exemplarily, the Kalman filter can be initialized first, and the specific operations include:

[0147] (1) Construct a state vector. For each sound source , there is: , where is the sound source position, is the sound source velocity (initial value is 0).

[0148] (2) Establish a state transition matrix according to the sound source motion model: . The sound source motion model may include, but is not limited to, a uniform motion model, a uniformly accelerated motion model, and a random walk model.

[0149] (3) Construct a process noise covariance matrix: .

[0150] (4) Construct an observation matrix: , and the covariance matrix of the observation noise: .

[0151] (5) Initialize the state estimation value and the error covariance matrix .

[0152] Exemplarily, according to the head pose information , the rotation matrix of the head coordinate system used for the head pose information relative to the world coordinate system can be calculated, and the first position is converted from the world coordinate system to the head coordinate system to obtain the second position as: .

[0153] Input the time delay estimation value and the second position into the Kalman filter for iterative filtering processing. Among them, for each sampling moment , for each microphone and each sound source , the specific process of the iterative filtering processing includes prediction calculation and result update.​

[0154] The prediction calculation may include the following steps:

[0155] (1) State prediction: ;

[0156] (2) Error covariance prediction: ;

[0157] (3) Calculate the observed value of the time difference of arrival: ;

[0158] (4) Calculate the Kalman gain: .

[0159] The result update may include the following steps:

[0160] (1) State update: ; where represents the state estimate value corresponding to the sound source.

[0161] (2) Error covariance update: , ([[]] is the identity matrix).

[0162] Exemplarily, through the above calculation formula, the state estimate value corresponding to each sound source can be calculated, and then the coordinate system conversion is performed on the state estimate value corresponding to each sound source to obtain the sound source position estimate value corresponding to each sound source.

[0163] In some embodiments, performing coordinate system conversion on the state estimate value corresponding to each sound source to obtain the sound source position estimate value corresponding to each sound source may include: determining the first position estimate value of each sound source in the head coordinate system according to the state estimate value corresponding to each sound source; converting the first position estimate value corresponding to each sound source from the head coordinate system to the world coordinate system according to the rotation matrix to obtain the second position estimate value corresponding to each sound source; and determining the sound source position estimate value corresponding to each sound source according to the second position estimate value corresponding to each sound source.

[0164] Exemplarily, the first position estimate value of each corresponding sound source in the head coordinate system can be extracted from the state estimate value corresponding to each sound source . According to the rotation matrix , the first position estimate value corresponding to each sound source is converted from the head coordinate system to the world coordinate system to obtain the second position estimate value corresponding to each sound source, and the second position estimate value can be expressed as . Finally, the second position estimate value corresponding to each sound source is determined as the sound source position estimate value corresponding to each sound source, that is, the sound source position estimate value can be expressed as 。

[0165] In some embodiments, the second position estimation values corresponding to each sound source may also be fused to obtain a sound source position estimation value corresponding to all sound sources , where represents the second position estimation value of the th sound source.

[0166] In the above embodiments, by performing sound source localization based on multi-channel time-frequency signals, head pose information, and position information, a sound source position estimation value corresponding to the sound source can be obtained, providing more accurate direction information for subsequent beamforming.

[0167] Step S402: Perform beamforming on the multi-channel time-frequency signals according to the head pose information, position information, and the sound source position estimation value corresponding to each sound source to obtain a beamformed signal.

[0168] It should be noted that the main function of beamforming is to perform spatial filtering, optimize the directivity of the array receiving system, and ensure that the beam direction is consistent with the designed direction to maximize the amplification of the signal in the observation direction. In the embodiments of the present application, through beamforming, the sound in the direction of the target sound source can be enhanced.

[0169] Please refer to Figure 10 , Figure 10 which is a schematic flowchart of a beamforming provided by an embodiment of the present application. As Figure 10 shown, performing beamforming on the multi-channel time-frequency signals according to the head pose information, position information, and the sound source position estimation value corresponding to each sound source in step S402 may include step S4021 and step S4022.

[0170] Step S4021: Calculate the sound source direction according to the head pose information and the sound source position estimation value corresponding to each sound source to obtain the target sound source direction corresponding to each sound source.

[0171] In some embodiments, the sound source position estimation value includes the second position estimation value corresponding to each sound source; calculating the sound source direction according to the head pose information and the sound source position estimation value to obtain the target sound source direction may include: calculating the front direction of the head for the head pose information to obtain the front direction of the head; determining the target sound source according to the second position estimation value corresponding to each sound source and the front direction of the head, where the target sound source is the sound source with an angle between the second position estimation value and the front direction of the head less than a preset angle; determining the target sound source direction according to the direction of the target sound source relative to the microphone array.

[0172] Exemplarily, the calculation of the front direction of the head may include the following steps:

[0173] (1) According to the head pose information , calculate the rotation matrix of the head coordinate system in the world coordinate system .

[0174] (2) The forward direction of the head in the head coordinate system is .

[0175] (3) Convert the forward direction of the head from the head coordinate system to the world coordinate system: .

[0176] Exemplarily, according to the second position estimate corresponding to each sound source and the forward direction of the head, determining the target sound source may include: calculating each sound source and the forward direction of the head of the included angle:

[0177] ,

[0178] find the sound source with the smallest included angle, which is the target sound source .

[0179] Calculate the direction of the target sound source relative to the microphone array to obtain the target sound source direction. For example, when the target sound source is the first microphone of the microphone array, the target sound source direction can be expressed as . Again, for example, when the target sound source is the second microphone of the microphone array, the target sound source direction can be expressed as .

[0180] Step S4022: Perform beam calculation on the multi-channel time-frequency signal according to the target sound source direction, position information, and covariance matrix estimation result to obtain a beamformed signal.

[0181] In the embodiments of the present application, as Figure 4 shown, after calculating the target sound source direction based on the head pose information and the sound source position estimate value, a beamforming weight value can be calculated based on an MVDR (Minimum Variance Distortionless Response) beamformer according to the target sound source direction, position information, and covariance matrix estimation result, and a beamformed signal can be determined according to the beamforming weight value and the multi-channel time-frequency signal. It should be noted that the beamforming weight value refers to the weight value corresponding to the multi-channel time-frequency signal during beamforming.

[0182] In some embodiments, beamforming calculation is performed on the multi-channel time-frequency signal according to the target sound source direction, position information, and covariance matrix estimation result to obtain a beamformed signal, which may include: calculating a steering vector according to the target sound source direction and position information; calculating a beamforming weight value according to the steering vector and the covariance matrix estimation result; performing matrix calculation on the beamforming weight value and the multi-channel time-frequency signal to obtain an output signal matrix; and determining the beamformed signal according to the output signal matrix.

[0183] Exemplarily, for each frame and each frequency , according to the target sound source direction and the position information of the microphone array , the steering vector is calculated as follows:

[0184] .

[0185] Based on the steering vector and the covariance matrix estimation result , weight calculation is performed to obtain the beamforming weight value, as shown below:

[0186]

[0187] In the formula, represents the transpose matrix of the steering vector , and represents the inverse matrix of the covariance matrix estimation result .

[0188] Based on the beamforming weight value and the multi-channel time-frequency signal, matrix calculation is performed to obtain an output signal matrix, and the output signal matrix can be expressed as ; based on the output signal matrix , it is determined as the beamformed signal, and the beamformed signal can be expressed as .

[0189] In the above embodiments, by performing beamforming according to the multi-channel time-frequency signal, head pose information, position information, and the estimated sound source position value corresponding to each sound source, spatial filtering of the multi-channel time-frequency signal can be achieved, the directivity of the array receiving system can be optimized, and only when the beam direction is consistent with the designed direction can the signal in the observation direction be maximally amplified, thereby enhancing the sound in the target sound source direction.

[0190] Step S403: Perform blind source separation on the multi-channel time-frequency signal to obtain an independent audio signal corresponding to each sound source.

[0191] It should be noted that the function of blind source separation is to analyze the unobserved original signals from multiple observed mixed signals. Usually, the observed mixed signals come from the outputs of multiple sensors, and the output signals of the sensors are independent (linearly uncorrelated). In the embodiments of the present application, by performing blind source separation on multi-channel time-frequency signals, it is possible to separate the interfering speech in the multi-channel time-frequency signals and suppress the background noise, thereby improving the speech quality.

[0192] Please refer to Figure 11 , Figure 11 which is a schematic flowchart of blind source separation provided by the embodiments of the present application. As Figure 11 shown, in step S403, blind source separation is performed on the multi-channel time-frequency signals to obtain independent audio signals corresponding to each sound source, which may include step S4031 and step S4033.

[0193] Step S4031: Extract features from the multi-channel time-frequency signals to obtain Mel-frequency cepstral coefficients.

[0194] In the embodiments of the present application, considering the auditory characteristics of the human ear, Mel-frequency cepstral coefficients (MFCC) are used for feature extraction.

[0195] Exemplarily, the specific process of extracting features from the multi-channel time-frequency signals is as follows:

[0196] (1) Calculate the energy spectrum of the multi-channel time-frequency signals. For each frame and each audio channel take the square of the modulus to obtain the energy spectrum .

[0197] (2) Mel-filter preprocessing:

[0198] a) The number of filters to be constructed is , the lowest frequency , the highest frequency ;

[0199] b) Convert the lowest and highest frequency points to Mel frequencies : ;

[0200] c) Divide the Mel frequency range of each filter into points;

[0201]

[0202]

[0203] d) Convert the Mel frequency points back to linear frequency points:

[0204]

[0205] (3) According to 、 、 , design the Mel filter bank , representing the response of the -th filter at frequency .

[0206] (4) Construct the Mel filter bank:

[0207] a) Initialize the filter bank as a zero matrix with shape , where is the number of points in the STFT (usually the frame length).

[0208] b) For each filter :

[0209] i. Select three frequency points ;

[0210] ii. For each frequency point :

[0211] 1. If : then

[0212] ;

[0213] 2. If : then

[0214] ;

[0215] 3. Otherwise: .

[0216] (5) Mel filtering:

[0217] For each time frame and each microphone channel , calculate the Mel filter bank energy .

[0218] .

[0219] (6) Logarithmic transformation:

[0220] Perform a logarithmic transformation on the result of Mel filtering to compress the dynamic range, making the features less sensitive to energy changes and improving the robustness of subsequent processing.

[0221] 。

[0222] (7) DCT Discrete Cosine Transform:

[0223] For each time frame and each audio channel , perform DCT (Discrete Cosine Transform) on the result of the logarithmic transform to finally obtain the Mel-frequency cepstral coefficients , where , is the number of required Mel-frequency cepstral coefficients (determined according to the feature dimension and the degree of information retention).

[0224] 。

[0225] Step S4032: Input the Mel-frequency cepstral coefficients into the trained deep learning model for mask estimation to obtain the first mask estimation value corresponding to each sound source.

[0226] In the embodiment of the present application, after performing feature extraction on the multi-channel time-frequency signal to obtain the Mel-frequency cepstral coefficients, mask estimation also needs to be performed based on the Mel-frequency cepstral coefficients.

[0227] It should be noted that the deep learning model can be trained with a large amount of data to learn complex noise patterns, so as to show better robustness in a noisy environment, which is particularly important for the application of smart glasses in noisy environments such as outdoors and public places. Among them, the deep learning model can include, but is not limited to, models such as CNN (Convolution Neural Network), RNN (Recurrent Neural Networks), and BM (Boltzmann machines). In the embodiment of the present application, taking the deep learning model as a convolutional neural network model as an example, how to perform mask estimation is described.

[0228] Please refer to Figure 12 , Figure 12 is a schematic structural diagram of a convolutional neural network model provided by the embodiment of the present application. As Figure 12 shown, the structure of the convolutional neural network model includes:

[0229] Input layer: Use MFCC features as the model input.

[0230] CNN layer (the number of layers is determined according to the actual effect and computational complexity), which can include a convolutional layer and a pooling layer.

[0231] Convolutional layer: Multiple convolutional layers are used to extract local features in the time and frequency domains. Usually, a batch normalization layer and a ReLU activation function are followed after each convolutional layer.

[0232] Pooling layer: After some convolutional layers, a pooling layer is used to reduce the dimension of the feature map and increase the robustness of the model.

[0233] Output layer:

[0234] Dense fully connected layer: Maps the features extracted by the CNN layer to the time-frequency masks of

[0235] Sigmoid activation function: Limits the output value between 0 and 1 to obtain a soft mask.

[0236] The Reshape layer adjusts the dimension of the output matrix to so as to correspond to the multi-channel audio signal correspondingly.

[0237] The loss function of the convolutional neural network model is used to calculate the mean square error between the estimated mask output by the model and the true mask. The loss function is as follows:

[0238] .

[0239] In the formula, is the number of sound sources, is the number of time frames for time-frequency analysis, is the number of filters in the Mel filter bank, is the number of audio channels (in this embodiment of the application, 3 audio channels are taken as an example), is the th sound source at time frame , frequency point , channel on the true mask, while is the estimated mask output by the convolutional neural network model, denoted as the first mask estimate value.

[0240] In this embodiment of the application, the convolutional neural network model can be iteratively trained in advance to obtain a trained convolutional neural network model. Among them, when preparing training data, a large amount of multi-channel mixed speech and clean speech data corresponding to each sound source can be collected according to the actual application scenario first; then, the mixed speech and clean speech data are processed through the foregoing steps to obtain the MFCC feature extraction result , the mask result output by the model , the feature extraction result of the clean speech ; perform STFT transformation on to obtain ; according to and Calculate the true mask After obtaining the training data, the convolutional neural network model can be iteratively trained based on the training data until the loss value of the loss function corresponding to the convolutional neural network model no longer decreases or is less than a preset threshold, and the trained convolutional neural network model is obtained.

[0241] Exemplarily, the Mel-frequency cepstral coefficients can be input into the trained convolutional neural network model for mask estimation to obtain a first mask estimation value corresponding to each sound source. The specific process of mask estimation can be referred to the above description and will not be elaborated here.

[0242] Step S4033: Based on the estimated number of sound sources, perform signal separation on the multi-channel time-frequency signal according to the first mask estimation value corresponding to each sound source to obtain an independent audio signal corresponding to each sound source.

[0243] In the embodiment of the present application, after inputting the Mel-frequency cepstral coefficients into the trained deep learning model for mask estimation to obtain a first mask estimation value corresponding to each sound source, it is necessary to perform signal separation on the multi-channel time-frequency signal according to the first mask estimation value corresponding to each sound source to obtain an independent audio signal corresponding to each sound source.

[0244] In some embodiments, based on the estimated number of sound sources, performing signal separation on the multi-channel time-frequency signal according to the first mask estimation value corresponding to each sound source to obtain an independent audio signal corresponding to each sound source may include: performing mask expansion on the first mask estimation value corresponding to each sound source to obtain a second mask estimation value after mask expansion for each sound source; superimposing the second mask estimation value corresponding to each sound source onto the multi-channel time-frequency signal to obtain a frequency-domain estimation signal of each sound source on each microphone channel; performing time-domain transformation on the frequency-domain estimation signal of each sound source on each microphone channel to obtain a multi-channel time-domain signal corresponding to each sound source; and determining an independent audio signal corresponding to each sound source according to the multi-channel time-domain signal corresponding to each sound source.

[0245] It should be noted that since the input of the convolutional neural network model is MFCC features and its frequency dimension is the number of Mel filter banks , while the multi-channel time-frequency signal and the beamforming signal have a frequency dimension of the number of frequency points after STFT transformation , it is necessary to expand the first mask estimation value output by the convolutional neural network model to the same dimension as and through linear interpolation to obtain a second mask estimation value after mask expansion Among them, for the specific process of expanding dimensions by the linear interpolation method, reference can be made to related technologies, which will not be elaborated here.

[0246] Exemplarily, for each sound source to be separated and each microphone channel perform the following processing: Using the soft masking processing method, superimpose the expanded time-frequency mask onto the multi-channel time-frequency signal to obtain the frequency-domain estimated signal of the sound source on the corresponding microphone channel , where, represents element-wise multiplication.

[0247] Exemplarily, the frequency-domain estimated signals of each sound source on each microphone channel can be subjected to ISTFT (Inverse Short-Time Fourier Transform) to obtain the multi-channel time-domain signal corresponding to each sound source ; According to the multi-channel time-domain signal corresponding to each sound source , determine the independent audio signal corresponding to each sound source, and the independent audio signal can be expressed as .

[0248] In the above embodiment, by performing blind source separation on the multi-channel time-frequency signal, it is possible to separate the interfering speech in the multi-channel time-frequency signal and suppress background noise, thereby improving the speech quality.

[0249] Step S404, perform sound enhancement based on the beamforming signal and the independent audio signal corresponding to each sound source to obtain the target audio signal.

[0250] In the embodiment of the present application, as Figure 4 shown, although the multi-channel time-frequency signal is processed to a certain extent through beamforming and blind source separation, in order to obtain better speech clarity and intelligibility, sound enhancement and suppression are still required to solve the interference from other directions that cannot be completely eliminated by the beamforming algorithm, as well as problems such as noise residue and signal distortion existing in the blind source separation algorithm.

[0251] Please refer to Figure 13 , Figure 13 which is a schematic flowchart of a sound enhancement provided by the embodiment of the present application. As Figure 13 shown, the sound enhancement performed in step S404 based on the beamforming signal and the independent audio signal corresponding to each sound source may include step S4041 and step S4045.

[0252] Step S4041, determine the target sound source from at least one sound source according to the head pose information and the estimated sound source position value corresponding to at least one sound source.

[0253] In some embodiments, the sound source direction can be calculated based on the head pose information and the estimated sound source position corresponding to each sound source, so as to obtain the target sound source direction corresponding to each sound source. Based on the target sound source direction corresponding to each sound source and the estimated sound source position corresponding to each sound source, the target sound source is determined.

[0254] Exemplarily, when the user hopes to hear the sound source directly in front, the target sound source is the sound source closest to the front of the head. First, the front direction of the head in the world coordinate system can be calculated , and the estimated sound source position with the smallest included angle with the front direction of the head is found , and the sound source corresponding to the estimated sound source position is used as the target sound source , where the specific calculation method is the same as the calculation method for sound source selection in step S4021 above. as the target sound source , where the specific calculation method is the same as the calculation method for sound source selection in the above step S4021.

[0255] Step S4042: Perform a frequency-domain transformation on the independent audio signal corresponding to each sound source to obtain the time-frequency signal of each sound source.

[0256] Exemplarily, the independent audio signal corresponding to each sound source can be subjected to a frequency-domain transformation to obtain the time-frequency signal of each sound source . Among them, the frequency-domain transformation refers to the STFT transformation, and the parameters used in the frequency-domain transformation are the same as those used in the frequency-domain transformation in step S201 above, and the specific process will not be elaborated here.

[0257] Step S4043: Perform gain calculation based on the beamformed signal and the time-frequency signal of each sound source to obtain the soft masking value of each sound source on each microphone channel.

[0258] In some embodiments, performing gain calculation based on the beamformed signal and the time-frequency signal of each sound source to obtain the soft masking value of each sound source on each microphone channel may include: inputting the beamformed signal and the time-frequency signal of each sound source into a trained deep learning model for masking estimation to obtain the time-frequency masking value of each sound source on each microphone channel; determining the soft masking value of each sound source on each microphone channel according to the multi-channel time-frequency signal corresponding to each sound source signal and the time-frequency masking value.

[0259] Exemplarily, for the time frame , frequency and microphone channel , a simplified deep learning model can be used to input the beamformed signal and the time-frequency signal of each sound source As input, predict the time-frequency masking values of each sound source on each microphone channel Among them, the model structure and training method of the deep learning model are similar to the CNN model used in blind source separation in steps S4031 to S4033. Parameters affecting power consumption such as the number of layers and the number of convolutions can be simplified according to actual applications.

[0260] When predicting the time-frequency masking values of each sound source on each microphone channel After that, based on the multi-channel time-frequency signal and the time-frequency masking value corresponding to each sound source signal, the soft masking value of each sound source on each microphone channel can be determined, where the soft masking value can be expressed as , denotes element-wise multiplication.

[0261] Step S4044: Synthesize the soft masking values of each sound source on each microphone channel to obtain a single-channel frequency-domain total signal.

[0262] Exemplarily, after obtaining the soft masking values of each sound source on each microphone channel, the soft masking values of each sound source on each microphone channel can be synthesized to obtain a single-channel frequency-domain total signal.

[0263] In some embodiments, synthesizing the soft masking values of each sound source on each microphone channel to obtain a single-channel frequency-domain total signal may include: summing the soft masking values of each sound source on each microphone channel to obtain a single-channel frequency-domain signal corresponding to each sound source; adding the single-channel frequency-domain signals corresponding to all sound sources to obtain a single-channel frequency-domain total signal.

[0264] Exemplarily, for each sound source , sum the soft masking values in the microphone channel dimension to obtain the single-channel frequency-domain signal of this sound source:[[]] . Add the single-channel frequency-domain signals of each sound source to obtain the final single-channel frequency-domain total signal, where the single-channel frequency-domain total signal can be expressed as:[[]]

[0265] .

[0266] Step S4045: Perform a time-domain transformation on the single-channel frequency-domain total signal to obtain the target audio signal.

[0267] Exemplarily, perform a time-domain transformation on each frame of the single-channel frequency-domain total signal. Among them, the time-domain transformation can be an inverse Fourier transform, and the inverse Fourier transform can be performed using the same parameters as the frequency-domain transformation in step 201 above to obtain the windowed time-domain signal:[[]]

[0268] 。

[0269] In the formula, represents the in-frame sample index. The windowed time-domain signal is superimposed on the corresponding position of the output signal to obtain the target audio signal , where 。

[0270] In the above embodiments, by performing sound enhancement based on the beam synthesis signal and the independent audio signal corresponding to each sound source, sound enhancement and suppression of the audio signal can be achieved, obtaining better speech clarity and intelligibility. At the same time, it can also solve the interference in other directions that cannot be completely eliminated by the beam synthesis algorithm, as well as problems such as noise residue and signal distortion existing in the blind source separation algorithm.

[0271] In the embodiments of the present application, when playing the target audio signal in step S103, in order to achieve stereo playback, the capabilities of the target audio signal can be allocated to the left and right channels of the wearable device according to the estimated sound source position and the head pose information.

[0272] In some embodiments, playing the target audio signal may include: allocating the target audio signal to the left and right channels of the wearable device according to the head pose information and the estimated sound source position corresponding to each sound source; respectively outputting the audio signal on the left channel and the audio signal on the right channel to the sound generating unit of the wearable device.

[0273] Exemplarily, the target audio signal may include the audio signals of at least one sound source. In some implementation manners, allocating the target audio signal to the left and right channels of the wearable device according to the head pose information and the estimated sound source position corresponding to each sound source may include: calculating the azimuth angle of the estimated sound source position corresponding to each sound source in the head coordinate system; determining the first gain on the left channel and the second gain on the right channel for each sound source according to the azimuth angle corresponding to each sound source; respectively multiplying the audio signal of each sound source by the corresponding first gain of each sound source and superimposing it on the left channel, and respectively multiplying the audio signal of each sound source by the corresponding second gain of each sound source and superimposing it on the right channel.

[0274] Exemplarily, for the estimated sound source position corresponding to each sound source, when calculating its azimuth angle in the head coordinate system respectively, first convert each estimated sound source position from the world coordinate system to the head coordinate system:

[0275] ,

[0276] where is the rotation matrix of the head coordinate system relative to the world coordinate system.

[0277] Exemplarily, for the position of the estimated value of the sound source position corresponding to each sound source in the head coordinate system , calculate the azimuth angle , where is the four-quadrant arctangent function, and the return value range is , which can be converted to other ranges according to actual application needs, and the unit is radians. The azimuth angle is as follows:

[0278] .

[0279] According to the calculation result of the azimuth angle , calculate the first gain of each sound source on the left channel , and the second gain on the right channel , as follows:

[0280] ;

[0281] Among them, the calculation method of the gain can be a linear method or a piecewise method. In the embodiments of the present application, a more flexible piecewise method can be adopted. It should be noted that the determination of the angle of the sound source being left or right can be determined according to the actual application situation (not necessarily ), and the gain value settings of the left and right channels in the same piecewise function should be selected according to the actual situation (it is also allowed for the user to set by themselves). Only an example is given here.

[0282] Exemplarily, in step S4045, after performing a time-domain transformation on the single-channel frequency-domain total signal to obtain the target audio signal, the target audio signal can be multiplied by the first gain corresponding to the left channel to obtain the first audio signal and the target audio signal can be multiplied by the second gain corresponding to the right channel to obtain the second audio signal ; then the first audio signal and the second audio signal are respectively superimposed on the left and right channels and output to the sound generating unit or the power amplifier unit corresponding to the sound generating unit. Among them, as Figure 5 shown, the sound generating unit can include sound emitting devices such as speakers, and the power amplifier unit can include a power amplifier.

[0283] Please refer to Figure 14 , Figure 14 is a schematic flowchart of the third sound processing method provided by the embodiments of the present application. As Figure 14 shown, it may include steps S501 to S505.

[0284] Step S501: Obtain the head pose information of the user, the original audio signal received by the microphone array, and the position information of the microphone array.

[0285] Exemplarily, the head pose information of the user can be collected by an inertial measurement unit, and the external audio signal received by the microphone array can be used as the original audio signal.

[0286] Step S502: Perform sound source localization based on the original audio signal, the head pose information, and the position information to obtain the estimated sound source position corresponding to at least one sound source.

[0287] Step S503: Perform beamforming on the original audio signal according to the head pose information, the position information, and the estimated sound source position corresponding to each sound source to obtain a beamformed signal.

[0288] Step S504: Perform blind source separation on the original audio signal to obtain the independent audio signal corresponding to each sound source.

[0289] Step S505: Perform sound enhancement based on the beamformed signal and the independent audio signal corresponding to each sound source to obtain the target audio signal.

[0290] It should be noted that before step S502, the original audio signal can be preprocessed to obtain a multi-channel time-frequency signal. The specific process of the preprocessing can refer to the above steps S201 to S202 and will not be elaborated here. In the subsequent steps S502, S503, and S504, sound source localization, beamforming, and blind source separation can be respectively performed on the multi-channel time-frequency signal. It can be understood that steps S502 to S504 are the same as steps S401 to S404 above, and the specific process can refer to the detailed description of the above embodiments and will not be elaborated here.

[0291] In the above embodiments, by performing sound source localization based on multi-channel time-frequency signals, head pose information, and position information, an estimated sound source position corresponding to the sound source can be obtained, providing more accurate direction information for subsequent beamforming. By performing beamforming based on multi-channel time-frequency signals, head pose information, position information, and the estimated sound source position corresponding to each sound source, spatial filtering of the multi-channel time-frequency signals can be achieved, optimizing the directivity of the array receiving system. Only when the beam direction is consistent with the designed direction can the signal in the observation direction be maximally amplified, thereby enhancing the sound in the direction of the target sound source. By performing blind source separation on the multi-channel time-frequency signals, interference speech in the multi-channel time-frequency signals can be separated and background noise can be suppressed, thereby improving the speech quality. By performing sound enhancement based on the beamformed signal and the independent audio signal corresponding to each sound source, sound enhancement and suppression of the audio signal can be achieved, obtaining better speech clarity and intelligibility. At the same time, it can also solve the interference from other directions that cannot be completely eliminated by the beamforming algorithm, as well as problems such as noise residue and signal distortion existing in the blind source separation algorithm.

[0292] In an embodiment of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. The processor executes the program instructions to implement any one of the sound processing methods provided in the embodiments of the present application. For example, when the computer program is loaded by the processor, the following steps can be executed:

[0293] Obtain the original audio signal received by the microphone array, where the original audio signal includes audio signals corresponding to at least one sound source; if it is detected that the user turns the head, obtain the head pose information of the user and the position information of the microphone array, and perform directional speech enhancement on the original audio signal according to the head pose information and the position information to obtain a target audio signal corresponding to the target sound source, where the target sound source is the sound source closest to the front of the user's head; play the target audio signal.

[0294] For another example, when the computer program is loaded by the processor, the following steps can be executed:

[0295] Obtain the head pose information of the user, the original audio signal received by the microphone array, and the position information of the microphone array; perform sound source localization according to the original audio signal, head pose information, and position information to obtain an estimated sound source position corresponding to at least one sound source; perform beamforming according to the original audio signal, head pose information, position information, and the estimated sound source position corresponding to each sound source to obtain a beamformed signal; perform blind source separation on the original audio signal to obtain an independent audio signal corresponding to each sound source; perform sound enhancement according to the beamformed signal and the independent audio signal corresponding to each sound source to obtain a target audio signal.

[0296] For the specific implementation of each of the above operations, reference may be made to the foregoing embodiments and will not be elaborated herein.

[0297] Among them, the computer-readable storage medium may be an internal storage unit of the wearable device in the foregoing embodiment, such as the hard disk or memory of the wearable device. The computer-readable storage medium may also be an external storage device of the wearable device, such as a plug-in hard disk equipped on the wearable device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.

[0298] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A sound processing method, characterized in that: Applied to a wearable device, the wearable device includes a microphone array; the method includes: Acquire an original audio signal received by the microphone array, wherein the original audio signal includes an audio signal corresponding to at least one sound source; If it is detected that the user turns his head, the head posture information of the user and the position information of the microphone array are obtained, and the original audio signal is subjected to directional speech enhancement according to the head posture information and the position information to obtain a target audio signal corresponding to a target sound source, where the target sound source is the sound source closest to the front of the user's head; Playing the target audio signal; The original audio signal includes an audio signal corresponding to at least one sound source received by at least one microphone channel; before performing directional speech enhancement on the original audio signal according to the head posture information and the position information to obtain a target audio signal corresponding to the target sound source, the method further includes: performing frequency domain transformation on the audio signal received by each of the microphone channels to obtain a time-frequency signal corresponding to each of the microphone channels; combining the time-frequency signals corresponding to each of the microphone channels to obtain a multi-channel time-frequency signal; The performing directional speech enhancement on the original audio signal according to the head posture information and the position information to obtain a target audio signal corresponding to a target sound source includes: performing sound source localization according to the multi-channel time-frequency signal, the head posture information and the position information to obtain a sound source position estimation value corresponding to at least one sound source; performing beam synthesis on the multi-channel time-frequency signal according to the head posture information, the position information and the sound source position estimation value corresponding to each sound source to obtain a beam synthesis signal; performing blind source separation on the multi-channel time-frequency signal to obtain an independent audio signal corresponding to each sound source; performing sound enhancement according to the beam synthesis signal and the independent audio signal corresponding to each sound source to obtain the target audio signal; The sound enhancement is performed according to the beam synthesis signal and the independent audio signal corresponding to each of the sound sources to obtain the target audio signal, including: performing frequency domain transformation on the independent audio signal corresponding to each of the sound sources to obtain the time-frequency signal of each of the sound sources; performing gain calculation according to the beam synthesis signal and the time-frequency signal of each of the sound sources to obtain the soft masking value of each of the sound sources on each of the microphone channels; performing signal synthesis on the soft masking value of each of the sound sources on each of the microphone channels to obtain a single-channel frequency domain total signal; performing time domain transformation on the single-channel frequency domain total signal to obtain the target audio signal; The step of performing gain calculation based on the beam synthesis signal and the time-frequency signal of each sound source to obtain a soft masking value of each sound source on each microphone channel includes: inputting the beam synthesis signal and the time-frequency signal of each sound source into a trained deep learning model for masking estimation to obtain a time-frequency masking value of each sound source on each microphone channel; and determining the soft masking value of each sound source on each microphone channel based on a multi-channel time-frequency signal corresponding to each sound source and the time-frequency masking value, wherein the soft masking value is obtained by performing element-by-element multiplication of the multi-channel time-frequency signal and the time-frequency masking value.

2. The sound processing method according to claim 1, characterized in that: Before performing directional speech enhancement on the original audio signal according to the head posture information and the position information to obtain a target audio signal corresponding to a target sound source, the method further includes: Performing frequency domain transformation on the audio signal received by each microphone channel to obtain a time-frequency signal corresponding to each microphone channel; Combining the time-frequency signals corresponding to each of the microphone channels to obtain a multi-channel time-frequency signal; Perform covariance matrix estimation on the multi-channel time-frequency signal to obtain a covariance matrix estimation result corresponding to the multi-channel time-frequency signal.

3. The sound processing method according to claim 1, characterized in that: The performing sound source localization according to the multi-channel time-frequency signal, the head posture information and the position information to obtain a sound source position estimation value corresponding to at least one sound source includes: Determine a delay estimation value corresponding to the arrival time difference of the multi-channel time-frequency signal; Based on the Kalman filter, sound source positioning is performed according to the time delay estimation value, the head posture information and the position information to obtain the sound source position estimation value corresponding to each sound source.

4. The sound processing method according to claim 3, characterized in that: The step of determining, according to the multi-channel time-frequency signal, a delay estimation value corresponding to the arrival time difference of the multi-channel time-frequency signal comprises: Constructing a cross-correlation function corresponding to the multi-channel time-frequency signal; constructing a generalized cross-correlation function according to the Fourier transform of the cross-correlation function; The delay estimation value is determined according to the peak position of the generalized cross-correlation function.

5. The sound processing method according to claim 3, characterized in that: The position information includes a first position of a microphone array; the method of performing sound source localization based on a Kalman filter according to the time delay estimation value, the head posture information, and the position information to obtain the sound source position estimation value corresponding to each sound source includes: Determine a rotation matrix of a head coordinate system relative to a world coordinate system used by the head posture information; According to the rotation matrix, converting the first position from the world coordinate system to the head coordinate system to obtain a second position; Inputting the time delay estimation value and the second position into the Kalman filter for loop filtering to obtain a state estimation value corresponding to each of the sound sources; The state estimation value corresponding to each of the sound sources is converted into a coordinate system to obtain the sound source position estimation value corresponding to each of the sound sources.

6. The sound processing method according to claim 5, characterized in that: The performing coordinate system conversion on the state estimation value corresponding to each of the sound sources to obtain the sound source position estimation value corresponding to each of the sound sources includes: Determine a first position estimation value of each sound source in the head coordinate system according to the state estimation value corresponding to each sound source; According to the rotation matrix, the first position estimation value corresponding to each of the sound sources is converted from the head coordinate system to the world coordinate system to obtain a second position estimation value corresponding to each of the sound sources; The sound source position estimation value corresponding to each sound source is determined according to the second position estimation value corresponding to each sound source.

7. The sound processing method according to claim 2, characterized in that: The performing beam synthesis on the multi-channel time-frequency signal according to the head posture information, the position information, and the sound source position estimation value corresponding to each sound source to obtain a beam synthesis signal includes: Calculate the sound source direction according to the head posture information and the sound source position estimation value corresponding to each sound source to obtain the target sound source direction corresponding to each sound source; The beamforming signal is obtained by performing beamforming calculation on the multi-channel time-frequency signal according to the target sound source direction, the position information and the covariance matrix estimation result corresponding to each sound source.

8. The sound processing method according to claim 7, characterized in that: The sound source position estimation value includes a second position estimation value corresponding to each sound source; The calculating the sound source direction according to the head posture information and the sound source position estimation value to obtain the target sound source direction includes: Calculating the front direction of the head based on the head posture information to obtain the front direction of the head; Determine the target sound source according to the second position estimation value corresponding to each sound source and the direction in front of the head, the target sound source being a sound source for which the angle between the second position estimation value and the direction in front of the head is less than a preset angle; The target sound source direction is determined according to the direction of the target sound source relative to the microphone array.

9. The sound processing method according to claim 7, characterized in that: The performing beam calculation on the multi-channel time-frequency signal according to the target sound source direction, the position information and the covariance matrix estimation result to obtain the beam synthesis signal includes: Calculating a steering vector according to the target sound source direction and the position information; Performing weight calculation according to the steering vector and the covariance matrix estimation result to obtain a beamforming weight value; Performing matrix calculation based on the beamforming weight value and the multi-channel time-frequency signal to obtain an output signal matrix; The beam synthesis signal is determined according to the output signal matrix.

10. The sound processing method according to claim 2, characterized in that: After performing covariance matrix estimation on the multi-channel time-frequency signal to obtain a covariance matrix estimation result corresponding to the multi-channel time-frequency signal, the method further includes: The number of sound sources is estimated according to the covariance matrix estimation result to obtain an estimated value of the number of sound sources.

11. The sound processing method according to claim 10, characterized in that: The step of performing blind source separation on the multi-channel time-frequency signal to obtain an independent audio signal corresponding to each sound source includes: Extracting features of the multi-channel time-frequency signal to obtain Mel-frequency cepstrum coefficients; Inputting the Mel-frequency cepstral coefficients into the trained deep learning model for masking estimation to obtain a first masking estimation value corresponding to each of the sound sources; Based on the estimated value of the number of sound sources, the multi-channel time-frequency signal is subjected to signal separation according to the first masking estimated value corresponding to each of the sound sources to obtain an independent audio signal corresponding to each of the sound sources.

12. The sound processing method according to claim 11, characterized in that: The step of performing signal separation on the multi-channel time-frequency signal based on the estimated value of the number of sound sources and according to the first masking estimated value corresponding to each of the sound sources to obtain an independent audio signal corresponding to each of the sound sources includes: Performing masking expansion on the first masking estimation value corresponding to each of the sound sources to obtain a second masking estimation value after masking expansion of each of the sound sources; Superimposing the second masking estimation value corresponding to each of the sound sources to the multi-channel time-frequency signal to obtain a frequency domain estimation signal of each of the sound sources on each microphone channel; Performing a time domain transformation on the frequency domain estimation signal of each sound source on each microphone channel to obtain a multi-channel time domain signal corresponding to each sound source; According to the multi-channel time domain signal corresponding to each of the sound sources, an independent audio signal corresponding to each of the sound sources is determined.

13. The sound processing method according to claim 1, characterized in that: The step of performing signal synthesis on the soft masking value of each sound source on each microphone channel to obtain a single-channel frequency domain total signal includes: Sum the soft masking values ​​of each sound source on each microphone channel to obtain a single-channel frequency domain signal corresponding to each sound source; The single-channel frequency domain signals corresponding to all the sound sources are added together to obtain the single-channel frequency domain total signal.

14. The sound processing method according to claim 1, characterized in that: The playing of the target audio signal comprises: Allocate the target audio signal to a left channel and a right channel of the wearable device according to the head posture information and the estimated value of the sound source position corresponding to each of the sound sources; The audio signal on the left channel and the audio signal on the right channel are respectively output to the sound-emitting unit of the wearable device.

15. The sound processing method according to claim 14, characterized in that: The target audio signal includes an audio signal of at least one sound source; and the step of allocating the target audio signal to a left channel and a right channel of the wearable device according to the head posture information and the sound source position estimation value corresponding to each sound source includes: Calculate the azimuth of the sound source position estimation value corresponding to each sound source in the head coordinate system; Determine, according to the azimuth angle corresponding to each of the sound sources, a first gain of each of the sound sources on the left channel and a second gain of each of the sound sources on the right channel; The audio signal of each sound source is multiplied by the first gain of each corresponding sound source and then added to the left channel, and the audio signal of each sound source is multiplied by the second gain of each corresponding sound source and then added to the right channel.

16. A sound processing method, characterized in that: Applied to a wearable device, the wearable device includes a microphone array; the method includes: Acquire the user's head posture information, the original audio signal received by the microphone array, and the position information of the microphone array; Performing sound source localization according to the original audio signal, the head posture information, and the position information to obtain a sound source position estimation value corresponding to at least one sound source; Performing beam synthesis on the original audio signal according to the head posture information, the position information, and the sound source position estimation value corresponding to each sound source to obtain a beam synthesis signal; Performing blind source separation on the original audio signal to obtain an independent audio signal corresponding to each sound source; Performing sound enhancement according to the beam synthesis signal and the independent audio signal corresponding to each of the sound sources to obtain a target audio signal; The sound enhancement is performed according to the beam synthesis signal and the independent audio signal corresponding to each of the sound sources to obtain the target audio signal, including: performing frequency domain transformation on the independent audio signal corresponding to each of the sound sources to obtain the time-frequency signal of each of the sound sources; performing gain calculation according to the beam synthesis signal and the time-frequency signal of each of the sound sources to obtain the soft masking value of each of the sound sources on each of the microphone channels; performing signal synthesis on the soft masking value of each of the sound sources on each of the microphone channels to obtain a single-channel frequency domain total signal; performing time domain transformation on the single-channel frequency domain total signal to obtain the target audio signal; The step of performing gain calculation based on the beam synthesis signal and the time-frequency signal of each sound source to obtain a soft masking value of each sound source on each microphone channel includes: inputting the beam synthesis signal and the time-frequency signal of each sound source into a trained deep learning model for masking estimation to obtain a time-frequency masking value of each sound source on each microphone channel; and determining the soft masking value of each sound source on each microphone channel based on a multi-channel time-frequency signal corresponding to each sound source and the time-frequency masking value, wherein the soft masking value is obtained by performing element-by-element multiplication of the multi-channel time-frequency signal and the time-frequency masking value.

17. A wearable device, characterized in that: The wearable device includes a memory, a processor, a microphone array, and an inertial measurement unit; The memory is used to store computer programs; The microphone array is used to receive audio signals; The inertial measurement unit is used to collect the user's head posture information; The processor is used to implement the sound processing method according to any one of claims 1 to 15 or the sound processing method according to claim 16 when executing the computer program.

18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the sound processing method according to any one of claims 1 to 15 or the sound processing method according to claim 16 is implemented.

Citation Information

Patent Citations

  • Multichannel voice separation method with unknown speaker number

    CN112116920A

  • Head-mounted computing device with microphone beam steering

    CN116636237A

  • Speech enhancement for an electronic device

    US20190272842A1