A method for estimating a direction of arrival of a sound source, an electronic device and a chip system
By combining reverberation dératio and blind source separation in a joint iterative process, the problem of low DOA estimation accuracy in multi-source and complex acoustic environments is solved, achieving high-precision DOA estimation, which is suitable for multi-source environments.
Patent Information
- Application Number
- CN202010643053.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-03
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2040-07-03
AI Technical Summary
Existing technologies have low DOA estimation accuracy in multiple sound sources and complex acoustic environments and are unable to effectively filter out noise and reverberation signals, resulting in insufficient estimation accuracy.
A combined iterative processing method of dereverberation and blind source separation is adopted. Noise and reverberation signals are removed through multiple iterations. Combined with the characteristics of acoustic vector sensors, the direction of arrival of each target sound source is separated and estimated.
Improving DOA estimation accuracy in multi-source and complex acoustic environments ensures accurate estimation of the direction of arrival (DOA) of target sound source signals, reduces computational load, and improves processing efficiency.
Smart Images

Figure CN113889135B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of audio processing, and in particular to a method for estimating the direction of arrival of a sound source, an electronic device, and a chip system. BACKGROUND
[0002] Direction of Arrival (DOA) estimation is a research hotspot in array signal processing. Through DOA estimation, the direction of a sound source can be determined, so that the target sound can be better picked up. Therefore, DOA estimation has become a key technology for human-computer interaction devices such as mobile phones, smart speakers, and remote conference screens.
[0003] Since the audio signal collected by the device is usually mixed by a sound source signal and noise, when performing DOA estimation, the noise is first treated as an interference signal, and the frequency point data that does not contain noise or has less noise influence is screened out from the audio signal. Then, DOA estimation is performed on the screened frequency point data. However, this DOA estimation method cannot be applied in a scene containing multiple sound sources when applied, and in a complex acoustic environment, the screened frequency point data is less and may contain other interference components, resulting in low accuracy in DOA estimation. SUMMARY
[0004] Embodiments of the present application provide a method for estimating the direction of arrival of a sound source, an electronic device, and a chip system, which can be applied in a multi-sound source acoustic environment and can improve the accuracy of DOA estimation in a complex acoustic environment.
[0005] To achieve the above-mentioned purpose, the technical solutions adopted by the present application are as follows:
[0006] In a first aspect, the embodiments of the present application provide a method for estimating the direction of arrival of a sound source, comprising: an electronic device acquires an audio signal, the audio signal comprising: noise, sound source signals of one or more target sound sources, and reverberation signals; the electronic device performs Nth de-reverberation processing on the audio signal to obtain an Nth prediction matrix and an Nth de-reverberation signal, the Nth de-reverberation signal comprising signals in the audio signal except for an Nth reverberation signal, the Nth reverberation signal being a reverberation signal removed in the Nth de-reverberation processing; the electronic device performs Nth blind source separation processing on the Nth de-reverberation signal to obtain an Nth de-mixing matrix and an Nth de-noise signal, the Nth de-noise signal being a signal in the audio signal from which noise obtained by the Nth blind source separation processing is removed; the electronic device continues to perform de-reverberation processing and blind source separation processing on the Nth de-noise signal; when N is a preset value or the Nth prediction matrix converges and the Nth de-mixing matrix converges, the electronic device obtains the sound source signals of the one or more target sound sources according to the Nth de-mixing matrix and the Nth de-reverberation signal; wherein N is a positive integer starting from 1; and the electronic device determines the direction of arrival of the sound source signals of the one or more target sound sources.
[0007] After the electronic device performs the joint iterative processing of dereverberation and blind source separation on the obtained audio signal, the noise, the reverberation signal and the sound source signal can be separated. That is, when the audio signal contains the sound source signal of one or more target sound sources, after the electronic device performs the joint iterative processing of dereverberation and blind source separation on the audio signal, the sound source signal of each target sound source can be obtained. Therefore, the electronic device can estimate the direction of arrival of the sound source signal of each target sound source, so that it can be applied in a multi-sound source acoustic environment; the reverberation signal is the signal collected by the microphone after the sound source signal of the target sound source is reflected and delayed, so the direction of the reverberation signal has changed; the direction of the noise is all around, therefore, the sound source signal of each target sound source obtained after the joint iterative processing contains little or almost no noise and reverberation signal that affects the accuracy of DOA estimation. Therefore, the electronic device has high DOA estimation accuracy when estimating the direction of arrival of the sound source signal of the target sound source; and even if the collection environment of the audio signal is a high-noise and high-reverberation acoustic environment, the electronic device still has high DOA estimation accuracy.
[0008] In a possible implementation of the first aspect, for the first target sound source, the direction of arrival of the sound source signal of the first target sound source includes: first direction information of the sound source signal of the first target sound source at different frequency points.
[0009] For ease of description, any one of the one or more target sound sources can be referred to as a first target sound source.
[0010] In a possible implementation of the first aspect, the electronic device performs joint processing of sound source separation and direction of arrival estimation on the sound source signal of the first target sound source according to the first direction information of the sound source signal of the first target sound source at different frequency points, to obtain second direction information of the sound source signal of the first target sound source at different frequency points, wherein the sound source separation processing includes: joint iterative processing of dereverberation and blind source separation.
[0011] The dereverberation and the blind source separation are both estimation algorithms. The sound source signal of the target sound source obtained after the joint iteration of the dereverberation and the blind source separation of the audio signal can still contain noise and / or reverberation signals. The first direction information obtained by the electronic device based on the sound source signal containing noise and / or reverberation signals can be less accurate. Therefore, the electronic device can continue to perform the sound source separation processing and the DOA estimation processing on the first target sound source based on the first direction information of the sound source signal of the first target sound source obtained at different frequency points. Since the sound source separation processing is constrained by the first direction information of the sound source signal of the first target sound source at different frequency points when the sound source separation processing is performed again, the electronic device can obtain more accurate sound source signal, demixing matrix and mixing matrix of the first target sound source after performing the sound source separation. The electronic device can obtain more accurate second direction information of the sound source signal of the first target sound source at different frequency points according to the more accurate sound source signal, demixing matrix or mixing matrix of the first target sound source. Of course, the second direction information of the sound source signal of the first target sound source at different frequency points obtained is more accurate than the first direction information of the sound source signal of the first target sound source at different frequency points.
[0012] In a possible implementation of the first aspect, the electronic device performs smoothing filtering processing or kernel density estimation processing on the first direction information of the sound source signal of the first target sound source at different frequency points to obtain third direction information of the first target sound source at different frequency points. The electronic device fuses the third direction information of the first target sound source at different frequency points to obtain the direction of the first target sound source.
[0013] In the embodiments of the present application, the smoothing filtering processing and the kernel density estimation processing are to remove some interference in the sound source signal of the first target sound source, so that the direction of the sound source signal of the first target sound source can be obtained according to the third direction information after removing some interference.
[0014] In a possible implementation of the first aspect, the electronic device continues to perform the dereverberation processing and the blind source separation processing on the N-th de-noised signal includes:
[0015] If the p-th prediction matrix obtained by the p-th dereverberation processing converges and the p-th demixing matrix obtained by the p-th blind source separation processing does not converge, the electronic device performs the p+i-th blind source separation processing until the p+i-th demixing matrix converges, or the electronic device alternately performs the p+i-th dereverberation processing and the p+i-th blind source separation processing until the p+i-th prediction matrix and the p+i-th demixing matrix converge simultaneously, where p is a positive integer and i is a positive integer starting from 1.
[0016] If the qth prediction matrix obtained by the qth dereverberation processing does not converge and the qth demixing matrix obtained by the qth blind source separation processing converges, the electronic device performs the q+i th dereverberation processing until the q+i th prediction matrix converges, or the electronic device alternately performs the q+i th dereverberation processing and the q+i th blind source separation processing until the q+i th prediction matrix and the q+i th demixing matrix converge simultaneously, where q is a positive integer and i is a positive integer starting from 1.
[0017] In an embodiment, when the electronic device performs the dereverberation processing and the blind source separation processing, in addition to being able to cyclically perform the dereverberation processing and the blind source separation processing until the prediction matrix and the demixing matrix converge, the electronic device can also independently iterate the matrix that does not converge until the prediction matrix and the demixing matrix converge when one of the matrices converges, thereby avoiding the process of repeatedly operating the converged matrix and improving processing efficiency.
[0018] In a possible implementation of the first aspect, performing the jth dereverberation processing by the electronic device includes a process of updating the prediction matrix m times, and performing the jth blind source separation processing by the electronic device includes a process of updating the demixing matrix n times, where j, m, and n are positive integers.
[0019] The electronic device can include a process of independently iterating the matrix in each execution of the dereverberation processing and the blind source separation processing, thereby reducing the amount of operation and improving processing efficiency.
[0020] In a possible implementation of the first aspect, the dereverberated signal obtained by the last dereverberation processing is used as the processing signal for the blind source separation, and the audio signal is used as the processing signal for the dereverberation; or the de-noised signal obtained by the last blind source separation processing is used as the processing signal for the dereverberation, and the audio signal is used as the processing signal for the blind source separation.
[0021] In a possible implementation of the first aspect, the dereverberated signal obtained by any one of the historical dereverberation processes is used as the processing signal for the blind source separation, and the de-noised signal obtained by any one of the historical dereverberation processes is used as the processing signal for the dereverberation.
[0022] When the electronic device performs the joint iteration of the dereverberation processing and the blind source separation processing, the parameter of the dereverberation processing process can affect the blind source separation processing, and the parameter of the blind source separation processing process can affect the dereverberation processing. Moreover, the parameter can be the dereverberated signal and / or the de-noised signal obtained by any one of the previous dereverberation processes.
[0023] In a possible implementation of the first aspect, the electronic device obtaining the audio signal includes: the electronic device collecting the audio signal through an acoustic vector sensor on the electronic device; or the electronic device receiving the audio signal collected by an acoustic vector sensor on another electronic device.
[0024] In a possible implementation manner of the first aspect, the electronic device determining the direction of arrival of the sound source signal of the one or more target sound sources comprises: the electronic device obtaining the direction of arrival of the sound source signal of a second target sound source according to one or more of the amplitudes of the sound source signal of the second target sound source on the multiple channels, a demixing matrix or a mixing matrix of the sound source signals of the one or more target sound sources, the second target sound source being any one of the one or more target sound sources; wherein the demixing matrix represents a conversion relationship when the audio signal is separated into the sound source signals of the one or more target sound sources, and the mixing matrix represents a conversion relationship when the sound source signals of the one or more target sound sources in the audio signal are mixed into the audio signal.
[0025] In the embodiments of the present application, in combination with the characteristic that the points of the multiple channels of the acoustic vector microphone are coincident, it is considered that the amplitude of the sound source signal of a pure target sound source in each channel direction is related to the direction of the target sound source after the sound source signal is collected by the acoustic vector microphone. Because the electronic device can obtain the direction of arrival of the sound source signal of the first target sound source according to the amplitudes of the sound source signal of the first target sound source on the multiple channels. In combination with the characteristic that the points of the multiple channels of the acoustic vector microphone are coincident, and the second mixing model of the multiple target sound sources, it is considered that the direction of the sound source signal of the target sound source is also contained in the conversion relationship when the sound source signal of the target sound source and the noise are mixed into the audio signal. Therefore, the electronic device can obtain the direction of arrival of any target sound source according to the mixing matrix or the demixing matrix of the sound source signals of the one or more target sound sources.
[0026] In a possible implementation manner of the first aspect, the electronic device obtaining the direction of arrival of the sound source signal of the second target sound source according to the mixing matrix of the sound source signals of the one or more target sound sources comprises: the electronic device determining a target column in the mixing matrix and a first target row and a second target row in the target column, wherein the target column is a column representing the sound source signal of the second target sound source, and the first target row and the second target row are rows related to the angle of the sound source signal of the second target sound source; and the electronic device obtaining the direction of arrival of the sound source signal of the second target sound source according to elements of the first target row and the second target row in the target column.
[0027] Due to the characteristic that the points of the acoustic vector sensors are coincident, the mixing matrix implicitly contains the ratio between the amplitudes of the sound source signals of each target sound source on the multiple channels, or it can be understood that the mixing matrix implicitly contains the angle relationship. The angle can be determined according to the elements on the first target row and the second target row in the column representing the sound source signal of the target sound source in the mixing matrix.
[0028] In a possible implementation manner of the first aspect, when the first target row represents a row of a first channel of the acoustic vector sensor, the second target row represents a row of a second channel of the acoustic vector sensor, and the direction of arrival of the sound source signal of the second target sound source comprises a horizontal angle of the sound source signal of the second target sound source, the horizontal angle being an angle in a coordinate system in which the acoustic vector sensor is located; and / or when the first target row represents a row of a third channel of the acoustic vector sensor, the second target row represents a row of an omnidirectional channel of the acoustic vector sensor, and the direction of arrival of the sound source signal of the second target sound source comprises a pitch angle of the sound source signal of the second target sound source, the pitch angle being an angle in the coordinate system in which the acoustic vector sensor is located.
[0029] In the implementation manner, the first column in the mixing matrix represents a column in which the sound source signal of the first target sound source is located, the second column in the mixing matrix represents a column in which the sound source signal of the second target sound source is located, and so on. Each row in the mixing matrix represents an omnidirectional channel, an X channel, a Y channel, and a Z channel (three-dimensional four-channel sound vector microphone). The electronic device can obtain, according to elements in the target column representing the first target sound source and elements in a row representing the X channel of the acoustic vector sensor, a horizontal angle of the sound source signal of the first target sound source, the horizontal angle being an angle in a coordinate system in which the acoustic vector sensor is located. The electronic device can obtain, according to elements in the target column representing the first target sound source and elements in a row representing the Z channel of the acoustic vector sensor, a pitch angle of the sound source signal of the first target sound source, the pitch angle being an angle in the coordinate system in which the acoustic vector sensor is located.
[0030] In a possible implementation manner of the first aspect, after the electronic device obtains the sound source signal of one or more target sound sources according to the Nth demixing matrix and the Nth de-noised signal, the electronic device performs first enhancement processing on the sound source signal of one or more target sound sources, where the first enhancement processing comprises: interference spectrum filtering processing and / or harmonic enhancement processing, the first target sound source being any one of the sound source signals of the one or more target sound sources; the interference spectrum filtering processing is used to filter out interference components mixed in the sound source signal of any one of the one or more target sound sources based on spectral energy of the sound source signal of the target sound source; and the harmonic enhancement processing is used to obtain a harmonic enhanced signal of the one or more target sound sources, the harmonic enhanced signal being a sound source signal containing harmonic components.
[0031] In the implementation manner, the interference spectrum filtering processing can be used to filter out interference components mixed in the sound source signal of the first target sound source based on spectral energy of the sound source signal of the first target sound source, so as to obtain a purer sound source signal; and the harmonic enhancement processing can enrich the sound heard by us or restore the real sound emitted by a musical instrument.
[0032] In a possible implementation of the first aspect, the electronic device performs second enhancement processing on the sound source signal of the first target sound source based on the first direction information of the sound source signal of the first target sound source at different frequency points, where the second enhancement processing includes interference direction filtering processing and / or beamforming directional enhancement processing; the interference direction filtering processing is used to filter out frequency points in the sound source signal of the first target sound source that are not within the expected angle range; and the beamforming directional enhancement processing is used to enhance the power of the sound source signal in the expected direction.
[0033] In the implementation, the interference direction filtering processing is used to filter out frequency points in the sound source signal of the first target sound source that are not within the expected angle range, and thus the sound in directions other than the direction of the first target sound source can be suppressed; and the beamforming directional enhancement processing is used to enhance the power of the sound source signal in the expected direction.
[0034] In a possible implementation of the first aspect, when N is a preset value or the Nth prediction matrix converges and the Nth demixing matrix converges, the electronic device further obtains noise and reverberation signals of one or more target sound sources from the audio signal; and the electronic device adjusts the proportional relationship among the noise, the sound source signal of the first target sound source, and the reverberation signal of the first target sound source, where the first target sound source is any one of the one or more target sound sources.
[0035] In the embodiments of the present application, the electronic device performs the step of adjusting the proportional relationship among the sound source signal of the first target sound source, the reverberation signal of the first target sound source, and the noise, and thus sound with different scene effects, such as KTV effects, concert hall effects, and open field effects, can be obtained.
[0036] In the second aspect, the embodiments of the present application provide an electronic device, which includes an audio signal acquisition unit configured to acquire an audio signal, where the audio signal includes noise, sound source signals of one or more target sound sources, and reverberation signals.
[0037] The audio signal acquisition unit is configured to acquire an audio signal, where the audio signal includes noise, sound source signals of one or more target sound sources, and reverberation signals.
[0038] The de-reverberation processing unit is configured to perform Nth de-reverberation processing on the audio signal to obtain an Nth prediction matrix and an Nth de-reverberation signal, where the Nth de-reverberation signal includes signals in the audio signal except for an Nth reverberation signal removed in the Nth de-reverberation processing, and the Nth reverberation signal is a reverberation signal removed in the Nth de-reverberation processing.
[0039] The blind source separation processing unit is configured to perform Nth blind source separation processing on the Nth de-reverberation signal to obtain an Nth demixing matrix and an Nth de-noise signal, where the Nth de-noise signal is a signal in the audio signal from which noise obtained in the Nth blind source separation processing is removed.
[0040] a sound source signal obtaining unit, configured to perform a de-reverberation processing and a blind source separation processing on the Nth de-noised signal; and obtain one or more sound source signals of one or more target sound sources according to the Nth de-mixing matrix and the Nth de-reverberation signal when N is a preset value or the Nth prediction matrix converges and the Nth de-mixing matrix converges; wherein N is a positive integer starting from 1
[0041] a sound source direction estimation unit, configured to determine a direction of arrival of the one or more sound source signals of the one or more target sound sources.
[0042] In a third aspect, an electronic device is provided, including a processor configured to execute a computer program stored in a memory to implement the method of any one of the first aspect.
[0043] In a fourth aspect, a chip system is provided, including a processor coupled with a memory, and the processor executes a computer program stored in the memory to implement the method of any one of the first aspect.
[0044] In a fifth aspect, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, and the computer program is executed by one or more processors to implement the method of any one of the first aspect.
[0045] In a sixth aspect, a computer program product is provided, and when the computer program product is executed on an electronic device, the electronic device executes the method of any one of the first aspect.
[0046] It can be understood that the beneficial effects of the second aspect to the sixth aspect can be referred to the related description of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0047] Figure 1 An application scenario of the method for estimating the direction of arrival of a sound source provided by the embodiments of the present application is shown in the figure;
[0048] Figure 2 A hardware structure of an electronic device for executing the method for estimating the direction of arrival of a sound source provided by the embodiments of the present application is shown in the figure;
[0049] Figure 3 A flowchart of the method for estimating the direction of arrival of a sound source provided by the embodiments of the present application is shown in the figure;
[0050] Figure 4 A flowchart of another method for estimating the direction of arrival of a sound source provided by the embodiments of the present application is shown in the figure;
[0051] Figure 5 A second mixing model of the sound source signals of the one or more target sound sources provided in the embodiments of the present application is shown in the figure;
[0052] Figure 6 A schematic diagram of a separation model for separating the sound source signal of each target sound source from an audio signal provided in an embodiment of the present application;
[0053] Figure 7 for Figure 4 A schematic flow chart of an implementation method of joint iterative processing of dereverberation and blind source separation in the illustrated embodiment;
[0054] Figure 8 A flowchart of another method for estimating the direction of arrival of a sound source provided in an embodiment of the present application;
[0055] Figure 9 A structural rendering of a three-dimensional four-channel acoustic vector sensor provided in an embodiment of the present application;
[0056] Figure 10 A schematic block diagram of a joint process including sound source separation, DOA estimation, and enhancement processing provided in an embodiment of the present application;
[0057] Figure 11 A schematic block diagram of a functional architecture module of an electronic device that performs a method for estimating the direction of arrival of a sound source provided in an embodiment of the present application. DETAILED DESCRIPTION
[0058] In the following description, specific details such as specific system structures and technologies are provided for illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details.
[0059] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0060] It should also be understood that in the embodiments of this application, "one or more" refers to one, two, or more than two; "and / or" describes the relationship between associated objects, indicating that three relationships can exist; for example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0061] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0062] Reference within the specification to "one embodiment" or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearances of the phrase "in one embodiment" or "in some embodiments" in various places within specified
[0063] Embodiments of the present application can be applied in an acoustic environment in which there is one or more target sound sources, as shown in Figure 1 , Figure 1 One application scenario of the method for estimating the direction of arrival of a sound source provided by embodiments of the present application can be that voice information emitted by one or more users is collected by a microphone array, as shown in Figure 1 The microphone array can include one or more microphones, and each microphone is used to collect voice information emitted by one or more users. Figure 1 In the example shown in FIG. 1, the microphone array includes four microphones, and the number of users is four. It can be understood that in actual processes, the number of microphones included in the microphone array can be more than four or less than four, and the number of users can be more than four or less than four.
[0064] For example, the four users are talking or singing from time to time (hereinafter, the sound emitted by a user can be referred to as voice information). In one time period, at least two users among the four users are talking at the same time, and it can be considered that there are at least two target sound sources. In another time period, only one user is talking, and it can be considered that there is one target sound source. Of course, in one time period, all four users are talking, and it can be considered that there are four target sound sources. If an acoustic environment in which all four users are talking is taken as an example, there are four target sound sources in the acoustic environment. There are four microphones in the microphone array, and the voice information emitted by each user can be collected by the four microphones. Similarly, the voice information emitted by each user can also be collected by each microphone. In addition to the voice information emitted by each user, the information collected by each microphone in the microphone array can also include noise (for example, mixed environmental noise, device noise), reverberation signals, and the like.
[0065] Since each microphone can collect not only the voice information emitted by each user, but also mixed environmental noise, device noise, and reverberation signals, in this embodiment of the application, all the information collected by the microphone can be referred to as an audio signal. In this embodiment of the application, mixed environmental noise and device noise are collectively referred to as noise.
[0066] The audio signal collected by each microphone in the microphone array is called a channel audio signal. Figure 1 When the microphone array in the application scenario shown collects audio signals, the audio signals collected by the microphone array are four-channel audio signals. The audio signals of one channel may include voice signals emitted by different users.
[0067] The electronic device can separate the voice information of each user from the audio signal containing reverberation signals and noise. The voice information of each user can be interpreted as the sound source signal of a target sound source. The electronic device can also obtain the direction of arrival of the sound source signal of each target sound source.
[0068] It should be understood that each microphone in the microphone array can collect voice information respectively emitted by the four users.
[0069] The present invention provides a method for estimating the direction of arrival of a sound source. The method can be applied to electronic devices, such as mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), and the like. The present invention does not limit the specific types of electronic devices.
[0070] Figure 2 The electronic device 200 may include a processor 210, an internal memory 221, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, and an earphone jack 270D.
[0071] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0072] The processor 210 can include one or more processing units, for example: the processor 210 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices, or can be integrated in one or more processors. For example, the processor 210 is configured to execute the method for estimating the sound source direction of arrival in the embodiments of the present application, for example, steps 301-302 described below.
[0073] The memory in the processor 210 can also be configured to store instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. The memory can save instructions or data that the processor 210 has just used or repeatedly uses. If the processor 210 needs to use the instructions or data again, it can be directly called from the memory. Avoiding repeated access reduces the waiting time of the processor 210, thereby improving the efficiency of the system.
[0074] In some embodiments, the processor 210 can include one or more interfaces. The interface can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0075] The wireless communication function of the electronic device 200 can be implemented by the antenna 1, the antenna 2, the mobile communication module 250, the wireless communication module 260, the modem processor, the baseband processor, and the like.
[0076] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 200 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna of a wireless local area network. In some other embodiments, the antennas can be used in combination with a tuning switch.
[0077] As an example, when the electronic device is not provided with a microphone array, the audio signal of another electronic device can be acquired through the antenna 1 and the mobile communication module 250, or the audio signal of another electronic device can be acquired through the antenna 2 and the wireless communication module 260.
[0078] The mobile communication module 250 can provide a solution including 2G / 3G / 4G / 5G wireless communication applied to the electronic device 200. The mobile communication module 250 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), and the like. The mobile communication module 250 can receive electromagnetic waves from the antenna 1, and perform filtering, amplification, and the like on the received electromagnetic waves, and transmit the processed signals to the modem processor for demodulation. The mobile communication module 250 can also amplify the signals modulated by the modem processor, and convert the signals into electromagnetic waves radiated through the antenna 1.
[0079] In some embodiments, at least part of the functional modules of the mobile communication module 250 can be disposed in the processor 210. In some embodiments, at least part of the functional modules of the mobile communication module 250 can be disposed in the same device as at least part of the modules of the processor 210.
[0080] The wireless communication module 260 can provide a solution for wireless communication, including wireless local area networks (WLAN) (e.g., wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc., which are applied to the electronic device 200. The wireless communication module 260 can be one or more devices that integrate at least one communication processing module. The wireless communication module 260 receives electromagnetic waves via the antenna 2, frequency-modulates and filters the electromagnetic wave signals, and transmits the processed signals to the processor 210. The wireless communication module 260 can also receive signals to be transmitted from the processor 210, frequency-modulate them, amplify them, and radiate them as electromagnetic waves via the antenna 2.
[0081] In some embodiments, the antenna 1 and the mobile communication module 250 of the electronic device 200 are coupled, and the antenna 2 and the wireless communication module 260 are coupled, so that the electronic device 200 can communicate with a network and other devices through wireless communication technology. The wireless communication technology can include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS can include a global positioning system (GPS), a global navigation satellite system (GLONASS), a beidu navigation satellite system (BDS), a quasi-zenith satellite system (QZSS), and / or a satellite based augmentation systems (SBAS).
[0082] The internal memory 221 can be used to store computer executable program codes including instructions. The processor 210 performs various functional applications and data processing of the electronic device 200 by executing the instructions stored in the internal memory 221. The internal memory 221 can include a program storage area and a data storage area. The program storage area can store an operating system and at least one application program (such as a sound play function, an image play function, etc.) required by a function. The data storage area can store data created during the use of the electronic device 200.
[0083] In addition, the internal memory 221 can include a high-speed random access memory, and can also include a non-volatile memory such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0084] The electronic device 200 can implement audio functions through the audio module 270, the speaker 270A, the receiver 270B, the microphone 270C, the earphone interface 270D, and the application processor, etc. For example, music playing, recording, etc.
[0085] The audio module 270 is configured to convert a digital audio signal into an analog audio signal for output, and to convert an analog audio input into a digital audio signal. The audio module 270 can also be configured to encode and decode audio signals. In some embodiments, the audio module 270 can be disposed in the processor 210, or some functional modules of the audio module 270 can be disposed in the processor 210.
[0086] The speaker 270A, also referred to as a “loudspeaker”, is configured to convert an audio electrical signal into a sound signal. The electronic device 200 can play the sound source signal obtained in the embodiments of the present application through the speaker 270A.
[0087] The receiver 270B, also referred to as a “earpiece”, is configured to convert an audio electrical signal into a sound signal. When the electronic device 200 is on a call or receiving a voice message, the receiver 270B can be placed close to a human ear to receive the voice, for example, a user receives the sound source signal obtained in the embodiments of the present application through the receiver in a hearing aid.
[0088] The microphone 270C, also referred to as a “microphone”, “sound transducer”, is configured to convert a sound signal into an electrical signal. When making a call or sending a voice message, a user can speak close to the microphone 270C to input the sound signal into the microphone 270C. The electronic device 200 can be provided with at least one microphone 270C. In other embodiments, the electronic device 200 can be provided with two microphones 270C, in addition to collecting sound signals, the electronic device 200 can also implement noise reduction functions. In other embodiments, the electronic device 200 can also be provided with three, four or more microphones 270C to form a microphone array provided in the embodiments of the present application, to implement sound signal collection, noise reduction, and also to identify the source of the sound, to implement directional recording functions, etc. For example, the microphone 270C can be used to collect the audio signal involved in the embodiments of the present application.
[0089] It should be noted that if the electronic device is a server, the server includes a processor and a communication interface.
[0090] In the embodiments of the present application, the specific structure of the execution subject of the method for estimating the direction of arrival of a sound source is not particularly limited, as long as the execution subject can communicate according to the method for estimating the direction of arrival of a sound source by running a program in which the code of the method for estimating the direction of arrival of a sound source is recorded. For example, the execution subject of the method for estimating the direction of arrival of a sound source provided by the embodiments of the present application can be a functional module capable of calling and executing a program in an electronic device, or a communication device applied to an electronic device, for example, a chip. The following embodiments are described by taking the execution subject of the method for estimating the direction of arrival of a sound source as an electronic device.
[0091] Referring to Figure 3 , Figure 3 The flowchart of the method for estimating the direction of arrival of a sound source provided by the embodiments of the present application is shown in the figure, and the method comprises the following steps:
[0092] In step 301, the electronic device acquires an audio signal, which comprises noise, sound source signals of one or more target sound sources, and reverberation signals.
[0093] In the embodiments of the present application, the audio signal can be a multi-channel audio signal or a single-channel audio signal, and the multi-channel audio signal indicates that the audio signal comes from multiple channels.
[0094] Since the electronic device has the function of collecting the audio signal or not, the implementation of step 301 is different, which will be introduced as follows:
[0095] Example 1: The electronic device has the function of collecting the audio signal.
[0096] In a possible implementation manner, step 301 can be implemented by the following manner: the electronic device collects the audio signal through the audio collection device (for example, a microphone array) arranged in the electronic device.
[0097] For example, the electronic device can be an electronic device in which the microphone array shown in the figure is arranged. Figure 1
[0098] Example 2: The electronic device does not have the function of collecting the audio signal.
[0099] In a possible implementation manner, step 301 can be implemented by the following manner: the electronic device receives the audio signal from other devices. The other devices are provided with a microphone array for collecting the audio signal.
[0100] For example, in example 2, the electronic device can be a server, a cloud platform, or the like.
[0101] It should be understood that, in the case that the electronic device has the function of collecting audio signals, the electronic device can also receive audio signals sent by other devices.
[0102] Most of the space in the natural state has environmental noise, and there may also be interference (referred to as device noise) in the signal regardless of the presence or absence of the device due to the occurrence, inspection, measurement or recording device. Therefore, the audio signal collected by the microphone array may contain environmental noise, device noise, collectively referred to as noise in the subsequent. In addition, since the environmental noise in the space is actually emitted by the sound source, in order to facilitate the distinction, the sound source in the space other than the noise emitting sound source can be recorded as the target sound source. Of course, in actual application, the number of target sound sources in the audio signal can be one or more.
[0103] Since the target sound source emits sound, the sound may be reflected by the ground, wall, etc. Therefore, in addition to containing noise and target sound source emitted sound source signals in the above-mentioned audio signal, it can also include the sound emitted by the target sound source after being reflected and collected by the microphone array. In the embodiments of the present application, the component collected by the microphone after the target sound source emits sound can be recorded as the sound source signal, and the component collected by the microphone after the target sound source emits sound after being reflected can be recorded as the reverberation signal.
[0104] The audio signal in the embodiments of the present application is a mixed signal obtained by noise, reverberation signal, one or more target sound source sound source signals. Alternatively, the audio signal in the embodiments of the present application is a mixed signal obtained by noise and one or more target sound source sound source signals.
[0105] Since the audio signal obtained by the electronic device can be a time domain signal or a frequency domain signal, but the action of the electronic device before performing step 302 is different in different cases, therefore, the following case will be introduced:
[0106] Case 1, the audio signal in the embodiments of the present application is a time domain signal.
[0107] Correspondingly, the method provided by the embodiments of the present application can further include: the electronic device can perform time-frequency transformation on the audio signal to obtain a frequency domain signal of the audio signal. It can be understood that the frequency domain signal corresponding to the audio signal is taken as the processing object in the subsequent steps.
[0108] For example, the above-mentioned time-frequency transformation can adopt Fourier transform, fast Fourier transform, wavelet transform, etc. The specific transformation method can be determined according to the actual application requirement, and the specific processing process of Fourier transform, fast Fourier transform and wavelet transform will not be repeated here. Of course, other time-frequency transformation methods can also be used in actual application, which is not limited here.
[0109] Case 2: The audio signal in the embodiment of the present application is a frequency domain signal.
[0110] In this case 2, the electronic device does not need to perform a step of time-frequency transform on the audio signal before performing joint processing on the audio signal.
[0111] Step 302: The electronic device performs joint processing on the audio signal to obtain sound source signals of one or more target sound sources and directions of arrival of the sound source signals of one or more target sound sources. The joint processing includes: dereverberation processing, blind source separation processing, and direction of arrival estimation processing.
[0112] In the embodiments of the present application, dereverberation and blind source separation may be collectively referred to as sound source separation. Sound source separation may be performed by first performing dereverberation and then performing blind source separation, or by first performing blind source separation and then performing dereverberation, or by performing a combined iterative process of dereverberation and blind source separation. For details on dereverberation, blind source separation, and the combined iterative process of dereverberation and blind source separation, please refer to the description of subsequent embodiments.
[0113] When an electronic device performs joint processing, it may first perform sound source separation processing, then perform direction of arrival estimation processing; it may also perform sound source separation processing and direction of arrival estimation processing in a loop; it may also perform direction of arrival estimation processing first, then perform sound source separation processing; it may also perform direction of arrival estimation processing and sound source separation processing in a loop. Therefore, electronic devices performing joint processing on audio signals include at least the following methods:
[0114] In the first method, the electronic device performs a combined iterative process of dereverberation processing and blind source separation processing on the audio signal to obtain a sound source signal of one or more target sound sources;
[0115] The electronic device determines the direction of arrival of sound source signals of one or more target sound sources.
[0116] As another implementation method of the first method, after the electronic device determines the arrival direction of the sound source signal of one or more target sound sources, the electronic device performs a joint iterative processing of dereverberation processing and blind source separation processing and arrival direction estimation processing on the audio signal based on the arrival direction of the sound source signal of the one or more target sound sources to obtain a purer sound source signal and a more accurate arrival direction of the one or more target sound sources.
[0117] In a second manner, the electronic device performs direction of arrival estimation processing on the audio signal to obtain a third signal of one or more target sound sources and a direction of arrival of the third signal of the one or more target sound sources, where the third signal is a sound source signal containing noise;
[0118] The electronic device performs joint iterative processing of dereverberation and blind source separation on the third signals of the one or more target sound sources according to the directions of arrival of the third signals of the one or more target sound sources to obtain sound source signals of the one or more target sound sources.
[0119] The electronic device takes the directions of arrival of the third signals of the one or more target sound sources as the directions of arrival of the sound source signals of the one or more target sound sources.
[0120] As another implementation manner of the second manner, after the electronic device performs joint iterative processing of dereverberation and blind source separation on the third signals of the one or more target sound sources according to the directions of arrival of the third signals of the one or more target sound sources to obtain sound source signals of the one or more target sound sources, the electronic device can continue to determine the directions of arrival of the sound source signals of the one or more target sound sources, instead of taking the directions of arrival of the third signals of the one or more target sound sources as the directions of arrival of the sound source signals of the one or more target sound sources.
[0121] As can be understood from the foregoing description, one difference between the first manner and the second manner in which the electronic device performs joint processing on the audio signals lies in whether the sound source separation processing is performed first or the direction of arrival estimation processing is performed first.
[0122] The present application describes the process in which the electronic device performs joint processing on the audio signals by taking the first manner as an example, and the dereverberation processing, the blind source separation processing, and the direction of arrival estimation processing in the second manner can refer to the description in the first manner. See Figure 4 , Figure 4 Another flowchart of a method for estimating the directions of arrival of sound sources provided by an embodiment of the present application is shown in the figure, which includes steps 401 to 403. The content of step 401 is the same as that of step 301, and thus is not described again.
[0123] In step 402, the electronic device performs joint iterative processing of dereverberation and blind source separation on the audio signals to obtain sound source signals of one or more target sound sources.
[0124] This step can refer to the description of subsequent embodiments for details, which are not described again here.
[0125] In step 403, the electronic device determines the directions of arrival of the sound source signals of the one or more target sound sources.
[0126] The electronic device can perform DOA estimation processing on the sound source signals of each target sound source in the one or more target sound sources to obtain the respective direction of arrival of each target sound source.
[0127] Taking an example in which the one or more target sound sources include a first target sound source, the electronic device performs DOA estimation processing on the sound source signals of the first target sound source to obtain the direction of arrival of the first target sound source.
[0128] As an example, the direction of arrival of each target sound source includes first direction information of the target sound source at different frequency points. For example, taking a first target sound source included in the plurality of target sound sources as an example, the direction of arrival of the first target sound source includes first direction information of the first target sound source at different frequency points (for example, frequency point 1, frequency point 2, and frequency point 3). As an example, when the audio signal is a frequency domain signal, the frequency value of the audio signal is within a certain range (for example, 500-3000 Hz), and each frequency value within the range can be recorded as a frequency point. For example, the first direction information of the first target sound source at frequency point 1 (taking horizontal angle and pitch angle as an example) can be the angle value (30°, 60°) of the first target sound source at frequency point 500 Hz. The first direction information of the first target sound source at frequency point 2 can be the angle value (31°, 59°) of the first target sound source at frequency point 501 Hz. The first direction information of the first target sound source at frequency point 3 can be the angle value (30°, 59°) of the first target sound source at frequency point 502 Hz. In actual application, the frequency values within the frequency range of the audio signal can also be divided into equally spaced frequency segments (for example, each interval of 5 Hz forms a frequency segment), and each frequency segment is recorded as a frequency point. Specific examples are not given.
[0129] It should be noted that the first target sound source described above is any one of the plurality of target sound sources and does not have indicative meaning. In addition, the frequency points corresponding to different target sound sources can be the same or different. For example, the different frequency points corresponding to the first target sound source are a plurality of frequency points within a first frequency value range (for example, 500-2500 Hz), and the different frequency points corresponding to a second target sound source in the plurality of target sound sources can be a plurality of frequency points within a second frequency value range (for example, 600-3000 Hz).
[0130] In a possible implementation, the electronic device determines the first direction information of the first target sound source at different frequency points. The electronic device obtains the direction of arrival of the first target sound source according to the first direction information of the first target sound source at different frequency points. Of course, the electronic device can also fuse the amplitude information of the first target sound source at different frequency points. Then, the electronic device calculates the direction angle of the first target sound source according to the fused amplitude information.
[0131] It should be noted that the above only takes how the electronic device calculates the direction of arrival of the first target sound source as an example, and the calculation method of the direction of arrival of the remaining target sound source(s) in the one or more target sound sources can refer to the calculation process of the direction of arrival of the first target sound source, which is not described herein.
[0132] The method for estimating the sound source direction of arrival provided in the embodiments of the present application can obtain the sound source signal of each target sound source without noise and reverberation signals affecting the DOA estimation accuracy or with a small amount of noise and reverberation signals affecting the DOA estimation accuracy. Therefore, the electronic device has high DOA estimation accuracy when estimating the direction of arrival of each target sound source. Even if the audio signal is collected in a high-noise and high-reverberation acoustic environment, the electronic device still has high DOA estimation accuracy.
[0133] It should be noted that the audio signal in the embodiments of the present application can be a signal without reverberation signals. When the audio signal without reverberation signals is processed by the steps of the joint iteration of the reverberation removal and blind source separation provided in the embodiments of the present application, the obtained reverberation signal can be 0. Considering the accuracy of the reverberation removal algorithm, the obtained reverberation signal can also be a very small proportion of the signal considered to be a reverberation signal obtained from the audio signal. Since even a very small proportion of the reverberation signal obtained from the audio signal without reverberation signals is equivalent to a very small proportion of the signal processed from the audio signal when blind source separation is used, the accuracy of the final DOA estimation is almost unaffected. Therefore, the audio signal in the embodiments of the present application can or can not contain reverberation signals. Whether the audio signal contains reverberation signals or not, it does not affect the implementation of the embodiments of the present application.
[0134] As a possible implementation manner, the step 402 in the embodiments of the present application can be implemented in the following manner:
[0135] The electronic device performs Nth reverberation removal processing on the audio signal to obtain an Nth prediction matrix and an Nth reverberation removed signal. The Nth reverberation removed signal includes signals in the audio signal except for an Nth reverberation signal. The Nth reverberation signal is a reverberation signal removed in the Nth reverberation removal processing. The electronic device performs Nth blind source separation processing on the Nth reverberation removed signal to obtain an Nth demixing matrix and an Nth noise removed signal. The Nth noise removed signal is a signal in the audio signal from which noise obtained by the Nth blind source separation processing is removed. The electronic device continues to perform the reverberation removal processing and the blind source separation processing on the Nth noise removed signal. When N is a preset value or the Nth prediction matrix converges and the Nth demixing matrix converges, the electronic device obtains the sound source signal of one or more target sound sources according to the Nth demixing matrix and the Nth noise removed signal. N is a positive integer starting from 1.
[0136] In the embodiments of the present application, the dereverberation processing performed by the electronic device can remove the reverberation signal in the audio signal to obtain a signal in which the reverberation signal is removed from the audio signal, which can be denoted as an Nth dereverberation signal, or as a dereverberation signal obtained in this time. The blind source separation processing performed by the electronic device can separate each sound source signal and noise. The Nth dereverberation signal is a signal in which the noise obtained in this time is removed from the audio signal, which can also be understood as a dereverberation signal obtained in this time. Therefore, theoretically, the sound source signal of the target sound source does not contain reverberation components and noise. However, considering the accuracy of the joint iterative processing of dereverberation and blind source separation, a small amount of reverberation signal and / or noise may still exist in each sound source signal of the target sound source that can be obtained. Therefore, the sound source signal of the target sound source described in the embodiments of the present application does not mean that it does not contain any reverberation signal and / or noise at all.
[0137] As an example of the dereverberation processing, the dereverberation processing is a WPE (Weighted Prediction Error) algorithm. The idea of the WPE algorithm is that the current received signal (which can also be understood as a processing signal in the dereverberation processing) is a linear combination of the current pure signal (which can be understood as a sound source signal) and the received signals of the past several frames (which can be understood as reverberation signals). In the process of dereverberation processing, the noise in the audio signal is ignored.
[0138] As an example, a first mixing model is obtained according to the idea of the above-mentioned WPE algorithm:
[0139]
[0140] wherein y l (t) represents the received signal of the microphone, represents the reverberation signal, i.e., the reverberation signal corresponding to the received signals of the past Δ frames to Δ+K l -1 frames, represents the pure signal, l represents a frequency point, and τ represents a frame number, is called the conjugate transpose matrix of the prediction matrix (also referred to as a linear coefficient).
[0141] From the above description, it can be understood that the purpose of the dereverberation processing is to estimate the prediction matrix, obtain the reverberation signal according to the prediction matrix, subtract the reverberation signal from the current received signal, and thus recover the current pure signal, i.e., the signal in which the reverberation signal is removed.
[0142] It should be noted that when the noise is ignored, the signal in which the reverberation signal is removed is the sound source signal obtained from the received signal. When the noise exists, the signal in which the reverberation signal is removed is the signal in which the estimated reverberation signal is removed from the received signal, which can also be understood as a mixture of the sound source signal and the noise.
[0143] Therefore, the electronic device performs the dereverberation process including the process of calculating the prediction matrix, and the electronic device can solve the prediction matrix in an iterative manner such that the processed signal of the dereverberation process passes through the matrix to separate the reverberation signal and the signal of the dereverberation signal as much as possible.
[0144] The process of solving the prediction matrix can be a process of optimizing the problem through a maximum likelihood function, and the optimization problem can be expressed as:
[0145] Step 1, initialization of the prediction matrix, where τ denotes a frame number, Δ ≤ τ ≤ Δ + K l -1.
[0146] Step 2, reverberation calculation
[0147]
[0148] Γ is a set of frame numbers of the received signal, and the other parameters are explained with reference to the first mixing model.
[0149] Step 3, spatial relation number estimation
[0150]
[0151] where E() is an expectation function, denotes a processed signal, denotes a conjugate transpose matrix of, and δ is a predetermined normal number.
[0152] Step 4, calculation of a weighted sample correlation matrix, assuming that the clean signal is subject to a Gaussian distribution, that is, then the following expression is obtained:
[0153]
[0154]
[0155] where is a conjugate transpose matrix of y l (t), ψ denotes a conjugate transpose matrix of ψ l (t - Δ), and N is the number of microphones in the microphone array.
[0156] Step 5, prediction matrix parameter update
[0157] The updated prediction matrix is obtained by rearranging the term
[0158] Step 6, judging whether the prediction matrix converges or not, if not, returning to step 2, if yes, ending.
[0159] The process of calculating the prediction matrix by the electronic device is described above, of course, in actual application, other iterative ways can also be used to calculate the prediction matrix, which is not limited here.
[0160] The embodiment of the present application illustrates the iterative process of the electronic device performing the dereverberation processing by example, it can be understood that the prediction matrix updated each time uses the prediction matrix updated last time and the processing signal (i.e. l (t))).
[0161] The above method is a post-reverberation suppression technique based on delay linear prediction, which can effectively suppress the post-reverberation signal (i.e. late reverberation signal), however, it also damages the short-time correlation of the speech, so it increases the independence between channels to some extent.
[0162] As an example of blind source separation processing, the method of the electronic device performing the blind source separation processing is to separate the sound source signal of each target sound source from the received signal (which can also be understood as the processing signal in the blind source separation processing), the blind source separation refers to the process of recovering each independent component (for example, the sound source signal of each target sound source) from the received signal in the blind source separation processing according to the statistical characteristics of the input sound source signal without knowing the parameters of the sound source signal and the transmission channel.
[0163] Referring to Figure 5 , Figure 5 The second mixing model of the sound source signals of multiple target sound sources is shown in FIG. 2, as shown in the figure, there can be multiple target sound sources in the environment, so there can be multiple sound source signals of target sound sources in the audio signal, there is also environmental noise in the environment, and the microphone array collecting the audio signal can also cause device noise in the audio signal due to its own reasons, therefore, the audio signal collected by the microphone array can be set to be mixed by the sound source signals of multiple target sound sources and noise, and the reverberation signal in the audio signal is ignored in the blind source separation processing, so the second mixing model can be expressed as: Figure 5
[0164] X=AS+N s
[0165] Wherein, X represents the processing signal in the blind source separation processing, A is the mixing matrix, S is the sound source signal, and N s is noise.
[0166] Of course, if the microphone array collecting the audio signal includes multiple microphones, the received signal is a multi-channel audio signal, and it is assumed that the number of channels of the processing signal is M.
[0167] The processing signal in the blind source separation processing is represented in the time domain as:
[0168] X(t) = [x1(t), x2(t), …, xM(t)] M T .
[0169] Suppose that the processing signal corresponds to N independent sound source signals, which are represented in the time domain as:
[0170] S(t) = [s1(t), s2(t), …, sN(t)] N T .
[0171] The second mixing model of the multiple sound source signal mixing system is:
[0172] X(t) = AS(t) + N s (t).
[0173] Wherein, X(t) is an M-dimensional observation vector, S(t) is an N-dimensional unknown sound source signal vector, N s (t) is an M-dimensional noise, and A is an M*N-dimensional mixing matrix.
[0174] Figure 6 The separation model diagram when each target sound source signal is separated from the processing signal is shown in the separation model diagram. According to the separation model diagram, the electronic device obtains the estimated value of each sound source information after performing blind source separation processing on the processing signal, which can also be regarded as a sound source signal. It can be represented as follows:
[0175]
[0176] Wherein, represents the estimated vector of the sound source signal, and W is the demixing matrix.
[0177] From the above formula, it can be understood that the blind source separation processing includes the process of solving the demixing matrix. In the embodiments of the present application, the electronic device can solve the demixing matrix in an iterative manner, so that the processing signal in the blind source separation processing passes through the matrix and separates each component as much as possible.
[0178] In performing blind source separation processing, the electronic device can use an independent vector analysis (IVA), independent component analysis (ICA), independent low-rank matrix analysis (ILRMA), or the like to separate the target sound source signal, and can perform separation based on signal independence maximization in the separation process. The cost function in the separation process can be a log-likelihood function of maximum likelihood estimation, and the specific separation method, cost function, and optimization algorithm are not limited.
[0179] For example, in an embodiment of the present application, the electronic device can use independent vector analysis. The essence of independent vector analysis is to extend independent component analysis technology to multiple data sets, fully utilize the statistical correlation between multiple data sets, and decompose the data sets using high-order statistics and second-order statistics. The goal of independent vector analysis is that each source in each data set is independent of each other, and a certain source in each data set is at most related to one source in other data sets.
[0180] To make the electronic device satisfy the above description when performing blind source separation processing on the processing signal, in an embodiment of the present application, the data at each frequency (frequency point) in the frequency domain signal can be referred to as a data set, and there are multiple data sets, each of which is linearly mixed by multiple independent sound source signals.
[0181] In independent vector analysis, a source component vector (SCV) is defined, which is composed of sound sources corresponding to different data sets. If the second mixing model is converted from the time domain representation described above to the frequency domain representation.
[0182] Independent vector analysis is actually to determine a cost function containing a demixing matrix and an optimization algorithm for solving the demixing matrix in the cost function. The cost function used needs to be based on a separation criterion of independence measure, such as non-Gaussian maximization criterion, mutual information minimization criterion, information maximization, maximum likelihood criterion, etc. The following is described from the perspective of the frequency domain.
[0183] For example, the cost function can be:
[0184]
[0185] where J(W) represents the mutual information in the SCV, E[ ] represents expectation, s k is the vector of the kth sound source, there are K sound sources, G() is a contrast function, and if G(sk ) = -log p(s k ), G() is the maximum entropy contrast function, p(s k ) is the marginal density function of SCV, ω represents frequency, there are N ω frequency points, and det() is the determinant of a matrix.
[0186] As can be seen from the above cost function, when the cost function is minimized, the entropy values between the sound source vectors are also minimized, and the mutual information between the SCVs is minimized.
[0187] For the above cost function, the process of minimizing the cost function is the process of iteratively solving the demixing matrix W.
[0188] As an example, the process of minimizing the cost function is as follows:
[0189] 1. Update the weighted covariance V k (ω):
[0190]
[0191] wherein E() represents an expectation function, For each frequency point ω, r k is common, represents the conjugate transpose matrix of the term w k (ω) in the demixing matrix, x(ω) is a received signal, and x h (ω) represents the conjugate transpose matrix of x(ω).
[0192] 2. Update the demixing matrix W:
[0193] w k (ω)←(W(ω)V k (ω)) -1 e k
[0194]
[0195] The explanations of the parameters are as described above.
[0196] Arrange w k (ω) to obtain the updated demixing matrix, and steps 1 and 2 are executed in a loop until the demixing matrix converges.
[0197] As can be understood from the above update process of the demixing matrix, the demixing matrix each time iteration update needs to use the demixing matrix and the processing signal (x(ω)) of the last iteration update.
[0198] Of course, in practical applications, batch processing algorithms, adaptive algorithms, successive extraction algorithms, gradient descent methods, and Newton-Raphson iterative algorithms can also be used to estimate the unmixing matrix of each data set.
[0199] As described above, the dereverberation process will increase the independence between channels when the electronic device performs the dereverberation process. Therefore, in an embodiment of the present application, the independent iterative process of the dereverberation process and the independent iterative process of the blind source separation process are combined, so that the dereverberation process and the blind source separation process are performed simultaneously. Moreover, when the electronic device performs the joint iterative processing, since the dereverberation processing signal is an audio signal with noise removed, and the blind source separation processing signal is an audio signal with reverberation components removed, the dereverberation processing and the blind source separation processing processes both meet their respective algorithm models, thereby being able to obtain more accurate separation results.
[0200] For example, the joint iterative process of dereverberation and blind source separation can adopt the process of "dereverberation, blind source separation, dereverberation, blind source separation, ... " or the process of "blind source separation, dereverberation, blind source separation, dereverberation, ... ". For ease of description, the embodiment of the present application denotes "dereverberation-blind source separation" as one joint iterative process. Of course, in actual applications, "blind source separation-dereverberation" can also be considered as one joint iterative process.
[0201] Of course, the dereverberation in one iteration of the joint iterative process does not mean that the prediction matrix is updated to convergence in one iteration, but means that the prediction matrix is updated in one or more iterations in the dereverberation process, that is, the prediction matrix obtained in the dereverberation in one iteration of the joint iterative process can not converge. Similarly, the blind source separation in one iteration of the joint iterative process does not mean that the demixing matrix is updated to convergence in one iteration, but means that the demixing matrix is updated in one or more iterations in the blind source separation process, that is, the demixing matrix obtained in the blind source separation in one iteration of the joint iterative process can not converge. In each iteration of the dereverberation, the electronic device can obtain the reverberation signal and the dereverberation signal in the current iteration according to the prediction matrix obtained in the current iteration. As the iteration process proceeds, the prediction matrix becomes more and more accurate, and the obtained reverberation signal and dereverberation signal become more and more accurate. In each iteration of the blind source separation, the electronic device can obtain the sound source signal of each target sound source and the noise in the current iteration according to the demixing matrix obtained in the current iteration. As the iteration process proceeds, the demixing matrix becomes more and more accurate, and the obtained sound source signal of each target sound source and the noise become more and more accurate. This cycle continues until the condition for stopping the cycle is met. The electronic device can determine whether the condition for stopping the cycle is met by determining whether the number of joint iterations reaches a preset number, or by determining whether the obtained prediction matrix and the obtained demixing matrix both converge. After the condition for stopping the cycle is met, the sound source signal of one or more target sound sources and the noise are obtained according to the demixing matrix obtained in the last iteration and the dereverberation signal obtained in the last iteration.
[0202] Referring to Figure 7 , Figure 7 A schematic diagram of the process of the joint iterative dereverberation and blind source separation of the electronic device is shown, taking a 3-time joint iteration process as an example.
[0203] In the first joint iteration process, in the dereverberation process, the electronic device calculates the prediction matrix 1 according to the audio signal. Then, the electronic device calculates the reverberation signal 1 and the dereverberation signal 1 (the dereverberation signal 1 includes signals in the audio signal other than the reverberation signal 1) according to the prediction matrix 1. In the blind source separation stage, the electronic device takes the dereverberation signal 1 as the processing signal, calculates the demixing matrix 1, and obtains the noise and the de-noise signal 1 (the de-noise signal 1 includes signals in the audio signal other than the noise signal 1) according to the demixing matrix 1.
[0204] During the second joint iterative processing, in the dereverberation processing stage, the electronic device updates prediction matrix 1 based on denoised signal 1 to obtain prediction matrix 2. Based on prediction matrix 2, the electronic device calculates reverberation signal 2 and dereverberation signal 2 (dereverberation signal 2 includes signals in the audio signal other than reverberation signal 2). During the blind source separation processing stage, the electronic device uses dereverberation signal 2 as the processing signal, calculates demixing matrix 2, and obtains noise and dereverberation signal 2 based on demixing matrix 2 (dereverberation signal 2 includes signals in the audio signal other than noise signal 2).
[0205] During the third joint iterative processing, in the dereverberation processing phase, the electronic device updates prediction matrix 2 based on denoised signal 2 to obtain prediction matrix 3. The electronic device calculates reverberation signal 3 and dereverberation signal 3 based on prediction matrix 3 (the dereverberation signal 3 includes signals in the audio signal other than reverberation signal 3). In the blind source separation processing phase, the electronic device uses dereverberation signal 3 as the processing signal, calculates to obtain demixing matrix 3, and obtains noise and dereverberation signal 3 based on demixing matrix 3 (the dereverberation signal 3 includes signals in the audio signal other than noise signal 3).
[0206] After the joint iterative processing is completed, the electronic device calculates the sound source signal of each target sound source based on the dereverberation signal obtained in the last joint iteration and the demixing matrix obtained in the last joint iteration. Of course, the electronic device can also calculate the sound source signal of each target sound source based on the dereverberation signal obtained in the second-to-last joint iteration and the demixing matrix obtained in the last calculation. This is because, at the end of the joint iteration, the prediction matrix and the demixing matrix are converged, that is, the difference between the prediction matrices of the last several times is small or the difference is within an acceptable range, and the difference between the demixing matrices of the last several times is small or the difference is within an acceptable range. Therefore, the electronic device can calculate the sound source signal of each target sound source based on any one of the dereverberation signals obtained in the last several joint iterations and any one of the demixing matrices obtained in the last several joint iterations.
[0207] As another example, during each joint iterative processing, during the dereverberation processing stage, when the electronic device calculates the prediction matrix obtained by this dereverberation, it may independently iterate the prediction matrix multiple times, and use the prediction matrix obtained by the last independent iteration as the prediction matrix obtained by this joint iterative dereverberation; during the blind source separation processing stage, when the electronic device calculates the demixing matrix obtained by this blind source separation, it may also independently iterate the demixing matrix multiple times, and use the demixing matrix obtained by the last independent iteration as the demixing matrix obtained by this joint iterative blind source separation. That is, the electronic device performing the j-th dereverberation processing includes the process of updating the prediction matrix m times, and the electronic device performing the j-th blind source separation processing includes the process of updating the demixing matrix n times, where j, m, and n are all positive integers.
[0208] As another example, in each joint iteration process, the dereverberation signal obtained in the last dereverberation process can be taken as the processing signal of the blind source separation, and the audio signal can be taken as the processing signal of the dereverberation. Of course, in each joint iteration process, the de-noise signal obtained in the last blind source separation process can be taken as the processing signal of the dereverberation, and the audio signal can be taken as the processing signal of the blind source separation.
[0209] As another example, in each joint iteration process, the dereverberation signal obtained in any one of the historical iteration processes can be taken as the processing signal of the blind source separation in this iteration process, and the de-noise signal obtained in any one of the historical iteration processes can be taken as the processing signal of the dereverberation in this iteration process.
[0210] As an example, if the pth prediction matrix obtained in the pth dereverberation process converges and the pth de-mixing matrix obtained in the pth blind source separation process does not converge, the electronic device performs the p+i th blind source separation process until the p+i th de-mixing matrix converges,
[0211] Or, the electronic device alternately performs the p+i th dereverberation process and the p+i th blind source separation process until the p+i th prediction matrix and the p+i th de-mixing matrix converge simultaneously, where p is a positive integer, and i is a positive integer starting from 1.
[0212] As another example, in the joint iteration process, if the qth prediction matrix obtained in the qth dereverberation process does not converge and the qth de-mixing matrix obtained in the qth blind source separation process converges, the electronic device performs the q+i th dereverberation process until the q+i th prediction matrix converges,
[0213] Or, the electronic device alternately performs the q+i th dereverberation process and the q+i th blind source separation process until the q+i th prediction matrix and the q+i th de-mixing matrix converge simultaneously, where q is a positive integer, and i is a positive integer starting from 1.
[0214] As shown in Figure 8 As shown in Figure 8 Another possible embodiment of the method for estimating the direction of arrival of the sound source provided by the present application is shown, which includes steps 801 to 804.
[0215] Steps 801 to 803 can refer to the description of steps 401 to 403 described above, and will not be described here.
[0216] In step 804, the electronic device performs joint processing of source separation and DOA estimation on the sound source signal of the first target sound source according to the first direction information of the sound source signal of the first target sound source at different frequency points, to obtain second direction information of the sound source signal of the first target sound source at different frequency points. The source separation processing includes joint iterative processing of dereverberation and blind source separation.
[0217] The step 804 can be implemented by the following steps:
[0218] A step, the electronic device performs source separation processing on the sound source signal of the first target sound source according to the first direction information of the sound source signal of the first target sound source at different frequency points, to obtain the sound source signal of the first target sound source after this joint processing.
[0219] B step, the electronic device performs DOA estimation processing on the sound source signal of the first target sound source after this joint processing, to obtain the first direction information of the first target sound source at different frequency points after this joint processing.
[0220] C step, the electronic device performs the A step to the B step for Q times, and records the first direction information of the first target sound source at different frequency points obtained at the last time as the second direction information of the first target sound source at different frequency points, where Q≥1 and Q is a positive integer.
[0221] In the embodiments of the present application, dereverberation and blind source separation are both estimation algorithms. After joint iterative processing of dereverberation and blind source separation, the sound source signal of the target sound source obtained by the joint iterative processing may still contain noise and / or reverberation signals, and the first direction information obtained by the electronic device based on the sound source signal containing noise and / or reverberation signals may not be very accurate. Therefore, the electronic device can continue to perform source separation processing and DOA estimation processing on the sound source signal of the first target sound source based on the first direction information of the sound source signal of the first target sound source at different frequency points obtained at the current time. Since the process of source separation processing is constrained by the first direction information of the sound source signal of the first target sound source at different frequency points when the source separation processing is performed again, the electronic device can obtain more accurate sound source signal of the first target sound source, demixing matrix and mixing matrix after performing source separation. The electronic device can obtain more accurate second direction information of the sound source signal of the first target sound source at different frequency points according to the more accurate sound source signal of the first target sound source, demixing matrix or mixing matrix. Of course, the second direction information of the sound source signal of the first target sound source at different frequency points obtained is more accurate than the first direction information of the sound source signal of the first target sound source at different frequency points.
[0222] Since the dereverberation processing, the blind source separation processing, and the DOA estimation process are all estimation processes based on a model, in the joint iterative processing of dereverberation and blind source separation performed by the electronic device, a small amount of reverberation signals can still be mixed in the last obtained dereverberation signal, and some noise and / or reverberation signals can still be mixed in the obtained sound source signal of the target sound source, or even a small amount of sound source signals of other target sound sources can be mixed in. Therefore, the electronic device performs the sound source separation processing again on the sound source signal of the first target sound source based on the first direction information of the first target sound source obtained in the step 704, which can make the obtained sound source signal of the first target sound source purer, thereby improving the accuracy of the subsequent calculation of the sound source signal of the first target sound source.
[0223] As a possible embodiment, the method for estimating the direction of arrival of a sound source provided by the embodiments of the present application further includes:
[0224] The electronic device performs smoothing filtering processing or kernel density estimation processing on the first direction information of the sound source signal of the first target sound source at different frequency points to obtain third direction information of the first target sound source at different frequency points; and the electronic device fuses the third direction information of the first target sound source at different frequency points to obtain the direction of the first target sound source.
[0225] For example, the electronic device performs smoothing filtering processing or kernel density estimation processing on the first direction information of the sound source signal of the first target sound source at frequency point 1, frequency point 2, …, and frequency point L (assuming that there are L frequency points) to obtain third direction information of the first target sound source at frequency point 1, frequency point 2, …, and frequency point L. The purpose of the smoothing filtering processing or the kernel density estimation processing can be to remove some interference, thereby obtaining more accurate third direction information of the sound source signal of the first target sound source at different frequency points. Finally, the electronic device fuses the third direction information of the sound source signal of the first target sound source at different frequency points together to determine a more accurate direction of the sound source signal of the first target sound source.
[0226] Of course, in actual applications, the electronic device can perform smoothing filtering processing or kernel density estimation processing on the first direction information of the sound source signal of the first target sound source at different frequency points, or perform smoothing filtering processing or kernel density estimation processing on the second direction information of the sound source signal of the first target sound source at different frequency points, thereby obtaining the third direction information of the sound source signal of the first target sound source at different frequency points. As can be understood from the above description, the second direction information of the sound source signal of the first target sound source at different frequency points is more accurate than the first direction information of the sound source signal of the first target sound source at different frequency points, and the third direction information of the sound source signal of the first target sound source at different frequency points is more accurate than the first direction information or the second direction information of the sound source signal of the first target sound source at different frequency points.
[0227] The first direction information (for example, the angle value of each frequency point) of the sound source signal of the first target sound source at each frequency point is composed into a first set, the second direction information of the sound source signal of the first target sound source at each frequency point is composed into a second set, and the third direction information of the sound source signal of the first target sound source at each frequency point is composed into a third set. The angle values in the first set may be relatively scattered, the angle values in the second set are more concentrated than the angle values in the first set, and the angle values in the third set are more concentrated than the angle values in the second set.
[0228] As another embodiment of the present application, the microphone array can be an Acoustic Vector Sensor (AVS), which can also be referred to as a sound vector microphone.
[0229] Since the DOA estimation algorithm is related to the size and arrangement of the microphone array, when the electronic device performs the DOA estimation process, the algorithm for DOA estimation needs to be adjusted according to the size and arrangement of the microphone array that collects the audio signal. Referring to Figure 9 A schematic diagram of a structure of a sound vector microphone is shown. The AVS is composed of one omnidirectional microphone and two to three orthogonal 8-shaped microphones. It can be considered that the microphones in the AVS are copoint, and the size of the AVS can be made smaller. For a sound source, the sound source signals received by each microphone channel of the sound source do not have phase differences. The amplitudes of the sound source signals received by each microphone channel of the same sound source are related to the direction of the sound source. Therefore, when the sound vector microphone is applied, the array size and arrangement do not need to be considered, and the sound vector microphone has a wider application scenario.
[0230] The sound source signal obtained after the audio signal is subjected to sound source separation processing is a signal in which noise and reverberation signals are removed. The microphone array is a copoint microphone. Therefore, for a sound source signal, the components of the sound source signal received by each channel do not have phase differences. Therefore, when a sound source signal is collected by two or three orthogonal microphone channels, the amplitude of the sound source signal on each channel is related to the direction, or the direction of arrival of the sound source signal is related to the amplitude of the sound source signal on each channel.
[0231] Taking a three-dimensional four-channel microphone array as an example, the three-dimensional four-channel microphone array includes one omnidirectional microphone and three orthogonal 8-shaped microphones. That is, the three-dimensional microphone array includes four microphones, and the audio signal collected by the microphone array is a four-channel audio signal. Assuming that the channel of the omnidirectional microphone is represented by W, the channels of the three orthogonal 8-shaped microphones are represented by X, Y, and Z respectively, the amplitude and direction of the microphone array are represented as follows:
[0232]
[0233] wherein W represents the amplitude of the signal collected by the omnidirectional microphone channel; X, Y and Z respectively represent the amplitudes of the signals collected by the channels in the three orthogonal directions of the Cartesian coordinate system, and represents the horizontal angle of the sound source in the Cartesian coordinate system, represents the pitch angle of the sound source in the Cartesian coordinate system, and F represents the amplitude of the signal collected by the omnidirectional channel.
[0234] For the sound source signal from which the reverberation signal and noise are removed, the above amplitude and direction relationship is satisfied, and therefore, the electronic device can obtain the direction of arrival of the second target sound source according to the relationship between the amplitudes of the sound source signal of the second target sound source in the directions of the channels of the acoustic vector sensor, the second target sound source being any one of the one or more target sound sources.
[0235] As an example, in a three-dimensional four-channel acoustic vector microphone, the ratio between the amplitudes in the X channel and the Y channel can obtain the inverse tangent value of the horizontal angle, and the ratio between the amplitudes in the Z channel and the W channel can obtain the inverse sine value of the pitch angle.
[0236] Of course, based on the relationship between the amplitudes in the channels of the acoustic vector microphone and the direction angle, other relationships between the amplitudes in the channels can also be derived to obtain the horizontal angle and the pitch angle, which will not be exemplified one by one.
[0237] Taking a two-dimensional three-channel microphone array as an example, the amplitudes and directions of the microphone array are represented as follows:
[0238]
[0239] wherein W represents the amplitude of the signal collected by the omnidirectional microphone channel; X, Y and Z respectively represent the amplitudes of the signals collected by the channels in the three orthogonal directions of the Cartesian coordinate system, and represents the horizontal angle of the sound source in the Cartesian coordinate system,
[0240] The ratio between the amplitudes of the sound source signal of the second target sound source in the X channel and the Y channel can obtain the inverse tangent value of the horizontal angle.
[0241] In addition, it should be noted that, taking the two-dimensional three-way sound vector microphone as an example, the relationship between the amplitude and the direction angle on each channel of the sound vector microphone described above is a first-order relationship. In order to obtain a more accurate DOA estimation result, when considering the relationship between the amplitude and the phase and the direction angle, there can also be a second-order relationship, or even a higher-order relationship. Of course, the microphone array can also be in other forms, such as a spherical microphone array or a ring-shaped microphone array. If the spherical microphone array or the ring-shaped microphone array is considered to be a case where each channel is at the same point, the phase difference on each channel can also not be considered. If it is necessary to consider the phase and the amplitude difference on each channel of the spherical microphone array or the ring-shaped microphone array, the relationship between the amplitude, the phase, and the direction angle on each channel can have a higher-order relationship, and the present application does not limit the order of the relationship between the amplitude, the phase, and the direction.
[0242] The above calculation process can be performed by the electronic device on the data of each frequency point of the sound source signal of the second target sound source, that is, the obtained horizontal angle and the elevation angle are the horizontal angle and the elevation angle of the second target sound source at each frequency point. The horizontal angle and the elevation angle of the second target sound source at each frequency point can be referred to as the DOA information (for example, the first direction information, the second direction information, and the third direction information described above) of the second target sound source.
[0243] In some embodiments, another way of obtaining the first direction information of the second target sound source is also provided.
[0244] In the description of the above second mixing model, since the amplitudes of the sound source signal of the second target sound source on each channel are mixed by the mixing matrix, and each channel of the microphone array is considered to be at the same point. That is, there is no phase difference between the components of the sound source signal of the second target sound source on each channel, and therefore, the mixing matrix implicitly contains the ratio between the amplitudes of the sound source signal of the second target sound source in the direction of each channel of the acoustic vector sensor. Therefore, the electronic device can obtain the direction of arrival of the sound source signal of the second target sound source according to the mixing matrix of the sound source signal of one or more target sound sources.
[0245] Taking a two-dimensional microphone array as an example, the second mixing model of the sound source signal is:
[0246]
[0247] wherein X W , X X , and X Y are the audio signals received by the three channels of the two-dimensional AVS, S1 and S2 are two sound source signals, and N = AN'.
[0248] Meanwhile, as described above, for one sound source signal, there is the following relationship between the amplitude and the horizontal angle of each channel:
[0249]
[0250] wherein, W represents the amplitude of the signal collected by the omnidirectional microphone channel; X and Y respectively represent the amplitudes of the signals collected by the channels in the two horizontal orthogonal directions of the Cartesian coordinate system, θ represents the horizontal angle of the sound source, and F represents the amplitude of the signal collected by the omnidirectional channel.
[0251] It can be understood from the above two formulas that the first column in the mixing matrix represents the column in which the sound source signal of the first target sound source is located, and the second column in the mixing matrix represents the column in which the sound source signal of the second target sound source is located. The first row in the mixing matrix represents the omnidirectional channel, the second row represents the X channel, and the third row represents the Y channel.
[0252] Based on the above description, the electronic device obtains the direction of arrival of the sound source signal of the second target sound source according to the mixing matrix of the sound source signal of one or more target sound sources, including: the electronic device determines a target column in the mixing matrix, and a first target row and a second target row in the target column, wherein the target column is a column representing the sound source signal of the second target sound source, and the first target row and the second target row are rows related to the angle of the sound source signal of the second target sound source; and the electronic device obtains the direction of arrival of the sound source signal of the second target sound source according to elements of the first target row and the second target row in the target column. When the first target row represents the row of the X channel of the acoustic vector sensor and the second target row represents the row of the Y channel of the acoustic vector sensor, the direction of arrival of the sound source signal of the second target sound source includes the horizontal angle of the sound source signal of the second target sound source, and the horizontal angle is the angle in the coordinate system in which the acoustic vector sensor is located.
[0253] As an example, the electronic device calculates the first direction information of each frequency point of the sound source by using the mixing matrix A, and when solving the first direction information of each frequency point of the sound source signal of the αth target sound source, any one of the following is used:
[0254]
[0255] wherein, A χα represents the element of the γth row and the αth column in the mixing matrix corresponding to the αth sound source, θ is the horizontal angle of the sound source signal at each frequency point, and γ = 1, 2, 3. That is, after the target column is determined, the elements on the rows corresponding to any two channels of the three channels of the two-dimensional three-channel acoustic vector microphone can be calculated to obtain the horizontal angle.
[0256] Taking the three-dimensional four-channel acoustic vector microphone as an example, the calculation manner of the horizontal angle refers to the description in the two-dimensional three-channel acoustic vector microphone, and when calculating the pitch angle, the target column needs to be determined first, and then the pitch angle is obtained according to the ratio between the element on the row where the Z channel is located and the element on the row where the omnidirectional channel is located in the mixing matrix.
[0257] For example, when solving the pitch angle of the sound source signal of the αth target sound source at each frequency point, any one of the following is adopted:
[0258]
[0259]
[0260] wherein A χα represents the element of the γth row and the αth column in the mixing matrix corresponding to the αth sound source, θ is the horizontal angle of the sound source signal at each frequency point, γ = 1, 2, 3, 4, and θ is the horizontal angle.
[0261] As described before, the mixing matrix and the demixing matrix are reciprocal matrices, so the electronic device can also use the demixing matrix W to calculate the first direction information of the sound source signal of the target sound source at each frequency point.
[0262] As a possible embodiment, the electronic device obtains the direction of arrival of the sound source signal of the second target sound source according to the demixing matrix of the sound source signal of one or more target sound sources, which includes: the electronic device obtains the horizontal angle of the sound source signal of the second target sound source according to the element representing the X channel of the acoustic vector sensor in the column of the element representing the Y channel of the acoustic vector sensor in the target row representing the second target sound source in the demixing matrix, the horizontal angle being the angle in the coordinate system in which the acoustic vector sensor is located; and the electronic device obtains the pitch angle of the sound source signal of the second target sound source according to the element representing the Z channel of the acoustic vector sensor in the column of the element representing the omnidirectional channel of the acoustic vector sensor in the target row representing the second target sound source in the demixing matrix, the pitch angle being the angle in the coordinate system in which the acoustic vector sensor is located.
[0263] Taking the two-dimensional three-channel acoustic vector microphone as an example, the electronic device first determines the target row, and then calculates the horizontal angle by using the ratio between the element in the column where the X channel is located and the element in the column where the Y channel is located in the target row in the demixing matrix. Y
[0264] For example, when solving the first direction information of the αth sound source at each frequency point, the electronic device adopts:
[0265]
[0266] wherein W γα The element in the αth row and γth column of the unmixing matrix corresponding to the αth sound source, γ = 1, 2, 3. As before, after determining the target row, the horizontal angle can be calculated from the elements in the columns corresponding to any two channels of the three channels of the two-dimensional three-channel sound vector microphone.
[0267] When the microphones are a three-dimensional, four-channel acoustic vector microphone array, the horizontal angle is calculated using the same method as described above. The pitch angle can be calculated by first determining the target row and then calculating the pitch angle based on the ratio between the elements in the column containing the Z channel and the elements in the column containing the omnidirectional channel in the unmixing matrix. Of course, in practical applications, many more calculation methods can be developed, and these are not listed here.
[0268] Through the above description process, it can be understood that, in fact, dereverberation processing and blind source separation processing are both estimation methods. The sound source signal obtained by the electronic device performing the above estimation method may also contain some interference factors. In addition, in order to be applied in different application scenarios, the electronic device can further perform other processing, such as the first enhancement processing and the second enhancement processing described later, and the processing of adjusting the proportional relationship between the sound source signal, the reverberation signal and the noise.
[0269] As a possible implementation, see Figure 10 , Figure 10 The illustrated figure includes dereverberation processing, blind source separation processing, DOA estimation, and enhancement processing. Dedereverberation processing, blind source separation processing, and DOA estimation processing can be described in the above embodiments and will not be repeated here. Enhancement processing can include two methods: one in which the electronic device performs a first enhancement processing on the sound source signal; the other in which the electronic device performs a second enhancement processing on the sound source signal based on the first, second, or third directional information of the sound source signal. Figure 10 The enhancement processing in the illustrated figures may include the second enhancement processing or the first enhancement processing. In the following description, the first target sound source and the first direction information are used as an example. The first target sound source is any one of the one or more target sound sources. In actual applications, the first enhancement processing, the second enhancement processing, and the proportional relationship adjustment processing described above may be performed for each target sound source.
[0270] The electronic device performs the first enhancement processing on the sound source signal of the first target sound source, where the first enhancement processing includes interference spectrum filtering processing and / or harmonic enhancement processing, the first target sound source is any one of the sound source signals of the one or more target sound sources; the interference spectrum filtering processing is used to filter out the interference components mixed in the sound source signal of the first target sound source based on the spectral energy of the sound source signal of the first target sound source; and the harmonic enhancement processing is used to obtain a harmonic enhanced signal of the first target sound source, where the harmonic enhanced signal is a sound source signal containing harmonic components.
[0271] In some embodiments, the interference spectrum filtering processing is a process in which the electronic device filters out the interference components mixed in the sound source signal of the first target sound source based on the spectral energy of the sound source signal of the first target sound source.
[0272] For example, the electronic device uses a Gaussian Mixture Model (GMM) to model the spectral energy of the sound source signal of the first target sound source, determines a main sound source spectral range according to the spectral energy, and then deletes the sound source signal corresponding to the spectral energy that is not within the main sound source spectral range in the sound source signal of the first target sound source.
[0273] After the electronic device performs the interference spectrum filtering processing, it can remove other interference signals in the sound source signal of the first target sound source that have a large difference in spectral energy from the main sound source, thereby obtaining a more pure sound source signal.
[0274] When the collection device of the audio signal converts the live sound into an audio signal, it generally does not fully record and convert the entire quality of the live sound, resulting in the audio signal not including many original harmonics. However, many sounds with tones or fundamental frequencies that people hear usually contain harmonics, for example, harmonics can produce the tone quality or sound quality emitted by a musical instrument. In order to enrich the sound we hear or to restore the real sound emitted by a musical instrument, it is necessary to add harmonic components to the audio signal. Harmonic enhancement processing is a technique in which the electronic device adds harmonics to the sound source signal. Of course, the sound source signal before the harmonic enhancement processing can or can not contain harmonic components. The purpose of the harmonic enhancement processing is to obtain a sound source signal containing harmonic components.
[0275] For example, the second enhancement processing based on the first direction information of the sound source signal of the first target sound source, the electronic device performs the second enhancement processing on the sound source signal of the first target sound source based on the first direction information of the sound source signal of the first target sound source at different frequency points, wherein the second enhancement processing includes interference direction filtering processing and / or beamforming directional enhancement processing; the interference direction filtering processing is used to filter out the frequency points in the sound source signal of the first target sound source whose direction angle is not within the expected angle range; and the beamforming directional enhancement processing is used to enhance the power of the sound source signal of the first target sound source in the expected direction.
[0276] In some embodiments, the interference direction filtering processing is used to filter out the frequency points in the sound source signal of the first target sound source whose direction angle is not within the expected angle range, so as to suppress the sound in other directions except the direction of the first target sound source. In specific implementation, the mask in the frequency domain can be performed on the first direction information of the sound source signal of the first target sound source, that is, the θ angle corresponding to the sound source signal is expanded by a certain range, the first direction information at each frequency point is compared with the range, and the component corresponding to the first direction information of the frequency point exceeding the range is removed.
[0277] For example, when the first target sound source is the sound source signal of the kth target sound source, the electronic device can set [θ k -δ k ,θ k +δ k as the mask range of the sound source signal of the kth target sound source when performing the mask processing. The first direction information of the sound source signal of the kth target sound source at each frequency point is compared with the mask range, and the component corresponding to the first direction information of the kth sound source signal which is not within the mask range is removed, k≥1, k is a positive integer.
[0278] The electronic device performs the interference direction filtering processing to remove the interference signal in other directions which is far away from the direction of the sound source signal of the first target sound source in the sound source signal of the first target sound source, so as to obtain a more pure sound source signal.
[0279] The beamforming directional enhancement processing is used to enhance the power of the sound source signal in the expected direction; in specific implementation, the electronic device adds the signals related to the sound source signal of the first target sound source to be enhanced, and does not add the sound source signals and interference of other target sound sources which are not related, so that the power of the sound source signal of the first target sound source to be enhanced is enhanced.
[0280] As an example, the electronic device can process the sound source signals by using a null steering method, which does not have a steering vector and uses the angles of different sound sources in the space to customize beamforming. Assuming there are four sound source signals with different angles, according to this method, only the signal in the direction of the first sound source signal is obtained, and the signals in other directions are suppressed.
[0281] In the embodiments of the present application, the sound source separation, DOA estimation and enhancement processing are combined, and after the audio signals are subjected to the joint iterative processing of the dereverberation processing and the blind source separation processing performed by the electronic device, more accurate sound source signals, reverberation signals and noises are obtained. The electronic device determines the direction of the target sound source by using the DOA estimation method, and further enhances the sound source signals by using the enhancement processing to weaken the interference components, so that the finally obtained sound source signals are more pure, have stronger power or have increased harmonic components, etc., so as to obtain a better hearing effect.
[0282] As another embodiment of the present application, in the process of obtaining the sound source signal of one or more target sound sources by using the joint iterative processing of dereverberation and blind source separation on the audio signals, the electronic device obtains the noise and the reverberation signal of one or more target sound sources from the audio signals.
[0283] The electronic device adjusts the proportional relationship among the noise, the sound source signal of the first target sound source and the reverberation signal of the first target sound source, and the first target sound source is any one of the one or more target sound sources.
[0284] As an example, when a signal of a KTV scene effect is needed, the electronic device can set the loudness proportional relationship among the sound source signal, the reverberation signal and the noise as sound source signal: reverberation signal: noise = β 11 : β 12 : β 13 When a signal of a music hall scene effect is needed, the electronic device can set the loudness proportional relationship among the sound source signal, the reverberation signal and the noise as sound source signal: reverberation signal: noise = β 21 : β 22 : β 23 When a signal of a field scene effect is needed, the electronic device can set the loudness proportional relationship among the sound source signal, the reverberation signal and the noise as sound source signal: reverberation signal: noise = β 31 : β 32 : β 33 .
[0285] When the electronic device adjusts the proportional relationship among the three, the electronic device can adjust the loudness proportional relationship, the power proportional relationship, etc. among the three. Of course, in actual applications, different adjustment parameters can be set according to specific scene effects.
[0286] The first enhancement processing, the second enhancement processing and the proportional relationship adjustment processing mentioned above are all post-processing methods. In actual applications, you can choose one of the post-processing methods, or you can choose to combine multiple post-processing methods. The specific post-processing method can be determined according to the specific application scenario.
[0287] As an example, when a mobile phone equipped with a microphone array is recording or videotaping, the microphone array collects multi-channel audio signals. The electronic device executes the method for estimating the direction of arrival of the sound source provided in the embodiment of the present application, and can separate the sound source signals of one or more target sound sources in the recording scene. It can also perform DOA estimation on the sound source signals of the separated target sound sources, and then realize automatic zooming of the target sound source or change the sound effect of the sound source signal of the target sound source through one or more post-processing.
[0288] As another application scenario, users can make video calls through electronic devices. Taking remote conferences as an example, the large screen used in remote conferences is equipped with a microphone array. The microphone array collects ambient sound to obtain audio signals. The electronic device executes any of the methods for estimating the direction of arrival of a sound source provided in the embodiments of the present application to obtain the sound source signals of one or more target sound sources and the direction of the sound source signal of each target sound source. The electronic device can also enhance the power of the signal in the direction of the speaker or filter out the interference signal in the direction of the speaker through post-processing. When all speakers are speaking and are close to each other, the voice signal of each speaker can be accurately separated and the position of each speaker can be accurately determined.
[0289] Of course, microphone arrays can also be used in hearing aids. In a complex acoustic environment, the microphone array collects ambient sound, and the hearing aid executes any of the above methods for estimating the direction of arrival of the sound source. This can increase the loudness ratio of the sound source signal and reduce the loudness ratio of the reverberation signal and noise, so that the user can obtain a clearer sound source signal when wearing a hearing aid.
[0290] The electronic device performs the above-mentioned processing on the frequency domain signal. After one or more combined processing steps are completed, the electronic device can also convert the frequency domain signal into a time domain signal and then transmit or play the time domain signal. The process of converting the frequency domain signal into the time domain signal by the electronic device can be understood as the inverse process of time-frequency conversion. The inverse process of time-frequency conversion can be referred to the existing frequency-time conversion method and will not be described in detail here.
[0291] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0292] The embodiments of the present application can divide the function modules of the electronic device according to the above method examples. For example, each function module can be divided according to each function, or two or more functions can be integrated in one processing module. The integrated module can be realized in the form of hardware or in the form of a software function module. It should be noted that the division of the modules in the embodiments of the present application is illustrative, and is only a logical function division. In actual implementation, another division mode can be used. The following will be described by taking the division of each function module according to each function as an example:
[0293] With reference to Figure 11 The electronic device 1100 includes:
[0294] An audio signal acquisition unit 1101 is configured to acquire an audio signal, the audio signal including noise, a sound source signal of one or more target sound sources, and a reverberation signal.
[0295] A reverberation removal processing unit 1102 is configured to perform Nth reverberation removal processing on the audio signal to obtain an Nth prediction matrix and an Nth reverberation-removed signal, the Nth reverberation-removed signal including a signal in the audio signal except for an Nth reverberation signal, and the Nth reverberation signal being a reverberation signal removed in the Nth reverberation removal processing.
[0296] A blind source separation processing unit 1103 is configured to perform Nth blind source separation processing on the Nth reverberation-removed signal to obtain an Nth demixing matrix and an Nth noise-removed signal, the Nth noise-removed signal being a signal in the audio signal from which noise obtained by the Nth blind source separation processing is removed.
[0297] A sound source signal obtaining unit 1104 is configured to perform reverberation removal processing and blind source separation processing on the Nth noise-removed signal, and when N is a preset value or the Nth prediction matrix converges and the Nth demixing matrix converges, obtain the sound source signal of the one or more target sound sources according to the Nth demixing matrix and the Nth noise-removed signal, where N is a positive integer starting from 1.
[0298] A sound source direction estimation unit 1105 is configured to determine a direction of arrival of the sound source signal of the one or more target sound sources.
[0299] As another embodiment of the present application, for the first target sound source, the first target sound source is any one of the one or more target sound sources,
[0300] The direction of arrival of the sound source signal of the first target sound source includes first direction information of the sound source signal of the first target sound source at different frequency points.
[0301] As another embodiment of the present application, the sound source direction estimation unit 1105 is further configured to:
[0302] perform joint processing of source separation and direction of arrival estimation on the source signal of the first target sound source according to the first direction information of the source signal of the first target sound source at different frequency points, to obtain second direction information of the source signal of the first target sound source at different frequency points, wherein the source separation processing includes joint iterative processing of dereverberation and blind source separation.
[0303] As another embodiment of the present application, the sound source direction estimation unit 1105 is further configured to:
[0304] perform smoothing filtering processing or kernel density estimation processing on the first direction information of the source signal of the first target sound source at different frequency points, to obtain third direction information of the first target sound source at different frequency points; and fuse the third direction information of the first target sound source at different frequency points to obtain the direction of the first target sound source.
[0305] As another embodiment of the present application, the sound source signal obtaining unit 1104 is further configured to:
[0306] If the pth prediction matrix obtained by the pth dereverberation processing converges and the pth demixing matrix obtained by the pth blind source separation processing does not converge, the electronic device performs the p+i th blind source separation processing until the p+i th demixing matrix converges, or the electronic device alternately performs the p+i th dereverberation processing and the p+i th blind source separation processing until the p+i th prediction matrix and the p+i th demixing matrix converge simultaneously, wherein p is a positive integer and i is a positive integer starting from 1.
[0307] If the qth prediction matrix obtained by the qth dereverberation processing does not converge and the qth demixing matrix obtained by the qth blind source separation processing converges, the electronic device performs the q+i th dereverberation processing until the q+i th prediction matrix converges, or the electronic device alternately performs the q+i th dereverberation processing and the q+i th blind source separation processing until the q+i th prediction matrix and the q+i th demixing matrix converge simultaneously, wherein q is a positive integer and i is a positive integer starting from 1.
[0308] As another embodiment of the present application, the sound source signal obtaining unit 1104 performs the jth dereverberation processing including the process of updating the prediction matrix m times, and performs the jth blind source separation processing including the process of updating the demixing matrix n times, wherein j, m and n are all positive integers.
[0309] As another embodiment of the present application, the audio signal obtaining unit 1101 is further configured to:
[0310] acquire the audio signal through an acoustic vector sensor on the electronic device;
[0311] receive the audio signal acquired by an acoustic vector sensor on another electronic device.
[0312] As another embodiment of the present application, the sound source direction estimation unit 1105 is further configured to:
[0313] The direction of arrival of the sound source signal of the second target sound source is obtained according to one or more of the amplitudes of the sound source signal of the second target sound source on the plurality of channels, a demixing matrix or a mixing matrix of the sound source signals of the one or more target sound sources, the second target sound source being any one of the one or more target sound sources; wherein the demixing matrix represents a conversion relationship when the audio signal is separated into the sound source signals of the one or more target sound sources, and the mixing matrix represents a conversion relationship when the sound source signals of the one or more target sound sources in the audio signal are mixed into the audio signal.
[0314] As another embodiment of the present application, the sound source direction estimation unit 1105 is further configured to:
[0315] The target column in the mixing matrix, and the first target row and the second target row in the target column are determined, wherein the target column is a column representing the sound source signal of the second target sound source, and the first target row and the second target row are rows related to the angle of the sound source signal of the second target sound source; and the direction of arrival of the sound source signal of the second target sound source is obtained according to the elements of the first target row and the elements of the second target row in the target column.
[0316] As another embodiment of the present application, the electronic device 1100 further comprises:
[0317] The post-processing unit 1106 is configured to perform first enhancement processing on the sound source signals of the one or more target sound sources, wherein the first enhancement processing comprises interference spectrum filtering processing and / or harmonic enhancement processing, the first target sound source being any one of the sound source signals of the one or more target sound sources; the interference spectrum filtering processing is configured to filter out interference components mixed in the sound source signal of any one of the one or more target sound sources based on the spectral energy of the sound source signal of the target sound source; and the harmonic enhancement processing is configured to obtain a harmonic enhanced signal of the one or more target sound sources, the harmonic enhanced signal being a sound source signal containing harmonic components.
[0318] As another embodiment of the present application, the post-processing unit 1106 is further configured to:
[0319] The post-processing unit 1106 is configured to perform second enhancement processing on the sound source signal of the first target sound source based on the first direction information of the sound source signal of the first target sound source at different frequency points, wherein the second enhancement processing comprises interference direction filtering processing and / or beamforming directional enhancement processing; the interference direction filtering processing is configured to filter out frequency points in the sound source signal of the first target sound source whose direction angle is not within an expected angle range; and the beamforming directional enhancement processing is configured to enhance the power of the sound source signal of the first target sound source in the expected direction.
[0320] As another embodiment of the present application, the sound source signal obtaining unit 1104 can also obtain: a reverberation signal of the noise and the one or more target sound sources;
[0321] The electronic device 1100 further includes:
[0322] The scene effect processing unit 1107 is configured to adjust a proportional relationship between the noise, the sound source signal of the first target sound source, and the reverberation signal of the first target sound source, the first target sound source being any one of the one or more target sound sources.
[0323] It should be noted that the information interaction between the above electronic devices / units, the execution process, and the like, based on the same concept as the method embodiments of the present application, the specific functions and the technical effects brought by the method embodiments, and the like, can be referred to the method embodiments part, and will not be repeated here.
[0324] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the electronic device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or software. In addition, the specific name of each functional unit and module is only for easy distinction, and does not limit the protection scope of the present application. The specific working process of the unit and module in the system can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0325] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps in each of the method embodiments.
[0326] The embodiments of the present application also provide a computer program product, which, when running on an electronic device, enables the electronic device to implement the steps in each of the method embodiments.
[0327] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The computer program can be stored in a computer readable storage medium, and the computer program can implement the steps of each method embodiment when executed by a processor. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms. The computer readable medium at least includes any entity or device capable of carrying the computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium. For example, a U disk, a mobile hard disk, a magnetic disk or an optical disk, etc. In some jurisdictions, according to legislation and patent practice, the computer readable medium can not be an electrical carrier signal and a telecommunications signal.
[0328] The embodiment of the present application also provides a chip system, which comprises a processor and a memory. The processor is coupled with the memory, and the processor executes a computer program stored in the memory to implement the steps of any method embodiment of the present application. The chip system can be a single chip or a chip module composed of multiple chips.
[0329] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.
[0330] Those skilled in the art can appreciate that the units and method steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0331] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A method of estimating a direction of arrival of a sound source, characterized by, The method comprises the following steps: An electronic device acquires an audio signal, the audio signal comprising: noise, sound source signals of one or more target sound sources, reverberation signals; The electronic device performs Nth de-reverberation processing on the audio signal to obtain an Nth prediction matrix and an Nth de-reverberation signal, the Nth de-reverberation signal comprising signals in the audio signal except for an Nth reverberation signal, the Nth reverberation signal being a reverberation signal removed in the Nth de-reverberation processing; the electronic device performs Nth blind source separation processing on the Nth de-reverberation signal to obtain an Nth de-mixing matrix and an Nth de-noise signal, the Nth de-noise signal being a signal in the audio signal from which noise obtained by the Nth blind source separation processing is removed; The electronic device continues to perform de-reverberation processing and blind source separation processing on the Nth de-noise signal; When N is a preset value or the Nth prediction matrix converges and the Nth de-mixing matrix converges, the electronic device obtains the sound source signals of the one or more target sound sources according to the Nth de-mixing matrix and the Nth de-reverberation signal; N is a positive integer starting from 1; The electronic device determines the directions of arrival of the sound source signals of the one or more target sound sources.
2. The method of claim 1, wherein, For a first target sound source, the first target sound source is any one of the one or more target sound sources, The direction of arrival of the sound source signals of the first target sound source comprises first direction information of the sound source signals of the first target sound source at different frequency points.
3. The method of claim 2, wherein, The method further comprises: The electronic device performs joint processing of sound source separation and direction of arrival estimation on the sound source signals of the first target sound source according to the first direction information of the sound source signals of the first target sound source at the different frequency points, to obtain second direction information of the sound source signals of the first target sound source at the different frequency points, wherein the sound source separation processing comprises joint iterative processing of de-reverberation and blind source separation.
4. The method of claim 2, wherein, The method further comprises: The electronic device performs smoothing filtering processing or kernel density estimation processing on the first direction information of the sound source signals of the first target sound source at the different frequency points to obtain third direction information of the sound source signals of the first target sound source at the different frequency points; The electronic device fuses the third direction information of the sound source signals of the first target sound source at the different frequency points to obtain the direction of the first target sound source.
5. The method of claim 1, wherein, The electronic device continues to perform de-reverberation processing and blind source separation processing on the Nth de-noise signal comprises: If an Nth prediction matrix obtained by Nth de-reverberation processing converges and an Nth de-mixing matrix obtained by Nth blind source separation processing does not converge, the electronic device performs Nth+i blind source separation processing until the Nth+i de-mixing matrix converges, or the electronic device alternately performs Nth+i de-reverberation processing and Nth+i blind source separation processing until the Nth+i prediction matrix and the Nth+i de-mixing matrix converge simultaneously, wherein N is a positive integer, and i is a positive integer starting from 1. If the qth prediction matrix obtained by the qth dereverberation processing does not converge and the qth demixing matrix obtained by the qth blind source separation processing converges, the electronic device performs q+i th dereverberation processing until the q+i th prediction matrix converges, or the electronic device alternately performs q+i th dereverberation processing and q+i th blind source separation processing until the q+i th prediction matrix and the q+i th demixing matrix simultaneously converge, wherein q is a positive integer, and i is a positive integer starting from 1.
6. The method of claim 1, wherein, The process of performing the jth dereverberation processing by the electronic device includes m times of updating the prediction matrix, and the process of performing the jth blind source separation processing by the electronic device includes n times of updating the demixing matrix, wherein j, m and n are all positive integers.
7. The method of claim 1, wherein, The electronic device acquires the audio signal by collecting the audio signal through an acoustic vector sensor on the electronic device. Or, the electronic device receives the audio signal collected by an acoustic vector sensor on another device.
8. The method of claim 7, wherein, The electronic device determines the direction of arrival of the sound source signal of the second target sound source, comprising: The electronic device obtains the direction of arrival of the sound source signal of the second target sound source according to one or more of the amplitude of the sound source signal of the second target sound source on multiple channels, the demixing matrix of the sound source signal of the one or more target sound sources, and the mixing matrix of the one or more target sound sources; The second target sound source is any one of the one or more target sound sources; Wherein, the demixing matrix represents the conversion relationship when the audio signal is separated into the sound source signal of the one or more target sound sources, and the mixing matrix represents the conversion relationship when the sound source signal of the one or more target sound sources in the audio signal is mixed into the audio signal.
9. The method of claim 8, wherein, The electronic device obtains the direction of arrival of the sound source signal of the second target sound source according to the mixing matrix of the sound source signal of the one or more target sound sources, comprising: The electronic device determines a target column in the mixing matrix, and a first target row and a second target row in the target column, wherein the target column is a column representing the sound source signal of the second target sound source, and the first target row and the second target row are rows related to the angle of the sound source signal of the second target sound source; The electronic device obtains the direction of arrival of the sound source signal of the second target sound source according to the elements of the first target row and the elements of the second target row in the target column.
10. The method of claim 9, wherein, When the first target row represents a row of the first channel of the acoustic vector sensor and the second target row represents a row of the second channel of the acoustic vector sensor, the direction of arrival of the sound source signal of the second target sound source includes the horizontal angle of the sound source signal of the second target sound source, and the horizontal angle is the angle in the coordinate system in which the acoustic vector sensor is located. And / or, When the first target row represents a row of the third channel of the acoustic vector sensor and the second target row represents a row of the omnidirectional channel of the acoustic vector sensor, the direction of arrival of the sound source signal of the second target sound source includes the pitch angle of the sound source signal of the second target sound source, and the pitch angle is the angle in the coordinate system in which the acoustic vector sensor is located.
11. The method according to any one of claims 1 to 10, wherein, After the electronic device obtains the sound source signal of the one or more target sound sources according to the Nth demixing matrix and the Nth de-noised signal, the method further comprises: The electronic device performs first enhancement processing on the sound source signal of the one or more target sound sources, wherein the first enhancement processing comprises interference spectrum filtering processing and / or harmonic enhancement processing; The interference spectrum filtering processing is configured to filter out interference components mixed in the sound source signal of any target sound source based on the spectral energy of the sound source signal of the any target sound source. The harmonic enhancement processing is configured to obtain a harmonic enhanced signal of the one or more target sound sources, wherein the harmonic enhanced signal is a sound source signal containing harmonic components.
12. The method of claim 2, 3, or 4, wherein, The method further comprises: The electronic device performs second enhancement processing on the sound source signal of the first target sound source based on the first direction information of the sound source signal of the first target sound source at different frequency points, wherein the second enhancement processing comprises interference direction filtering processing and / or beamforming directional enhancement processing; The interference direction filtering processing is configured to filter out frequency points in the sound source signal of the first target sound source whose direction angles are not within an expected angle range. The beamforming directional enhancement processing is configured to enhance the power of the sound source signal of the first target sound source in an expected direction.
13. The method according to any one of claims 1 to 10, wherein When N is a preset value or the Nth prediction matrix converges and the Nth demixing matrix converges, the method further comprises: The electronic device obtains the noise and the reverberation signal of the one or more target sound sources from the audio signal; The electronic device adjusts the proportional relationship among the noise, the sound source signal of the first target sound source, and the reverberation signal of the first target sound source, wherein the first target sound source is any target sound source in the one or more target sound sources.
14. An electronic device, comprising: The method comprises: An audio signal acquisition unit configured to acquire an audio signal, wherein the audio signal comprises noise, sound source signals of one or more target sound sources, and reverberation signals; An Nth de-reverberation processing unit configured to perform Nth de-reverberation processing on the audio signal to obtain an Nth prediction matrix and an Nth de-reverberation signal, wherein the Nth de-reverberation signal comprises signals in the audio signal except for an Nth reverberation signal, and the Nth reverberation signal is a reverberation signal removed in the Nth de-reverberation processing; A blind source separation processing unit configured to perform Nth blind source separation processing on the Nth de-reverberation signal to obtain an Nth demixing matrix and an Nth de-noised signal, wherein the Nth de-noised signal is a signal in the audio signal from which noise obtained by the Nth blind source separation processing is removed; A sound source signal obtaining unit configured to perform de-reverberation processing and blind source separation processing on the Nth de-noised signal; When N is a preset value or the Nth prediction matrix converges and the Nth demixing matrix converges, the sound source signal of the one or more target sound sources is obtained according to the Nth demixing matrix and the Nth de-reverberation signal, wherein N is a positive integer starting from 1; A sound source direction estimation unit configured to determine the direction of arrival of the sound source signal of the one or more target sound sources.
15. An electronic device, comprising: The electronic device comprises a processor for running a computer program stored in a memory to implement the method of any one of claims 1 to 13.
16. A chip system, characterized by The chip system comprises a processor coupled with a memory, the processor being configured to run a computer program stored in the memory to implement the method of any one of claims 1 to 13.
Citation Information
Patent Citations
Signal processing apparatus for dereverberating a number of input audio signals
CN106233382A
Target speaker estimation method and system applied to multi-person speech mixture
CN108766459A