Voice signal extraction method and device, electronic equipment, chip and medium

By using the sound source probability density and independent vector analysis algorithm trained by the model in vehicle voice separation, the target signal that meets the preset target angle is extracted, solving the problems caused by noise interference and seat position particularity in vehicle voice separation, and achieving efficient and robust signal extraction effect.

CN120199267APending Publication Date: 2025-06-24BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311791456.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In vehicle voice separation, there is noise interference and special distribution of seat positions in the car, resulting in increased difficulty in extracting target signals (such as the direction of the main driver). The existing technology has poor robustness and separation performance in scenarios with low signal-to-noise.

Method used

The sound source probability density obtained through model training combined with independent vector analysis and signal estimation algorithm to extract target signals that meet the preset target angle, solving the problem of channel aliasing, and has a certain fault tolerance for target angle offset and is highly robust.

Benefits of technology

It realizes efficient extraction of target signals in scenarios with relatively low signal-to-noise, improves the robustness and separation performance of vehicle voice separation, and adapts to complex in-vehicle acoustic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199267A_ABST
    Figure CN120199267A_ABST
Patent Text Reader

Abstract

The invention provides a voice signal extraction method and device, electronic equipment, a chip and a medium, and relates to the technical field of Internet of Vehicles. The method comprises the following steps: acquiring a mixed audio signal; based on the mixed audio signal, using a trained sound source probability density model to obtain sound source probability density weights of the plurality of sound source signals; based on the sound source probability density weight, performing independent vector analysis on the mixed audio signal to obtain an estimated signal of each sound source signal and a covariance of the amplitude of each sound source signal; and determining a target voice signal conforming to a preset target angle from the mixed audio signal based on the estimated signal and the covariance. According to the method provided by the invention, the sound source probability density weight obtained by model training is combined with the independent vector analysis to separate the mixed audio signal, the robustness is high, the method can adapt to a scene with a low signal-to-noise ratio, the obtained covariance matrix is used as the input of the estimation signal, the target voice signal is extracted, the problem of channel aliasing is solved, and the accuracy of the channel aliasing is improved. And the method has a certain error-tolerant rate for target angle deviation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of vehicle networking, and in particular, to a method, device, electronic device, chip and medium for extracting voice signals. Background Art

[0002] Voice interaction has become the main control method in the intelligent cockpit. Usually, one or more microphone arrays are used in the intelligent cockpit for sound pickup, and the main sound area is extracted according to the direction of arrival of the array signal beam. However, in vehicle-mounted voice separation, there is noise interference and the distribution of seat positions in the vehicle has particularity. Therefore, extracting the target signal in a specific direction (such as the driver's seat direction) is of great significance for intelligent voice interaction. Summary of the Invention

[0003] The present disclosure provides a method, device, electronic device, chip and medium for extracting voice signals. By combining the source probability density obtained through model training with the independent vector analysis and the estimated signal algorithm, the target signal that meets the preset target angle is extracted, solving the problem of channel aliasing, having a certain tolerance rate for target angle deviation, high robustness, and being able to adapt to scenarios with a relatively low signal-to-noise ratio.

[0004] A first aspect embodiment of the present disclosure proposes a method for extracting voice signals. The method includes: obtaining a mixed audio signal, where the mixed audio signal includes multiple source signals, and the multiple source signals correspond to multiple channels of a microphone array; based on the mixed signal, using a trained source probability density model to obtain the source probability density weights of the multiple source signals, where the source probability density weights are the probabilities of obtaining each source signal from the mixed audio signal; based on the source probability density weights, performing independent vector analysis on the mixed audio signal to obtain the estimated signal of each source signal and the covariance of the amplitude of each source signal; based on the estimated signal and the covariance, determining the target voice signal that meets the preset target angle from the mixed audio signal.

[0005] In some embodiments of the present disclosure, performing independent vector analysis on the mixed audio signal based on the source probability density weights to obtain the estimated signal of each source signal and the covariance of the amplitude of each source signal includes: performing a short-time Fourier transform on the mixed audio signal to obtain the frequency-domain signal of the mixed audio signal; performing independent vector analysis on the frequency-domain signal to obtain the estimated signal of each source signal and the covariance of the amplitude of each source signal.

[0006] In some embodiments of the present disclosure, determining a target voice signal that conforms to a preset target angle from a mixed audio signal based on an estimated signal and covariance includes: estimating the covariance to obtain the incident angle of each estimated signal; determining the channel where the incident angle that conforms to the first preset condition is located as the target channel, where the first preset condition is that the difference between the incident angle and the preset target angle is less than or equal to the first threshold; determining the estimated signal of the target channel as the target signal; and performing a short-time inverse Fourier transform on the target signal to obtain the target voice signal.

[0007] In some embodiments of the present disclosure, estimating the covariance to obtain the incident angle of each estimated signal includes: performing fixed-frame sampling on the covariance to obtain multiple fixed-frame covariances of each estimated signal; performing eigenvalue decomposition on the multiple fixed-frame covariances to obtain the eigenvalues and eigenvectors of the multiple fixed-frame covariances; extracting the eigenvector corresponding to the maximum eigenvalue of the multiple fixed-frame covariances as the signal subspace of the estimated signal; and determining the direction of arrival of the signal subspace as the incident angle of the estimated signal.

[0008] In some embodiments of the present disclosure, determining the channel where the incident angle that conforms to the first preset condition is located as the target channel includes: when there are multiple incident angles of the estimated signals that conform to the first preset condition, comparing the voiceprint features of the multiple estimated signals; and determining the channel corresponding to the estimated signal whose voiceprint feature conforms to the second preset condition as the target channel.

[0009] In some embodiments of the present disclosure, the second preset condition includes: the voiceprint feature matches the target voiceprint feature.

[0010] In some embodiments of the present disclosure, the method further includes: using a simulated voice signal to train a sound source probability density model to obtain a trained sound source probability density model.

[0011] An embodiment of the second aspect of the present disclosure provides a voice signal extraction device, including: a separation module and an extraction module. The separation module is configured to obtain a mixed audio signal, where the mixed audio signal includes multiple sound source signals, and the multiple sound source signals correspond to multiple channels of a microphone array; based on the mixed audio signal, use the trained sound source probability density model to obtain the sound source probability density weights of the multiple sound source signals, where the sound source probability density weights are the probabilities of obtaining each sound source signal from the mixed audio signal; and based on the sound source probability density weights, perform independent vector analysis on the mixed audio signal to obtain the estimated signal of each sound source signal and the covariance of the amplitude of each sound source signal. The extraction module is configured to determine a target voice signal that conforms to a preset target angle from the mixed audio signal based on the estimated signal and the covariance.

[0012] A third aspect embodiment of the present disclosure provides an electronic device, including: a processor and a memory for storing a computer program that can run on the processor. When the processor is used to run the computer program, it executes the method described in any one of the embodiments of the first aspect of the present disclosure.

[0013] A fourth aspect embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause a computer to execute the method described in any one of the embodiments of the first aspect of the present disclosure.

[0014] A fifth aspect embodiment of the present disclosure provides a chip, including at least one processor and a communication interface. The communication interface is used to receive signals input to the chip or signals output from the above chip. The processor communicates with the communication interface and implements the method described in any one of the embodiments of the first aspect of the present disclosure through logic circuits or by executing code instructions.

[0015] A sixth aspect embodiment of the present disclosure provides a vehicle, including the voice signal extraction device described in the second aspect embodiment or the electronic device described in the third aspect embodiment.

[0016] In summary, the voice signal extraction method, device, electronic device, chip and medium provided by the present disclosure include: obtaining a mixed audio signal, where the mixed audio signal includes multiple sound source signals, and the multiple sound source signals correspond to multiple channels of a microphone array; based on the mixed signal, using a trained sound source probability density model to obtain the sound source probability density weights of the multiple sound source signals, where the sound source probability density weights are the probabilities of obtaining each sound source signal from the mixed audio signal; based on the sound source probability density weights, performing independent vector analysis on the mixed audio signal to obtain the estimated signal of each sound source signal and the covariance of the amplitude of each sound source signal; based on the estimated signal and the covariance, determining a target voice signal that meets a preset target angle from the mixed audio signal.

[0017] The method provided by the present disclosure can separate the mixed voice signal by combining the sound source probability density weights obtained through model training with independent vector analysis, has high robustness, can adapt to scenarios with relatively low signal-to-noise ratios, and uses the obtained covariance matrix as the input of the estimated signal algorithm, thereby extracting the target signal, solving the problem of channel aliasing, and having a certain tolerance for target angle offset.

[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation to the present disclosure.

[0020] Figure 1 It is a scene graph provided for the embodiments of the present disclosure;

[0021] Figure 2 It is a schematic flowchart of a method for extracting a voice signal proposed in the embodiments of the present disclosure;

[0022] Figure 3 It is a flowchart of a method for determining a target voice signal proposed in the embodiments of the present disclosure;

[0023] Figure 4 It is a schematic diagram for screening a target channel according to a preset target angle provided by the embodiments of the present disclosure;

[0024] Figure 5 It is a schematic diagram of the incident angles of two sound sources and the preset target angle in the actual road test provided by the embodiments of the present disclosure;

[0025] Figure 6 It is a flowchart of a method for estimating covariance proposed in the embodiments of the present disclosure;

[0026] Figure 7 It is a flowchart of a method for determining a target channel proposed in the embodiments of the present disclosure;

[0027] Figure 8A It is a schematic flowchart of a method for extracting a voice signal provided by the embodiments of the present disclosure;

[0028] Figure 8B It is a general flowchart of the front-end voice processing provided by the embodiments of the present disclosure;

[0029] Figure 8C It is a distribution diagram of a four-microphone array provided by the embodiments of the present disclosure;

[0030] Figure 8D It is a schematic flowchart of signal extraction using the ESPRIT algorithm provided by the embodiments of the present disclosure;

[0031] Figure 9 It is a schematic structural diagram of a voice signal extraction device proposed in the embodiments of the present disclosure;

[0032] Figure 10 It is a schematic structural diagram of an electronic device provided by the embodiments of the present disclosure;

[0033] Figure 11 It is a schematic structural diagram of a chip provided by the embodiments of the present disclosure. Detailed implementation manners

[0034] Embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present disclosure, and should not be construed as a limitation of the present disclosure.

[0035] As one of the most common and simplest interaction methods for humans, voice interaction has become the main control method in intelligent cockpits. Through the driver's voice, the interaction task requirements of the driver during driving can be quickly, accurately, and intelligently realized, which can reduce unnecessary, inconvenient, and unsafe hand movements during driving and achieve operations such as navigation, music playback, and window opening and closing. Due to the complex in-vehicle acoustic environment, except for the relatively fixed position of the driver, the source directions and composition components of other interfering voices are not determined. Therefore, using real-time voice extraction instructions to extract the driver's voice from the mixed voice signal has become a crucial step in the in-vehicle voice intelligent interaction process. Currently, in intelligent cockpits, one or more microphone arrays are usually used for sound pickup, and then the main sound area is extracted according to the direction of arrival of the array signal beam. However, due to the presence of wind noise, tire noise, engine noise, and interference from multiple speakers in in-vehicle voice separation, the voice directly collected by the microphone array often has a low signal-to-noise ratio and intelligibility and cannot be correctly recognized. Therefore, first, front-end voice signal processing technologies (such as blind source separation, echo cancellation, noise reduction, etc.) need to be used to suppress interference and enhance the target voice, and then the target voice signal is extracted and input into the recognition system for text conversion and semantic understanding, and finally, the corresponding instructions are executed on the terminal. Due to the special distribution of seat positions in the car, that is, the functions of the driver's seat, co-driver's seat, and rear seats are different, extracting the voice signal of the target in a specific direction (such as the driver's direction) is of great significance for intelligent voice interaction.

[0036] Currently, there is a technical solution that estimates the source direction through the phase difference and then separates the speakers; the second technical solution uses a directional microphone for directional sound pickup, and the overall area inside the vehicle is covered by increasing the number of directional microphones. After receiving the audio at the directional position, it is further processed, and then the feedback audio information is transmitted to the position where the directional sound source is located; the third technical solution uses a sound pickup device to pick up the mixed voice, directly extracts the voiceprint characteristics as the input for intelligent interaction, and prompts for re-recording when the voice signal is unclear or does not match the stored voiceprint characteristics.

[0037] Among the above three related technical solutions, Solution 1 assumes that the speaker is in a fixed position, and the frequency points are extracted based on the phase difference twice. However, the speaker's position often shifts. Especially for the speakers in the back row, they are often at the edge of the sound area, and there is also the situation of mutual leakage between sound areas in the in-vehicle voice environment. The signal-to-noise ratio of the sound signals extracted by sound area is relatively low, and the algorithm robustness and separation performance are poor. Solution 2 selects a more expensive directional microphone, which has a higher cost. In addition, when there are multiple signal sources at the same time, the signal sources in other directions will interfere with the target signal source, and the signal-to-noise ratio obtained by direct sound collection is relatively low. Solution 3 pre-stores the voiceprint characteristics of the target speaker as matching information, and cannot be normally executed when the voiceprints do not match. Its main information is targeted at the speaker rather than the position of the speech. Therefore, it is not applicable when the position of the driver's seat in the vehicle is fixed but the speaker in the driver's seat is not determined. In addition, in a low signal-to-noise ratio scenario, the voiceprint characteristics are not obvious and it is difficult to execute the instructions.

[0038] In summary, to solve the technical problems in the related technologies, the embodiments of the present disclosure provide a method for extracting voice signals. By obtaining a mixed audio signal, the mixed audio signal includes multiple sound source signals, and the multiple sound source signals correspond to multiple channels of a microphone array; based on the mixed signal, using a trained sound source probability density model, the sound source probability density weights of the multiple sound source signals are obtained, and the sound source probability density weight is the probability of obtaining each sound source signal from the mixed audio signal; based on the sound source probability density weights, independent vector analysis is performed on the mixed audio signal to obtain the estimated signal of each sound source signal and the covariance of the amplitude of each sound source signal; based on the estimated signal and the covariance, the target voice signal that meets the preset target angle is determined from the mixed audio signal. The purpose of this technical solution is to fuse the sound source probability density obtained through model training into the independent vector analysis algorithm to obtain the estimated signal and the covariance matrix, and obtain the direction of arrival of each estimated signal through the covariance matrix, and determine the estimated signal in the channel that meets the preset target angle as the target signal, which solves the problem of channel aliasing, has a certain tolerance for target angle offset, and has high robustness and can adapt to scenarios with a relatively low signal-to-noise ratio.

[0039] The method for extracting in-vehicle voice signals provided by the present application will be introduced in detail below with reference to the accompanying drawings.

[0040] Figure 1 For the application scenario diagram of the present disclosure, as Figure 1 shown, the embodiments of the present disclosure are applied to a four-array voice interaction scenario. The voice signal can be divided into a "target signal" and an "interference signal". The angle of the target signal in the present disclosure is known. According to the target angle, a clean, accurate, and signal closest to the target angle is extracted from the mixed voice signal. The interference signal refers to other sound signals except the target signal, including voice signals and noises at other angles.

[0041] Figure 2 The flowchart of a method for extracting a voice signal proposed by an embodiment of the present disclosure is shown as follows. Figure 2 As shown, this method can be executed by a terminal, specifically, it can be a vehicle. This method may include the following steps.

[0042] Step 201: Obtain a mixed audio signal.

[0043] In some embodiments, the mixed audio signal includes multiple sound source signals, and the multiple sound source signals correspond to multiple channels.

[0044] Exemplarily, in a four-seat vehicle, by setting four microphones at equal intervals (spacing is △) as the sound pickup devices, each microphone corresponds to a channel and a corresponding sound source signal.

[0045] In some embodiments, the number of microphones is equal to the number of channels, and the number of sound sources is less than or equal to the number of microphones. Exemplarily, if the number of sound sources is N and the number of microphones is K, then the number of channels is K, and K≥N.

[0046] Exemplarily, assuming that the pure original signal is s, the mixed signal received by the microphone array is x, and the separated estimated signal is y. The embodiment of the present disclosure is to obtain the separated estimated signal y through the obtained mixed signal x, and screen the target signal that meets the preset conditions with the preset target angle from the estimated signal.

[0047] Step 202: Based on the mixed audio signal, use the trained sound source probability density model to obtain the sound source probability density weights of the multiple sound source signals.

[0048] In some embodiments, the sound source probability density model is trained using simulated voice signals.

[0049] In some embodiments, the sound source probability density weights can be obtained through other models. Exemplarily, it can be a Laplace model, and the present disclosure is not limited thereto.

[0050] Step 203: Based on the sound source probability density weights, perform independent vector analysis on the mixed audio signal to obtain the estimated signal of each sound source signal and the covariance of the amplitude of each sound source signal.

[0051] In some embodiments, before the independent vector analysis (Independent Vector Analysis, IVA), it further includes performing a short-time Fourier transform on the mixed audio signal to obtain a frequency-domain signal as the input of the independent vector analysis algorithm.

[0052] In some embodiments, after performing the short-time Fourier transform on the mixed audio signal, interference signals such as noise can be removed to obtain a voice signal in the frequency domain.

[0053] In some embodiments, independent vector analysis is performed on the mixed audio signal to separate the mixed signal and obtain an estimated signal of the speech signal of each sound source.

[0054] In some embodiments, based on the sound source probability density weight weight, an auxiliary variable in the independent vector analysis algorithm can be obtained. Taking the mixed audio signal as the input of the algorithm, an estimated signal and a covariance matrix can be obtained.

[0055] Exemplarily, in the scenario of a four-microphone array, after the mixed audio signal undergoes short-time Fourier transform, the sound source signal can be expressed as s(t,f) = [s1(t,f), s2(t,f), s3(t,f), s4(t,f)] T ; the mixed audio signal acquired by the microphone is x(t,f) = [x1(t,f), x2(t,f), x3(t,f), x4(t,f)] T ; the estimated signal is y(t,f) = [y1(t,f), y2(t,f), y3(t,f), y4(t,f)] T ; where the frame label t ∈ {1, 2,..., T}, the frequency bin label f ∈ {1, 2,..., F}, independent vector analysis is performed on x to obtain the estimated source signal y, and ρ(r) is an auxiliary variable obtained from the weight weight: Among them, f is the number of frequency bands, ξ is a very small value to avoid division by zero, k is the number of channels, t is the time frame, H is the complex conjugate transpose symbol, and the covariance matrix can be obtained based on the following formula

[0056] Among them, the mixing matrix and the separation matrix are A(f) and W(f) respectively. Then the mixed audio signal and the estimated signal can be expressed as: x(t,f) = A(f)s(t,f), y(t,f) = W(f)x(t,f), where W(f) = [W1(f), W2(f), W3(f), W4(f)] T , then the separation matrix W(f) can be iteratively updated in the following manner:

[0057] w(k,f) ← (W(f)V(k,f)) -1 e n ;

[0058] Among them, e n is a K×1 unit vector, the nth element is 1, and the other elements are 0.

[0059] The estimated signal can be obtained through y(k,t,f) = W(k,f) × x(k,t,f).

[0060] In some embodiments, independent vector analysis estimates the mixing coefficients and independent components by minimizing mutual information. All frequency components of each source signal are combined as a multivariate probability density, which not only preserves the internal dependencies between the frequencies of each source signal but also maximizes the independence between different source signals to avoid the permutation problem.

[0061] In the above embodiments, independent vector analysis is performed on the mixed audio signal using the source probability density weight, and the obtained estimated signal is more suitable for the in-vehicle voice scenario and has a better separation effect.

[0062] Step 204: Determine the target voice signal that meets the preset target angle based on the estimated signal and covariance.

[0063] In some embodiments, the estimated signal and covariance of each source signal obtained in step 103 are used as the input of the Estimation of Signal Parameters Via Rotational Invariance Techniques (ESPRIT) algorithm. The main signal source direction of each channel is deduced based on the covariance matrix, the channel where the target is located is determined according to the target direction, and the estimated signal of this channel is extracted as the target signal.

[0064] In some embodiments, the ESPRIT algorithm uses the rotational invariance between sub-arrays to achieve the Direction of Arrival (DOA) estimation of the array.

[0065] In some embodiments, the direction of arrival of each channel obtained by the ESPRIT algorithm is the incident angle of the estimated signal of each channel.

[0066] In some embodiments, the incident angle of the estimated signal of each obtained channel is compared with the target angle, the eligible incident angles are screened, and the estimated signal of the channel where the incident angle is located is used as the target voice signal.

[0067] In summary, in the above embodiments of the present application, a mixed audio signal is obtained; based on the mixed audio signal, using a trained sound source probability density model, the sound source probability density weights of multiple sound source signals are obtained; based on the sound source probability density weights, independent vector analysis is performed on the mixed audio signal to obtain the covariance of the estimated signal of each sound source signal and the amplitude of each sound source signal; based on the estimated signal and the covariance, a target voice signal that meets a preset target angle is determined. The independent vector analysis is performed on the mixed audio signal using the sound source probability density weights obtained through model training, and the covariance matrix and the estimated signal are combined with the preset target angle to extract the target voice signal, solving the problem of channel aliasing, having a certain tolerance for target angle offset, high robustness, and being able to adapt to scenarios with relatively low signal-to-noise ratios.

[0068] Figure 3 is a flowchart of a method for determining a target voice signal proposed in an embodiment of the present disclosure. Based on Figure 2 the embodiments shown, Figure 3 is a further description of Figure 2 step 204. Figure 3 The embodiments shown may include the following steps.

[0069] Step 301, estimate the covariance to obtain the incident angle of each estimated signal.

[0070] In some embodiments, the covariance of each channel where the separated estimated signal is located is estimated using the Estimation of Signal Parameters Via Rotational Invariance Techniques (ESPRIT) algorithm to obtain the direction of arrival, i.e., the incident angle, of each estimated signal.

[0071] In some embodiments, the estimation of the covariance can be performed through the ESPRIT algorithm. According to the maximum eigenvalue of the covariance matrix of each channel, the most main sound source direction of the channel where the estimated signal is located is judged, and then the incident angle of each estimated signal is obtained.

[0072] Exemplarily, after sampling the covariance matrix of each channel according to a fixed frame, eigenvalue decomposition is performed, and the eigenvector corresponding to the maximum eigenvalue is taken as the signal subspace of the estimated signal of the channel. The direction of arrival of the signal subspace is calculated as the incident angle of the channel.

[0073] Step 302, determine the channel where the incident angle that meets the first preset condition is located as the target channel.

[0074] In some embodiments, the first preset condition may be that the difference between the incident angle and the preset target angle is less than or equal to the first threshold.

[0075] Exemplarily, as Figure 4 shown, θ_target is a preset target angle. The channel where the line segment with the smallest included angle with the line segment of the preset target angle is located can be determined as the target channel.

[0076] Exemplarily, the first threshold can be an infinitesimal angle or the minimum difference value is taken according to the comparison result. As Figure 5 shown, a is the preset target angle, b is the incident angle of the first voice signal, c is the incident angle of the second voice signal, and the difference between b and a is much smaller than the difference between c and a. Therefore, the channel where b is located is determined as the target channel.

[0077] Exemplarily, the preset target angle is 45°, b is 40°, c is 60°. Then the difference between b and the preset target angle is 5°, the difference between c and the preset target angle is 15°, and the first threshold is 15°. Therefore, the second channel where b is located is determined as the target channel.

[0078] Step 303: Determine the estimated signal of the target channel as the target signal.

[0079] In some embodiments, based on the target channel determined in the foregoing step 302, determining the estimated signal of the target channel as the target signal can obtain an estimated signal of the sound source signal closer to the target angle.

[0080] In some embodiments, for example, y(k = 1) is determined as the target signal, that is, the estimated signal of the 1st channel is determined as the target signal.

[0081] Step 304: Perform a short-time inverse Fourier transform on the target signal to obtain the target voice signal.

[0082] In some embodiments, the estimated signal of the determined target channel is a signal in the frequency domain. Performing a short-time inverse Fourier transform on it can reconstruct the corresponding time-domain signal, that is, the target voice signal that meets the preset target angle can be obtained.

[0083] In the embodiments of the present disclosure, the covariance is used as the input of the ESPRIT algorithm. By the rotational invariance of the subarray and the array relationship, the direction of the main signal source of each channel is deduced, so as to determine the incident angle of the estimated signal of each channel. Since the preset target angle is known, comparing the incident angles can achieve the purpose of screening out the channel where the target is located.

[0084] Figure 6 The flowchart of the method for estimating the covariance proposed in the embodiments of the present disclosure. Based on Figure 3 the embodiments shown, Figure 6 For Figure 3 step 301 in the embodiments is further described. As Figure 6 shown, it includes the following steps:

[0085] Step 601: Perform fixed-frame sampling on the covariance to obtain multiple fixed-frame covariances of each estimated signal.

[0086] In some embodiments, perform fixed-frame sampling on the covariance matrix of each estimated signal, extract N τ frames as the guiding sampling frames, mark the sampling frames as τ, and obtain multiple fixed-frame covariances V(k,τ) of each estimated signal.

[0087] Step 602: Perform eigen-decomposition on the multiple fixed-frame covariances to obtain the eigenvalues and eigenvectors of the multiple fixed-frame covariances.

[0088] In some embodiments, perform eigen-decomposition on the multiple fixed-frame covariances of each estimated signal obtained in Step 601 to obtain its eigenvalues and eigenvectors.

[0089] Exemplarily, perform eigen-decomposition on V(k,τ), λ Ev (k,τ), Ev(k,τ) = eig(V(k,τ)) to obtain the eigenvalue λ Ev and the eigenvector Ev.

[0090] Step 603: Extract the eigenvector corresponding to the maximum eigenvalue of the multiple fixed-frame covariances as the signal subspace of the estimated signal.

[0091] In some embodiments, in each channel where each estimated signal is located, take the eigenvector corresponding to the maximum eigenvalue among the multiple fixed-frame covariances as the signal subspace of the corresponding channel. Exemplarily, take 1 Ev of the maximum eigenvalues in λ(k), and the corresponding eigenvector is the signal subspace Es(k) corresponding to the estimated signal. Regard it as the most main signal in the k-th channel, then a signal subspace is determined for each channel.

[0092] Step 604: Determine the direction of arrival of the signal subspace as the incident angle of the estimated signal.

[0093] In some embodiments, by calculating the signal subspace of each channel, the direction of arrival of each speech signal can be obtained, and take the direction of arrival of each estimated signal as its incident angle.

[0094] Exemplarily, split the signal subspace Es of each channel into two sub-array spaces Ea and Eb, where Ea takes the first K - 1 rows of Es, and Eb takes the last K - 1 rows of Es: The sub-array spaces Ea and Eb are concatenated by columns into Eab, and perform singular value decomposition on Eab to obtain the eigenmatrix E: E, ∼, ∼ = SVD(Eab H Eab), sort according to the singular values from large to small, and split E into four small matrices Take the right singular vectors corresponding to the N smallest singular values According to φ(k,τ) = -F0(k,τ) / F1 -1 (k,τ), Obtain the delay phase φ(k), Combined with the microphone array structure, solve to obtain the direction of arrival θ(k): θ(k) is the incident angle of the most dominant signal source in the k-th channel to the array. Here, the symbol ~ represents the irrelevant parameters omitted in the solution result. H is the complex conjugate transpose symbol, λ is the wavelength, and Δ represents the spacing between microphones.

[0095] In the above embodiment, by using the covariance as the input of the ESPRIT algorithm and solving to obtain the incident angle of the most dominant sound source signal in each channel, the incident angle of each channel can be compared with the preset target angle to determine the target channel, which can solve the problem of channel aliasing and has a certain tolerance rate when there is an offset in the target angle.

[0096] Figure 7 This is the flowchart of the method for determining the target channel proposed in the embodiments of the present disclosure. Based on Figure 3 the embodiments shown Figure 7 to Figure 3 further describe step 302 in the embodiment. As Figure 7 shown, it includes the following steps:

[0097] Step 701, when the incident angles of multiple estimated signals meet the first preset condition, compare the voiceprint features of the multiple estimated signals.

[0098] In some embodiments, when the differences between the incident angles of multiple estimated signals and the preset target angle are all less than or equal to the first threshold, that is, the channels where the multiple estimated signals are located may all be determined as the target channels, and it is necessary to extract the voiceprint features of the channels where the estimated signals are located for comparison to obtain the target channel closest to the target sound source.

[0099] Exemplarily, as Figure 5 shown, assuming that the included angle between the incident angle d of the third channel and the preset target angle is 5°, then it is necessary to extract the voiceprint features of the first channel where b is located and the third channel where d is located for comparison.

[0100] Step 702, determine the channel corresponding to the estimated signal whose voiceprint feature meets the second preset condition as the target channel.

[0101] In some embodiments, the second preset condition may be that the voiceprint feature matches the target voiceprint feature.

[0102] Exemplarily, the target voiceprint feature is the voiceprint feature of a preset target angle obtained in advance. Compare the voiceprint feature of the estimated signal of the third channel where d is located with the voiceprint feature of the estimated signal of the first channel where b is located. If the voiceprint feature of the third channel where d is located matches the target voiceprint feature, then the third channel where d is located is determined as the target channel.

[0103] In the above embodiment, using the matching of voiceprint features as an auxiliary verification means for preset target angle matching can avoid the influence on the final result when the interfering signal source is too close to the target signal source.

[0104] The embodiments of the present disclosure have the following beneficial effects:

[0105] 1. The independent vector analysis algorithm combines the source probability density weights obtained through model training to separate audio signals, which is more suitable for in-vehicle voice scenarios, resulting in better in-vehicle voice separation effect, high robustness, and the ability to adapt to scenarios with relatively low signal-to-noise ratio.

[0106] 2. Using the covariance of each separated channel as the input of the ESPRIT algorithm can extract the voice signal in the target direction, solve the problem of channel aliasing, and has a certain tolerance rate when there is an offset in the target angle.

[0107] Figure 8A It is a schematic flowchart of a voice signal extraction method of the present disclosure. Figure 8B It is a general flowchart of front-end voice processing, including STFT noise reduction preprocessing, IVA separation, ESPRIT localization, and ISTFT postprocessing.

[0108] As Figure 8A shown, the voice signal extraction method includes the following steps:

[0109] Step 1: Simulate in-vehicle data for DNN pre-training, and obtain the weight weight of the source probability density based on the mixed audio signal.

[0110] In some embodiments, the source probability density model can be trained based on the simulated voice signal. Input the mixed audio signal into the source probability density model to obtain the source probability density weight.

[0111] Step 2: The four-microphone array acquires the mixed audio signal x. After short-time Fourier transform, it is used as the input of the real-time IVA algorithm, and the estimated signal and covariance matrix are output.

[0112] In some embodiments, the mixed audio signal is used as the input of the IVA algorithm. After denoising by short-time Fourier transform, it is converted into a frequency-domain signal, and the output is the estimated signal and covariance matrix in the separated frequency domain.

[0113] Step 3: Extract the covariance matrix V once every fixed number of frames.

[0114] In some embodiments, sampling is performed at fixed frames to obtain multiple fixed-frame covariances V of each estimated signal.

[0115] Step 4: Perform eigen-decomposition on the multiple fixed-frame covariances V for each channel, take the eigenvector corresponding to the 1 largest eigenvalue to obtain the main signal subspace Es, calculate the eigenvalues φ corresponding to each channel according to the TLS-ESPRIT method, take the average of φ for all sampling frames, perform eigen-decomposition on the average eigenvalue, and use the eigenvalues to form a matrix of delay phases Combined with the array structure, solve to obtain the direction of arrival θ(k).

[0116] In some embodiments, taking the covariance matrix as the input of the ESPRIT algorithm can obtain the incident angle of each channel corresponding to the estimated signal.

[0117] Step 5: Sort θ(k), and select the channel closest to the target direction as the target channel. If there is an uncertain situation, enter the voiceprint feature matching to select the target signal.

[0118] In some embodiments, by comparing the incident angle of each estimated signal with a preset target angle, obtaining the incident angle that satisfies the first preset condition, determining the channel where it is located as the target channel, and taking the estimated signal of the target channel as the target signal, performing a short-time inverse Fourier transform on the target signal can obtain the target voice signal that satisfies the preset target angle.

[0119] In some embodiments, the uncertain situation can be that there are multiple incident angles that satisfy the first preset condition. The voiceprint feature of the estimated signal of the channel where the satisfied incident angle is located can be compared with the target voiceprint feature, select the channel where the matching voiceprint feature is located as the target channel, take the estimated signal of the target channel as the target signal, and similarly, perform a short-time inverse Fourier transform on the target signal to obtain the target voice signal that satisfies the preset target angle.

[0120] The principle of voice separation by the IVA algorithm is introduced as follows:

[0121] Place a set of equidistant four-microphone (spacing is Δ) microphone arrays in the front row of the vehicle to pick up sound, separate the voices of four-channel speakers through real-time IVA, and extract the target channel through the array method by the target angle.

[0122] Specifically, assume the number of microphones is K, where K = 4 in this embodiment, and the number of sound sources is N, where K is greater than or equal to N and K is greater than or equal to 2. The pure original signal is s, the mixed signal received by the microphone array is x, and the separated estimated signal is y. After short-time Fourier transform, the four sound source signals, microphone observation signals, and estimated source signals in the cockpit can be respectively expressed as:

[0123] (1) s(t,f) = [s1(t,f), s2(t,f), s3(t,f), s4(t,f)] T

[0124] (2) x(t,f) = [x1(t,f), x2(t,f), x3(t,f), x4(t,f)] T

[0125] (3) y(t,f) = [y1(t,f), y2(t,f), y3(t,f), y4(t,f)] T

[0126] Among them, the frame label t ∈ {1, 2,..., T}, and the frequency point label f ∈ {1, 2,..., F}. Assume the mixing matrix and separation matrix are A(f) and W(f) respectively, then the models of the mixing system and separation system can be respectively expressed as:

[0127] (4) x(t,f) = A(f)s(t,f)

[0128] (5) y(t,f) = W(f)x(t,f) where W(f) = [W1(f), W2(f), W3(f), W4(f)] T

[0129] IVA believes that each sound source is independent of each other. By minimizing the mutual information, it estimates the mixing coefficients and independent components. IVA combines all frequency components of each source signal as a multivariate probability density, which not only retains the internal dependence between the frequencies of each source signal but also maximizes the independence between different sound source signals to avoid the permutation problem. Therefore, it theoretically ensures that the separated signals are consistent across the entire frequency band. The following is the objective function of IVA:

[0130] (6)

[0131] where, y(k,t) = [y(k,t,1), y(k,t,2),..., y(k,t,f),..., y(k,t,F)] T , 1 ≤ t ≤ T, 1 ≤ f ≤ F.

[0132] Among them, W refers to the demixing matrix, i.e., the separation matrix, and the estimated speech signal y is obtained through equation (5);

[0133] p(y) refers to the probability density function (PDF) of the speech signal y in the frequency domain. In this embodiment, the PDF is time-varying and different among various sound sources.

[0134] detW(f) is to find the determinant of W at frequency band f;

[0135] const is a constant that does not change and is not considered in the iterative descent process.

[0136] In the real-time version of IVA, V(k,t,f) should be calculated frame by frame for each time frame t. An adaptive algorithm can be used to recursively obtain V(k,t,f). V(k,t,f) is an auxiliary variable and is part of the update rule of W(f) in the IVA method. V(k,t,f) in the real-time version of AuxIVA can be understood as the covariance matrix of the weighted mixed signal x that is continuously transformed in the time domain.

[0137] W(f) is iteratively updated according to the following rule: w(k,f)←(W(f)V(k,f)) -1 e n ;

[0138]

[0139] where e n is a K×1 unit vector, the nth element is 1, and other elements are 0.

[0140]

[0141] where α∈[0,1) is a forgetting factor and ξ is a very small value to avoid division by zero.

[0142] Assume that the sound source signal follows a non-stationary Gaussian distribution, and define σ 2 (k,t) as the time-varying variance, then there is:

[0143] (8)

[0144] where σ 2 (k,t) refers to the time-varying variance of the estimated signal y received at the kth channel at time t, and ||y(k,t)||2 refers to the L2 norm of the y signal.

[0145] The principle of the ESPRIT algorithm is introduced as follows:

[0146] Such as Figure 8CAs shown in the figure, assume that there are two identical sub-arrays in the array itself. Two identical sub-arrays are obtained through certain transformations. There is a fixed spacing between adjacent sub-arrays, and this spacing reflects a fixed relationship between adjacent sub-arrays, that is, the rotational invariance between sub-arrays. The ESPRIT algorithm precisely utilizes the rotational invariance between sub-arrays to achieve DOA estimation of the array. Assume that the sub-array 1 is within the red box, and the received acoustic signal is x1. The sub-array 2 is within the blue box, and the received acoustic signal is x2. is the delay phase between the two sub-arrays. Combining with formula (4), we have:

[0147] (9)

[0148] (10)

[0149] (11)λ = c x / f x , where c x is the speed of sound of the voice signal x, and f x is the frequency of the voice signal x.

[0150] Take the covariance matrix Vz of z:

[0151] Perform SVD decomposition on Vz: Vz = U s ΛU s H

[0152] is consistent with the subspace of U s , then there must exist a unique non-singular matrix Q such that Then U s can be decomposed as: Then

[0153] Therefore, by obtaining φ, the delay phase can be obtained

[0154]

[0155] According to the Total Least Square (TLS), the above formula can be transformed into a generalized eigenvalue problem with a smaller dimension, and then the calculation of φ can be realized.

[0156] The process of using the IVA algorithm for voice separation in this embodiment is as follows:

[0157] 1. According to the weight weight of the probability density of the sound source obtained during the pre-training process, the updated auxiliary variable can be obtained: where the initial value of w(k, f) is 0, F is the number of frequency bands, and H is the symbol of complex conjugate transpose.

[0158] 2. The separation matrix is updated once per iteration, i.e.:

[0159] w(k,f)←(W(f)V(k,f)) -1 e n

[0160]

[0161] where e n is a K×1 unit vector, with the nth element being 1 and the other elements being 0.

[0162] 3. Based on the above, the covariance matrix is obtained

[0163] 4. The estimated source signal y(k,t,f) is obtained using formula (5), y(k,t,f) = W(k,f) × x(k,t,f).

[0164] The covariance matrix and the estimated source signal obtained above through the IVA algorithm are used as the input of the ESPRIT algorithm.

[0165] Figure 8D For the signal extraction process schematic diagram of the ESPRIT algorithm, as Figure 8D shown, the signal extraction process includes the following steps:

[0166] 1. Extract N τ frames as the steering sampling frames, with the sampling frames labeled as τ, and perform fixed-frame sampling on the covariance matrix.

[0167] 2. Perform eigenvalue decomposition on the covariance matrix V(k,t,f) separated by the IVA algorithm to obtain its eigenvalues λ Ev and eigenvectors Ev: λ Ev (k,τ), Ev(k,τ) = eig(V(k,τ)).

[0168] 3. Take 1 largest eigenvalue from λ Ev (k), and the corresponding eigenvector is the corresponding signal subspace Es(k), which is regarded as the main signal of the k channel.

[0169] 4. Split the signal subspace Es into two subarray spaces Ea and Eb, where Ea takes the first K - 1 rows of Es and Eb takes the last K - 1 rows of Es: The subarray spaces Ea and Eb are concatenated by columns into Eab, and singular value decomposition is performed on Eab to obtain the eigenmatrix E:

[0170] E = SVD(Eab H Eab)

[0171] Sorted in descending order of singular values, and E is split into four small matrices:

[0172] Take the right singular vectors corresponding to the N smallest singular values:

[0173] Then the eigenvalue of the k-th channel is φ(k,τ) = -F0(k,τ) / F1 -1 (k,τ). The average value of the eigenvalues of all frames of the k-th channel is Perform eigenvalue decomposition on the average value to obtain:

[0174] 5. Combining with the microphone array structure, solve to obtain the direction of arrival θ(k), and θ(k) is the incident angle of the most main signal source of the k-th channel to the array:

[0175] 6. Determine the channel where the target voice signal is located by comparing θ(k).

[0176] Figure 9 It is a schematic structural diagram of a voice signal extraction device 900 according to an embodiment of the present disclosure. As Figure 9 shown, the device includes:

[0177] A separation module 901, configured to obtain a mixed audio signal, where the mixed audio signal includes multiple sound source signals, and the multiple sound source signals correspond to multiple channels of a microphone array; based on the mixed audio signal, use a trained sound source probability density model to obtain the sound source probability density weights of the multiple sound source signals, and the sound source probability density weights are the probabilities of obtaining each sound source signal from the mixed audio signal; based on the sound source probability density weights, perform independent vector analysis on the mixed audio signal to obtain the estimated signal of each sound source signal and the covariance of the amplitude of each sound source signal;

[0178] An extraction module 902, configured to determine a target voice signal that meets a preset target angle from the mixed audio signal based on the estimated signal and the covariance.

[0179] In some embodiments, the separation module is configured to: perform short-time Fourier transform on the mixed audio signal to obtain the frequency-domain signal of the mixed audio signal; perform independent vector analysis on the frequency-domain signal to obtain the estimated signal of each sound source signal and the covariance of the amplitude of each sound source signal.

[0180] In some embodiments, the extraction module is further configured to: estimate the covariance to obtain the incident angle of each estimated signal; determine the channel where the incident angle that meets the first preset condition is located as the target channel, where the first preset condition is that the difference between the incident angle and the preset target angle is less than or equal to the first threshold; determine the estimated signal of the target channel as the target signal; perform a short-time inverse Fourier transform on the target signal to obtain the target voice signal.

[0181] In some embodiments, the extraction module is further configured to: perform fixed-frame sampling on the covariance to obtain multiple fixed-frame covariances of each estimated signal; perform eigen-decomposition on the multiple fixed-frame covariances to obtain the eigenvalues and eigenvectors of the multiple fixed-frame covariances; extract the eigenvector corresponding to the maximum eigenvalue of the multiple fixed-frame covariances as the signal subspace of the estimated signal; determine the direction of arrival of the signal subspace as the incident angle of the estimated signal.

[0182] In some embodiments, the extraction module is further configured to: when there are incident angles of multiple estimated signals that meet the first preset condition, compare the voiceprint features of the multiple estimated signals; determine the channel corresponding to the estimated signal whose voiceprint feature meets the second preset condition as the target channel.

[0183] In some embodiments, the second preset condition includes: the voiceprint feature matches the target voiceprint feature.

[0184] In some embodiments, the separation module is further configured to: use the simulated voice signal to train the sound source probability density model to obtain the trained sound source probability density model.

[0185] Figure 10 FIG. 1000 is a schematic structural diagram of an electronic device 1000 for implementing the above voice signal extraction method according to an exemplary embodiment.

[0186] Referring to Figure 10 , the electronic device 1000 may include one or more of the following components: a processing component 1002, a memory 1004, a power supply component 1006, a multimedia component 1008, an audio component 1010, an input / output (I / O) interface 1012, a sensor component 1014, and a communication component 1016.

[0187] The processing component 1002 generally controls the overall operation of the electronic device 1000, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 1002 may include one or more processors 1020 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 1002 may include one or more modules to facilitate the interaction between the processing component 1002 and other components. For example, the processing component 1002 may include a multimedia module to facilitate the interaction between the multimedia component 1008 and the processing component 1002.

[0188] The memory 1004 is configured to store various types of data to support the operation of the electronic device 1000. Examples of such data include instructions for any application or method operating on the electronic device 1000, contact data, phone book data, messages, pictures, videos, etc. The memory 1004 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0189] The power component 1006 provides power to various components of the electronic device 1000. The power component 1006 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 1000.

[0190] The multimedia component 1008 includes a screen that provides an output interface between the electronic device 1000 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 1008 includes a front camera and / or a rear camera. When the electronic device 1000 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0191] The audio component 1010 is configured to output and / or input audio signals. For example, the audio component 1010 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 1000 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 1004 or transmitted via the communication component 1016. In some embodiments, the audio component 1010 further includes a speaker for outputting audio signals.

[0192] The I / O interface 1012 provides an interface between the processing component 1002 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0193] The sensor component 1014 includes one or more sensors for providing an assessment of various aspects of the status of the electronic device 1000. For example, the sensor component 1014 can detect the on / off state of the electronic device 1000, the relative positioning of components, such as the display and keypad of the electronic device 1000. The sensor component 1014 can also detect a change in the position of the electronic device 1000 or a component of the electronic device 1000, the presence or absence of user contact with the electronic device 1000, the orientation or acceleration / deceleration of the electronic device 1000, and the temperature change of the electronic device 1000. The sensor component 1014 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 1014 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 1014 can further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0194] The communication component 1016 is configured to facilitate communication between the electronic device 1000 and other devices in a wired or wireless manner. The electronic device 1000 can access a wireless network based on communication standards, such as WiFi, 2G or 3G, 4G LTE, 5G NR (New Radio), or a combination thereof. In an exemplary embodiment, the communication component 1016 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1016 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0195] In an exemplary embodiment, the electronic device 1000 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described method.

[0196] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1004 including instructions, and the above instructions can be executed by a processor 1020 of the electronic device 1000 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0197] An embodiment of the present disclosure also proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the voice signal extraction method described in the above embodiments of the present disclosure.

[0198] An embodiment of the present disclosure also proposes a computer program product, including a computer program, and the computer program executes the voice signal extraction method described in the above embodiments of the present disclosure when being executed by a processor.

[0199] An embodiment of the present disclosure also proposes a vehicle including Figure 9 the voice signal extraction device described as Figure 10 or the electronic device described as

[0200] Figure 11 FIG. 20 is a schematic structural diagram of a chip 1100 for implementing the above voice signal extraction method according to an exemplary embodiment. Referring to Figure 11 , the chip 1100 includes at least one communication interface 1101 and a processor 1102. The communication interface 1101 is configured to receive signals input to the chip 1100 or signals output from the chip 1100, and the processor 1102 communicates with the communication interface 1101 and implements the in-vehicle voice signal extraction method described in the above embodiments of the present disclosure through logic circuits or by executing code instructions.

[0201] It should be noted that the terms "first", "second", etc. in the description of the present disclosure, the claims and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0202] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples" or "some examples", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0203] Any process or method description, whether in a flowchart or otherwise described herein, can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present disclosure includes additional implementations, where the functions can be executed in a manner other than shown or discussed, including in a substantially simultaneous manner or in a reverse order according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure pertain.

[0204] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processing module, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (control method), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0205] It should be understood that various parts of the embodiments of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0206] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0207] In addition, each functional unit in various embodiments of the present disclosure may be integrated into a processing module, may exist separately physically as individual units, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, or the like.

[0208] Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.

Claims

1. A method for extracting a voice signal, characterized in that, The method includes: Obtaining a mixed audio signal, where the mixed audio signal includes a plurality of sound source signals, and the plurality of sound source signals correspond to a plurality of channels of a microphone array; Based on the mixed audio signal, using a trained sound source probability density model, obtaining the sound source probability density weights of the plurality of sound source signals, where the sound source probability density weights are the probabilities of obtaining each sound source signal from the mixed audio signal; Based on the sound source probability density weights, performing independent vector analysis on the mixed audio signal to obtain an estimated signal of each sound source signal and a covariance of the amplitude of each sound source signal; Based on the estimated signal and the covariance, determining a target voice signal that meets a preset target angle from the mixed audio signal.

2. The method according to claim 1, wherein The performing independent vector analysis on the mixed audio signal based on the sound source probability density weights to obtain an estimated signal of each sound source signal and a covariance of the amplitude of each sound source signal includes: Performing a short-time Fourier transform on the mixed audio signal to obtain a frequency-domain signal of the mixed audio signal; Performing independent vector analysis on the frequency-domain signal to obtain an estimated signal of each sound source signal and a covariance of the amplitude of each sound source signal.

3. The method according to claim 1, wherein The determining a target voice signal that meets a preset target angle from the mixed audio signal based on the estimated signal and the covariance includes: Estimating the covariance to obtain the incident angle of each estimated signal; Determining the channel where the incident angle that meets a first preset condition is located as the target channel, where the first preset condition is that the difference between the incident angle and the preset target angle is less than or equal to a first threshold; Determining the estimated signal of the target channel as the target signal; Performing a short-time inverse Fourier transform on the target signal to obtain the target voice signal.

4. The method according to claim 3, characterized in that, The estimating the covariance to obtain the incident angle of each estimated signal includes: Performing fixed-frame sampling on the covariance to obtain a plurality of fixed-frame covariances of each estimated signal; Performing eigenvalue decomposition on the plurality of fixed-frame covariances to obtain eigenvalues and eigenvectors of the plurality of fixed-frame covariances; Extracting the eigenvector corresponding to the maximum eigenvalue of the plurality of fixed-frame covariances as the signal subspace of the estimated signal; Determining the direction of arrival of the signal subspace as the incident angle of the estimated signal.

5. The method according to claim 3, wherein The determining the channel where the incident angle that meets a first preset condition is located as the target channel includes: When there are incident angles of multiple estimated signals that meet the first preset condition, comparing the voiceprint features of the multiple estimated signals; Determining the channel corresponding to the estimated signal whose voiceprint feature meets a second preset condition as the target channel.

6. The method according to claim 5, where the second preset condition includes: The voiceprint feature matches the target voiceprint feature.

7. The method according to claim 1, characterized in that The method further includes: Using a simulated voice signal to train the sound source probability density model to obtain the trained sound source probability density model.

8. A voice signal extraction device, characterized in that, Including: A separation module and an extraction module, The separation module is used to obtain a mixed audio signal, where the mixed audio signal includes a plurality of sound source signals, and the plurality of sound source signals correspond to a plurality of channels of a microphone array; Based on the mixed audio signal, using the trained sound source probability density model, obtain the sound source probability density weights of the multiple sound source signals, where the sound source probability density weights are the probabilities of obtaining each sound source signal from the mixed audio signal; Based on the sound source probability density weights, perform independent vector analysis on the mixed audio signal to obtain the estimated signal of each sound source signal and the covariance of the amplitude of each sound source signal; The extraction module is configured to determine a target voice signal that meets a preset target angle from the mixed audio signal based on the estimated signal and the covariance.

9. An electronic device, characterized in that, Comprising: a processor and a memory for storing a computer program that can run on the processor, wherein, when the processor is used to run the computer program, it executes the method according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.

11. A chip, characterized in that, Comprising at least one processor and a communication interface; the communication interface is used to receive signals input to the chip or signals output from the chip, and the processor communicates with the communication interface and implements the method according to any one of claims 1-7 through logic circuits or by executing code instructions.

12. A vehicle, characterized in that, Comprising the voice signal extraction device according to claim 8 or the electronic device according to claim 9.