Vehicle control method, vehicle, and storage medium

Through a two-stage separation processing method, combined with beam design and blind source separation model, the problem of poor audio signal separation performance in the vehicle cabin in the existing technology is solved, and efficient and accurate sound zone signal separation is achieved, which is suitable for various vehicles.

CN116364071BActive Publication Date: 2025-10-24GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310334599.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-29
Publication Date
2025-10-24
Estimated Expiration
2043-03-29

AI Technical Summary

Technical Problem

Among the existing audio signal separation methods in vehicle cabins, the blind source separation method based on distributed arrays degrades in performance under cutting-edge mixing conditions, while the deep learning method has high computational complexity and lacks universality, resulting in poor separation performance.

Method used

A two-stage separation processing method is adopted. First, pre-separation enhancement is performed through beam design, and then further processing is performed using the blind source separation model, including determining the beam coefficients and blind source separation coefficient matrix, to achieve accurate separation of the audio signal.

Benefits of technology

The accuracy and universality of signal separation are improved, the amount of calculation is reduced, and more vehicles with insufficient performance can effectively separate audio signals in different sound zones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116364071B_ABST
    Figure CN116364071B_ABST
Patent Text Reader

Abstract

The application discloses a vehicle control method, comprising: acquiring an audio signal in a vehicle cabin; performing first separation processing on the audio signal based on audio zone information in the vehicle cabin to determine a first audio signal corresponding to each audio zone; performing second separation processing on each first audio signal to determine a second audio signal corresponding to each audio zone in the vehicle cabin, so as to control the vehicle according to the second audio signal. Through the two separation processes, the acquired audio signal is separated and enhanced for different audio zones, and the second audio signal based on different audio zones can be obtained. Each group of second audio signals enhances the signal in the corresponding audio zone and isolates the signal in the non-corresponding audio zone. Meanwhile, the calculation amount of each stage is effectively reduced, the universality of the signal separation method is improved, and more vehicles with insufficient performance to support a large amount of calculation can also realize the separation of the audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent vehicles, in particular to a vehicle control method, a vehicle and a computer readable storage medium. BACKGROUND

[0002] In the related art, in order to realize the separation and positioning of different sound zones in the vehicle cabin, a blind source separation method based on a distributed array or a deep learning method is generally used. However, for the blind source separation method based on the distributed array, the derivation process is generally implemented when all sound sources are active. However, in actual scenarios, the voice signal is mostly in a sharp mixing state, and the pure noise segment or the segment in which less than the total number of sound sources are active will interfere with the estimation of parameters and the convergence of filters, resulting in a decrease in separation performance. For the deep learning method, the calculation amount is large, and the performance is extremely dependent on real vehicle data, and the universality is not strong. SUMMARY

[0003] The present application provides a vehicle control method, a vehicle and a computer readable storage medium.

[0004] The vehicle control method provided by the present application embodiment comprises:

[0005] obtaining an audio signal in a vehicle cabin;

[0006] performing first separation processing on the audio signal based on sound zone information in the vehicle cabin to determine a first audio signal corresponding to each sound zone;

[0007] performing second separation processing on each first audio signal to determine a second audio signal corresponding to each sound zone in the vehicle cabin, and controlling the vehicle according to the second audio signal.

[0008] In this way, the present application can obtain multiple groups of second audio signals based on different sound zones by performing different separation and enhancement on the obtained audio signal for different sound zones through the first and second separation processing, and each group of second audio signals enhances the signal in the sound zone corresponding thereto and isolates the signal in the sound zone not corresponding thereto. At the same time, by dividing the separation process into two stages, the calculation amount of each stage is effectively reduced, the universality of the signal separation method is improved, and more vehicles with insufficient performance to support the deep learning method and other methods with large calculation amount can also realize the separation of audio signals in different sound zones.

[0009] The first separation processing on the audio signal based on the sound zone information in the vehicle cabin comprises:

[0010] The audio signals are pre-separated and enhanced based on the sound zone information in the vehicle cabin to determine first audio signals corresponding to each sound zone.

[0011] Thus, the application can realize pre-separated enhancement of the acquired audio signals, thereby realizing preliminary screening and separation of the audio signals into signals corresponding to each sound zone and pre-enhancing the signals in the corresponding sound zone for subsequent further processing.

[0012] The pre-separated enhancement of the audio signals based on the sound zone information in the vehicle cabin to determine first audio signals corresponding to each sound zone includes:

[0013] According to a preset first function relationship and a preset constraint condition, a plurality of beam coefficients corresponding to a plurality of different sound zones in the vehicle cabin are determined;

[0014] According to a preset second function relationship, a plurality of first audio signals are determined.

[0015] Thus, the application can realize pre-separated enhancement of the audio signals based on beam design.

[0016] The determination of the plurality of beam coefficients corresponding to the plurality of different sound zones in the vehicle cabin according to the preset first function relationship and the preset constraint condition includes:

[0017] According to the preset first function relationship composed of the audio signals, preset ideal beam parameters and a preset steering vector matrix, and the preset constraint condition composed of the beam coefficients, steering vectors in the preset steering vector matrix and a preset white noise gain threshold, a plurality of beam coefficients corresponding to a plurality of different sound zones in the vehicle cabin are determined;

[0018] According to the preset second function relationship composed of the audio signals, the beam coefficients, time information and time delay information of a radio component used to acquire the audio signals, a plurality of first audio signals are determined.

[0019] Thus, the application can realize design of the beam coefficients based on various parameters and corresponding relationships to isolate audio signals of non-corresponding sound zones for pre-separated enhancement of the audio signals.

[0020] The second separation processing of each first audio signal to determine second audio signals corresponding to different sound zones in the vehicle cabin includes:

[0021] According to a plurality of first audio signals and the preset blind source separation model, a blind source separation coefficient matrix is determined.

[0022] According to the first audio signals and the blind source separation coefficient matrix, the second audio signals are determined.

[0023] Thus, the application can enhance the filtering of audio signals in the irrelevant sound area through pre-separation, avoid the signal in the intermittent mixing state, and improve the accuracy and reliability of blind source separation.

[0024] The blind source separation coefficient matrix is determined according to the first audio signals and the preset blind source separation model, including:

[0025] According to the preset blind source separation model, an acoustic representation matrix is determined.

[0026] According to the first audio signals and the acoustic representation matrix, a weight covariance matrix is determined.

[0027] According to the weight covariance matrix, the blind source separation coefficient matrix is determined through iterative processing.

[0028] Thus, the application can realize blind source separation through multiple parameter matrices.

[0029] The acoustic representation matrix is determined according to the preset blind source separation model, including:

[0030] According to the preset blind source separation model, a non-negative matrix factorization process is performed to determine a basis matrix and a coefficient matrix.

[0031] According to the basis matrix and the coefficient matrix, the acoustic representation matrix is determined.

[0032] Thus, the application limits the generation method of the acoustic representation matrix.

[0033] The weight covariance matrix is determined according to the first audio signals and the acoustic representation matrix, including:

[0034] According to a preset third function relationship constituted by the first audio signals, the acoustic representation matrix and a unit matrix, the weight covariance matrix is determined.

[0035] Thus, the application limits the generation method of the weight covariance matrix.

[0036] The application also provides a vehicle including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the vehicle control method according to any one of the above embodiments.

[0037] The application further provides a computer readable storage medium, which stores a computer program, and the computer program, when executed by one or more processors, implements the vehicle control method according to any one of the above embodiments.

[0038] Additional aspects and advantages of the embodiments of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS

[0039] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, taken in conjunction with the following drawings in which:

[0040] Figure 1 Flowchart of the vehicle control method of the present application;

[0041] Figure 2 Application scenario diagram of the vehicle control method of the present application;

[0042] Figure 3 Flowchart of the vehicle control method of the present application;

[0043] Figure 4 Application data heat diagram of the vehicle control method of the present application;

[0044] Figure 5 Flowchart of the vehicle control method of the present application;

[0045] Figure 6 Flowchart of the vehicle control method of the present application;

[0046] Figure 7 Flowchart of the vehicle control method of the present application;

[0047] Figure 8 Flowchart of the vehicle control method of the present application. DETAILED DESCRIPTION

[0048] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the drawings, in which the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below are exemplary and are used only to explain the embodiments of the present application, and cannot be understood as limiting the embodiments of the present application.

[0049] As shown in Figure 1 The present application provides a vehicle control method, comprising:

[0050] 01: acquiring an audio signal in a vehicle cabin;

[0051] 02: Based on the sound zone information in the vehicle cabin, perform a first separation process on the audio signal to determine a first audio signal uniquely corresponding to each sound zone;

[0052] 03: Perform a second separation process on each first audio signal to determine a second audio signal uniquely corresponding to each sound zone in the vehicle cabin, so as to control the vehicle according to the second audio signal.

[0053] The present application also provides a vehicle, including a memory and a processor, the memory storing a computer program, the processor being used to obtain an audio signal in a vehicle cabin, and to perform a first separation process on the audio signal based on sound zone information in the vehicle cabin to determine a first audio signal uniquely corresponding to each sound zone, and to perform a second separation process on each first audio signal to determine a second audio signal uniquely corresponding to each sound zone in the vehicle cabin, so as to control the vehicle according to the second audio signal.

[0054] Specifically, in order to separate the audio signal in the vehicle cabin, the audio signal in the vehicle cabin should first be obtained through relevant equipment. Figure 2 As shown, in order to obtain the audio signal in the vehicle cabin more accurately and comprehensively, a pair of sound collecting devices can be respectively set at the left rear door, the right rear door and the central rearview mirror inside the vehicle cabin. Figure 2 The three groups of side-by-side circles in the figure indicate that the collected audio signals are divided into six waveforms according to the different sound collection devices. Figure 2 In the figure, each star represents the center position of the corresponding seat, which is the approximate sound source. The two straight lines in the figure divide the interior of the vehicle cabin into four different sound zones according to the distribution of the center position of the seats in the vehicle cabin.

[0055] Then, the collected audio signals are subjected to two separate separation processes, one for each audio zone, to convert the audio signals into second audio signals corresponding to different sound zones within the vehicle cabin. In certain embodiments, the vehicle cabin includes four seats, including a driver, a front passenger, and two rear seats, each corresponding to a sound zone. Therefore, the vehicle cabin is divided into four sound zones. The first separation process converts the audio signals into four first audio signals corresponding to the four sound zones. A second separation process is then performed on the four first audio signals to further enhance the portions of the four first audio signals in the corresponding sound zones and isolate the portions in other sound zones, ultimately determining four groups of second audio signals corresponding to the sound zones. This method allows the mixed audio signals within the vehicle cabin to be separated by sound zones, enabling the in-vehicle system to more easily identify the user's voice commands within the vehicle cabin and more accurately control the vehicle to perform different actions.

[0056] Thus, the application can obtain multiple groups of second audio signals based on different sound zones by performing two separation processes to enhance the acquired audio signals for different sound zones, and each group of second audio signals is enhanced in the sound zone corresponding thereto and isolated from the sound zone not corresponding thereto. Meanwhile, by dividing the separation process into two stages, the calculation amount of each stage is effectively reduced, the universality of the signal separation method is improved, and more vehicles with insufficient performance to support deep learning methods and other methods with large calculation amount can also realize separation of audio signals in different sound zones.

[0057] Step 02 comprises:

[0058] Based on the sound zone information in the vehicle cabin, the audio signal is pre-separated and enhanced to determine the first audio signal corresponding to each sound zone.

[0059] The processor is configured to pre-separate and enhance the audio signal based on the sound zone information in the vehicle cabin to determine the first audio signal corresponding to each sound zone.

[0060] Specifically, for the first separation process, the main purpose is to initially separate the acquired audio signal according to the sound zone, and to enhance the signal corresponding to the sound zone in the audio signal to a certain extent, so that the subsequent second separation process can more conveniently and effectively realize the enhancement or isolation of different parts of the first audio signal, thereby realizing the complete correspondence between the second audio signal and different sound zones.

[0061] Thus, the application can realize pre-separation and enhancement of the acquired audio signal, thereby realizing initial screening and separation of the audio signal into signals corresponding to each sound zone, and pre-enhancing the signal in the corresponding sound zone to facilitate further processing.

[0062] As shown in the first aspect, Figure 3 the application pre-separates and enhances the audio signal based on the sound zone information in the vehicle cabin to determine the first audio signal corresponding to each sound zone, comprising:

[0063] 021: determining multiple beam coefficients corresponding to multiple different sound zones in the vehicle cabin according to a preset first function relationship and a preset constraint condition;

[0064] 022: determining multiple first audio signals according to a preset second function relationship.

[0065] The processor is configured to determine a plurality of beam coefficients respectively uniquely corresponding to a plurality of different sound zones in the vehicle cabin according to a preset first function relationship and a preset constraint condition, and to determine the plurality of first audio signals according to a preset second function relationship.

[0066] Specifically, in some examples, the first separation processing adopts a beam design manner, which utilizes the directivity and null of the beam to filter out the audio signals in a specific direction and retain the audio signals in other directions. Specifically, it determines the plurality of first audio signals by designing beam coefficients and applying the beam coefficients to the audio signals.

[0067] In some examples, Figure 4 A beam heat map corresponding to the beam coefficients corresponding to the main driver sound zone of the vehicle cabin as shown in Figure 2 The beam heat map corresponding to the beam coefficients corresponding to the main driver sound zone of the vehicle cabin as shown in Figure 4 It can be seen from Figure 4 The situation of the other three sound zones is similar to that shown in Figure 2 According to the above method, the corresponding number of beam coefficients can be designed according to the number of sound zones, which are respectively applied to the audio signals obtained by the plurality of sound collecting devices to obtain a plurality of pre-separated enhanced first audio signals. As shown in

[0068] Thus, the present application can achieve pre-separation and enhancement of audio signals based on beam design.

[0069] As shown in Figure 5 Step 021 comprises:

[0070] 0211: determining a plurality of beam coefficients respectively uniquely corresponding to a plurality of different sound zones in the vehicle cabin according to a preset first function relationship constituted by an audio signal, a preset ideal beam parameter and a preset steering vector matrix, and a preset constraint condition constituted by a beam coefficient, a steering vector in the preset steering vector matrix and a preset white noise gain threshold;

[0071] 0212: determining a plurality of first audio signals according to a preset second function relationship constituted by an audio signal, a beam coefficient, time information and time delay information of a sound collecting component used to obtain the audio signal.

[0072] The processor is configured to determine a plurality of beam coefficients respectively uniquely corresponding to a plurality of different sound zones in the vehicle cabin according to a preset first function relationship constituted by the audio signal, the preset ideal beam parameter, and the preset steering vector matrix, and a preset constraint condition constituted by the beam coefficient, the steering vector in the preset steering vector matrix, and the preset white noise gain threshold, and to determine a plurality of first audio signals according to a preset second function relationship constituted by the audio signal, the beam coefficient, the time information, and the time delay information of the radio component used to obtain the audio signal.

[0073] Specifically, the design process of the beam coefficient is described in detail as follows. First, in order to determine the beam coefficient, the objective function and the constraint condition need to be determined.

[0074] In some examples, the preset first objective function is as follows:

[0075]

[0076] Meanwhile, for the preset first objective function, the preset constraint condition is as follows:

[0077]

[0078] In the above relationship, ω p represents the frequency point of the audio signal collected by the sound collecting device numbered p, is the ideal beam parameter matrix, G(ω p ) is the steering vector matrix, d(ω p ) is the steering vector in the steering vector matrix G(ω p ), γ is the lower limit of the preset white noise gain, w f (ω p ) is the beam coefficient to be solved, and its form is a beam coefficient matrix.

[0079] According to the first objective function and the constraint condition as described above, the required beam coefficient matrix can be determined. The number of the required beam coefficient matrix is the same as the number of the sound zones in the vehicle cabin. In some examples, as shown in the vehicle cabin Figure 2 , there are four required beam coefficient matrices corresponding to the four sound zones in the vehicle cabin. The beam coefficient matrix obtained by the above method is expressed in the frequency domain. If it is to be applied to the audio signal to determine the first audio signal, it needs to be converted into the time domain.

[0080] After the beam coefficient matrix is converted into the time domain, the audio signal is subjected to the following second objective function, and the first audio signal corresponding to the sound zone corresponding to the beam coefficient matrix can be obtained:

[0081]

[0082] wherein y(t) is the first audio signal, P is the total number of sound collection devices, in some examples 6, p is the number of sound collection devices (a positive integer ranging from 0 to (P-1)), w p is the time domain form of the beam coefficient matrix, x p is the beam coefficient matrix w p corresponds to the audio signal, t is the time length, and τ is the time delay length of the corresponding sound collection device, which is determined by the attribute of the sound collection device itself.

[0083] Since the beam coefficient matrix w p has a total of 4, there are also 4 first audio signals generated, and each expression of the first audio signal is in the time domain form, and each expression has 6 terms, each of which corresponds to an audio signal obtained by a sound collection device.

[0084] In this way, the present application can realize the design of the beam coefficient to isolate the audio signal of the non-corresponding sound area based on the parameters and the corresponding relationship, so as to enhance the pre-separation of the audio signal.

[0085] As Figure 6 shown, step 03 includes:

[0086] 031: determining a blind source separation coefficient matrix according to the plurality of first audio signals and the preset blind source separation model;

[0087] 032: determining a plurality of second audio signals according to the plurality of first audio signals and the blind source separation coefficient matrix.

[0088] The processor is configured to determine a blind source separation coefficient matrix according to the plurality of first audio signals and the preset blind source separation model, and to determine a plurality of second audio signals according to the plurality of first audio signals and the blind source separation coefficient matrix.

[0089] Specifically, for the second separation processing, in some examples, a blind source separation method is adopted. Since the first separation processing, the signal of the non-corresponding sound area part in the first audio signal has been suppressed or even filtered out due to the influence of the beam coefficient matrix, and the signal of the corresponding sound area part is retained, the intermittent mixing state of the first audio signal is suppressed, and it can be approximately considered that the sound source of the total sound source number is in the sound state, the interference caused by the estimation of the parameter and the convergence of the filtering is limited, thereby improving the separation characteristics of the blind source separation. At the same time, since the first audio signal has a clear correspondence with the sound area, the probability of the signal separated by the blind source separation being the expected signal is very high, and it is not easy to cause sound area crosstalk and leakage. Therefore, after the first separation processing by the beam design, the second separation processing can be directly performed by the blind source separation method, and the signal separation effect can be further improved.

[0090] Further, in order to realize blind source separation, first, a blind source separation matrix for acting on the first audio signal is determined by taking the first audio signal and a preset blind source separation model as a reference, the number of elements of the matrix is related to the number of the first audio signals, then the blind source separation matrix is applied to the plurality of first audio signals respectively, a plurality of second audio signals corresponding to the first audio signals are determined, the number of the second audio signals is the same as that of the first audio signals, and the second audio signals correspond one-to-one to different sound areas in the vehicle cabin.

[0091] In this way, the present application can enhance the filtering of the audio signal in the irrelevant sound area through pre-separation, avoid the signal in the intermittent mixing state, and improve the accuracy and reliability of the blind source separation.

[0092] Step 031 includes:

[0093] 0311: determining an acoustic representation matrix according to a preset blind source separation model;

[0094] 0312: determining a weight covariance matrix according to the first audio signal and the acoustic representation matrix;

[0095] 0313: determining a blind source separation coefficient matrix through iterative processing according to the weight covariance matrix.

[0096] The processor is configured to determine an acoustic representation matrix according to a preset blind source separation model, determine a weight covariance matrix according to the first audio signal and the acoustic representation matrix, and determine a blind source separation coefficient matrix through iterative processing according to the weight covariance matrix.

[0097] Specifically, for the second separation processing, the process of determining the blind source separation coefficient matrix according to the preset blind source separation model is roughly divided into three steps. In some examples, for a plurality of first audio signals that have been determined, a real-time ILRMA method is generally adopted to perform blind source separation. For the ILRMA method, first, the harmonic information of each first audio signal is highlighted by using non-negative matrix factorization (NMF), and an acoustic representation matrix can be obtained according to the non-negative matrix factorization. Then, based on the acoustic representation matrix and in combination with each first audio signal, a plurality of weight covariance matrices corresponding to the first audio signals can be obtained. Then, based on each weight covariance matrix, a blind source separation coefficient matrix for applying an effect to the first audio signals to achieve a separation effect can be obtained, and the number of elements in each row of the blind source separation coefficient matrix is equal to the number of first audio signals, and the number of elements in each column is also equal to the number of first audio signals. Such a setting can make the action between the matrices legal on a mathematical level.

[0098] Thus, the application can achieve blind source separation through a plurality of parameter matrices.

[0099] As shown in Figure 8 , step 0311 includes:

[0100] 03111: performing non-negative matrix factorization processing according to a preset blind source separation model to determine a basis matrix and a coefficient matrix;

[0101] 03112: determining an acoustic representation matrix according to the basis matrix and the coefficient matrix.

[0102] The processor is configured to perform non-negative matrix factorization processing according to a preset blind source separation model to determine a basis matrix and a coefficient matrix, and to determine an acoustic representation matrix according to the basis matrix and the coefficient matrix.

[0103] Specifically, for the generation of the acoustic representation matrix, the following is described in detail: in some examples, according to the non-negative matrix factorization and according to the model adopted by the ILRMA method, a basis matrix T n and a coefficient matrix V n are solved, and then the coefficient matrix V n is right multiplied by the basis matrix T n to obtain a matrix, which is the acoustic representation matrix R n , that is:

[0104] R n = n V n

[0105] Thus, the application limits the generation method of the acoustic representation matrix.

[0106] Step 0312 includes:

[0107] The weight covariance matrix is determined according to a preset third function relationship constituted by the first audio signals, the acoustic characteristic matrix, and a unit matrix.

[0108] The processor is configured to determine the weight covariance matrix according to a preset third function relationship constituted by the first audio signals, the acoustic characteristic matrix, and a unit matrix.

[0109] Specifically, after the acoustic characteristic matrix R n is determined, the acoustic characteristic matrix R n may be used to solve the weight covariance matrix V m () in combination with multiple first audio signals. In some examples, the number of the first audio signals and the number of the weight covariance matrices are both 4, and the weight covariance matrix V m () is determined by the following function:

[0110]

[0111] wherein E is a unit matrix, k is the number of the audio zone, and x(ω, t) is the first audio signal corresponding to the aforementioned audio zone.

[0112] In addition, in some examples, after the weight covariance matrix V m (k) is determined, the weight covariance matrix V m (k) is iterated to finally obtain a K×K deconvolution matrix, i.e., the blind source separation coefficient matrix to be solved, where K is the number of the first audio signals, and in some examples, K is 4. The iteration is performed in the order shown in the following relationship until k reaches its maximum value K:

[0113] w k (ω, t)←(W(ω, t)V m (k)) -1 e k

[0114]

[0115] wherein w k (ω, t) is the separation coefficient corresponding to the audio zone numbered k determined according to the ILRMA algorithm, W(ω, t) is the deconvolution matrix, i.e., the blind source separation coefficient matrix to be solved, W(ω, t) = [w1(ω, t) w2(ω, t)...w K (ω, t)] H K is the number of the first audio signals, i.e., the number of the audio zones, and e k is a column vector, and column vector e kThe kth element of is 1 and all other elements are 0.

[0116] Furthermore, in some examples, based on the above implementation, the blind source separation coefficient matrix W(ω, t) to be determined can be determined according to the above process. Then, the obtained blind source separation coefficient matrix is ​​applied to the first audio signal to obtain K final separated second audio signals. For convenience of representation, the specific relationship is expressed in a matrix as follows:

[0117] Y(ω,t)=W(ω,t)X(ω,t)

[0118] Among them, X(ω,t)=[x1(ω,t)x2(ω,t)...x K (ω, t)], that is, a matrix composed of K first audio signals, in some examples K = 4, Y (ω, t) = [y1 (ω, t) y2 (ω, t) ... y K (ω, t)], that is, a matrix composed of K second audio signals, in some examples K = 4. In this way, second audio signals corresponding to multiple sound zones are obtained, thereby completing the separation of the audio signals in the vehicle cabin.

[0119] The present application also provides a computer-readable storage medium storing a computer program, which implements the vehicle control method in the above embodiment when the computer program is executed by one or more processors.

[0120] In the description of this specification, the reference terms "certain embodiments", "in an example", "exemplarily", etc. mean that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are mutually inconsistent.

[0121] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0122] Although the embodiments of the present application have been shown and described above, it is understood that the above-described embodiments are exemplary and are not to be construed as limiting the present application, and that changes, modifications, substitutions and variations can be made by those skilled in the art without departing from the scope of the present application.

Claims

1. A vehicle control method characterized by, The method comprises: acquiring an audio signal in a vehicle cabin; performing first separation processing on the audio signal based on audio zone information in the vehicle cabin to determine a first audio signal corresponding to each audio zone; performing second separation processing on each first audio signal to determine a second audio signal corresponding to each audio zone in the vehicle cabin to control the vehicle according to the second audio signal; wherein the second separation processing on each first audio signal to determine a second audio signal corresponding to each audio zone in the vehicle cabin comprises: determining a blind source separation coefficient matrix according to a plurality of first audio signals and a preset blind source separation model; determining a plurality of second audio signals according to a plurality of first audio signals and the blind source separation coefficient matrix.

2. The method of claim 1, wherein, The first separation processing on the audio signal based on the audio zone information in the vehicle cabin to determine a first audio signal corresponding to each audio zone comprises: performing pre-separation enhancement processing on the audio signal based on the audio zone information in the vehicle cabin to determine a first audio signal corresponding to each audio zone.

3. The method of claim 2, wherein, The pre-separation enhancement processing on the audio signal based on the audio zone information in the vehicle cabin to determine a first audio signal corresponding to each audio zone comprises: determining a plurality of beam coefficients corresponding to a plurality of different audio zones in the vehicle cabin according to a preset first functional relationship and a preset constraint condition; determining a plurality of first audio signals according to a preset second functional relationship.

4. The method of claim 3, wherein, The determination of a plurality of beam coefficients corresponding to a plurality of different audio zones in the vehicle cabin according to a preset first functional relationship and a preset constraint condition comprises: determining a plurality of beam coefficients corresponding to a plurality of different audio zones in the vehicle cabin according to the preset first functional relationship composed of the audio signal, a preset ideal beam parameter and a preset steering vector matrix, and the preset constraint condition composed of the beam coefficient, a steering vector in the preset steering vector matrix and a preset white noise gain threshold; determining a plurality of first audio signals according to the preset second functional relationship composed of the audio signal, the beam coefficient, time information and time delay information of a radio component used to acquire the audio signal.

5. The method of claim 1, wherein, The determination of a blind source separation coefficient matrix according to a plurality of first audio signals and a preset blind source separation model comprises: determining an acoustic representation matrix according to the preset blind source separation model; determining a weight covariance matrix according to the first audio signal and the acoustic representation matrix; determining the blind source separation coefficient matrix through iterative processing according to the weight covariance matrix.

6. The method of claim 5, wherein, The determination of an acoustic representation matrix according to the preset blind source separation model comprises: performing non-negative matrix factorization processing according to the preset blind source separation model to determine a basis matrix and a coefficient matrix; determining the acoustic representation matrix according to the basis matrix and the coefficient matrix.

7. The method of claim 5, wherein, The determining the weight covariance matrix according to the first audio signal and the acoustic characterization matrix comprises: The weight covariance matrix is determined according to a preset third function relationship constituted by the first audio signal, the acoustic characterization matrix and a unit matrix.

8. A vehicle characterized by comprising: The vehicle comprises a memory and a processor, the memory stores a computer program, and the computer program causes the processor to execute the vehicle control method according to any one of claims 1-7 when the computer program is executed by the processor.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program causes one or more processors to implement the vehicle control method according to any one of claims 1-7 when the computer program is executed by the one or more processors.

Citation Information

Patent Citations

  • Voice signal processing method and device, storage medium, electronic equipment and vehicle

    CN114783458A