Speech separation method, apparatus, system, electronic device, and storage medium

By deploying multiple microphone arrays inside the vehicle, using spectrum and phase information to separate speech within and between sound zones, and combining this with narrow-beam enhancement technology, the problem of poor separation performance of in-vehicle speech separation systems in harsh scenarios has been solved, improving wake-up rate and positioning accuracy, and optimizing user experience.

CN119785815BActive Publication Date: 2025-11-18IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411759439.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-11-18
Estimated Expiration
2044-12-03

AI Technical Summary

Technical Problem

Existing in-vehicle voice separation systems perform poorly in harsh environments, failing to accurately separate and locate voice, which affects subsequent interaction processes and leads to a decline in user experience.

Method used

Multiple microphone arrays are used for signal acquisition. By acquiring the spectral and phase information of speech signals from multiple sound regions, speech is separated between and within sound regions. The separation results between and within sound regions are then fused, and narrowband enhancement technology is used to improve the separation effect.

Benefits of technology

It achieves accurate voice separation in harsh environments, improves the wake-up rate and positioning accuracy of the in-vehicle interactive system, and optimizes the user's voice interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785815B_ABST
    Figure CN119785815B_ABST
Patent Text Reader

Abstract

The application provides a speech separation method, device, system, electronic equipment and storage medium, wherein the method comprises: based on the frequency spectrum information and phase information of the speech signals of multiple sound areas on a target vehicle, performing inter-sound-area speech separation and intra-sound-area speech separation on the speech signals of the multiple sound areas; based on the inter-sound-area separation result obtained by inter-sound-area speech separation and the intra-sound-area separation result obtained by intra-sound-area speech separation, determining the speech separation result of the multiple sound areas, overcoming the defect that the vehicle-mounted speech separation effect is poor in a traditional scheme under a harsh scene, and through signal collection by multiple microphone arrays, spatial information can be better acquired, and on this basis, speech separation of two different precisions is performed, and the results are fused, more accurate speech separation is realized, accurate and reliable speech separation results can be obtained, the wake-up rate of a vehicle-mounted interaction system and the positioning accuracy are effectively improved, and therefore the speech interaction experience of a user is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a speech separation method, apparatus, system, electronic device, and storage medium. Background Technology

[0002] With the development of automotive intelligence, in-vehicle interactive systems have gradually become an important part of modern automobiles, providing drivers and passengers with a more convenient and intelligent interactive experience. This system controls various in-vehicle functions by recognizing passengers' wake words and commands. However, in practical applications, the in-vehicle environment is complex and variable, including interference factors such as vehicle noise, wind noise, and multiple people talking simultaneously, all of which pose significant challenges to speech separation for in-vehicle speakers.

[0003] Currently, most vehicle voice separation methods rely on a single microphone array, meaning that voice separation is achieved through a single microphone array deployed within the vehicle. This method often performs poorly in harsh environments with low signal-to-noise ratios or other severe interference. Specifically, in low signal-to-noise ratio environments, the signal quality of a single microphone array deteriorates significantly, making accurate voice separation impossible. Even if separation is achieved, the separated signal is prone to distortion, affecting subsequent wake-up and recognition processes, and ultimately leading to a significant decline in the user's voice interaction experience. Summary of the Invention

[0004] This invention provides a voice separation method, apparatus, system, electronic device, and storage medium to solve the problem of poor voice separation performance in vehicle-mounted systems under harsh conditions in the prior art, which makes it impossible to accurately separate and locate voice information and affects subsequent interaction processes. The invention improves the voice separation performance under harsh conditions, achieves accurate voice separation, and helps to improve interaction efficiency.

[0005] This invention provides a speech separation method, comprising:

[0006] Acquire speech signals from multiple registers on the target vehicle;

[0007] Based on the spectral and phase information of the speech signals from the multiple voice regions, speech separation between voice regions and speech separation within voice regions are performed on the speech signals from the multiple voice regions.

[0008] Based on the speech separation results obtained from inter-speech interval separation and the speech separation results obtained from intra-speech interval separation, the speech separation results of the multiple speech regions are determined.

[0009] According to a speech separation method provided by the present invention, the step of performing inter-tone speech separation and intra-tone speech separation on the speech signals of the multiple tone regions based on the spectral and phase information of the speech signals of the multiple tone regions includes:

[0010] Based on the spectral and phase information of the speech signals from the multiple voice regions, the speech features of the multiple voice regions are determined, and based on the speech features of the multiple voice regions, speech separation between the multiple voice regions is performed.

[0011] Each of the multiple sound regions is divided into sub-regions to obtain multiple sub-sound regions under each sound region. Based on the speech features of the multiple sub-sound regions, intra-sound separation is performed on the speech signal of the corresponding sound region.

[0012] According to a speech separation method provided by the present invention, the inter-tone separation result includes the separation result of each of the plurality of tone regions, and the intra-tone separation result includes the separation result of each of the plurality of sub-tone regions under the corresponding tone region;

[0013] The determination of speech separation results for the multiple speech regions based on the inter-regional speech separation results and the intra-regional speech separation results includes:

[0014] The separation results of the target sub-region are obtained by filtering the separation results within each region.

[0015] Based on the separation results of the target sub-sound region and the separation results of the corresponding sound region, the speech separation results of the corresponding sound region are determined.

[0016] According to a speech separation method provided by the present invention, determining the speech separation result of a corresponding speech region based on the separation result of the target sub-speech region and the separation result of the corresponding speech region of the target sub-speech region includes:

[0017] Based on the separation results of the target consonant region and the separation results of the corresponding consonant region, the target separation result of the corresponding consonant region is determined;

[0018] Narrowband enhancement is performed based on the target separation results of the corresponding sound region to obtain the speech separation results of the corresponding sound region.

[0019] According to a speech separation method provided by the present invention, the step of performing narrowband enhancement based on the target separation result of the corresponding voice region to obtain the speech separation result of the corresponding voice region includes:

[0020] Based on the target separation results and speech signals of the corresponding sound regions, determine the covariance matrix of the corresponding sound regions;

[0021] Based on the covariance matrix of the corresponding tone range, the narrowband signal of the corresponding tone range is determined;

[0022] Narrowband enhancement is performed based on the narrowband signal and target separation results of the corresponding voice region to obtain the speech separation results of the corresponding voice region.

[0023] According to a speech separation method provided by the present invention, determining the speech features of the multiple speech regions based on the spectral and phase information of the speech signals of the multiple speech regions includes:

[0024] Fixed beamforming is performed on the speech signals of the multiple voice regions to obtain fixed beams for the multiple voice regions;

[0025] Based on the spectral and phase information corresponding to the fixed beams of the multiple voice regions, the speech features of the multiple voice regions are determined.

[0026] The present invention also provides a speech separation device, comprising:

[0027] The acquisition unit is used to acquire speech signals from multiple audio zones on the target vehicle.

[0028] The separation unit is used to perform inter-tone speech separation and intra-tone speech separation on the speech signals of the multiple tone regions based on the spectral information and phase information of the speech signals of the multiple tone regions.

[0029] The fusion unit is used to determine the speech separation results of the multiple speech regions based on the speech separation results obtained from speech separation between speech regions and speech separation within speech regions.

[0030] The present invention also provides an in-vehicle interactive system, including multiple microphone arrays, a processor and a display; the multiple microphone arrays are respectively deployed in multiple sound zones on the target vehicle;

[0031] The processor is used to acquire speech signals from multiple audio regions collected by the multiple microphone arrays on the target vehicle, and to perform inter-audio speech separation and intra-audio speech separation on the speech signals from the multiple audio regions based on the spectral information and phase information of the speech signals from the multiple audio regions; to determine the speech separation results of the multiple audio regions based on the inter-audio speech separation results and the intra-audio speech separation results, and to perform voice wake-up and voice interaction based on the speech separation results of the multiple audio regions to obtain response content, and to control the display to display the response content.

[0032] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the speech separation method as described above.

[0033] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the speech separation method as described above.

[0034] The speech separation method, apparatus, system, electronic device, and storage medium provided by this invention extract the spectral and phase information of speech signals from multiple sound zones in a target vehicle. Based on this, it performs inter-sound zone speech separation and intra-sound zone speech separation on the speech signals from multiple sound zones. Based on the inter-sound zone separation results obtained from inter-sound zone speech separation and the intra-sound zone separation results obtained from intra-sound zone speech separation, it determines the speech separation results for multiple sound zones. This overcomes the shortcomings of traditional solutions, such as poor speech separation performance in harsh environments, inability to accurately separate and locate speech, and impact on subsequent interaction processes. By using multiple microphone arrays for signal acquisition, spatial information can be better obtained. On this basis, two speech separations with different precisions are performed, and the results are fused to achieve more accurate speech separation. This results in accurate and reliable speech separation results, effectively improving the wake-up rate and positioning accuracy of the in-vehicle interactive system, thereby optimizing the user's voice interaction experience. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0036] Figure 1 This is a flowchart illustrating the speech separation method provided by the present invention;

[0037] Figure 2 This is a flowchart of the overall speech separation method provided by the present invention;

[0038] Figure 3 This is a schematic diagram of the speech separation device provided by the present invention;

[0039] Figure 4 This is a schematic diagram of the structure of the in-vehicle interactive system provided by the present invention;

[0040] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0041] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0042] Currently, most in-vehicle voice separation methods rely on a single microphone array distributed throughout the vehicle. While this approach can achieve voice separation, its effectiveness is often poor in harsh environments with low signal-to-noise ratios or significant interference. Furthermore, the separated signal is prone to distortion, leading to wake-up failures and severely degrading the user's voice interaction experience.

[0043] To address this issue, the present invention provides a speech separation method that aims to acquire signals through multiple microphone arrays to better obtain spatial information. The method utilizes speech signals from multiple vocal ranges acquired by the multiple microphone arrays to perform two speech separations with different levels of precision, and then fuses the two separation results to achieve more accurate speech separation. This method can obtain accurate speech separation results, solve the problem of poor speech separation performance in in-vehicle systems under harsh conditions, and effectively improve the wake-up rate and positioning rate of in-vehicle interactive systems.

[0044] Figure 1 This is a flowchart illustrating the speech separation method provided by the present invention, as shown below. Figure 1 As shown, this method can be applied to in-vehicle interaction systems, and the specific execution entity can be the processor within the in-vehicle interaction system. The method includes:

[0045] Step 110: Acquire voice signals from multiple audio zones on the target vehicle;

[0046] Step 120: Based on the spectral and phase information of the speech signals in multiple voice regions, perform inter-voice speech separation and intra-voice speech separation on the speech signals in multiple voice regions.

[0047] Step 130: Based on the inter-speech separation results obtained from inter-speech separation and the intra-speech separation results obtained from intra-speech separation, determine the speech separation results for multiple speech regions.

[0048] Specifically, considering that current speech separation using a single array is not feasible in harsh scenarios with low signal-to-noise ratios or many interference factors, the separation effect is poor, affecting the subsequent speech interaction process and thus leading to a decline in user experience.

[0049] Based on this, in this embodiment of the invention, it is proposed that multiple microphone arrays be deployed on the target vehicle, signals be collected through microphone arrays with multiple sound zones, and the collected speech signals from multiple sound zones be used for two speech separations. The separation results are then fused, which effectively improves the speech separation effect and can obtain accurate and reliable speech separation results. This can improve the wake-up rate of the voice interaction process and thus enhance the user's voice interaction experience.

[0050] Understandably, in practical applications, before performing speech separation, it is necessary to first determine the speech signal to be separated. Considering the problems existing in current speech separation based on a single array, this embodiment of the invention deploys multiple microphone arrays to perform speech separation using the speech signals corresponding to the multiple microphone arrays. Specifically, this can be achieved by pre-deploying microphone arrays in multiple sound zones on the target vehicle, collecting speech signals from inside the vehicle through these arrays, thus obtaining speech signals from multiple sound zones on the target vehicle. It should be noted that because vehicles are typically not soundproof, and sound zones are not completely isolated, the speech signals collected by the microphone arrays in each sound zone do not perfectly correspond to their respective sound zones. They often contain sound signals from other sound zones, such as the voice signals of speakers in other sound zones, noise signals, etc. It is precisely for this reason that this embodiment of the invention requires speech separation of the collected speech signals to obtain accurate separated signals for each sound zone (excluding signals from other sound zones).

[0051] Here, the audio zones in the target vehicle can be divided according to seat positions. For example, for a typical four-seater sedan, it can be divided into a driver's seat zone, a passenger's seat zone, a driver's seat rear zone, and a passenger's seat rear zone. For five-seater and seven-seater sedans, it can be divided into a driver's seat zone, a passenger's seat zone, a center zone, a left rear zone, and a right rear zone. Alternatively, it can be divided according to the vehicle's structure, interior environment, layout, etc., and this embodiment of the invention does not impose specific limitations on this. One microphone array or multiple microphone arrays can be deployed within each audio zone, and this embodiment of the invention does not impose specific limitations. Preferably, to save costs, improve processing efficiency, and ensure separation effectiveness, this embodiment of the invention deploys one microphone array within each audio zone to collect signals and obtain the speech signal for each zone.

[0052] Compared to traditional arrays, the distributed microphone array (a collection of multiple microphone arrays) used in this embodiment of the invention is not constrained by array configuration. Not only can each array independently process the received audio signal, but the overall performance will not degrade due to the failure of one or a few array elements. Using a distributed microphone array for speech separation, due to the participation of multiple arrays, enables fast and accurate speech separation and sound source localization even in harsh environments.

[0053] Furthermore, after obtaining speech signals from multiple voice regions, speech separation can be performed based on these signals in this embodiment of the invention to obtain separation results. That is, speech separation is performed on the speech signals from multiple voice regions using the spectral and phase information of the speech signals from multiple voice regions, thereby obtaining separation results for multiple voice regions. To achieve more accurate speech separation and obtain more reliable speech separation results, the speech separation process here can be divided into two levels: inter-voice speech separation and intra-voice speech separation. Inter-voice speech separation involves separating speech between different voice regions to obtain inter-voice separation results; intra-voice speech separation involves separating speech within each voice region to obtain intra-voice separation results.

[0054] Specifically, this can begin by extracting information from the speech signal in each vocal range to obtain information about the spectrum and phase, such as logarithmic power spectrum and phase difference. This yields the spectral and phase information of the speech signal in each vocal range. Next, speech separation can be performed based on this spectral and phase information. Specifically, using the spectral and phase information of the speech signals from multiple vocal ranges, inter-vocal speech separation and intra-vocal speech separation are performed separately to separate the signals from different vocal ranges, thus obtaining the separated signals for each vocal range, i.e., the inter-vocal separation results. Simultaneously, signals from different regions / positions within each vocal range are separated, thus obtaining the separated signals for each region / position within each vocal range, i.e., the intra-vocal separation results.

[0055] Following this, the final separation result can be determined based on the inter-regional speech separation results and the intra-regional speech separation results. That is, based on the inter-regional and intra-regional speech separation results obtained from the inter-regional and intra-regional speech separation, signal fusion is performed to obtain the final separation result for multiple regions. This separated signal is the final speech separation result for multiple regions. Specifically, this involves fusing the inter-regional and intra-regional separation results obtained from the two speech separation processes. By fusing the separated signals, more accurate speech separation can be achieved, ultimately yielding speech separation results for multiple regions.

[0056] Furthermore, after obtaining the speech separation results for multiple voice zones, in this embodiment of the invention, in-vehicle voice interaction can be performed based on these speech separation results, and a response output from the in-vehicle interaction system can be obtained. That is, voice wake-up can be performed based on the speech separation results for multiple voice zones, and sound source localization can be performed based on energy, thereby enabling targeted voice interaction to respond to the input speech, generate response information, and output and display it, so that the corresponding personnel in the target vehicle can be promptly informed of the system's response information. This achieves voice interaction in the target vehicle, improves the wake-up rate of the system backend, and optimizes the user's voice interaction experience.

[0057] The following example uses a target vehicle equipped with a distributed array consisting of four microphone arrays to illustrate the speech separation process of its speech signals:

[0058] The target vehicle is equipped with four microphone arrays, positioned above the four seats (driver's seat, front passenger seat, rear of the driver's seat, and rear of the front passenger seat). The distributed microphone arrays can capture multi-channel speech signals, i.e., speech signals from multiple frequency ranges. Speech separation can then be performed using these multiple frequency range signals. Specifically, inter-frequency range speech separation is first performed based on the multi-frequency range signals, while intra-frequency range speech separation is performed using the speech signals acquired by a single array. The inter-frequency range separation results and intra-frequency range separation results are then fused to obtain accurate and reliable separation signals for multiple frequency ranges. This enables accurate speech separation even in harsh environments, extracting accurate separation signals for each frequency range from complex in-vehicle environments for subsequent voice wake-up and sound source localization, thus improving in-vehicle voice interaction.

[0059] The speech separation method provided by this invention extracts the spectral and phase information of speech signals from multiple sound zones in a target vehicle. Based on this, it performs inter-sound zone speech separation and intra-sound zone speech separation. The speech separation results for multiple sound zones are determined based on the inter-sound zone separation results and the intra-sound zone separation results. This overcomes the shortcomings of traditional solutions, such as poor speech separation performance in harsh environments, inaccurate speech separation and localization, and impact on subsequent interaction processes. By using multiple microphone arrays for signal acquisition, spatial information can be better obtained. Based on this, two speech separations with different precisions are performed, and the results are fused to achieve more accurate speech separation. This results in accurate and reliable speech separation results, effectively improving the wake-up rate and localization accuracy of the in-vehicle interactive system, thereby optimizing the user's voice interaction experience.

[0060] Based on the above embodiments, step 120 includes:

[0061] Based on the spectral and phase information of speech signals from multiple vocal regions, the speech features of multiple vocal regions are determined, and based on the speech features of multiple vocal regions, inter-vocal speech separation is performed on the speech signals from multiple vocal regions.

[0062] Each of the multiple sound regions is divided into sub-regions to obtain multiple sub-sound regions under each sound region. Based on the speech features of multiple sub-sound regions, intra-sound separation is performed on the speech signal of the corresponding sound region.

[0063] Specifically, in step 120, the process of performing inter-tone speech separation and intra-tone speech separation on the speech signals of multiple tone regions based on the spectral and phase information of the speech signals of multiple tone regions may include:

[0064] First, speech features of multiple speech regions can be determined based on the spectral and phase information of the speech signals from multiple regions. Specifically, the spectral and phase information of the speech signal can be directly used as the speech features of the corresponding region, or the spectral and phase information can be further encoded, fused, or processed, and the processed features can be used as speech features. This embodiment of the invention does not impose specific limitations on this. Next, the speech features of multiple regions can be used to perform inter-regional speech separation on the speech signals of multiple regions to separate signals from different regions in the mixed signal, thereby obtaining accurate separated signals in each region that do not contain signals from other regions. These separated signals are the inter-regional speech separation results.

[0065] Simultaneously, each pitch range can be subdivided into multiple sub-regions, resulting in multiple sub-regions within each pitch range. Here, the rules for subdividing different pitch ranges and the number of sub-regions obtained can be the same or different. For example, the driver's and passenger's pitch ranges can be divided into left and right sub-regions, while the rear driver's and passenger's pitch ranges can be divided into left and right sub-regions, or they can be divided into three sub-regions (left, center, and right). This embodiment of the invention does not specifically limit this division.

[0066] After completing the division of the sound regions and obtaining multiple sub-sound regions under each sound region, in this embodiment of the invention, speech separation can be performed on the speech signal of the corresponding sound region based on the speech characteristics of each sub-sound region. That is, intra-sound region speech separation is performed to separate the signals of different sub-sound regions within that sound region, thereby obtaining the separated signal of each sub-sound region. This separated signal is the intra-sound region separation result. The speech characteristics of each sub-sound region can be determined by the spectral information and phase information of the speech signal collected by the array elements (array elements in the microphone array) within each sub-sound region. The specific process is basically the same as the determination of the speech characteristics of the sound region.

[0067] It should be noted that both inter-regional and intra-regional speech separation processes can be implemented using models. Specifically, speech features from multiple regions can be input into an inter-regional separation model, allowing the model to separate speech signals from multiple regions based on the input features and output the inter-regional separation result. Similarly, speech features from multiple sub-regions can be input into an intra-regional separation model, allowing the model to separate speech signals from the corresponding regions based on the input features and output the intra-regional separation result.

[0068] It is worth noting that, before inputting the speech features into the model, this embodiment of the invention also requires pre-training of the interval separation model and the intra-interval separation model. That is, the interval separation model and the intra-interval separation model can be pre-trained using the sample speech signal and the corresponding sample separation results.

[0069] Specifically, this can involve pre-collecting sample speech signals for training. These sample speech signals can be actual speech signals recorded in a vehicle-mounted scenario or simulated signals; this embodiment of the invention does not impose specific limitations on this. If simulated signals are used as sample speech, the image method can be used to simulate and generate room impact responses in different areas of the vehicle beforehand, and then convolve these responses with a clean speech signal. Noise signals are then added according to different signal-to-noise ratios to generate noisy signals, which are the simulated sample speech signals.

[0070] Here, the frequency domain representation of the noisy signal is:

[0071]

[0072] In the formula, It is a noisy signal. The signal after convolution processing. This is a noise signal. For frequency point subscripts, Indicates a time frame.

[0073] When adding noise signals, the signal-to-noise ratio of the noise signals should cover the actual vehicle noise scenario as much as possible.

[0074] After obtaining the sample speech signals, in this embodiment of the invention, information extraction is performed based on the sample speech signals collected by multiple arrays during recording to obtain the spectral and phase information of the sample speech signals in multiple sample voice regions. Based on this, the sample speech features are determined. Then, the sample speech features of multiple sample voice regions are input into an initial voice region separation model to perform speech separation and output predicted voice region separation results, i.e., the speech masks of each sample voice region. Subsequently, the initial voice region separation model can be trained based on the sample voice region separation results and the predicted voice region separation results to learn the mapping relationship between input and output, ultimately resulting in a trained voice region separation model.

[0075] The sample speech interval separation result can be calculated using the following formula after determining the sample speech signal:

[0076]

[0077] In the formula, The time-frequency mask for each sample tone region, i.e., the result of sample tone region separation, is calculated by taking the amplitude spectrum of Y and S. exist The proportion it occupies is obtained.

[0078] Correspondingly, the training of the intra-regional separation model is basically the same as that of the inter-regional separation model. The sample speech signals used for training can be either real-world recordings or simulated signals. During training, the sample speech features of multiple sample sub-regions are used as input, and the speech masks (predicted intra-regional separation results) of multiple sample sub-regions are used as output. Based on the predicted output and the corresponding labels of the samples, the initial intra-regional separation model is trained to obtain the intra-regional separation model.

[0079] Here, the loss function used during training of the pitch interval separation model and the pitch interval separation model is the mean squared error loss function, and the learning rate is set to 0.001. The initial pitch interval separation model and the initial pitch interval separation model can be built on the basis of a convolutional neural network or a recurrent neural network.

[0080] In this embodiment of the invention, when performing speech separation, speech separation between pitch regions is first performed based on the speech characteristics of each pitch region, then speech signal within each single array is separated within the pitch region, and finally the two separation results are fused. Through more detailed region division, more accurate speech separation is achieved, which helps to improve the subsequent positioning effect and optimize the interactive experience.

[0081] Based on the above embodiments, speech features of multiple speech regions are determined based on the spectral and phase information of speech signals from multiple regions, including:

[0082] Fixed beamforming is performed on speech signals from multiple vocal registers to obtain fixed beams for multiple vocal registers;

[0083] Based on the spectral and phase information corresponding to fixed beams in multiple voice regions, the speech features of multiple voice regions are determined.

[0084] Specifically, the process of determining the speech features of multiple speech regions using the spectral and phase information of speech signals from multiple regions can include:

[0085] To improve the signal-to-noise ratio and enhance the speech signal, in this embodiment of the invention, after acquiring speech signals from multiple vocal registers, beamforming can be performed on the speech signals from each vocal register to increase the proportion of effective components in the signal, thereby obtaining an enhanced signal. Specifically, this can involve applying fixed beamforming to the speech signals from each vocal register, thereby improving the signal-to-noise ratio.

[0086] Furthermore, the spectral and phase information of the fixed beams for each audio region can be calculated. That is, after applying fixed beams to the speech signals of each region, the spectral information of the fixed beams for each region is calculated. This spectral information can be the logarithmic power spectrum. For example, when the target vehicle has four audio regions, the logarithmic power spectra of the four regions are calculated and denoted as LPS_fix1, LPS_fix2, LPS_fix3, and LPS_fix4, respectively. Simultaneously, the phase information, i.e., the phase difference between adjacent microphones within each region, is calculated and denoted as IPD (Inter-channel Phase Difference). Then, the speech characteristics of each region can be determined based on this logarithmic power spectrum and phase difference. Preferably, in this embodiment of the invention, the logarithmic power spectrum and phase difference are directly used as the speech features.

[0087] Based on the above embodiments, the pitch interval separation result includes the separation result of each of the multiple pitch regions, and the pitch region separation result includes the separation result of each of the multiple sub-pitch regions under the corresponding pitch region;

[0088] Step 130 includes:

[0089] The separation results of the target sub-region are obtained by filtering the separation results within each region.

[0090] Based on the separation results of the target sub-sound region and the separation results of the corresponding sound region, the speech separation results of the corresponding sound region are determined.

[0091] Specifically, step 130, the process of determining the speech separation results for multiple speech regions based on the inter-regional speech separation results and the intra-regional speech separation results, may include:

[0092] After two speech separation operations, we can obtain the separation signals for each vocal region and the separation signals for each sub-vocal region within each vocal region. This results in the separation of multiple vocal regions and the separation results for each sub-vocal region within each vocal region. Then, we can fuse the two separation results. Specifically, we can first select the largest separation signal from the separation results within each vocal region, and denot this as the separation result for the target sub-vocal region within that vocal region. For example, if there are four vocal regions, each with three sub-vocal regions, inter-vocal speech separation can yield the separation signals for each vocal region, denoted as... Speech separation within a vocal range yields the separated signals for each sub-vocal range, denoted as... Then for each It can be seen from its three corresponding Select the largest one from the list, and denote it as... That is, the separation result of the target subphone region.

[0093] Next, the separation result of the target sub-region can be compared with the separation result of the corresponding region to select the largest separation signal as the final separation signal for the corresponding region. That is, the separation result of the target sub-region can be compared with the separation result of the corresponding region. With the corresponding The two values ​​are compared, and the maximum value is taken as the final separation signal for the corresponding vocal region. This method yields the final separation signal for each vocal region, i.e., the speech separation result for multiple vocal regions.

[0094] Based on the above embodiments, the speech separation result of the corresponding speech region is determined based on the separation result of the target sub-speech region and the separation result of the corresponding speech region of the target sub-speech region, including:

[0095] Based on the separation results of the target consonant region and the separation results of the corresponding region, the target separation result of the corresponding region is determined.

[0096] Narrowband enhancement is performed based on the target separation results of the corresponding sound region to obtain the speech separation results of the corresponding sound region.

[0097] Specifically, the process of determining the speech separation result of the corresponding speech region based on the separation result of the target subphone region and the separation result of the corresponding speech region of the target subphone region includes:

[0098] Considering the poor performance of speech separation in harsh environments, and the tendency for signals obtained through forced separation to be distorted and ineffective for voice wake-up, this embodiment of the invention, in order to further improve the quality of the signal after speech separation, thereby increasing the wake-up rate of the backend and optimizing the user experience, firstly, when fusing the results of the two speech separations to obtain the final speech separation result, the result fusion can be performed. The fusion process has been described in detail above and will not be repeated here. Next, the target separation result obtained by fusion is enhanced to strengthen the separated signal, thus obtaining the final separated signal, i.e., the speech separation result.

[0099] In detail, this can be achieved by using an adaptive narrowband enhancement method to enhance the signal of the target separation result. Specifically, narrowband enhancement is applied to the target separation result for the corresponding voice region, resulting in an enhanced target separation result for that region. This enhanced result is the speech separation result for that voice region. Afterward, backend wake-up and localization can be performed based on the speech separation result, significantly improving the wake-up rate and thus optimizing the user's voice interaction experience.

[0100] Based on the above embodiments, narrowband enhancement is performed based on the target separation results of the corresponding sound region to obtain the speech separation results of the corresponding sound region, including:

[0101] Based on the target separation results and speech signals of the corresponding sound regions, determine the covariance matrix of the corresponding sound regions;

[0102] Based on the covariance matrix of the corresponding tone range, the narrowband signal of the corresponding tone range is determined;

[0103] Narrowband enhancement is performed based on the narrowband signal and target separation results of the corresponding voice region to obtain the speech separation results of the corresponding voice region.

[0104] Specifically, the process of performing narrowband enhancement based on the target separation results of the corresponding sound region to obtain the speech separation results of the corresponding sound region may include:

[0105] First, the parameters of the beamformer for narrowband enhancement can be calculated based on the target separation results of the corresponding voice region. Specifically, the covariance matrix of the beamformer can be updated using the target separation results. Here, the covariance matrix includes the speech covariance matrix and the noise covariance matrix.

[0106] Then, the narrowband signal for the corresponding tone range can be determined based on the updated covariance matrix. Specifically, this involves applying a beamformer with updated parameters, using the target separation results for the corresponding tone range as guidance to determine the narrowband signal. The beamformer here can be understood as any adaptive beamforming algorithm.

[0107] Then, based on the narrowband signal of the corresponding pitch region and the target separation result of the corresponding pitch region, the speech separation result of the corresponding pitch region can be obtained. Specifically, after passing through the adaptive narrowband beam, the narrowband signal of the corresponding pitch region can be obtained. Multiplying the narrowband signal by the corresponding separation signal yields the final separation signal.

[0108] Based on the above embodiments, the speech separation result of the corresponding audio region can be expressed by the following formula:

[0109]

[0110] In the formula, For the first Speech separation results for each vocal region For the first Target separation results for each sound region For the first Narrow-band signal in each frequency range.

[0111] The speech covariance matrix and noise covariance matrix are updated as follows:

[0112]

[0113]

[0114] In the formula, The speech covariance matrix, The noise covariance matrix is... For the first Speech signals in each vocal register This is the conjugate transpose.

[0115] In this embodiment of the invention, fixed beamforming is first performed on the speech signal of each voice region to improve the signal-to-noise ratio. The enhanced speech signal is then used for inter-voice speech separation and intra-voice speech separation. The two separation results are then fused, and a narrow beam is used to enhance the target separation result. This can greatly improve the accuracy and precision of speech separation, and ensure that the separated signal is not easily distorted. It can effectively wake up the backend, thereby improving the wake-up rate and positioning rate of the backend, and greatly optimizing the user's voice interaction experience.

[0116] Based on the above embodiments, Figure 2 This is a flowchart of the overall speech separation method provided by the present invention, as follows: Figure 2 As shown, the method includes:

[0117] First, acquire voice signals from multiple audio zones on the target vehicle.

[0118] Subsequently, fixed beamforming is performed on the speech signals of multiple voice regions to obtain fixed beams for multiple voice regions. Based on the spectral and phase information corresponding to the fixed beams of multiple voice regions, the speech features of multiple voice regions are determined.

[0119] Subsequently, based on the speech features of multiple sound regions, speech separation between sound regions is performed on the speech signals of multiple sound regions, and each sound region in multiple sound regions is divided into sub-regions to obtain multiple sub-sound regions under each sound region. Based on the speech features of multiple sub-sound regions, speech separation within the corresponding sound region is performed on the speech signals of the corresponding sound region.

[0120] Next, from the separation results within each pitch region, the separation results of the target sub-pitch region are obtained. Based on the separation results of the target sub-pitch region and the separation results of the corresponding pitch region, the target separation result for the corresponding pitch region is determined. The inter-pitch separation results include the separation results of multiple pitch regions individually, and the intra-pitch separation results include the separation results of multiple sub-pitch regions within the corresponding pitch region.

[0121] Finally, based on the target separation results and speech signals of the corresponding sound regions, the covariance matrix of the corresponding sound regions is determined; based on the covariance matrix of the corresponding sound regions, the narrowband signal of the corresponding sound regions is determined; based on the narrowband signal of the corresponding sound regions and the target separation results, narrowband enhancement is performed to obtain the speech separation results of the corresponding sound regions.

[0122] The method provided in this invention extracts the spectral and phase information of speech signals from multiple sound zones on a target vehicle. Based on this, it performs inter-sound zone speech separation and intra-sound zone speech separation on the speech signals from multiple sound zones. Based on the inter-sound zone separation results obtained from inter-sound zone speech separation and the intra-sound zone separation results obtained from intra-sound zone speech separation, the speech separation results of multiple sound zones are determined. This overcomes the shortcomings of traditional solutions, such as poor in-vehicle speech separation performance in harsh scenarios, inability to accurately separate and locate speech, and impact on subsequent interaction processes. By acquiring signals through multiple microphone arrays, spatial information can be better obtained. On this basis, two speech separations with different precisions are performed, and the results are fused to achieve more accurate speech separation. This results in accurate and reliable speech separation results, effectively improving the wake-up rate and positioning accuracy of the in-vehicle interactive system, thereby optimizing the user's voice interaction experience.

[0123] The speech separation device provided by the present invention is described below. The speech separation device described below and the speech separation method described above can be referred to in correspondence.

[0124] Figure 3 This is a schematic diagram of the speech separation device provided by the present invention, as shown below. Figure 3 As shown, the device includes:

[0125] Acquisition unit 310 is used to acquire speech signals from multiple audio zones on the target vehicle;

[0126] The separation unit 320 is used to perform inter-tone speech separation and intra-tone speech separation on the speech signals of the multiple tone regions based on the spectral information and phase information of the speech signals of the multiple tone regions.

[0127] The fusion unit 330 is used to determine the speech separation results of the multiple speech regions based on the speech separation results obtained from speech separation between speech regions and the speech separation results obtained from speech separation within speech regions.

[0128] The speech separation device provided by this invention extracts the spectral and phase information of speech signals from multiple sound zones in a target vehicle. Based on this, it performs inter-sound zone speech separation and intra-sound zone speech separation on the speech signals from multiple sound zones. Based on the inter-sound zone separation results obtained from inter-sound zone speech separation and the intra-sound zone separation results obtained from intra-sound zone speech separation, the speech separation results of multiple sound zones are determined. This overcomes the shortcomings of traditional solutions, such as poor in-vehicle speech separation performance in harsh scenarios, inability to accurately separate and locate speech, and impact on subsequent interaction processes. By acquiring signals through multiple microphone arrays, spatial information can be better obtained. On this basis, two speech separations with different precisions are performed, and the results are fused to achieve more accurate speech separation. This results in accurate and reliable speech separation results, effectively improving the wake-up rate and positioning accuracy of the in-vehicle interactive system, thereby optimizing the user's voice interaction experience.

[0129] Based on the above embodiments, the separation unit 320 is used for:

[0130] Based on the spectral and phase information of the speech signals from the multiple voice regions, the speech features of the multiple voice regions are determined, and based on the speech features of the multiple voice regions, speech separation between the multiple voice regions is performed.

[0131] Each of the multiple sound regions is divided into sub-regions to obtain multiple sub-sound regions under each sound region. Based on the speech features of the multiple sub-sound regions, intra-sound separation is performed on the speech signal of the corresponding sound region.

[0132] Based on the above embodiments, the pitch interval separation result includes the separation result of each of the plurality of pitch regions, and the pitch region intra-region separation result includes the separation result of each of the plurality of sub-pitch regions under the corresponding pitch region; the fusion unit 330 is used for:

[0133] The separation results of the target sub-region are obtained by filtering the separation results within each region.

[0134] Based on the separation results of the target sub-sound region and the separation results of the corresponding sound region, the speech separation results of the corresponding sound region are determined.

[0135] Based on the above embodiments, the fusion unit 330 is used for:

[0136] Based on the separation results of the target consonant region and the separation results of the corresponding consonant region, the target separation result of the corresponding consonant region is determined;

[0137] Narrowband enhancement is performed based on the target separation results of the corresponding sound region to obtain the speech separation results of the corresponding sound region.

[0138] Based on the above embodiments, the fusion unit 330 is used for:

[0139] Based on the target separation results and speech signals of the corresponding sound regions, determine the covariance matrix of the corresponding sound regions;

[0140] Based on the covariance matrix of the corresponding tone range, the narrowband signal of the corresponding tone range is determined;

[0141] Narrowband enhancement is performed based on the narrowband signal and target separation results of the corresponding voice region to obtain the speech separation results of the corresponding voice region.

[0142] Based on the above embodiments, the separation unit 320 is used for:

[0143] Fixed beamforming is performed on the speech signals of the multiple voice regions to obtain fixed beams for the multiple voice regions;

[0144] Based on the spectral and phase information corresponding to the fixed beams of the multiple voice regions, the speech features of the multiple voice regions are determined.

[0145] Figure 4 This is a structural schematic diagram of the in-vehicle interactive system provided by the present invention, as shown below. Figure 4 As shown, the system includes multiple microphone arrays 410, a processor 420, and a display 430; the multiple microphone arrays 410 are respectively deployed in multiple sound zones on the target vehicle;

[0146] The processor 420 is used to acquire speech signals from multiple sound zones collected by the multiple microphone arrays 410 on the target vehicle, and to perform inter-sound zone speech separation and intra-sound zone speech separation based on the spectral and phase information of the speech signals from the multiple sound zones; to determine the speech separation results of the multiple sound zones based on the inter-sound zone separation results and the intra-sound zone separation results, and to perform voice wake-up and voice interaction based on the speech separation results of the multiple sound zones to obtain response content, and to control the display 430 to display the response content.

[0147] The in-vehicle interactive system provided by this invention includes multiple microphone arrays, a processor, and a display. The processor acquires speech signals from multiple microphone arrays on a target vehicle across multiple sound zones, and performs inter-sound zone speech separation and intra-sound zone speech separation based on the spectral and phase information of the speech signals from multiple sound zones. Based on the inter-sound zone separation results and intra-sound zone separation results, the system determines the speech separation results for multiple sound zones, performs voice wake-up and voice interaction based on the speech separation results for multiple sound zones, obtains the response content, and controls the display to show the response content. This overcomes the shortcomings of traditional solutions, such as poor in-vehicle speech separation performance in adverse scenarios, inability to accurately separate and locate speech, and impact on subsequent interaction processes. It achieves more accurate speech separation, obtains accurate and reliable speech separation results, effectively improves the wake-up rate and positioning accuracy of the in-vehicle interactive system, and thus optimizes the user's voice interaction experience.

[0148] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device may include a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, communications interface 520, and memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a speech separation method, which includes: acquiring speech signals from multiple sound zones on a target vehicle; performing inter-sound zone speech separation and intra-sound zone speech separation on the speech signals from the multiple sound zones based on the spectral and phase information of the speech signals from the multiple sound zones; and determining the speech separation results of the multiple sound zones based on the inter-sound zone separation results obtained from the inter-sound zone speech separation and the intra-sound zone separation results obtained from the intra-sound zone speech separation.

[0149] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0150] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer is able to execute the speech separation method provided by the above methods, the method comprising: acquiring speech signals of multiple voice zones on a target vehicle; performing inter-voice speech separation and intra-voice speech separation on the speech signals of the multiple voice zones based on the spectral information and phase information of the speech signals of the multiple voice zones; and determining the speech separation result of the multiple voice zones based on the inter-voice separation result obtained by inter-voice speech separation and the intra-voice separation result obtained by intra-voice speech separation.

[0151] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the speech separation method provided by the above methods. The method includes: acquiring speech signals from multiple sound zones on a target vehicle; performing inter-sound zone speech separation and intra-sound zone speech separation on the speech signals from the multiple sound zones based on the spectral information and phase information of the speech signals from the multiple sound zones; and determining the speech separation result of the multiple sound zones based on the inter-sound zone separation result obtained by inter-sound zone speech separation and the intra-sound zone separation result obtained by intra-sound zone speech separation.

[0152] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0153] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A speech separation method, characterized in that, include: Acquire speech signals from multiple registers on the target vehicle; Based on the spectral and phase information of the speech signals from the multiple voice regions, speech separation between voice regions and speech separation within voice regions are performed on the speech signals from the multiple voice regions. Based on the inter-tone speech separation results obtained from inter-tone speech separation and the intra-tone speech separation results obtained from intra-tone speech separation, the speech separation results of the multiple tone regions are determined; the inter-tone speech separation is based on the speech signals of multiple tone regions collected by multiple microphone arrays on the target vehicle, and the intra-tone speech separation is based on the speech signals collected by a single microphone array in the corresponding tone region; The interval separation result includes the separation result of each of the multiple intervals, and the interval separation result includes the separation result of each of the multiple sub-intervals under the corresponding interval; The determination of speech separation results for the multiple speech regions based on the inter-regional speech separation results and the intra-regional speech separation results includes: The separation results of the target sub-region are obtained by filtering the separation results within each region. Based on the separation results of the target sub-sound region and the separation results of the corresponding sound region, the speech separation results of the corresponding sound region are determined.

2. The speech separation method according to claim 1, characterized in that, The process of performing inter-tone speech separation and intra-tone speech separation on the speech signals of the multiple tone regions based on the spectral and phase information of the speech signals of the multiple tone regions includes: Based on the spectral and phase information of the speech signals from the multiple voice regions, the speech features of the multiple voice regions are determined, and based on the speech features of the multiple voice regions, speech separation between the multiple voice regions is performed. Each of the multiple sound regions is divided into sub-regions to obtain multiple sub-sound regions under each sound region. Based on the speech features of the multiple sub-sound regions, intra-sound separation is performed on the speech signal of the corresponding sound region.

3. The speech separation method according to claim 1, characterized in that, The step of determining the speech separation result of the corresponding speech region based on the separation result of the target sub-speech region and the separation result of the corresponding speech region of the target sub-speech region includes: Based on the separation results of the target consonant region and the separation results of the corresponding consonant region, the target separation result of the corresponding consonant region is determined; Narrowband enhancement is performed based on the target separation results of the corresponding sound region to obtain the speech separation results of the corresponding sound region.

4. The speech separation method according to claim 3, characterized in that, The narrowband enhancement based on the target separation results of the corresponding voice region, to obtain the speech separation results of the corresponding voice region, includes: Based on the target separation results and speech signals of the corresponding sound regions, determine the covariance matrix of the corresponding sound regions; Based on the covariance matrix of the corresponding tone range, the narrowband signal of the corresponding tone range is determined; Narrowband enhancement is performed based on the narrowband signal and target separation results of the corresponding voice region to obtain the speech separation results of the corresponding voice region.

5. The speech separation method according to any one of claims 2 to 4, characterized in that, The determination of speech features of the multiple speech regions based on the spectral and phase information of the speech signals from the multiple speech regions includes: Fixed beamforming is performed on the speech signals of the multiple voice regions to obtain fixed beams for the multiple voice regions; Based on the spectral and phase information corresponding to the fixed beams of the multiple voice regions, the speech features of the multiple voice regions are determined.

6. A speech separation device, characterized in that, include: The acquisition unit is used to acquire speech signals from multiple audio zones on the target vehicle. The separation unit is used to perform inter-tone speech separation and intra-tone speech separation on the speech signals of the multiple tone regions based on the spectral information and phase information of the speech signals of the multiple tone regions. The fusion unit is used to determine the speech separation results of the multiple sound regions based on the inter-speech separation results obtained from inter-speech separation and the intra-speech separation results obtained from intra-speech separation; the inter-speech separation is based on the speech signals of multiple sound regions collected by multiple microphone arrays on the target vehicle, and the intra-speech separation is based on the speech signals collected by a single microphone array in the corresponding sound region; The interval separation result includes the separation result of each of the multiple intervals, and the interval separation result includes the separation result of each of the multiple sub-intervals under the corresponding interval; The determination of speech separation results for the multiple speech regions based on the inter-regional speech separation results and the intra-regional speech separation results includes: The separation results of the target sub-region are obtained by filtering the separation results within each region. Based on the separation results of the target sub-sound region and the separation results of the corresponding sound region, the speech separation results of the corresponding sound region are determined.

7. An in-vehicle interactive system, characterized in that, It includes multiple microphone arrays, a processor, and a display; the multiple microphone arrays are respectively deployed in multiple sound zones on the target vehicle; The processor is used to acquire speech signals from multiple sound zones collected by the multiple microphone arrays on the target vehicle, and to perform inter-sound zone speech separation and intra-sound zone speech separation based on the spectral and phase information of the speech signals from the multiple sound zones, obtaining inter-sound zone separation results and intra-sound zone separation results; the inter-sound zone separation results include the separation results of each of the multiple sound zones, and the intra-sound zone separation results include the separation results of each of the multiple sub-sound zones under the corresponding sound zone; from the intra-sound zone separation results of each sound zone, the separation results of the target sub-sound zone are selected; based on the separation results of the target sub-sound zone and the separation results of the corresponding sound zone of the target sub-sound zone, the speech separation results of the corresponding sound zone are determined, and voice wake-up and voice interaction are performed based on the speech separation results of the multiple sound zones to obtain response content, and the display is controlled to display the response content; The inter-tone speech separation is based on speech signals from multiple microphone arrays on the target vehicle, while the intra-tone speech separation is based on speech signals from a single microphone array in the corresponding tone region.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the speech separation method as described in any one of claims 1 to 5.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the speech separation method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and device for positioning sound area, storage medium and electronic equipment

    CN113380267A

  • Vehicle control method, vehicle and storage medium

    CN116364071A