Speech separation method and apparatus, system, device, storage medium, and program product

By extracting speech features and directional features independent of microphone array configuration, the problem of requiring customized separation models in traditional in-vehicle interaction systems is solved, achieving complete sound isolation and independent interaction in the vehicle, improving system delivery efficiency and reducing costs.

WO2026107884A1PCT designated stage Publication Date: 2026-05-28IFLYTEK CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2024-12-09
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Traditional in-vehicle interactive systems require customized separate models for each vehicle model due to differences in microphone array configurations. This results in low efficiency and high costs, and each time the vehicle model changes, customization is required again, affecting the efficiency of large-scale delivery.

Method used

Speech features and directional features independent of microphone array configuration are extracted, and speech separation is performed based on these features. A speech separation model is then used for speech separation, and the training model does not depend on array configuration information.

Benefits of technology

It achieves complete isolation of sound in the vehicle, ensuring independent interaction between different sound zones, improving the efficiency of large-scale system delivery, reducing costs, and minimizing resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024137921_28052026_PF_FP_ABST
    Figure CN2024137921_28052026_PF_FP_ABST
Patent Text Reader

Abstract

A speech separation method and apparatus, a system, a device, a storage medium, and a program product. The method comprises: acquiring a speech signal from a target vehicle, wherein the speech signal is collected by a microphone array mounted on the target vehicle (110); extracting a speech feature and a directional feature of the speech signal, wherein both the speech feature and the directional feature are independent of the configuration of the microphone array (120); and performing speech separation on the basis of the speech feature and the directional feature so as to obtain a speech separation result corresponding to the speech signal (130). The complete isolation of sound in the vehicle can be realized without depending on the configuration information of the microphone array, so as to provide a guarantee for the independent interaction of speech zones, thereby avoiding the problem, in conventional solutions, that speech separation in vehicles relying on a separation model obtained by training configuration information, leading to "one-model-per-vehicle" limitation. The speech separation is performed on the basis of the configuration-independent features , which improves the large-scale delivery efficiency of the system, reduces costs, and reduces the waste of resources.
Need to check novelty before this filing date? Find Prior Art

Description

Speech separation methods, apparatuses, systems, devices, storage media, and software products

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 2024116519093, filed on November 19, 2024, entitled “Speech Separation Method, Apparatus, System, Device, Storage Medium and Program Product”, which is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure relates to the field of artificial intelligence technology, and in particular to a speech separation method, apparatus, system, device, storage medium, and program product. Background Technology

[0004] With the development of automotive intelligence, in-vehicle intelligent voice interaction systems have gradually become an important part of modern automobiles, providing drivers and passengers with a more convenient and intelligent interactive experience. This system controls various in-vehicle functions by recognizing passengers' wake words and commands. However, in practical applications, due to differences in vehicle size and interior layout, the microphone array configurations of different car models vary, posing a challenge to the large-scale delivery of in-vehicle voice interaction systems.

[0005] Traditional in-vehicle interaction systems mostly rely on microphone array layout information for voice separation to achieve sound isolation and avoid interference between different sound zones. However, this approach requires customizing a separation model for each vehicle model. This "one-vehicle-one-customization" approach is not only inefficient but also costly. Whenever a new model is launched, or even if there are minor changes in the microphone array layout of the same model, a new separation model needs to be customized, consuming a lot of time and resources and reducing the efficiency of large-scale delivery. Summary of the Invention

[0006] This disclosure provides a voice separation method, apparatus, system, device, storage medium, and program product to solve the problem in related technologies where in-vehicle interactive systems are prone to falling into the "one vehicle, one customization" dilemma when achieving sound isolation so that each sound zone is independent and does not interfere with each other.

[0007] This disclosure provides a speech separation method, including:

[0008] Acquire voice signals from the target vehicle, which are obtained through a microphone array mounted on the target vehicle;

[0009] The speech features and directional features of the speech signal are extracted, and both the speech features and the directional features are independent of the array configuration of the microphone array;

[0010] Speech separation is performed based on the speech features and the directional features to obtain the speech separation result corresponding to the speech signal.

[0011] According to a speech separation method provided in this disclosure, the step of performing speech separation based on the speech features and the directional features to obtain the speech separation result corresponding to the speech signal includes:

[0012] Based on the speech features and the directional features, a speech separation model is applied to perform speech separation, and the speech separation result corresponding to the speech signal is obtained.

[0013] The speech separation model is trained based on microphone signals collected by microphones at different intervals in the sample scene, and the corresponding sample speech separation results.

[0014] According to a speech separation method provided in this disclosure, the speech separation model is trained based on the following steps:

[0015] Acquire the multi-microphone signals in the sample scenario;

[0016] Each microphone signal is encoded, and the features of the encoded microphone signals are fused to obtain sample microphone features;

[0017] Single-channel speech separation is performed on each microphone signal to obtain the time-frequency mask of each clean microphone signal. Based on the time-frequency mask of each clean microphone signal and the frequency domain phase of each microphone signal, the sample directional features are determined.

[0018] The sample microphone features and the sample directional features are input into the initial speech separation model to obtain the predicted speech separation result output by the initial speech separation model;

[0019] Based on the predicted speech separation results and the sample speech separation results, the parameters of the initial speech separation model are iterated to obtain the speech separation model.

[0020] According to a speech separation method provided in this disclosure, determining the sample directional features based on the time-frequency mask of each clean microphone signal and the frequency domain phase of each microphone signal includes:

[0021] Based on the time-frequency mask of each pure microphone signal, determine the spatial covariance matrix of each pure microphone signal;

[0022] Based on the spatial covariance matrix of each pure microphone signal and the frequency domain phase of each microphone signal, the directional characteristics of the sample are determined.

[0023] According to a speech separation method provided in this disclosure, determining the sample directional features based on the spatial covariance matrix of each clean microphone signal and the frequency domain phase of each microphone signal includes:

[0024] The spatial covariance matrix of each pure microphone signal is decomposed by eigenvalue decomposition to obtain the steering vector of each pure microphone signal.

[0025] Based on the steering vector of each pure microphone signal and the frequency domain phase of each microphone signal, the directional characteristics of the sample are determined.

[0026] According to a speech separation method provided in this disclosure, determining the sample directional features based on the steering vector of each clean microphone signal and the frequency domain phase of each microphone signal includes:

[0027] Based on the steering vector of each pure microphone signal, the time delay difference between the two microphones in each microphone pair is determined;

[0028] Based on the frequency domain phase of each microphone signal, the phase difference between the two microphones in each microphone pair is determined;

[0029] Based on the time delay difference and phase difference of each microphone pair, the cosine distance of each microphone pair in the direction of the microphone signal source is determined, and the sample directional characteristics are determined based on the cosine distance of each microphone pair in the direction of the microphone signal source.

[0030] According to a speech separation method provided in this disclosure, determining the spatial covariance matrix of each clean microphone signal based on the time-frequency mask of each clean microphone signal includes:

[0031] Based on the time-frequency masks of each pure microphone signal, determine the median time-frequency mask;

[0032] Based on the median time-frequency mask and the microphone signals, the spatial covariance matrix of each clean microphone signal is determined.

[0033] According to a speech separation method provided in this disclosure, the steps of encoding each microphone signal and fusing the features of the encoded microphone signals to obtain sample microphone features include:

[0034] Each microphone signal is encoded, and the characteristics of each encoded microphone signal are nonlinearly activated to obtain the microphone features of each channel.

[0035] Average pooling is performed on the microphone features of each channel to obtain global microphone features of each channel, and the global microphone features of each channel are fused to obtain sample microphone features.

[0036] According to a speech separation method provided in this disclosure, the step of fusing the global microphone features of each channel to obtain sample microphone features includes:

[0037] The global microphone features and corresponding microphone features of each channel are fused to obtain the target microphone features of each channel;

[0038] The microphone features of each target are fused to obtain the sample microphone features.

[0039] This disclosure also provides a speech separation device, including:

[0040] A signal acquisition unit is used to acquire voice signals from the target vehicle, the voice signals being collected by a microphone array mounted on the target vehicle;

[0041] The feature extraction unit is used to extract the speech features and directional features of the speech signal, wherein the speech features and the directional features are independent of the array configuration of the microphone array.

[0042] The speech separation unit is used to perform speech separation based on the speech features and the directional features to obtain the speech separation result corresponding to the speech signal.

[0043] This disclosure also provides an in-vehicle interactive system, including a microphone array, a processor, and an audio player;

[0044] The processor is used to acquire the voice signal collected by the microphone array on the target vehicle, extract the voice features and directional features of the voice signal, perform voice separation based on the voice features and directional features to obtain the voice separation result corresponding to the voice signal, perform signal response based on the voice separation result, generate reply content, and control the audio player to play the reply content.

[0045] The speech features and the directional features are independent of the array configuration of the microphone array.

[0046] This disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the program to implement any of the above-described speech separation methods.

[0047] This disclosure also provides a non-transitory computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the speech separation method as described above.

[0048] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements any of the speech separation methods described above.

[0049] The speech separation method, apparatus, system, device, storage medium, and program products disclosed herein extract array-independent speech features and directional features from the speech signals collected by the microphone array on the target vehicle, and perform speech separation based on these features. This enables complete isolation of sound on the vehicle without relying on the array information of the microphone array, ensuring independent interaction between different sound zones. It avoids the problem of traditional solutions relying on separation models trained with array information, which leads to a "one vehicle, one customization" dilemma. Speech separation based on array-independent features greatly improves the efficiency of large-scale system delivery, reduces costs, and minimizes resource waste. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 is a flowchart illustrating the speech separation method provided in this disclosure;

[0052] Figure 2 is a general framework diagram of the training process of the speech separation model provided in this disclosure;

[0053] Figure 3 is an overall flowchart of the sample directional feature determination process provided in this disclosure;

[0054] Figure 4 is a general framework diagram of the sample microphone feature determination process provided in this disclosure;

[0055] Figure 5 is a schematic diagram of the speech separation device provided in this disclosure;

[0056] Figure 6 is a structural schematic diagram of the in-vehicle interactive system provided in this disclosure;

[0057] Figure 7 is a schematic diagram of the structure of the electronic device provided in this disclosure. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of this disclosure clearer, the technical solutions of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0059] In the field of intelligent vehicles, a highlight of in-vehicle intelligent voice interaction systems (IVS) is the achievement of independent sound zones, enabling each seat to have a unique interactive experience. This means that people in different positions (passengers and drivers) can independently conduct voice interactions without interference from other sound zones. To achieve complete sound isolation, current methods often use spatial information of known array configurations to train separation models. However, this approach leads to a "one-vehicle-one-customization" dilemma. That is, a separation model needs to be customized for each vehicle model. Furthermore, whenever a new model is launched, or even if only a minor change is made to the microphone array layout of the same model, a new separation model needs to be created. This poses a significant challenge to the large-scale delivery of the system.

[0060] In response, this disclosure provides a speech separation method that aims to extract two types of array-independent features from the speech signal and perform speech separation based on these features to obtain signals for each voice region, thereby achieving complete sound isolation and ensuring independent interaction between each voice region. This avoids the current problem of speech separation in vehicles relying on separation models trained through array information, which can lead to a "one vehicle, one customization" dilemma and greatly improves system delivery efficiency.

[0061] Figure 1 is a flowchart illustrating the speech separation method provided in this disclosure. As shown in Figure 1, this method can be applied to in-vehicle interactive systems, and the specific execution entity can be the processor in the in-vehicle interactive system. The method includes:

[0062] Step 110: Acquire the voice signal from the target vehicle. The voice signal is acquired through a microphone array mounted on the target vehicle.

[0063] Step 120: Extract speech features and directional features from the speech signal. Both speech features and directional features are independent of the microphone array configuration.

[0064] Step 130: Perform speech separation based on speech features and directional features to obtain the speech separation result corresponding to the speech signal.

[0065] Considering that current in-vehicle interactive systems often use the spatial information of the vehicle's microphone array to customize a separate model for the vehicle in order to achieve independent interaction between different audio zones, that is, passengers and drivers in different positions can conduct independent voice interaction without interference from other areas. However, this approach is not only time-consuming and labor-intensive, but also requires a large investment of resources to customize the new model, which is costly and causes unnecessary waste.

[0066] Based on this, this disclosure proposes a separation method independent of microphone array configuration. By extracting configuration-independent features from the speech signal for speech separation, the speech separation process can be separated from the configuration of the microphone array on the vehicle. This makes the speech separation on the vehicle no longer dependent on the configuration information of the microphone array, thus effectively avoiding the problem of "one vehicle, one customization". It can achieve accurate separation and localization of different sound sources without being affected by differences in microphone array configuration, meeting the needs of in-vehicle interactive systems in practical applications and providing a guarantee for the rapid, efficient, and low-cost delivery of in-vehicle interactive systems for different vehicle models.

[0067] Understandably, in practical applications, before performing voice separation, it is first necessary to determine the voice signal to be separated, which is the mixed signal picked up by the vehicle's microphone. Specifically, this can be achieved when the target vehicle is started and the in-vehicle interactive system is working. The voices of the occupants can be picked up by the microphone array installed in the target vehicle, and a mixed signal is obtained through multi-channel pickup. This signal is the voice signal to be separated from the target vehicle.

[0068] After obtaining the speech signal to be separated, in order to achieve independent voice regions and ensure that each voice region does not interfere with each other and can interact independently, this embodiment of the disclosure further requires speech separation of the speech signal to obtain the speech separation result, i.e., the speech signal of each voice region. Considering that the current speech separation on vehicles relies on a separation model trained with spatial information of a known array pattern, which leads to a "one vehicle, one customization" problem and reduces delivery efficiency, this embodiment of the disclosure selects to extract features independent of the array pattern when performing speech separation on the speech signal of the target vehicle, and performs speech separation based on these features, thus completely solving the vehicle customization problem.

[0069] Specifically, the speech signal can first be processed to obtain speech features. This can be achieved through operations such as encoding, activation, and pooling to extract effective features from the speech signal, and then fusing the effective features extracted from multiple signals to obtain the speech features. Next, considering that speech features only guarantee some information about the signal itself, speech separation based on this is somewhat weak and cannot guarantee a good separation effect. Therefore, in this embodiment, it is necessary to further extract directional features that can characterize the signal direction, so as to combine the speech features and directional features for speech separation, achieving accurate sound source localization and precise separation of speech signals in each vocal region. Specifically, this can be achieved by processing the speech signal using signal processing algorithms to extract information related to the signal direction, especially information related to the direction of the signal source, to obtain the directional features of the speech signal.

[0070] It should be noted that, to avoid falling into the misconception of "one-vehicle-one-customization," the speech features and directional features extracted in this embodiment are features unrelated to the array configuration of the microphone array on the target vehicle. "Unrelated to array configuration" does not mean that the entire feature extraction process does not involve array-related information, but rather that array-related information is converted into array-independent information during the extraction process, thus making the final extracted speech features and directional features independent of the microphone array configuration and unaffected by it.

[0071] Furthermore, after obtaining the speech separation result, in this embodiment of the disclosure, interaction can be performed based on the speech separation result, thereby realizing independent interaction of each voice region. That is, the speech signals of each separated voice region can be responded to separately, and reply information can be generated and displayed, so that the corresponding person in the target vehicle can know the system's reply information in a timely manner, realizing that the interaction of each seat does not interfere with each other and enjoys a unique interactive experience.

[0072] In this embodiment of the disclosure, speech separation is performed by combining two types of array-independent features obtained from signal processing and encoding. This enables accurate localization and separation of different sound sources regardless of differences in microphone array array configuration, thereby completely solving the "one vehicle, one customization" problem and greatly improving delivery efficiency.

[0073] The speech separation method disclosed herein extracts array-independent speech features and directional features from the speech signals collected by the microphone array on the target vehicle, and performs speech separation based on these features. This method can achieve complete isolation of sound on the vehicle without relying on the array information of the microphone array, ensuring independent interaction between different sound zones. It avoids the problem of traditional solutions relying on separation models trained with array information, which leads to a "one vehicle, one customization" dilemma. Speech separation based on array-independent features greatly improves the efficiency of large-scale system delivery, reduces costs, and minimizes resource waste.

[0074] Based on the above embodiments, step 130 includes:

[0075] Based on speech features and directional features, a speech separation model is applied to perform speech separation, and the speech separation results corresponding to the speech signal are obtained.

[0076] The speech separation model is trained based on microphone signals collected by microphones at different intervals in the sample scene, and the corresponding sample speech separation results.

[0077] Specifically, in step 130, the process of performing speech separation based on speech features and directional features to obtain the speech separation result corresponding to the speech signal can be implemented through a speech separation model. Here, specifically, it can be achieved by applying a speech separation model to perform speech separation on the speech signal based on the extracted array-independent speech features and directional features, and then outputting the separated signal. In other words, the extracted speech features and directional features of the speech signal are input into the speech separation model, so that the speech separation model performs the speech separation task according to the input features and outputs the separated signal, thus obtaining the speech separation result.

[0078] It is worth noting that before applying this speech separation model for speech separation, it is usually necessary to pre-train the speech separation model to enable it to have better speech separation capabilities, accurately separate the speech signals of each voice region from the mixed signal, and then perform high-precision speech separation on the speech signals obtained from multiple voice pickups in practical applications.

[0079] In detail, the speech separation model can be pre-trained using application samples and labels. However, considering the current problem that the training of the separation model depends on the spatial information of the known array, in this embodiment of the disclosure, in order to make the trained speech separation model unrestricted by the array and applicable to various arrays of in-vehicle interactive systems, it is possible to accurately separate the speech signals collected by microphones of different arrays. Here, multiple microphones with different spacings can be selected to collect signals and label the separation signals in the corresponding scenarios. That is, the microphone signals collected by multiple microphones with different spacings in the sample scenario are obtained, and the corresponding sample speech separation results in the sample scenario are determined. Based on this, the model is trained so that the model can learn the correspondence between input and output during the training process, thereby having the ability to accurately separate signals of different sound regions based on the array-independent features in the microphone signals. Thus, in practical applications, it can efficiently and accurately locate and separate different sound sources without being affected by the array.

[0080] Based on the above embodiments, the speech separation model is trained using the following steps:

[0081] Acquire multi-channel microphone signals in a sample scenario;

[0082] Each microphone signal is encoded, and the features of the encoded microphone signals are fused to obtain sample microphone features;

[0083] Single-channel speech separation is performed on each microphone signal to obtain the time-frequency mask of each clean microphone signal. Based on the time-frequency mask of each clean microphone signal and the frequency domain phase of each microphone signal, the directional features of the sample are determined.

[0084] Input the sample microphone features and sample directional features into the initial speech separation model to obtain the predicted speech separation result output by the initial speech separation model;

[0085] Based on the predicted speech separation results and the sample speech separation results, the parameters of the initial speech separation model are iterated to obtain the speech separation model.

[0086] Specifically, the training process of the speech separation model may include the following steps:

[0087] Figure 2 is the overall framework diagram of the speech separation model training process provided in this disclosure. As shown in Figure 2, the speech signal in the sample scene needs to be determined first, that is, the speech signals collected by multiple microphones with different spacing in the sample scene are obtained, which are referred to here as microphone signals, that is, the multiple microphone signals y1, y2, ..., y3 in the sample scene. P P represents the total number of microphones.

[0088] Next, feature extraction can be performed on the multi-channel microphone signals, and the array-independent features extracted from each microphone signal can be fused to obtain the speech features and directional features corresponding to the microphone signals, which are called sample microphone features and sample directional features.

[0089] Specifically, this can involve encoding the microphone signals from each channel using convolutional layers to extract effective features from the signals. These encoded features can then be fused to obtain an overall speech feature corresponding to the multi-channel microphone signals, i.e., sample microphone features. This fusion can be addition, concatenation, etc., and this embodiment does not impose specific limitations.

[0090] Simultaneously, the microphone signals can be processed to extract features related to signal direction, thereby obtaining sample directional features. Considering the noise in microphone signals obtained through multiple microphones, i.e., they may contain background noise, other environmental factors, etc., to improve the accuracy of speech separation, in this embodiment, coarse separation can be performed first using a single-channel speech separation method. That is, single-channel speech separation can be performed on each microphone signal to obtain the time-frequency mask of the clean signal corresponding to each microphone signal, i.e., the time-frequency mask of each pure microphone signal. Then, based on this time-frequency mask and the frequency domain phase of the microphone signal, the signal direction features can be determined. Specifically, based on the time-frequency mask of each pure microphone signal and the frequency domain phase of each microphone signal, features that are independent of the array configuration but related to the source direction of the microphone signal can be solved to obtain sample directional features.

[0091] Then, based on the sample microphone features and sample directional features, the initial speech separation model can be applied to perform speech separation, thereby obtaining the speech separation result corresponding to the sample scene, i.e., the predicted speech separation result. Specifically, the sample microphone features and sample directional features are input into the initial speech separation model, so that the initial speech separation model performs speech separation on the microphone signal in the sample scene based on the input sample microphone features and sample directional features, and outputs the separated signal, thereby obtaining the output predicted speech separation result. It should be noted that the initial speech separation model here is a pre-built initial model for the training process, which can be built on the basis of an existing multi-channel speech separation model. The specific model structure can be set according to the actual situation and specific needs, etc., and this embodiment does not impose specific limitations.

[0092] Then, based on the predicted speech separation results and the sample speech separation results, the parameters of the initial speech separation model can be iterated to obtain the final speech separation model. That is, based on the results of the speech separation task output by the model and the labels of the samples, the loss of the initial speech separation model on the speech separation task can be determined, and the parameters of the initial speech separation model can be adjusted according to this loss so that the predicted speech separation results output by the model after parameter adjustment are as close as possible to the sample speech separation results, and finally, a trained speech separation model can be obtained.

[0093] Here, the loss function used to measure the loss of the initial speech separation model is SISNR (Scale-Invariant Signal-to-Noise Ratio) loss, and the final loss uses the PIT (Permutation Invariant Training) strategy, which is to arrange the minimum loss values ​​among various combinations.

[0094] Based on the above embodiments, the sample directional characteristics are determined based on the time-frequency mask of each pure microphone signal and the frequency domain phase of each microphone signal, including:

[0095] Based on the time-frequency mask of each pure microphone signal, the spatial covariance matrix of each pure microphone signal is determined;

[0096] Based on the spatial covariance matrix of each pure microphone signal and the frequency domain phase of each microphone signal, the directional characteristics of the sample are determined.

[0097] Specifically, the process of determining the sample directional characteristics based on the time-frequency mask of each pure microphone signal and the frequency domain phase of each microphone signal includes:

[0098] First, based on the time-frequency masks of the clean microphone signals obtained from single-channel speech separation, the spatial covariance matrix of each clean microphone signal is determined. This matrix reflects information about the signal source direction, facilitating the determination of sample directional characteristics. Specifically, this can be achieved by combining the time-frequency masks of each clean microphone signal to select a time-frequency mask that ensures stable arrangement of the single-channel speech separation results. Based on this, the spatial covariance matrix of each clean microphone signal is calculated, thus reflecting the information of the corresponding signal in the signal source direction.

[0099] Subsequently, the directional characteristics of the samples can be determined based on the spatial covariance matrix of each pure microphone signal and the frequency domain phase of each microphone signal. That is, information related to the direction of the signal source is calculated based on the time-frequency mask and the frequency domain phase, thereby obtaining the directional characteristics of the samples.

[0100] Based on the above embodiments, the sample directional characteristics are determined based on the spatial covariance matrix of each pure microphone signal and the frequency domain phase of each microphone signal, including:

[0101] The spatial covariance matrix of each pure microphone signal is decomposed by eigenvalue decomposition to obtain the steering vector of each pure microphone signal;

[0102] Based on the steering vector of each pure microphone signal and the frequency domain phase of each microphone signal, the directional characteristics of the sample are determined.

[0103] Specifically, the process of determining the sample directional characteristics based on the spatial covariance matrix of each pure microphone signal and the frequency domain phase of each microphone signal can include:

[0104] First, the vector information in the direction of the signal source can be calculated based on the spatial covariance matrix of each pure microphone signal, thereby obtaining the steering vector of each pure microphone signal. Specifically, the eigenvalue decomposition of the spatial covariance matrix of each pure microphone signal is performed to obtain the steering vector of each pure microphone signal.

[0105] Then, the sample directional characteristics can be determined based on this steering vector and the frequency domain phase of each microphone signal. That is, based on the information of each microphone signal in its source direction indicated by the steering vector, and the phase information of the corresponding frequency domain signal of each microphone signal reflected by the frequency domain phase, the directional information independent of the array configuration can be solved to obtain the sample directional characteristics.

[0106] Based on the above embodiments, the guide vector can be calculated using the following formula:

[0107] In the formula, The guide vector representing the pure microphone signal. This represents the spatial covariance matrix of the pure microphone signal. This represents eigenvalue decomposition.

[0108] Based on the above embodiments, the sample directional characteristics are determined based on the steering vector of each pure microphone signal and the frequency domain phase of each microphone signal, including:

[0109] Based on the steering vector of each pure microphone signal, the time delay difference between the two microphones in each microphone pair is determined;

[0110] Based on the frequency domain phase of each microphone signal, determine the phase difference between the two microphones in each microphone pair;

[0111] Based on the time delay difference and phase difference of each microphone pair, the cosine distance of each microphone pair in the direction of the microphone signal source is determined, and the sample directional characteristics are determined based on the cosine distance of each microphone pair in the direction of the microphone signal source.

[0112] Specifically, the process of determining the sample directional characteristics based on the steering vector of each pure microphone signal and the frequency domain phase of each microphone signal may include:

[0113] Figure 3 is an overall block diagram of the sample directionality feature determination process provided in this disclosure. As shown in Figure 3, after single-channel speech separation and steering vector calculation, in this embodiment of the disclosure, multiple microphone signals can be grouped into pairs to form multiple microphone pairs. The time delay difference between the two microphones can be calculated based on the steering vectors of the pure microphone signals corresponding to the two microphones in each microphone pair. The phase difference between the two microphones can also be calculated based on the frequency domain phase of the microphone signals corresponding to the two microphones in each microphone pair. Thus, the signal source direction information corresponding to each microphone pair is obtained.

[0114] Specifically, the time delay difference between the two microphones in the microphone pair can be calculated based on the steering vector of one pure microphone signal and the steering vector of the pure microphone signal of the other microphone in the microphone pair. At the same time, the phase difference between the two microphones in the microphone pair can be determined based on the phase of the frequency domain signals of the microphone signals of the two microphones in the microphone pair.

[0115] Subsequently, considering that the delay difference and phase difference of the microphone pairs calculated in the previous step are related to the microphone array and are array-dependent information, in this embodiment of the disclosure, after obtaining the delay difference and phase difference of each microphone pair, it is necessary to process them to convert the array-dependent information into array-independent distances. That is, based on the delay difference and phase difference of each microphone pair, the cosine distance of each microphone pair in the direction of the microphone signal source is calculated. The sample directional characteristics are determined by the array-independent cosine distance. In other words, the directional characteristics of each microphone pair can be determined by the cosine distance of each microphone pair in the direction of the microphone signal source. Based on this, the directional characteristics of each microphone are solved, and finally, the directional characteristics of each microphone are spliced ​​together to obtain the sample directional characteristics.

[0116] It should be noted that, compared to the array-related delay and phase differences, the cosine distance along the microphone signal source direction calculated in this embodiment is only related to the relative position between the signal source direction and the microphones; that is, whether the line connecting them forms a specific angle with the signal source direction. It does not involve the specific distance between the microphones or the shape of the entire array, and is unrelated to the array configuration. Conversely, signal delay and phase differences are array-related and are affected by the distance between microphones. Under the same signal source direction, different microphone array configurations (such as linear arrays and circular arrays) will result in different distances between microphones, thus producing different delay differences.

[0117] In short, since the cosine distance is only related to the direction of the signal source and the relative position between the microphones, and is not related to the specific distance or array of the microphones involved in the time delay difference, the cosine distance will not change even if the array of the microphones is changed (thus changing the time delay difference), thus fundamentally solving the "one car, one customization" problem.

[0118] Based on the above embodiments, the sample directionality feature is determined using the following formula:

[0119] In the formula, This represents the directional characteristics of the microphone signal corresponding to the p-th microphone, where P represents the total number of microphones, and P-1 is the number of microphone pairs formed by the p-th microphone and the remaining microphones.<q′,p> Let represent the microphone pair consisting of the q′-th microphone and the p-th microphone, Ω be the set of all microphone pairs including the p-th microphone, and ∠Y q′ (t,f) and ∠Y p (t,f) represent the phase of the frequency domain signal of the microphone signal corresponding to the q′-th microphone and the phase of the frequency domain signal of the microphone signal corresponding to the p-th microphone, respectively. and Let represent the guide vectors of the clean microphone signal corresponding to the q′-th microphone and the p-th microphone, respectively.

[0120] Based on the above embodiments, the spatial covariance matrix of each clean microphone signal is determined based on the time-frequency mask of each clean microphone signal, including:

[0121] Determine the median time-frequency mask based on the time-frequency masks of each pure microphone signal;

[0122] Based on the median time-frequency mask and the microphone signals, the spatial covariance matrix of each clean microphone signal is determined.

[0123] Specifically, the process of determining the spatial covariance matrix of each pure microphone signal based on the time-frequency mask of each pure microphone signal may include:

[0124] Considering that the signal arrangement of single-channel speech separation is not fixed, which is not conducive to the determination of directional features, in this embodiment of the present disclosure, after obtaining the time-frequency mask of each pure microphone signal, the median time-frequency mask can be determined by the median solution strategy, that is, the median of the time-frequency mask of multiple pure microphone signals is determined and used as the median time-frequency mask in subsequent calculations, thereby improving the accuracy of the calculation.

[0125] After this, the spatial covariance matrix of each clean microphone signal can be calculated based on the median time-frequency mask and the signals from each microphone channel. That is, based on the median time-frequency mask determined in the previous step, the spatial covariance matrix of each clean microphone signal is calculated.

[0126] Based on the above embodiments, the spatial covariance matrix of the pure microphone signal can be calculated using the following formula:

[0127] In the formula, η represents the spatial covariance matrix of the pure microphone signal. (c) Y(t,f) is the median time-frequency mask of the t-th frame, Y(t,f) represents the microphone signal of the t-th frame, T is the total number of frames of the signal, and H is the transpose.

[0128] The median time-frequency mask can be determined using the following formula:

[0129] In the formula, η (c) This is the median time-frequency mask, where `median` represents calculating the median. This is the time-frequency mask for the clean microphone signal corresponding to the multiple microphone signals, where P represents the total number of microphones.

[0130] Based on the above embodiments, each microphone signal is encoded, and the features of the encoded microphone signals are fused to obtain sample microphone features, including:

[0131] Each microphone signal is encoded, and the characteristics of each encoded microphone signal are nonlinearly activated to obtain the microphone features of each channel.

[0132] Average pooling is performed on the microphone features of each channel to obtain the global microphone features of each channel, and the global microphone features of each channel are then fused to obtain the sample microphone features.

[0133] Specifically, the process of encoding each microphone signal and fusing the features of the encoded microphone signals to obtain sample microphone features may include:

[0134] Figure 4 is the overall framework diagram of the sample microphone feature determination process provided in this disclosure. As shown in Figure 4, for the input multi-channel microphone signals, they are first encoded, and the features of each encoded microphone signal are nonlinearly activated to obtain the microphone features of each channel. Next, average pooling can be performed on each microphone feature to reduce the feature dimensionality and obtain the output features, which are used as global variables independent of the array configuration, i.e., the global microphone features of each channel. Afterwards, the global microphone features and the individual microphone features can be fused to obtain the sample microphone features.

[0135] Based on the above embodiments, the global microphone features of each channel are fused to obtain sample microphone features, including:

[0136] The global microphone features and corresponding microphone features of each channel are fused to obtain the target microphone features of each channel;

[0137] The microphone features of each target are fused to obtain the sample microphone features.

[0138] Specifically, the process of fusing the global microphone features from various channels to obtain the sample microphone features may include:

[0139] First, the global microphone features and the microphone features of each channel can be fused together, that is, the global microphone features and the microphone features of each channel can be concatenated to obtain the target microphone features of each channel. Next, the target microphone features of each channel can be fused together, that is, the target microphone features of each channel can be concatenated to obtain the sample microphone features.

[0140] The speech separation apparatus provided in this disclosure is described below. The speech separation apparatus described below can be referred to in correspondence with the speech separation method described above.

[0141] Figure 5 is a schematic diagram of the speech separation device provided in this disclosure. As shown in Figure 5, the device includes:

[0142] The signal acquisition unit 510 is used to acquire voice signals from the target vehicle, which are obtained by a microphone array mounted on the target vehicle.

[0143] The feature extraction unit 520 is used to extract the speech features and directional features of the speech signal, and the speech features and directional features are independent of the array configuration of the microphone array.

[0144] The speech separation unit 530 is used to perform speech separation based on the speech features and the directional features to obtain the speech separation result corresponding to the speech signal.

[0145] The speech separation device disclosed herein extracts array-independent speech features and directional features from the speech signals collected by the microphone array on the target vehicle, and performs speech separation based on these features. It can achieve complete isolation of sound on the vehicle without relying on the array information of the microphone array, thus ensuring independent interaction of each sound zone. This avoids the problem of traditional solutions relying on separation models trained with array information, which leads to a "one vehicle, one customization" dilemma. Speech separation based on array-independent features greatly improves the efficiency of large-scale system delivery, reduces costs, and minimizes resource waste.

[0146] Based on the above embodiments, the speech separation unit 530 is used for:

[0147] Based on the speech features and the directional features, a speech separation model is applied to perform speech separation, and the speech separation result corresponding to the speech signal is obtained.

[0148] The speech separation model is trained based on microphone signals collected by microphones at different intervals in the sample scene, and the corresponding sample speech separation results.

[0149] Based on the above embodiments, the device further includes a model training unit, used for:

[0150] Acquire the multi-microphone signals in the sample scenario;

[0151] Each microphone signal is encoded, and the features of the encoded microphone signals are fused to obtain sample microphone features;

[0152] Single-channel speech separation is performed on each microphone signal to obtain the time-frequency mask of each clean microphone signal. Based on the time-frequency mask of each clean microphone signal and the frequency domain phase of each microphone signal, the sample directional features are determined.

[0153] The sample microphone features and the sample directional features are input into the initial speech separation model to obtain the predicted speech separation result output by the initial speech separation model;

[0154] Based on the predicted speech separation results and the sample speech separation results, the parameters of the initial speech separation model are iterated to obtain the speech separation model.

[0155] Based on the above embodiments, the model training unit is used for:

[0156] Based on the time-frequency mask of each pure microphone signal, determine the spatial covariance matrix of each pure microphone signal;

[0157] Based on the spatial covariance matrix of each pure microphone signal and the frequency domain phase of each microphone signal, the directional characteristics of the sample are determined.

[0158] Based on the above embodiments, the model training unit is used for:

[0159] The spatial covariance matrix of each pure microphone signal is decomposed by eigenvalue decomposition to obtain the steering vector of each pure microphone signal.

[0160] Based on the steering vector of each pure microphone signal and the frequency domain phase of each microphone signal, the directional characteristics of the sample are determined.

[0161] Based on the above embodiments, the model training unit is used for:

[0162] Based on the steering vector of each pure microphone signal, the time delay difference between the two microphones in each microphone pair is determined;

[0163] Based on the frequency domain phase of each microphone signal, the phase difference between the two microphones in each microphone pair is determined;

[0164] Based on the time delay difference and phase difference of each microphone pair, the cosine distance of each microphone pair in the direction of the microphone signal source is determined, and the sample directional characteristics are determined based on the cosine distance of each microphone pair in the direction of the microphone signal source.

[0165] Based on the above embodiments, the model training unit is used for:

[0166] Based on the time-frequency masks of each pure microphone signal, determine the median time-frequency mask;

[0167] Based on the median time-frequency mask and the microphone signals, the spatial covariance matrix of each clean microphone signal is determined.

[0168] Based on the above embodiments, the model training unit is used for:

[0169] Each microphone signal is encoded, and the characteristics of each encoded microphone signal are nonlinearly activated to obtain the microphone features of each channel.

[0170] Average pooling is performed on the microphone features of each channel to obtain global microphone features of each channel, and the global microphone features of each channel are fused to obtain sample microphone features.

[0171] Based on the above embodiments, the model training unit is used for:

[0172] The global microphone features and corresponding microphone features of each channel are fused to obtain the target microphone features of each channel;

[0173] The microphone features of each target are fused to obtain the sample microphone features.

[0174] Figure 6 is a schematic diagram of the structure of the in-vehicle interactive system provided in this disclosure. As shown in Figure 6, the system includes a microphone array 610, a processor 620, and an audio player 630.

[0175] The processor 620 is used to acquire the voice signal collected by the microphone array 610 on the target vehicle, extract the voice features and directional features of the voice signal, perform voice separation based on the voice features and directional features to obtain the voice separation result corresponding to the voice signal, and perform signal response based on the voice separation result to generate reply content, and control the audio player 630 to play the reply content.

[0176] The speech features and the directional features are independent of the array configuration of the microphone array.

[0177] The in-vehicle interactive system disclosed herein collects voice signals from a target vehicle via a microphone array. A processor extracts array-independent voice features and directional features from the voice signals collected by the microphone array, and performs voice separation based on these features. The system then generates response content based on the voice separation results and controls an audio player to play the response content. This system enables independent interaction across different audio zones without relying on the array's array configuration information. This avoids the problem of traditional solutions where voice separation in vehicles depends on a separation model trained on array information, leading to a "one vehicle, one customization" dilemma. Voice separation based on array-independent features significantly improves the efficiency of large-scale system delivery, reduces costs, and minimizes resource waste.

[0178] Figure 7 illustrates a schematic diagram of the physical structure of an electronic device. As shown in Figure 7, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. The processor 710, communication interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute a speech separation method. This method includes: acquiring a speech signal from a target vehicle, the speech signal being collected by a microphone array mounted on the target vehicle; extracting speech features and directional features from the speech signal, the speech features and the directional features being independent of the array configuration of the microphone array; and performing speech separation based on the speech features and the directional features to obtain a speech separation result corresponding to the speech signal.

[0179] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0180] On the other hand, this disclosure also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, which, when executed by a computer, enable the computer to execute the speech separation method provided by the above methods. The method includes: acquiring a speech signal from a target vehicle, the speech signal being collected by a microphone array mounted on the target vehicle; extracting speech features and directional features from the speech signal, the speech features and the directional features being independent of the array configuration of the microphone array; and performing speech separation based on the speech features and the directional features to obtain a speech separation result corresponding to the speech signal.

[0181] In another aspect, this disclosure also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform the speech separation methods provided by the methods described above. The method includes: acquiring a speech signal from a target vehicle, the speech signal being collected by a microphone array mounted on the target vehicle; extracting speech features and directional features from the speech signal, the speech features and the directional features being independent of the array configuration of the microphone array; and performing speech separation based on the speech features and the directional features to obtain a speech separation result corresponding to the speech signal.

[0182] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0183] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit them. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure.

Claims

1. A speech separation method, comprising: Acquire voice signals from the target vehicle, which are obtained through a microphone array mounted on the target vehicle; The speech features and directional features of the speech signal are extracted, and both the speech features and the directional features are independent of the array configuration of the microphone array; Speech separation is performed based on the speech features and the directional features to obtain the speech separation result corresponding to the speech signal.

2. The speech separation method according to claim 1, wherein, The process of performing speech separation based on the speech features and the directional features to obtain the speech separation result corresponding to the speech signal includes: Based on the speech features and the directional features, a speech separation model is applied to perform speech separation, and the speech separation result corresponding to the speech signal is obtained. The speech separation model is trained based on microphone signals collected by microphones at different intervals in the sample scene, and the corresponding sample speech separation results.

3. The speech separation method according to claim 2, wherein, The speech separation model is trained based on the following steps: Acquire the multi-microphone signals in the sample scenario; Each microphone signal is encoded, and the features of the encoded microphone signals are fused to obtain sample microphone features; Single-channel speech separation is performed on each microphone signal to obtain the time-frequency mask of each clean microphone signal. Based on the time-frequency mask of each clean microphone signal and the frequency domain phase of each microphone signal, the sample directional features are determined. The sample microphone features and the sample directional features are input into the initial speech separation model to obtain the predicted speech separation result output by the initial speech separation model; Based on the predicted speech separation results and the sample speech separation results, the parameters of the initial speech separation model are iterated to obtain the speech separation model.

4. The speech separation method according to claim 3, wherein, The determination of sample directional characteristics based on the time-frequency mask of each pure microphone signal and the frequency domain phase of each microphone signal includes: Based on the time-frequency mask of each pure microphone signal, determine the spatial covariance matrix of each pure microphone signal; Based on the spatial covariance matrix of each pure microphone signal and the frequency domain phase of each microphone signal, the directional characteristics of the sample are determined.

5. The speech separation method according to claim 4, wherein, The determination of sample directional characteristics based on the spatial covariance matrix of each pure microphone signal and the frequency domain phase of each microphone signal includes: The spatial covariance matrix of each pure microphone signal is decomposed by eigenvalue decomposition to obtain the steering vector of each pure microphone signal. Based on the steering vector of each pure microphone signal and the frequency domain phase of each microphone signal, the directional characteristics of the sample are determined.

6. The speech separation method according to claim 5, wherein, The determination of sample directional characteristics based on the steering vector of each pure microphone signal and the frequency domain phase of each microphone signal includes: Based on the steering vector of each pure microphone signal, the time delay difference between the two microphones in each microphone pair is determined; Based on the frequency domain phase of each microphone signal, the phase difference between the two microphones in each microphone pair is determined; Based on the time delay difference and phase difference of each microphone pair, the cosine distance of each microphone pair in the direction of the microphone signal source is determined, and the sample directional characteristics are determined based on the cosine distance of each microphone pair in the direction of the microphone signal source.

7. The speech separation method according to claim 4, wherein, The determination of the spatial covariance matrix of each clean microphone signal based on the time-frequency mask of each clean microphone signal includes: Based on the time-frequency masks of each pure microphone signal, determine the median time-frequency mask; Based on the median time-frequency mask and the microphone signals, the spatial covariance matrix of each clean microphone signal is determined.

8. The speech separation method according to any one of claims 3 to 7, wherein, The process of encoding each microphone signal and fusing the features of the encoded microphone signals to obtain sample microphone features includes: Each microphone signal is encoded, and the characteristics of each encoded microphone signal are nonlinearly activated to obtain the microphone features of each channel. Average pooling is performed on the microphone features of each channel to obtain global microphone features of each channel, and the global microphone features of each channel are fused to obtain sample microphone features.

9. The speech separation method according to claim 8, wherein, The process of fusing the global microphone features from each path to obtain sample microphone features includes: The global microphone features and corresponding microphone features of each channel are fused to obtain the target microphone features of each channel; The microphone features of each target are fused to obtain the sample microphone features.

10. A speech separation device, comprising: A signal acquisition unit is used to acquire voice signals from the target vehicle, the voice signals being collected by a microphone array mounted on the target vehicle; The feature extraction unit is used to extract the speech features and directional features of the speech signal, wherein the speech features and the directional features are independent of the array configuration of the microphone array. The speech separation unit is used to perform speech separation based on the speech features and the directional features to obtain the speech separation result corresponding to the speech signal.

11. The speech separation device according to claim 10, wherein, The process of performing speech separation based on the speech features and the directional features to obtain the speech separation result corresponding to the speech signal includes: Based on the speech features and the directional features, a speech separation model is applied to perform speech separation, and the speech separation result corresponding to the speech signal is obtained. The speech separation model is trained based on microphone signals collected by microphones at different intervals in the sample scene, and the corresponding sample speech separation results.

12. The speech separation device according to claim 11, wherein, The speech separation model is trained based on the following steps: Acquire the multi-microphone signals in the sample scenario; Each microphone signal is encoded, and the features of the encoded microphone signals are fused to obtain sample microphone features; Single-channel speech separation is performed on each microphone signal to obtain the time-frequency mask of each clean microphone signal. Based on the time-frequency mask of each clean microphone signal and the frequency domain phase of each microphone signal, the sample directional features are determined. The sample microphone features and the sample directional features are input into the initial speech separation model to obtain the predicted speech separation result output by the initial speech separation model; Based on the predicted speech separation results and the sample speech separation results, the parameters of the initial speech separation model are iterated to obtain the speech separation model.

13. The speech separation device according to claim 12, wherein, The determination of sample directional characteristics based on the time-frequency mask of each pure microphone signal and the frequency domain phase of each microphone signal includes: Based on the time-frequency mask of each pure microphone signal, determine the spatial covariance matrix of each pure microphone signal; Based on the spatial covariance matrix of each pure microphone signal and the frequency domain phase of each microphone signal, the directional characteristics of the sample are determined.

14. The speech separation device according to claim 13, wherein, The determination of sample directional characteristics based on the spatial covariance matrix of each pure microphone signal and the frequency domain phase of each microphone signal includes: The spatial covariance matrix of each pure microphone signal is decomposed by eigenvalue decomposition to obtain the steering vector of each pure microphone signal. Based on the steering vector of each pure microphone signal and the frequency domain phase of each microphone signal, the directional characteristics of the sample are determined.

15. The speech separation device according to claim 14, wherein, The determination of sample directional characteristics based on the steering vector of each pure microphone signal and the frequency domain phase of each microphone signal includes: Based on the steering vector of each pure microphone signal, the time delay difference between the two microphones in each microphone pair is determined; Based on the frequency domain phase of each microphone signal, the phase difference between the two microphones in each microphone pair is determined; Based on the time delay difference and phase difference of each microphone pair, the cosine distance of each microphone pair in the direction of the microphone signal source is determined, and the sample directional characteristics are determined based on the cosine distance of each microphone pair in the direction of the microphone signal source.

16. The speech separation device according to claim 13, wherein, The determination of the spatial covariance matrix of each clean microphone signal based on the time-frequency mask of each clean microphone signal includes: Based on the time-frequency masks of each pure microphone signal, determine the median time-frequency mask; Based on the median time-frequency mask and the microphone signals, the spatial covariance matrix of each clean microphone signal is determined.

17. The speech separation method according to any one of claims 12 to 16, wherein, The process of encoding each microphone signal and fusing the features of the encoded microphone signals to obtain sample microphone features includes: Each microphone signal is encoded, and the characteristics of each encoded microphone signal are nonlinearly activated to obtain the microphone features of each channel. Average pooling is performed on the microphone features of each channel to obtain global microphone features of each channel, and the global microphone features of each channel are fused to obtain sample microphone features.

18. The speech separation method according to claim 17, wherein, The process of fusing the global microphone features from each path to obtain sample microphone features includes: The global microphone features and corresponding microphone features of each channel are fused to obtain the target microphone features of each channel; The microphone features of each target are fused to obtain the sample microphone features.

19. An in-vehicle interactive system, comprising a microphone array, a processor, and an audio player; The processor is used to acquire the voice signal collected by the microphone array on the target vehicle, extract the voice features and directional features of the voice signal, perform voice separation based on the voice features and directional features to obtain the voice separation result corresponding to the voice signal, perform signal response based on the voice separation result, generate reply content, and control the audio player to play the reply content. in, Both the speech features and the directional features are independent of the array configuration of the microphone array.

20. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor, when executing the program, implements the speech separation method as described in any one of claims 1 to 9.

21. A non-transitory computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the speech separation method as described in any one of claims 1 to 9.

22. A computer program product comprising a computer program that, when executed by a processor, implements the speech separation method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Voice separation method, model training method and electronic equipment

    CN111429937A

  • Method and device for separating target voice from multiple speakers

    CN113808610A

  • Sound source separation method and sound source separation device

    CN115116465A

  • Single-channel multi-speaker voice separation method and system

    CN115831137A

  • Multi-channel voice separation method based on attention weighting in reverberation environment

    CN118675542A