Voice separation method, system, medium, product, equipment and vehicle

Through audio signal filtering and deep learning model based on area position information, the problem of low speech signal separation accuracy in complex environments is solved, and the speech separation effect with higher accuracy and quality is achieved.

CN120472927APending Publication Date: 2025-08-12BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510391880.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art has low accuracy in speech signal separation in complex environments and insufficient noise resistance, making it difficult to effectively separate the voice of different speakers.

Method used

Based on the audio signal and area position information of multiple regions, the audio signal is filtered by determining mask data, and combined with graph convolution and deep learning models, the spatial characteristics of the audio signal are extracted for precise separation.

Benefits of technology

It improves the extraction accuracy and signal quality of the target audio signal, and enhances the application effect in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472927A_ABST
    Figure CN120472927A_ABST
Patent Text Reader

Abstract

The invention relates to a voice separation method and system, a medium, a product, equipment and a vehicle, and the method comprises the steps: determining a target audio signal after the audio signals are filtered based on the audio signals of a plurality of regions and the region position information of the plurality of regions; according to the method, based on audio signals of multiple areas and position information of the areas, mixed audio signals of different areas can be accurately separated by combining an audio processing algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio processing technology, and in particular to a speech separation method, system, medium, product, equipment and vehicle. Background Art

[0002] In practical applications, voice interaction systems often need to process multiple sound source signals and accurately separate and recognize the voices of different speakers in noisy environments. However, existing technologies suffer from low separation accuracy and insufficient noise immunity when processing voice signals in complex environments. Summary of the Invention

[0003] The embodiments of the present application provide a speech separation method, system, medium, product, device and vehicle to solve the above-mentioned problems.

[0004] In order to achieve the above object, according to a first aspect of the present application, a speech separation method is provided, the method comprising:

[0005] Based on the audio signals of the multiple regions and the region position information of the multiple regions, a target audio signal after the audio signals are filtered is determined.

[0006] Optionally, determining the target audio signal after filtering the audio signal based on the audio signals of the multiple regions and the region position information of the multiple regions includes:

[0007] determining mask data of the audio signal based on the audio signal and the region position information of the plurality of regions;

[0008] The audio signal is filtered based on the mask data to determine a target audio signal after the audio signal is filtered.

[0009] Optionally, determining mask data of the audio signal based on the audio signal and the region position information of the multiple regions includes:

[0010] Determining weight features of the multiple regions based on the regional location information of the multiple regions, wherein the weight features are used to characterize interactions between different regions;

[0011] Based on the weight features, the audio signals of the multiple regions are processed to obtain mask data of the audio signals.

[0012] Optionally, determining the weight features of the multiple regions based on the regional location information of the multiple regions includes:

[0013] Determining relative position information of the multiple areas based on the area position information of the multiple areas;

[0014] The relative position information of the multiple regions is used as an initial weight feature, and the initial weight feature is optimized by a preset method to obtain the weight features of the multiple regions.

[0015] Optionally, the method of obtaining the mask data based on the audio signal and the region position information of the multiple regions includes:

[0016] Performing feature extraction on the audio signal to obtain a first audio feature of the audio signal;

[0017] The mask data is obtained based on the first audio feature and the weight feature.

[0018] Optionally, obtaining the mask data based on the first audio feature and the weight feature includes:

[0019] Obtaining a hidden audio feature of the audio signal based on the first audio feature and the weight feature;

[0020] The mask data is obtained based on the hidden audio feature.

[0021] Optionally, obtaining the hidden audio feature of the audio signal based on the first audio feature and the weight feature includes:

[0022] The hidden audio feature of a second time period is obtained based on the first audio feature and the weight feature of a first time period, where the second time period is after the first time period.

[0023] Optionally, obtaining the hidden audio feature of the second time period based on the first audio feature and the weight feature of the first time period includes:

[0024] Performing feature fusion based on the first audio feature of the first time period and the weight feature to obtain a first fused feature of the first time period;

[0025] Performing feature fusion based on the preset hidden audio feature and the weight feature to obtain a second fused feature for the first time period;

[0026] Feature extraction is performed based on the first fusion feature of the first time period and the second fusion feature of the first time period to obtain the hidden audio feature of the second time period.

[0027] Optionally, obtaining the mask data based on the hidden audio feature includes:

[0028] The mask data is obtained based on a target hidden audio feature among the plurality of hidden audio features.

[0029] Optionally, the performing feature extraction on the audio signal to obtain a first audio feature of the audio signal includes:

[0030] Performing frequency domain analysis on the audio signal to obtain a second audio feature of the audio signal;

[0031] Feature extraction is performed based on the second audio feature to obtain the first audio feature.

[0032] Optionally, filtering the audio signal based on the mask data to determine a target audio signal after filtering the audio signal includes:

[0033] The target audio signal is obtained based on the product of the mask data and the audio signal.

[0034] Optionally, the target audio signal after the audio signal is filtered is obtained by processing the audio signals of the multiple regions and the region position information of the multiple regions using a preset separation model.

[0035] Optionally, the step of processing the audio signals of the multiple regions and the region position information of the multiple regions using a preset separation model to obtain the target audio signal includes:

[0036] Obtaining mask data of the audio signal based on the audio signals of the multiple regions and the region position information of the multiple regions using the separation model;

[0037] The audio signal is filtered based on the mask data to obtain the target audio signal.

[0038] Optionally, obtaining mask data of the audio signal based on the audio signals of the multiple regions and the region position information of the multiple regions by using the separation model includes:

[0039] performing feature extraction on the audio signal based on the first network layer of the separation model to obtain a first audio feature of the audio signal;

[0040] Performing feature extraction on the regional location information of the multiple regions based on the second network layer of the separation model to obtain weight features of the multiple regions;

[0041] performing feature fusion on the first audio feature and the weight feature based on the third network layer of the separation model to obtain a hidden audio feature of the audio signal;

[0042] The hidden audio features are processed based on the fourth network layer of the separation model to obtain the mask data.

[0043] Optionally, the separation model is trained through the following steps:

[0044] Inputting the sample audio signals of the plurality of regions and the sample region position information of the plurality of regions into the initial model to obtain sample mask data of the sample audio signals;

[0045] filtering the sample audio signal based on the sample mask data to obtain a predicted audio signal;

[0046] Determining a loss value of the initial model based on the predicted audio signal and the labeled data of the sample audio signal;

[0047] The parameters of the initial model are trained according to the loss value to obtain the separation model.

[0048] Optionally, the method further includes:

[0049] The sample weight features of the sample region position information of the multiple regions are trained based on the loss value.

[0050] Optionally, the regional location information of the plurality of regions is graph structure data, and the graph structure data includes nodes and edges;

[0051] A node in the regional location information of the plurality of regions corresponds to at least one of the regions, and a node feature of the node is related to the location information of the corresponding region;

[0052] The edges in the region position information of the multiple regions are related to the associated regions, and the edge features of the edges include weight features corresponding to the relative position information between the associated regions.

[0053] Optionally, the plurality of areas correspond to physical spaces within a cabin of a vehicle.

[0054] Optionally, the multiple areas include: a main driver's seat area, a co-driver's seat area, a first rear row area and a second rear row area.

[0055] According to a second aspect of the present application, an embodiment of the present application further provides a speech separation system, the system comprising:

[0056] The separation device is used to determine a target audio signal after filtering the audio signal based on the audio signals of multiple areas and the area position information of the multiple areas.

[0057] Optionally, the system further includes a plurality of sound pickup devices, which are distributed in the plurality of areas and are used to acquire audio signals from the plurality of areas.

[0058] According to a third aspect of the present application, an embodiment of the present application further provides an electronic device, including:

[0059] a memory having a computer program stored thereon;

[0060] A processor is used to execute the computer program in the memory to implement the steps of any one of the methods provided in the embodiments of the present application.

[0061] According to the fourth aspect of the present application, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the methods provided in the embodiments of the present application are implemented.

[0062] According to the fifth aspect of the present application, an embodiment of the present application further provides a computer program product, comprising a computer program or instructions, which, when executed by a processor, implement the steps of any one of the methods provided in the embodiments of the present application.

[0063] According to the sixth aspect of the present application, an embodiment of the present application also provides a vehicle, including the speech separation system, or, an electronic device as described, or, executing the steps of any one of the methods provided in the embodiments of the present application.

[0064] Some embodiments of this specification include at least the following beneficial effects: Based on the audio signals of multiple regions and the location information of the regions, by combining audio processing algorithms, the technical problem of the difficulty in accurately extracting the target audio signal from complex mixed signals in traditional audio signal processing is overcome. When faced with multi-region mixed audio signals, traditional methods are often unable to effectively distinguish the signal sources, resulting in low extraction accuracy of the target audio signal. By introducing regional location information and combining the spatial characteristics of the audio signal, this technology can accurately separate the mixed audio signals of different regions. Not only does it improve the extraction accuracy of the target audio signal, it also significantly improves the quality and availability of the audio signal, which is conducive to application in complex environments.

[0065] Other features and advantages of the present application will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.

[0067] In order to more completely understand the present application and its beneficial effects, the following description will be given in conjunction with the accompanying drawings, wherein the same drawing numbers represent the same parts in the following description.

[0068] Figure 1 This is a diagram of an application scenario of the speech separation method according to some embodiments of this specification;

[0069] Figure 2 is an exemplary flow chart of a speech separation method according to some embodiments of this specification;

[0070] Figure 3 is an exemplary schematic diagram of graph structure data according to some embodiments of this specification;

[0071] Figure 4 is an exemplary schematic diagram of determining a target audio signal according to some embodiments of this specification;

[0072] Figure 5 is an exemplary schematic diagram of determining hidden audio features according to some embodiments of this specification;

[0073] Figure 6 is an exemplary schematic diagram of a training model according to some embodiments of this specification;

[0074] Figure 7 is a structural diagram of a speech separation system according to some embodiments of this specification;

[0075] Figure 8 is a schematic structural diagram of an electronic device according to some embodiments of this specification;

[0076] Figure 9 is an exemplary schematic diagram of a vehicle according to some embodiments of the present specification. DETAILED DESCRIPTION

[0077] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0078] In order to facilitate understanding of the implementation scheme provided in the embodiment of the present application, the relevant application background of the speech separation method provided in the embodiment of the present application is first explained.

[0079] Speech separation is a key technology in modern intelligent cockpit voice assistants. However, while a vehicle is in motion, multiple sound sources, such as engine noise, wind noise, road noise, and passenger chatter, can significantly interfere with speech recognition performance. While existing speech processing technologies can suppress noise to a certain extent, they struggle to effectively separate the voice signals of different speakers in multi-source scenarios, resulting in low speech recognition accuracy and poor separation performance.

[0080] In view of this, some embodiments of this specification provide a speech separation method, which combines the audio signals of each area according to the positional relationship of each area, can capture the relative position information between different areas, so as to better understand the spatial distribution of the audio signal. Combined with the graph convolution method, the robustness and interpretability of the separation model can be further improved, so that the separation model can still accurately separate the target audio signal in a complex noise environment.

[0081] Figure 1 This is an application scenario diagram of the speech separation method shown in some embodiments of this specification.

[0082] like Figure 1 As shown, the application scenario diagram of the speech separation method may include an electronic device and a server, and the electronic device can communicate with the server.

[0083] In some embodiments, Figure 1 The application scenarios shown may also include: storage devices, signal sources, etc. In addition, Figure 1 An electronic device and a server are shown as an example, but other numbers of electronic devices and servers may actually be included, and this application does not impose any limitation on this.

[0084] In some embodiments, Figure 1 The server in the "server" can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. This application does not impose any restrictions on this.

[0085] Optionally, in the embodiment of the present application, the electronic device may be a mobile phone, a tablet computer, a wearable device, an in-vehicle device, or other electronic device with a display function. The embodiment of the present application does not impose any restrictions on the specific type of the electronic device.

[0086] In some embodiments, as Figure 1The electronic device shown can install and run an application with a voice separation function. The application can be, for example, a voice assistant application or an application with a remote conferencing function. The application can also have functions such as data recording, audio and video playback, and data query. When the application is running on the electronic device, it can exchange data with the server.

[0087] In some embodiments, the server is used to provide background services for applications with voice separation capabilities. Optionally, the server performs primary voice separation processing, and the electronic device performs secondary voice separation processing; alternatively, the electronic device performs primary voice separation processing, and the server performs secondary voice separation processing; or, alternatively, the server or the electronic device can each perform voice separation processing.

[0088] In some embodiments, the electronic device is disposed on a vehicle, for example, the electronic device is mounted (or fixed) to the vehicle via a bracket. In some embodiments, the vehicle is communicatively connected to the electronic device. The vehicle can send information to the electronic device or receive information from the electronic device. For example, if the vehicle and the electronic device are connected via Bluetooth, the voice, music, and other data of the electronic device can be transmitted to the vehicle via Bluetooth and played through the vehicle's speakers (or horns).

[0089] The speech separation method provided in the embodiments of the present application can be applied to products such as vehicle-mounted terminals, speech recognition products, voiceprint recognition products, intelligent voice assistants, and smart speakers. It can be applied to the front end of the above products, and can also be implemented through the interaction between electronic devices and servers.

[0090] For example, in vehicle applications, the audio signal in the vehicle is a mixed voice signal from audio signals in multiple areas. When performing voice separation on the mixed voice signal, the voice separation method provided in the embodiment of the present application can be used to separate the voice signal of a single sound source.

[0091] In the embodiment of the present application, the audio signals of multiple areas may include different speakers, and may also include other pronunciation sources (i.e., non-speakers) in the voice signal collection scene, such as other pronunciation sources such as audio that produces background music.

[0092] In the embodiment of the present application, it can be used to realize speech separation of speech signals including speech sources of a speaker and a non-speaker, and can also be used to realize speech separation of speech signals including speech sources of multiple speakers.

[0093] It is worth noting that the application scenarios of the speech separation method are provided for illustrative purposes only and are not intended to limit the scope of this specification. Those skilled in the art can make various changes and modifications based on the description of this specification. For example, the application scenarios may also include databases, information sources, etc. For another example, the application scenarios may be implemented on other devices to achieve similar or different functions. However, these changes and modifications do not deviate from the scope of this specification.

[0094] Figure 2 2 is an exemplary flow chart of a speech separation method according to some embodiments of this specification. In some embodiments, process 200 can be executed based on an electronic device. Figure 2 As shown, the process 200 includes the following steps.

[0095] Step 210: Acquire audio signals of multiple areas.

[0096] Multi-zone audio signals refer to sound data collected in different physical or logical areas.

[0097] Typically, audio signals recorded across multiple locations contain both the target sound source and other sound sources. These other sound sources can interfere with the identification of the target sound source. Therefore, it's necessary to separate the audio signals to clearly identify the target sound source.

[0098] In some embodiments, the audio signals of the multiple regions include sub-audio signals collected by multiple sound pickup devices, each of which may be located in a different region, and the collected sub-audio signals reflect the sound information of the different regions.

[0099] In some embodiments, the audio signal to be processed can be actively acquired by live recording or other methods. After acquiring the audio signal to be processed, the electronic device needs to encode the audio signal to be processed and convert the audio signal to other forms so as to perform audio separation on the audio signal to be processed.

[0100] In some embodiments, the audio signals of multiple zones include: a target sound source and other sound sources. The target sound source refers to the sound source that is desired to be separated from the audio signals of multiple zones, and the other sound sources refer to sound sources in the audio signals of multiple zones that interfere with the identification or reception of the target sound source. In other embodiments, the audio signals of multiple zones include a single sound source (the target sound source for audio separation) as well as noise or echo.

[0101] Each area can be a physically separated space or a virtual area divided by a specific method (such as signal processing, filtering, etc.).

[0102] In some embodiments, zones refer to different physical spatial locations or environments. For example, a room can be divided into several sub-zones, such as a living room, dining room, and hallway. Another example is that within a vehicle, the driver's seat, front passenger seat, first rear seat, and second rear seat areas can be considered different zones. Another example is that in a large conference room, the podium area, audience area, and discussion area can be considered different zones.

[0103] In some embodiments, a sound pickup device can be used to capture sound signals within a space. A space can refer to any physical environment or area. For example, indoor locations such as enclosed or semi-enclosed rooms, conference rooms, classrooms, and recording studios; outdoor environments such as squares, parks, and streets; or specific small areas such as the performance area on a stage or the interior of a vehicle.

[0104] Step 220 : Determine a target audio signal after filtering the audio signals based on the audio signals of the multiple regions and the region position information of the multiple regions.

[0105] The regional location information of multiple regions refers to the specific locations of each region in physical space and their relationships with each other. For example, the regional location information of multiple regions may include the absolute location of each region, the relative location information between different regions, the boundaries and ranges of the regions, etc.

[0106] The target audio signal is the audio signal of the target sound source to be identified. For example, in a conference room with multiple speakers, the target audio signal could be the speech of a specific speaker. Another example is the driver's voice commands in a noisy vehicle interior.

[0107] In some embodiments, a regional layout map of the space can be created in advance, and the location information of each area can be marked, or based on a microphone array and signal processing algorithm, the time difference and intensity difference of sound waves reaching different microphones can be analyzed to calculate the location information of the area where the audio signal corresponding to the sound source is located.

[0108] In some embodiments, based on the audio signals of multiple regions and the regional location information of the multiple regions, a target audio signal after the audio signal filtering is determined in various ways. For example, the audio signals of multiple regions and the regional location information of multiple regions can be processed based on a deep learning model (such as a convolutional neural network (CNN) or a recurrent neural network (RNN)) to obtain the target audio signal.

[0109] In some embodiments of this specification, by fusing the position information of different areas and capturing the spatial relationship between different areas, the interaction between the areas is better understood, the relationship between the audio signals captured by different microphones is clarified, and the target audio signal can be separated more accurately.

[0110] In some embodiments, determining a target audio signal after filtering the audio signals based on the audio signals of the multiple regions and the region position information of the multiple regions includes:

[0111] determining mask data of the audio signal based on the audio signal and region position information of the plurality of regions;

[0112] The audio signal is filtered based on the mask data to determine a target audio signal after the audio signal is filtered.

[0113] The mask data of the audio signal is a weight or filter calculated based on the relationship between the target audio signal and the interference signal. For example, the mask data can be an ideal ratio mask, an ideal amplitude mask, etc.

[0114] The mask data is used to modify the amplitude information of the spectral characteristics of the standard audio signal without modifying the phase information of the spectral characteristics. In other words, based on real-valued masking, the amplitude information corresponding to each time-frequency unit can be filtered. In some embodiments, real-valued masking is referred to as real-valued time-frequency masking.

[0115] In some embodiments, the mask data can be represented in the form of a real-valued masking matrix. The real-valued masking matrix is a multidimensional matrix composed of at least two real-valued masks. There is a corresponding relationship between the real-valued masks in the real-valued masking matrix and the time-frequency units. In some embodiments, the real-valued masks in the real-valued masking information are in a one-to-one correspondence with the time-frequency units, that is, one time-frequency unit in the spectrum feature corresponds to one real-valued mask, and at the same time, one real-valued mask in the real-valued masking information corresponds to one time-frequency unit in the spectrum feature. If the spectrum feature includes n time-frequency units, then the real-valued masking information includes n real-valued masks.

[0116] When obtaining the spectral features, such as obtaining the spectral features through short-time Fourier transform, the original audio signal is divided into multiple time periods, and the spectral features in each time period may include one or more time-frequency units.

[0117] In some embodiments, the audio signal and the region position information of the multiple regions may be processed based on a deep learning model or other methods (such as Wiener filtering) to estimate mask data.

[0118] In some embodiments, the estimated mask data may be multiplied by the spectrum of the audio signal to obtain the spectrum of the target audio signal; and the spectrum of the target audio signal may be restored to the target audio signal in the time domain by methods such as inverse short-time Fourier transform.

[0119] In some embodiments of the present specification, mask data is determined by using the location information of multiple areas where audio signals exist and the corresponding multiple audio signals to separate the target audio signal. This allows for separation and processing of audio signals in most scenarios where aliasing sounds exist, thereby improving the accuracy of audio signal separation.

[0120] In some embodiments, determining mask data of the audio signal based on the audio signal and region position information of the plurality of regions includes:

[0121] Based on the regional location information of the multiple regions, weight features of the multiple regions are determined, and the weight features are used to characterize the interaction between different regions;

[0122] Based on the weighted features, the audio signals of multiple regions are processed to obtain mask data of the audio signals.

[0123] A weighted feature represents a property or metric that connects two regions. For example, it can represent the degree of interaction or influence between different regions. For example, a weighted feature can be related to the distance or angle between regions, the similarity or correlation of audio signals in different regions, the energy ratio or phase difference between signals in different regions, and so on.

[0124] In some embodiments, the weight feature corresponding to the relative position information may be a weight matrix, where each element Wij in the weight matrix represents the degree of influence of region i on region j.

[0125] In some embodiments, different spatial regions have corresponding weight features that are different.

[0126] In some embodiments of the present specification, by determining weight features corresponding to relative position information, the degree of influence of audio signals in different regions on a target audio signal can be determined, which helps to more accurately separate the target audio signal.

[0127] In some embodiments, determining weight features of the plurality of regions based on the region location information of the plurality of regions includes:

[0128] Determining relative position information of the plurality of regions based on the regional position information of the plurality of regions;

[0129] The relative position information of the multiple regions is used as an initial weight feature, and the initial weight feature is optimized by a preset method to obtain the weight features of the multiple regions.

[0130] The relative position information of multiple regions refers to the spatial position relationship of each region in space. For example, the relative position information of multiple regions may include the straight-line distance between two regions, the angle or azimuth between two regions, etc.

[0131] In some embodiments, in a fixed-layout environment (such as a vehicle or a conference room), relative position information can be directly defined based on the known physical layout, or the relative positions between areas can be estimated by analyzing the propagation delay and intensity difference of the audio signal.

[0132] In some embodiments, the preset method may be a statistical or machine learning method, etc.

[0133] In some embodiments of this specification, determining the relative positions of different regions facilitates subsequent fusion of audio features of different regions, thereby improving separation effect.

[0134] In some embodiments, mask data is obtained based on an audio signal and region position information of a plurality of regions, and the method includes:

[0135] Extracting features from the audio signal to obtain a first audio feature of the audio signal;

[0136] Mask data is obtained based on the first audio feature and the region feature.

[0137] The first audio feature refers to a feature or vector extracted from the audio signal that describes the characteristics of the audio signal. For example, the first audio feature may include time domain features (such as amplitude features, phase features, signal energy distribution, and zero-crossing rate), frequency domain features (such as spectrum, spectrum bandwidth, and Mel-frequency cepstral coefficients), and the like.

[0138] In some embodiments, the audio features and the regional features are fused to obtain fused features. For example, the audio features and the regional features are fused using a neural network (such as a Transformer); and mask data is generated through a machine learning model (such as a deep learning model).

[0139] The mask data may be a data structure having the same dimension as the audio signal, and different values of the mask data may be used to indicate a target audio signal or background noise or other interfering signals.

[0140] In some embodiments, the region location information of the plurality of regions is graph structure data, and the graph structure data includes nodes and edges;

[0141] A node in the regional location information of the plurality of regions corresponds to at least one region, and a node feature of the node is related to the location information of the corresponding region;

[0142] The edges in the region position information of the multiple regions are related to the associated regions, and the edge features of the edges include weight features corresponding to the relative position information between the associated regions.

[0143] Figure 3 This is an exemplary schematic diagram of graph structure data according to some embodiments of this specification.

[0144] like Figure 3 As shown, a node represents an area in a space. Node features can reflect information related to the corresponding area. For example, the node features can include the location information of the area.

[0145] There is an edge between any two regions. Edge features can reflect the degree of influence between different regions. For example, edge features can include weight features between different regions.

[0146] In some embodiments, graph structure data can be processed based on a machine learning model (such as a graph neural network) to extract weight features.

[0147] In some embodiments, graph convolution may be performed on graph structure data based on a Laplacian operator to obtain weight features.

[0148] Graph convolution can further enhance the separation model's understanding of the spatial relationship between multiple regions. By performing convolution operations on the graph structure to achieve feature extraction, the separation model can effectively integrate feature information from different sound zones.

[0149] In some embodiments, the region location information of the plurality of regions may be a causal graph describing the causal relationship between the regions.

[0150] In some embodiments of this specification, by considering the relative position relationship between different sound zones and their potential mutual influence, the audio features of different sound zones can be effectively fused through graph modeling and graph convolution, which helps to improve the separation effect.

[0151] In some embodiments, obtaining mask data based on the first audio feature and the weight feature includes:

[0152] Obtaining a hidden audio feature of the audio signal based on the first audio feature and the weight feature;

[0153] Based on the hidden audio features, mask data is obtained.

[0154] Hidden audio features refer to potential features extracted from the first audio features and weighted features. Hidden audio features can capture the deep structure and semantic information of audio signals.

[0155] In some embodiments, hidden audio features can be extracted from audio features and weight features through a deep learning model (such as a convolutional neural network, a recurrent neural network, or a Transformer, etc.)

[0156] In some embodiments, the hidden audio features may be used as input to a deep learning network to generate mask data.

[0157] In some embodiments of this specification, by hiding audio features, high-level features are extracted from audio features and position information of multiple regions, which can capture complex information and structure in audio signals; by generating mask data through hidden audio features, the robustness and effectiveness of feature representation can be further enhanced.

[0158] In some embodiments, obtaining a hidden audio feature of the audio signal based on the first audio feature and the weight feature includes:

[0159] Based on the first audio feature and the weight feature of the first time period, a hidden audio feature of the second time period is obtained, where the second time period is after the first time period.

[0160] In some embodiments, the second time period is a time period subsequent to the first time period.

[0161] In some embodiments, each time an audio signal of a time period is obtained, the first audio feature of the audio signal of the time period is extracted to obtain the first audio features of multiple time periods.

[0162] In some embodiments, the time period is a fixed time period, or the time period is a dynamically changing time period. In some embodiments, the electronic device may set an initial time period for extracting the audio signal from the buffer; each time an audio signal of a time period is obtained, feature extraction is performed on the audio signal to obtain a first audio feature.

[0163] In some embodiments, before extracting hidden audio features, it is usually necessary to pre-process the original audio signal, including: dividing the original audio signal into multiple time periods, each time period containing an audio signal within a certain time range.

[0164] In some embodiments of the present specification, the hidden audio features obtained after processing by a recursive neural network (such as a recurrent neural network, a long short-term memory network, or a gated recurrent unit) can capture the long-term dependencies and complex temporal relationships in the audio signal, are more expressive than the original audio features (such as Mel-frequency cepstral coefficients), and can better reflect the intrinsic structure of the audio signal.

[0165] In some embodiments, obtaining the hidden audio feature of the second time period based on the first audio feature and the weight feature of the first time period includes:

[0166] Performing feature fusion based on the first audio feature and the weight feature of the first time period to obtain a first fused feature of the first time period;

[0167] Perform feature fusion based on the preset hidden audio feature and the weight feature to obtain a second fused feature for the first time period;

[0168] Feature extraction is performed based on the first fusion feature of the first time period and the second fusion feature of the first time period to obtain the hidden audio feature of the second time period.

[0169] In some embodiments, the hidden audio feature of the last time period may be determined based on at least one round of iteration, the at least one round of iteration comprising: performing feature fusion based on the first audio feature of the current time period and the weight feature to obtain a first fused feature of the current time period;

[0170] Perform feature fusion based on the hidden audio features and weight features of the current time period to obtain a second fused feature of the current time period;

[0171] Feature extraction is performed based on the first fusion feature of the current time period and the second fusion feature of the current time period to obtain the hidden audio feature of the next time period.

[0172] In some embodiments, the input to an iteration is related to the iteration round. For the first iteration round, the input to the iteration includes the first audio feature and weight feature of the first time period; for subsequent iteration rounds, the input to the iteration includes the first audio feature, weight feature, and hidden audio feature of the current time period.

[0173] In some embodiments, in each iteration, the output of the iteration is the hidden audio features for the next iteration.

[0174] In some embodiments, in a first round of iteration, feature fusion is performed based on the first audio feature of the first time period and the weight feature to obtain a first fused feature of the first time period;

[0175] Perform feature fusion based on the preset hidden audio features and weight features to obtain the second fused features of the first time period;

[0176] Feature extraction is performed based on the first fusion feature of the first time period and the second fusion feature of the first time period to obtain the hidden audio feature of the next time period, and it is determined whether the preset conditions are met; in response to not meeting the conditions, the next round of iteration is continued.

[0177] In some embodiments, in at least one non-last iteration of iteration, the first audio feature of the current time period, the weight feature, and the hidden audio feature of the current time period output from the previous iteration are obtained; feature fusion is performed based on the first audio feature of the current time period and the weight feature to obtain the first fused feature of the current time period; feature fusion is performed based on the hidden audio feature of the current time period and the weight feature to obtain the second fused feature of the current time period; feature extraction is performed based on the first fused feature of the current time period and the second fused feature of the current time period to obtain the hidden audio feature of the next time period, and the hidden audio feature of the next time period is used as the hidden audio feature of the next iteration, and the next iteration is continued. In the last iteration, it can be determined whether the preset conditions are met based on the last round; in response to the first audio feature corresponding to the last round being the last time period, the hidden audio feature of the last time period is determined, and the iteration is terminated.

[0178] The preset condition is a judgment condition for evaluating whether the iteration is terminated. For example, the preset condition may include the absence of an unprocessed first audio feature.

[0179] In some embodiments, obtaining mask data based on hidden audio features includes:

[0180] Mask data is obtained based on a target hidden audio feature among the multiple hidden audio features.

[0181] The target hidden audio feature refers to the hidden audio feature corresponding to the audio signal in the last time period.

[0182] Figure 4 is an exemplary schematic diagram of determining a target audio signal according to some embodiments of this specification.

[0183] In some embodiments, as Figure 4 As shown, graph convolution is performed on the causal graph to obtain weight features, and feature extraction is performed on the second audio features in sequence through multiple convolution kernels to obtain first audio features. Feature fusion is performed on the first audio features and the weight features, and the result of the feature fusion is input into the forward network to obtain mask data. The audio signals of multiple regions are post-processed based on the mask data to obtain the target audio signal.

[0184] Figure 5 is an exemplary schematic diagram of determining hidden audio features according to some embodiments of this specification.

[0185] In some embodiments, as Figure 5As shown, the original audio signal is divided into several time periods in chronological order. In time period t, information is transferred between the weight feature and the first audio feature of time period t obtained by the convolution kernel. For example, the first product between the two matrices of the weight feature and the first audio feature of time period t is calculated. At the same time, the second product of the weight feature and the hidden audio feature of time period t is calculated. Finally, the first product and the second product are multiplied by the gated loop unit and activation mechanism of the recurrent neural network (RNN) to obtain the hidden audio feature of the next time period t+1. The hidden audio feature finally output is the hidden audio feature of the last time period.

[0186] The hidden audio features of the last time period are passed through a feedforward network to obtain the final mask data. The feedforward network is used to convert the features into mask data to separate the target audio signal from background noise or other interfering signals.

[0187] In some embodiments of this specification, by calculating the attention between different regions, the separation model can combine the spatial relationship between different regions with the first audio features in the current time period, thereby better capturing the spatiotemporal features; through the gated recurrent unit of the recurrent neural network, the separation model can effectively memorize and transmit information, thereby improving the continuity and consistency of feature fusion.

[0188] In some embodiments, extracting features from the audio signal to obtain a first audio feature includes:

[0189] Performing frequency domain analysis on the audio signal to obtain a second audio feature of the audio signal;

[0190] Feature extraction is performed based on the second audio feature to obtain the first audio feature.

[0191] Frequency domain analysis refers to the process of converting audio signals from the time domain to the frequency domain to extract the spectral characteristics of the signal.

[0192] The second audio feature refers to a spectrum feature obtained through frequency domain analysis.

[0193] In some embodiments, the original audio signal is divided into multiple time periods, and Fourier transform is performed on the audio signal of each time period to obtain the second audio feature of the time period.

[0194] For example, the original audio signal can be divided into frames based on the length of the time period to obtain the audio signal of each frame. Subsequently, the spectral features corresponding to each frame of the audio signal can be extracted using discrete Fourier transform, wavelet transform, or Mel-frequency cepstrum transform, where the spectral features mainly include amplitude features and phase features.

[0195] In addition, a certain overlap rate can be set, for example, 50%, to ensure the continuity of each frame of audio signal. At the same time, in order to ensure that each frame of audio signal has the same frame length, the audio signal at the end of the frame or the insufficient frame length can be padded with zeros to ensure that the frame length of each audio signal is the same.

[0196] The first audio feature refers to a higher-level feature extracted from the second audio feature.

[0197] In some embodiments, the input audio signal is preprocessed. For example, the electronic device can collect analog signals from different regions through a sound pickup device and convert the obtained analog signals into digital audio signals for subsequent processing; based on the digital audio signals, the electronic device can detect and remove silent segments to reduce interference from invalid data; and the audio signals are normalized to ensure that the audio signals in different regions are on the same scale. For example, filtering and echo cancellation techniques can be used to reduce the impact of background noise and improve the quality of the audio signal. The preprocessed audio signal can be subjected to fast Fourier transform to extract the first audio feature. The fast Fourier transform can convert the time domain signal into the frequency domain signal, facilitating subsequent feature extraction and analysis.

[0198] In some embodiments, the first audio feature can be applied to several convolution kernels to extract correlation features between each point in the sequence of the first audio feature and several adjacent points. The convolution operation using the convolution kernels can extract local correlations in the audio signal, thereby better capturing key features in the audio signal. This leverages local contextual information to enhance the model's ability to separate audio signals from different sound sources.

[0199] In some embodiments, the region position information of multiple regions may be processed by a Laplacian operator to obtain a result of graph convolution, that is, a weight feature.

[0200] In some embodiments, the second audio feature may be input into the first audio feature obtained by the neural network model.

[0201] In some embodiments of this specification, hierarchical feature extraction can better capture the deep structure and semantic information of the audio signal, which helps to obtain feature representations that are more suitable for subsequent tasks (such as speech separation, etc.).

[0202] In some embodiments, filtering the audio signal based on the mask data to determine the target audio signal after filtering the audio signal includes:

[0203] Based on the product of the mask data and the audio signal, a target audio signal is obtained.

[0204] In some embodiments, the electronic device can determine the mask data of the audio signal to be processed at one time based on the separation model, multiply the mask data by the frequency domain characteristics of the audio signal element by element to obtain the frequency domain characteristics of the target audio signal, and perform time-frequency analysis on the frequency domain characteristics of the target audio signal, such as inverse short-time Fourier transform, to obtain the target audio signal originally in the time domain.

[0205] In some embodiments of the present specification, a target audio signal can be effectively extracted from a complex audio signal by performing a multiplication operation on mask data and an audio signal, thereby improving the performance and effect of audio processing.

[0206] In some embodiments, the target audio signal after the audio signal is filtered is obtained by processing the audio signals of multiple regions and the region position information of the multiple regions using a preset separation model.

[0207] The preset separation model may be a preset neural network model.

[0208] In some embodiments, the input of the preset neural network model may include audio signals of multiple regions and regional position information of multiple regions. In some embodiments, the output of the preset neural network model may include the target audio signal, or mask data of the audio signal.

[0209] In some embodiments, regional location information of multiple regions is displayed via a terminal. The terminal may include a vehicle display terminal or a user terminal. A vehicle display terminal refers to a display device installed inside a vehicle.

[0210] A user terminal refers to one or more terminal devices or software used by a user. In some embodiments, a user terminal may be used by one or more users, including users who directly use the relevant service or other related users. In some embodiments, a user terminal may be a mobile device, tablet computer, laptop computer, desktop computer, or any combination of other devices with input and / or output functions.

[0211] In some embodiments, the step of processing the audio signals of multiple regions and the region position information of the multiple regions using a preset separation model to obtain a target audio signal includes:

[0212] Based on the audio signals of the multiple regions and the regional position information of the multiple regions, mask data of the audio signals are obtained through a separation model;

[0213] The audio signal is filtered based on the mask data to obtain a target audio signal.

[0214] The separation model is a model or algorithm used to determine mask data.

[0215] In some embodiments, the separation model is a machine learning model. For example, the separation model may include any one or combination of a convolutional neural network (CNN) model, a neural network (NN) model, or other customized model structures.

[0216] In some embodiments, the input of the separation model includes audio signals of multiple regions and region position information of the multiple regions, and the output may include mask data.

[0217] In some embodiments, the separation model can be trained based on a large number of training samples with labeled data through various feasible methods. For example, parameter updates can be performed based on the gradient descent method. An exemplary training process includes: inputting multiple labeled training samples into the initial separation model, constructing a loss function based on the labels and the results of the initial separation model, and iteratively updating the parameters of the initial separation model based on the loss function through gradient descent or other methods. When the preset conditions are met, the model training is completed and a trained separation model is obtained. The preset conditions may be the convergence of the loss function, the number of iterations reaching a threshold, etc.

[0218] In some embodiments, the training samples include at least sample audio signals of multiple regions and sample region position information of the multiple regions. The training samples can be obtained based on historical data.

[0219] In some embodiments, the label data may include actual mask data corresponding to the training sample. The label data may be obtained through automatic or manual annotation.

[0220] In some embodiments of this specification, by introducing a causal graph corresponding to the regional location information of multiple regions, the separation model not only improves robustness but also enhances interpretability. The causal graph can intuitively display the causal relationship between different audio signals, making the decision-making process of the separation model more transparent. The enhanced interpretability enables users to better understand the working principle of the separation model, making it easier to debug and optimize. For example, by analyzing the distribution of weight features in the causal graph, it is possible to discover the degree to which different regions affect the separation effect, and then perform targeted optimization.

[0221] In some embodiments, obtaining mask data of the audio signal based on the audio signals of the multiple regions and the region position information of the multiple regions by using a separation model includes:

[0222] Extracting features of the audio signal based on the first network layer of the separation model to obtain a first audio feature of the audio signal;

[0223] The second network layer based on the separation model extracts features of the regional location information of multiple regions to obtain weight features of the multiple regions;

[0224] Performing feature fusion on the first audio feature and the weight feature based on the third network layer of the separation model to obtain a hidden audio feature of the audio signal;

[0225] The fourth network layer based on the separation model processes the hidden audio features to obtain mask data.

[0226] In some embodiments, the separation model may include a first network layer, a second network layer, a third network layer, and a fourth network layer.

[0227] The first network layer is a model for determining a first audio feature.

[0228] In some embodiments, the first network layer may be a machine learning model, for example, the first network layer may be a convolutional neural network (CNN).

[0229] In some embodiments, the input to the first network layer may include the second audio feature; and the output may include the first audio feature.

[0230] The second network layer is a model for determining weight features.

[0231] In some embodiments, the second network layer may be a machine learning model, for example, the second network layer may be a graph neural network model (GNN).

[0232] In some embodiments, the input of the second network layer may include regional location information of multiple regions; and the output may include weight features.

[0233] The third network layer is a model for determining hidden audio features.

[0234] In some embodiments, the third network layer may be a machine learning model, for example, the third network layer may be a convolutional neural network (CNN).

[0235] In some embodiments, the input of the third network layer may include the first audio feature and the weight feature; the output may include the hidden audio feature.

[0236] The fourth network layer is a model for determining mask data.

[0237] In some embodiments, the fourth network layer may be a machine learning model, for example, the fourth network layer may be a convolutional neural network (CNN).

[0238] In some embodiments, the input to the fourth network layer may include target hidden audio features; the output may include mask data.

[0239] In some embodiments, the separation model can be obtained by jointly training the first network layer, the second network layer, the third network layer, and the fourth network layer based on a large amount of training samples with labeled data. The training samples used for joint training include sample audio signals of multiple regions and sample region location information of multiple regions. The training samples can be obtained based on historical data.

[0240] An exemplary joint training process includes: inputting sample audio signals from multiple regions of the training sample into an initial first network layer, and inputting sample region location information from multiple regions into an initial second network layer; inputting the outputs of the initial first and second network layers into an initial third network layer; inputting the outputs of the initial third network layer into an initial fourth network layer; constructing a loss function based on the output and label of the initial fourth network layer, and simultaneously updating the parameters of the initial first, second, third, and fourth network layers until a preset condition is met and training is completed. The preset condition may be that the loss function is less than a threshold, converges, or the training cycle reaches a threshold.

[0241] In some embodiments of the present specification, the joint training of the initial first network layer, the initial second network layer, the initial third network layer, and the initial fourth network layer is beneficial to solving the problem of difficulty in obtaining labels when training the adjustment model alone, improving the training efficiency of the adjustment model, and reducing the training difficulty.

[0242] In some embodiments, the separation model is trained by the following steps:

[0243] Inputting the sample audio signals of the plurality of regions and the sample region position information of the plurality of regions into the initial model to obtain sample mask data of the sample audio signals;

[0244] Filtering the sample audio signal based on the sample mask data to obtain a predicted audio signal;

[0245] Determine the loss value of the initial model based on the labeled data of the predicted audio signal and the sample audio signal;

[0246] The parameters of the initial model are trained according to the loss value to obtain the separation model.

[0247] Figure 6is an exemplary schematic diagram of a training model according to some embodiments of this specification.

[0248] In some embodiments, as Figure 6 As shown, the separation model is trained in the following way:

[0249] Step 601: Randomly initialize model parameters.

[0250] In some embodiments, the parameters of the separation model can be randomly initialized using the Xavier random initialization strategy to ensure that the outputs of each layer in the neural network model have the same variance, thereby alleviating the gradient vanishing problem.

[0251] Step 602: Initialize the initial weight feature to the relative distance between regions.

[0252] In some embodiments, the initial weight features of the edges in the causal graph are initialized to the distances between associated regions.

[0253] Step 603: Input the training samples into the initial model.

[0254] In some embodiments, the preprocessed first audio features and the initial weight features corresponding to the causal graph are input into an already initialized model (e.g., an initialized deep learning model), and forward propagation is performed through the various layers of the deep learning model. During the forward propagation process, the deep learning model calculates intermediate features based on the current weight parameters and ultimately outputs predicted sample mask data. The sample audio signal in the training sample is filtered based on the predicted sample mask data to obtain a predicted audio signal.

[0255] Step 604: Calculate the loss value according to the loss function.

[0256] In some embodiments, a SISNR error function is used to calculate the error between the predicted audio signal and the labeled data to obtain a loss value.

[0257] Step 605: Optimize the parameters of the initial model according to the loss value.

[0258] In some embodiments, the model parameters are optimized using stochastic gradient descent based on the loss value. For example, the chain rule is used to calculate the gradient of the loss function with respect to each layer parameter, starting from the output layer and continuing through the network layer until the input layer. The value of the new parameter is the value of the old parameter minus n times the gradient of the loss function with respect to the old parameter, where n represents the learning rate. For example, the learning rate is set to 1e-4.

[0259] Step 606: Determine whether convergence has occurred.

[0260] In some embodiments, the error trend on the training samples is statistically analyzed, i.e., a trend chart of the error change over the training rounds is plotted to see if there is a clear downward trend. If the error fluctuates within a small range and no longer decreases significantly, the separation model is considered to have reached a good training state and the training process can be stopped. Otherwise, it is considered that there is still room for significant error reduction and the training process is continued.

[0261] Step 607: Optimized separation model.

[0262] In some embodiments, the training process ends when the convergence condition is reached, and the parameters of the separation model obtained at this time are considered to be the optimal solution and can be used for subsequent testing and practical applications.

[0263] In some embodiments, the method further comprises:

[0264] The sample weight features of the sample area position information of multiple regions are trained based on the loss value.

[0265] By automatically adjusting the sample weight features in the causal graph to learn the causal relationships between different audio signals, the separation model can better cope with complex noisy environments and improve robustness. Furthermore, the separation model can dynamically adjust the mutual influence between different audio signals, thereby improving the accuracy and stability of the separation.

[0266] In some embodiments of this specification, weight features are learnable using graph-structured data, with their initial values being the relative distances between regions. This graph-structured data enables the separation model to capture the spatial relationships between regions. By leveraging the nodes and edges in the graph-structured data, the separation model can better understand the interactions between regions and perform feature fusion and separation.

[0267] In some embodiments, the plurality of zones corresponds to physical spaces within a cabin within a vehicle.

[0268] In some embodiments, the plurality of areas include: a driver's seat area, a passenger seat area, a first rear row area, and a second rear row area.

[0269] The main driving area refers to the area where the driver is located.

[0270] The passenger area refers to the area where the co-pilot is located.

[0271] The first rear area and the second rear area are the right rear area of the vehicle and the left rear area of the vehicle.

[0272] In some embodiments, the division of areas can be adjusted according to the specific layout of the vehicle. For example, the division logic can be defined according to the seat position.

[0273] In some embodiments, by mapping multiple areas to the physical space of the vehicle cabin and separating the audio signals of different areas, the audio separation effect and voice interaction experience in the vehicle cabin can be significantly improved.

[0274] It should be noted that the above description of the relevant processes is for illustration and purpose only and does not limit the scope of application of this specification. For those skilled in the art, various modifications and changes can be made to the processes under the guidance of this specification. However, such modifications and changes are still within the scope of this specification.

[0275] Figure 7 It is a structural diagram of a speech separation system according to some embodiments of this specification.

[0276] like Figure 7 As shown, one or more embodiments of this specification also provide a structural diagram of a speech separation system. The speech separation system may include a separation device 701 for determining a target audio signal after filtering the audio signal based on the audio signals of multiple regions and the regional position information of the multiple regions.

[0277] In some embodiments, the system further includes a plurality of sound pickup devices 702 , which are distributed in a plurality of areas and are used to acquire audio signals from the plurality of areas.

[0278] The separation device 701 can process data and / or information obtained from other devices or system components. The processor can execute program instructions based on this data, information, and / or processing results to perform one or more functions described in this specification. For example, the separation device 701 can determine a target audio signal after filtering the audio signals based on the audio signals of multiple regions and the regional location information of the multiple regions. For more details, please refer to the relevant description above.

[0279] In some embodiments, the separation device 701 can be a single server or a group of servers, such as centralized or distributed. In some embodiments, the separation device 701 can be local or remote. In some embodiments, the separation device 701 can be implemented on a cloud platform. By way of example only, the cloud platform can include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an internal cloud, a multi-layer cloud, or any combination thereof.

[0280] In some embodiments, the separation device 701 can be integrated into a vehicle terminal. For example, the separation device 701 can be a computing device installed in a vehicle (e.g., an onboard computer). The vehicle can include, but is not limited to, various types of vehicles, such as gasoline-powered vehicles, electric vehicles, and trucks.

[0281] The sound pickup device 702 is used to obtain the user's audio signal to pick up the user's sound. In some embodiments, the sound pickup device 702 can be fixedly arranged at one or more locations in a space. For example, the sound pickup device 702 is arranged on the interior of the vehicle to improve the driver's driving safety and make the vehicle interior more beautiful and neat.

[0282] In some embodiments, the sound pickup device 702 may include multiple microphones arranged in a space. After the multiple microphones pick up sounds, the captured audio signals in the area are sent to the electronic device for processing to obtain the target audio signal.

[0283] For example, the sound pickup device 702 can be installed on the vehicle's interior, such as the center console, A / B pillars, front seats, rear seats, and the vehicle ceiling, to improve the sound pickup device's ability to capture sound signals, thereby improving the effectiveness of subsequent sound separation. The location of the sound pickup device is merely an example. In other embodiments, the sound pickup device can also be installed elsewhere in the vehicle or in other locations, which will not be detailed here.

[0284] Among them, the separation device 701 can be used to respectively execute the steps in the embodiments corresponding to the above-mentioned speech separation method. For the specific implementation methods of these modules and more details, please refer to the corresponding method part, which will not be repeated here.

[0285] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0286] Figure 8 This is a schematic diagram of the structure of an electronic device according to some embodiments of this specification. Figure 8 As shown, the electronic device 800 may include: a processor 801, a memory 802. The electronic device 800 may also include one or more of a multimedia component 803, an input / output (I / O) component 804, and a communication component 805. In this embodiment, the electronic device 800 may be a device that implements the speech separation method provided in this embodiment.

[0287] The processor 801 is used to control the overall operation of the electronic device 800 to complete all or part of the steps in the above-mentioned speech separation method. The memory 802 is used to store various types of data to support the operation of the electronic device 800. For example, this data may include instructions for any application or method operating on the electronic device 800, as well as application-related data, such as contact information, sent and received messages, pictures, audio, video, etc. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 802 or transmitted via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. The I / O component 804 provides an interface between the processor 801 and other interface modules, which may include a keyboard, a mouse, buttons, etc. These buttons may be virtual or physical buttons. The communication component 805 is used for wired or wireless communication between the electronic device 800 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, Narrow Band Internet of Things (NB-IOT), Enhanced Machine Type Communication (eMTC), or other 5G technologies, or a combination thereof, is not limited here. Therefore, the corresponding communication component 805 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.

[0288] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-mentioned speech separation method.

[0289] In another exemplary embodiment, a computer-readable storage medium is provided, which stores a computer program. When the program instructions are executed by a processor, the steps of the above-mentioned speech separation method are implemented. For example, the computer-readable storage medium may be the memory 802 containing the program instructions. The program instructions may be executed by the processor 801 of the electronic device 800 to implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application.

[0290] Alternatively, when the instructions are executed by a computer, they implement or execute the methods, steps, and logic diagrams disclosed in the embodiments of this application.

[0291] In another exemplary embodiment, a computer program product is also provided, including a computer program or instructions, which, when executed by a processor, implement the steps of the above-mentioned speech separation method. For example, the computer program product may be the aforementioned memory 802 including the computer program, which may be executed by the processor 801 of the electronic device 800 to implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0292] Alternatively, when the instructions are executed by a computer, they implement or execute the methods, steps, and logic diagrams disclosed in the embodiments of this application.

[0293] Figure 9 is an exemplary schematic diagram of a vehicle according to some embodiments of the present specification.

[0294] like Figure 9 As shown, the present application also provides a vehicle, which is equipped with the electronic device provided by any of the above embodiments, and the electronic device is used to perform the speech separation method provided by any of the above embodiments. The vehicle can be a fuel vehicle, a plug-in hybrid vehicle, or a new energy vehicle, etc., which is not specifically limited in this specification.

[0295] In one embodiment, a vehicle can be configured for a fully or partially autonomous driving mode. For example, while in autonomous driving mode, the vehicle can control itself and, through human interaction, determine the current state of the vehicle and its surroundings, determine the possible behavior of at least one other vehicle in the surroundings, and determine a confidence level corresponding to the likelihood that the other vehicle will perform the possible behavior, and control the vehicle based on this information. While in autonomous driving mode, the vehicle can be configured to operate without human interaction.

[0296] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.

[0297] The embodiments, implementation methods and related technical features of the present application can be combined and replaced with each other without conflict.

[0298] The above are only preferred embodiments of the present application and do not constitute any form of limitation to the present application. Although the descriptions of each embodiment in the embodiments of the present application have different focuses, for parts that are not described in detail in a certain embodiment, please refer to the relevant embodiments of other embodiments. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the technical solution of the present application.

Claims

1. A speech separation method, characterized in that: The method comprises: Based on the audio signals of the multiple regions and the region position information of the multiple regions, a target audio signal after the audio signals are filtered is determined.

2. The method according to claim 1, characterized in that The step of determining the target audio signal after filtering the audio signal based on the audio signals of the plurality of regions and the region position information of the plurality of regions includes: determining mask data of the audio signal based on the audio signal and the region position information of the plurality of regions; The audio signal is filtered based on the mask data to determine a target audio signal after the audio signal is filtered.

3. The method according to claim 2, characterized in that The determining, based on the audio signal and the region position information of the plurality of regions, mask data of the audio signal includes: Determining weight features of the multiple regions based on the regional location information of the multiple regions, wherein the weight features are used to characterize interactions between different regions; Based on the weight features, the audio signals of the multiple regions are processed to obtain mask data of the audio signals.

4. The method according to claim 3, characterized in that The determining, based on the regional location information of the multiple regions, weight features of the multiple regions includes: Determining relative position information of the multiple areas based on the area position information of the multiple areas; The relative position information of the multiple regions is used as an initial weight feature, and the initial weight feature is optimized by a preset method to obtain the weight features of the multiple regions.

5. The method according to claim 3, characterized in that The method of processing the audio signals of the multiple regions based on the weight features to obtain mask data of the audio signals includes: Performing feature extraction on the audio signal to obtain a first audio feature of the audio signal; The mask data is obtained based on the first audio feature and the weight feature.

6. The method according to claim 5, characterized in that The obtaining the mask data based on the first audio feature and the weight feature includes: Obtaining a hidden audio feature of the audio signal based on the first audio feature and the weight feature; The mask data is obtained based on the hidden audio feature.

7. The method according to claim 6, characterized in that The obtaining of the hidden audio feature of the audio signal based on the first audio feature and the weight feature includes: The hidden audio feature of a second time period is obtained based on the first audio feature and the weight feature of a first time period, where the second time period is after the first time period.

8. The method according to claim 7, characterized in that The obtaining, based on the first audio feature and the weight feature of the first time period, the hidden audio feature of the second time period includes: Performing feature fusion based on the first audio feature of the first time period and the weight feature to obtain a first fused feature of the first time period; Performing feature fusion based on the preset hidden audio feature and the weight feature to obtain a second fused feature for the first time period; Feature extraction is performed based on the first fusion feature of the first time period and the second fusion feature of the first time period to obtain the hidden audio feature of the second time period.

9. The method according to claim 6, characterized in that The obtaining of the mask data based on the hidden audio feature includes: The mask data is obtained based on a target hidden audio feature among the plurality of hidden audio features.

10. The method according to claim 6, characterized in that The extracting features from the audio signal to obtain a first audio feature of the audio signal includes: Performing frequency domain analysis on the audio signal to obtain a second audio feature of the audio signal; Feature extraction is performed based on the second audio feature to obtain the first audio feature.

11. The method according to claim 4, characterized in that The filtering the audio signal based on the mask data to determine the target audio signal after the audio signal is filtered includes: The target audio signal is obtained based on the product of the mask data and the audio signal.

12. The method according to claim 1, characterized in that The target audio signal after the audio signal is filtered is obtained by processing the audio signals of the multiple regions and the region position information of the multiple regions through a preset separation model.

13. The method according to claim 12, characterized in that The step of processing the audio signals of the multiple regions and the region position information of the multiple regions using a preset separation model to obtain the target audio signal includes: Obtaining mask data of the audio signal based on the audio signals of the multiple regions and the region position information of the multiple regions using the separation model; The audio signal is filtered based on the mask data to obtain the target audio signal.

14. The method according to claim 13, characterized in that The obtaining, based on the audio signals of the multiple regions and the region position information of the multiple regions, mask data of the audio signals by the separation model includes: performing feature extraction on the audio signal based on the first network layer of the separation model to obtain a first audio feature of the audio signal; Performing feature extraction on the regional location information of the multiple regions based on the second network layer of the separation model to obtain weight features of the multiple regions; performing feature fusion on the first audio feature and the weight feature based on the third network layer of the separation model to obtain a hidden audio feature of the audio signal; The hidden audio features are processed based on the fourth network layer of the separation model to obtain the mask data.

15. The method according to claim 13, characterized in that The separation model is trained by the following steps: Inputting the sample audio signals of the plurality of regions and the sample region position information of the plurality of regions into the initial model to obtain sample mask data of the sample audio signals; filtering the sample audio signal based on the sample mask data to obtain a predicted audio signal; Determining a loss value of the initial model based on the predicted audio signal and the labeled data of the sample audio signal; The parameters of the initial model are trained according to the loss value to obtain the separation model.

16. The method according to claim 15, characterized in that The method further comprises: The sample weight features of the sample region position information of the multiple regions are trained based on the loss value.

17. The method according to claim 1, wherein The regional location information of the plurality of regions is graph structure data, and the graph structure data includes nodes and edges; A node in the region location information of the plurality of regions corresponds to at least one of the regions, and a node feature of the node is related to the location information of the corresponding region; The edges in the region position information of the multiple regions are related to the associated regions, and the edge features of the edges include weight features corresponding to the relative position information between the associated regions.

18. The method according to claim 1, wherein The plurality of regions correspond to physical spaces within a cabin within the vehicle.

19. The method according to claim 18, characterized in that The multiple areas include: a main driver area, a co-driver area, a first rear area and a second rear area.

20. A speech separation system, characterized in that: The system comprises: The separation device is used to determine a target audio signal after filtering the audio signal based on the audio signals of multiple areas and the area position information of the multiple areas.

21. The system according to claim 20, wherein: The system further includes a plurality of sound pickup devices, which are distributed in the plurality of areas and are used to acquire audio signals from the plurality of areas.

22. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 19.

23. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 19 are implemented.

24. A computer program product, characterized in that The method comprises a computer program or instructions, which implement the steps of the method according to any one of claims 1 to 19 when executed by a processor.

25. A vehicle, characterized in that: Comprising the speech separation system as claimed in claim 20, or the electronic device as claimed in claim 22.