Sound area separation method, device and equipment for audio in vehicle, storage medium and vehicle

By using the target end-to-end network model in the car for sound source attribute detection and suppression, the problem of low accuracy of sound separation in the car is solved, and the accuracy of voice interaction and user experience are improved.

CN119943083APending Publication Date: 2025-05-06BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311466107.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Due to the small space in the car and the high microphone sensitivity, it is difficult to effectively separate the sound area on the car, resulting in a reduced accuracy of voice interaction and affecting the user experience.

Method used

By obtaining the multi-channel audio signals collected by multiple sound sensors in the car, the pre-trained target end-to-end network model is used to detect sound source attributes, distinguish between real-person sound sources and non-car media sound sources, suppressing the latter to reduce interference and achieving sound separation.

Benefits of technology

It improves the accuracy of sound area recognition in the car, reduces the interference of non-car media sound sources on voice interaction, and improves the user's voice interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943083A_ABST
    Figure CN119943083A_ABST
Patent Text Reader

Abstract

The invention discloses a sound area separation method, device and equipment for in-vehicle audio, a storage medium and a vehicle. The method comprises the steps that multiple channels of audio signals collected by a plurality of sound sensors in a vehicle are obtained, the multiple channels of audio signals comprise sound signals generated by a plurality of sound sources with different sound source attributes, and different sound sensors are arranged in different sound areas; based on a pre-trained target end-to-end network model, performing sound source attribute detection on the sound signals of each sound source in the multiple paths of audio signals to obtain sound source attributes corresponding to each sound signal in the multiple paths of audio signals; the sound source attribute comprises a real person sound source attribute and a non-vehicle end media sound source attribute; suppressing a sound signal of which the sound source attribute is a non-vehicle end media sound source attribute in the multiple paths of audio signals, and obtaining a sound signal corresponding to the real person sound source attribute; and identifying a sound area of the sound signal corresponding to the real person sound source attribute in the vehicle, and performing sound area separation on the sound signal corresponding to the real person sound source attribute according to an identification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of vehicle technology, and in particular, relates to a method, device, equipment, storage medium and vehicle for separating sound zones of in-vehicle audio. Background Art

[0002] With the continuous development of vehicle technology, vehicles have more and more functions. In particular, passengers can interact with the vehicle terminal by inputting voice information. When passengers interact with the vehicle end by voice, in order to ensure that the vehicle end recognizes the voice of each passenger without being interfered by the voice of other people in the car, it is necessary to separate the received voice information according to the position in the car to realize the voice interaction between each passenger and the vehicle terminal. However, due to the small space in the car and the high sensitivity of the microphone in the car, the vehicle end collects all the sounds in the car, including but not limited to normal human voices and media sounds played by non-car terminals such as mobile phones, etc. In particular, the media sounds played by non-car terminals are very similar to real people speaking, which will interfere with the vehicle end's separation of voice information, resulting in reduced accuracy of voice zone recognition, and thus affecting the voice interaction experience of passengers in the car. Summary of the invention

[0003] The embodiments of the present application provide a method, device, equipment, storage medium and vehicle for separating sound zones of in-vehicle audio, which can improve the accuracy of in-vehicle sound zone recognition and thus enhance the user's voice interaction experience.

[0004] In a first aspect, an embodiment of the present application provides a method for separating sound zones of in-vehicle audio, the method comprising:

[0005] Acquire multi-channel audio signals collected by multiple sound sensors in the vehicle, wherein the multi-channel audio signals include sound signals generated by multiple sound sources with different sound source properties, and different sound sensors are arranged in different sound zones;

[0006] Based on the pre-trained target end-to-end network model, sound source attribute detection is performed on the sound signal of each sound source in the multi-channel audio signal to obtain the sound source attribute corresponding to each sound signal in the multi-channel audio signal; the sound source attribute includes the attribute of a real person sound source and the attribute of a non-vehicle media sound source;

[0007] Suppressing the sound signal whose sound source attribute is not the sound source attribute of the vehicle-side media in the multi-channel audio signal, and obtaining the sound signal corresponding to the sound source attribute of the real person;

[0008] The sound zone in the car of the sound signal corresponding to the real person's sound source attribute is identified, and the sound zone of the sound signal corresponding to the real person's sound source attribute is separated according to the identification result.

[0009] In a second aspect, an embodiment of the present application provides a sound zone separation device for in-car audio, the device comprising:

[0010] A first acquisition module is used to acquire multi-channel audio signals collected by multiple sound sensors in the vehicle, wherein the multi-channel audio signals include sound signals generated by multiple sound sources with different sound source properties, and different sound sensors are arranged in different sound zones;

[0011] A detection module, configured to perform sound source attribute detection on the sound signal of each sound source in the multi-channel audio signal based on a pre-trained target end-to-end network model, and obtain the sound source attribute corresponding to each sound signal in the multi-channel audio signal; the sound source attribute includes a real person sound source attribute and a non-vehicle media sound source attribute;

[0012] A processing module, used for suppressing the sound signal whose sound source attribute is a non-vehicle-side media sound source attribute in the multi-channel audio signal, and obtaining the sound signal corresponding to the real person sound source attribute;

[0013] The recognition module is used to recognize the sound zone in the car of the sound signal corresponding to the real person sound source attribute, and perform sound zone separation on the sound signal corresponding to the real person sound source attribute according to the recognition result.

[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the method for separating sound zones of in-vehicle audio as described in any one of the above is implemented.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method for separating sound zones of in-vehicle audio as described in any one of the above is implemented.

[0016] In a fifth aspect, an embodiment of the present application provides a vehicle, comprising: an electronic device, wherein the electronic device is used to implement the sound zone separation method for in-vehicle audio as described in any one of the above items.

[0017] The in-car audio sound zone separation method, apparatus, device, storage medium and vehicle of the embodiments of the present application can detect the sound source attributes of the sound signals of each sound source in the multi-channel audio signals based on the pre-trained target end-to-end network model according to the multi-channel audio signals collected by multiple sound sensors in the car, and obtain the sound source attributes corresponding to each sound signal in the multi-channel audio signals; the sound source attributes include real person sound source attributes and non-vehicle media sound source attributes; suppress the sound signals with sound source attributes of non-vehicle media sound source attributes in the multi-channel audio signals, and obtain the sound signals corresponding to the real person sound source attributes; identify the sound zone in the car for the sound signal corresponding to the real person sound source attributes, and perform sound zone separation on the sound signal corresponding to the real person sound source attributes according to the identification result. In this way, the embodiment of the present application can determine whether the sound source attribute corresponding to each sound signal in the multi-channel audio signal is a real person sound source attribute or a non-vehicle media sound source attribute based on the multi-channel audio signals collected by multiple sound sensors in the car and based on the pre-trained target end-to-end network model, and suppress the sound signal with the non-vehicle media sound source attribute to reduce the interference of the sound signal corresponding to the non-vehicle media sound source attribute in the sound zone recognition in the car, thereby improving the accuracy of the sound zone recognition in the car, and further enhancing the user's voice interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 It is a flowchart of a method for separating audio zones of in-car audio provided by an embodiment of the present application;

[0020] Figure 2 is a structural schematic diagram of a sound zone separation device for in-vehicle audio provided by another embodiment of the present application;

[0021] Figure 3 It is a structural schematic diagram of an electronic device provided in yet another embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0023] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0024] With the continuous development of vehicle technology, vehicles have more and more functions. In particular, passengers can interact with the vehicle terminal by inputting voice information. When passengers interact with the vehicle end by voice, in order to ensure that the vehicle end recognizes the voice of each passenger without being interfered by the voice of other people in the car, it is necessary to separate the received voice information according to the position in the car to realize the voice interaction between each passenger and the vehicle terminal. However, due to the small space in the car and the high sensitivity of the microphone in the car, the vehicle end collects all the sounds in the car, including but not limited to normal human voices and media sounds played by non-car terminals such as mobile phones, etc. In particular, the media sounds played by non-car terminals are very similar to real people speaking, which will interfere with the vehicle end's separation of voice information, resulting in reduced accuracy of voice zone recognition, and thus affecting the voice interaction experience of passengers in the car.

[0025] In order to solve the problems of the prior art, the embodiments of the present application provide a method, device, equipment, storage medium and vehicle for separating the sound zones of in-vehicle audio. The following first introduces the method for separating the sound zones of in-vehicle audio provided by the embodiments of the present application.

[0026] Figure 1 FIG. 1 is a flow chart of a method for separating audio zones of in-car audio provided by an embodiment of the present application. Figure 1 As shown, a method for separating audio zones of in-vehicle audio may include the following steps S101 to S104:

[0027] S101, acquiring multi-channel audio signals collected by multiple sound sensors in the vehicle, where the multi-channel audio signals include sound signals generated by multiple sound sources with different sound source properties, and different sound sensors are set in different sound zones;

[0028] S102, based on the pre-trained target end-to-end network model, performing sound source attribute detection on the sound signal of each sound source in the multi-channel audio signal to obtain the sound source attribute corresponding to each sound signal in the multi-channel audio signal; the sound source attribute includes the sound source attribute of a real person and the sound source attribute of a non-vehicle media;

[0029] S103, suppressing the sound signal whose sound source attribute is not the sound source attribute of the vehicle-side media in the multi-channel audio signal, and obtaining the sound signal corresponding to the sound source attribute of the real person;

[0030] S104, identifying the sound zone in the car for the sound signal corresponding to the real person's voice source attribute, and performing sound zone separation for the sound signal corresponding to the real person's voice source attribute according to the identification result.

[0031] The sound zone separation method for in-car audio of an embodiment of the present application can detect the sound source attributes of the sound signals of each sound source in the multi-channel audio signals based on the pre-trained target end-to-end network model according to the multi-channel audio signals collected by multiple sound sensors in the car, and obtain the sound source attributes corresponding to each sound signal in the multi-channel audio signals; the sound source attributes include real person sound source attributes and non-car-end media sound source attributes; suppress the sound signals with sound source attributes of non-car-end media sound source attributes in the multi-channel audio signals, and obtain the sound signals corresponding to the real person sound source attributes; identify the sound zone in the car for the sound signal corresponding to the real person sound source attributes, and perform sound zone separation on the sound signal corresponding to the real person sound source attributes according to the identification result. In this way, the embodiment of the present application can determine whether the sound source attribute corresponding to each sound signal in the multi-channel audio signal is a real person sound source attribute or a non-vehicle media sound source attribute based on the multi-channel audio signals collected by multiple sound sensors in the car and based on the pre-trained target end-to-end network model, and suppress the sound signal with the non-vehicle media sound source attribute to reduce the interference of the sound signal corresponding to the non-vehicle media sound source attribute in the sound zone recognition in the car, thereby improving the accuracy of the sound zone recognition in the car, and further enhancing the user's voice interaction experience.

[0032] The above-mentioned audio zone separation method for in-vehicle audio is mainly applied to vehicle terminals. The present application is not limited thereto, and can also be a terminal device with data processing capability such as a vehicle central control system.

[0033] The specific implementation methods of the above steps are introduced below.

[0034] In S101, the sound sensor may be an energy conversion device that converts a sound signal into an electrical signal, and may be, for example, a microphone.

[0035] The above-mentioned sound zones may be obtained by dividing the interior space of the vehicle into different position spaces according to the seats, and each position space corresponds to a sound zone.

[0036] Since the multiple sound sensors correspond to the sound zones one by one, the multiple sound sensors can be arranged at different seats in the vehicle.

[0037] The above multi-channel audio signals are obtained by collecting the sound signals of at least one sound source by multiple sound sensors. Each audio signal corresponds to a sound sensor, so the multi-channel audio signals include the sound signals of at least one sound source in the vehicle. The object that is making sound is called the sound source.

[0038] The at least one sound source mentioned above may be all sound sources in the vehicle, including real-person sound sources, non-vehicle media sound sources, noise, etc.

[0039] In S102, the above-mentioned real-person sound source indicates that the object making the sound is a real person; the above-mentioned non-vehicle-side media sound source indicates that the object making the sound is non-vehicle-side media sound. Non-vehicle-side devices are mobile terminals such as mobile phones, tablets and laptops. Non-vehicle-side media sound may include WeChat phone calls, mobile phone navigation, mobile phone Douyin, mobile phone music, etc.

[0040] The above-mentioned sound source attribute detection is performed on the sound signal of each sound source in the multi-channel audio signal to obtain the sound source attribute corresponding to each sound signal in the multi-channel audio signal. Exemplarily, it can be through a target end-to-end network model to detect the complete information of the sound signal in the multi-channel audio signal of each sound source in a preset frequency band, and the complete information is used to indicate whether the sound signal in the multi-channel audio signal of each sound source is complete in the preset frequency band; according to the complete information of each sound source, it is determined that each sound source is a real person sound source or a non-vehicle-end media sound source, wherein, if the sound signal in the multi-channel audio signal is complete in the preset frequency band, the sound source attribute corresponding to the sound signal includes the real person sound source attribute; if the sound signal in the multi-channel audio signal is incomplete in the preset frequency band, the sound source attribute corresponding to the sound signal includes the non-vehicle-end media sound source attribute.

[0041] The above end-to-end means that the model directly outputs the target waveform after inputting the source waveform, without the need for Fourier transform to convert the time domain signal into the frequency domain. The above audio separation means separating two or more sound sources in a mixed speech.

[0042] In S103, the sound signal whose sound source attribute is the non-vehicle-side media sound source attribute in the multi-channel audio signal is suppressed to obtain the sound signal corresponding to the real person sound source attribute. Exemplarily, the sound signal of each non-vehicle-side media sound source in the multi-channel audio signal may be muted to obtain the target multi-channel audio signal, so that the target multi-channel audio signal does not include the sound signal of the non-vehicle-side media sound source, so that the sound signal of the non-vehicle-side media sound source is suppressed. Alternatively, the attribute label processing may be performed on each sound source in the multi-channel audio signal, and the target multi-channel audio signal includes the sound signal of at least one sound source and the attribute label of each sound source, and the attribute label is used to indicate that the sound source attribute corresponding to the sound source is the sound source attribute corresponding to the real person sound source or the sound source attribute corresponding to the non-vehicle-side media sound source. In this way, the subsequent sound zone recognition processing only processes the sound signal whose sound source label indicates that it is a real person sound source in the target multi-channel audio signal, so that the sound signal of the non-vehicle-side media sound source is suppressed.

[0043] In S104, the sound zone recognition is performed on the multi-channel audio signals to obtain the sound zone of each real person sound source in at least one sound source. Exemplarily, the sound zone recognition of the real person sound source can be performed based on the difference between the sound signals of each sound source in the multi-channel audio signals to obtain the sound zone of each real person sound source in at least one sound source. The sound zone of each real person sound source can be the sound zone corresponding to the sound sensor of the optimal audio signal in the multi-channel audio signals of each real person sound source, wherein the optimal audio signal is the audio signal with the earliest sound signal reception time or the strongest sound signal reception energy in the multi-channel audio signals.

[0044] In some embodiments, the above S102 may specifically include:

[0045] Detecting complete information of a sound signal in a preset frequency band in the multi-channel audio signal of each sound source, wherein the complete information is used to indicate whether the sound signal in the multi-channel audio signal of each sound source is complete in the preset frequency band;

[0046] Based on the complete information of each sound source, it is determined that each sound source is a real person sound source or a non-vehicle-end media sound source, wherein, if the sound signal in the multi-channel audio signal is complete in the preset frequency band, the sound source attributes corresponding to the sound signal include real person sound source attributes; if the sound signal in the multi-channel audio signal is incomplete in the preset frequency band, the sound source attributes corresponding to the sound signal include non-vehicle-end media sound source attributes.

[0047] Due to the hardware characteristics of non-vehicle devices, there will be significant differences in statistical characteristics between the audio played and the sound produced by a real person. For example, the sound signal collected by the microphone in the low frequency band [0-1kHz] will be missing part of the media sound signal, while the real person's sound signal is complete.

[0048] The above complete information can be used to indicate whether the sound signals in the multi-channel audio signals of each sound source are complete in the preset frequency band. The preset frequency band can be, for example, a low frequency band [0-1kHz].

[0049] In this embodiment, by detecting the complete information of the sound signal in the preset frequency band of the multi-channel audio signal of each sound source of at least one sound source, it is possible to accurately determine whether each sound source is a real sound source or a non-vehicle media sound source, thereby improving the accuracy of in-vehicle sound zone recognition and further enhancing the user's voice interaction experience.

[0050] As an implementation of the present application, in order to obtain a target end-to-end network model, the training process of the target end-to-end network model may include:

[0051] Obtain a target training set, the target training set includes multiple training samples, the multiple training samples include a first training sample, a second training sample and a third training sample, the first training sample includes sound signals of multiple historical real sound sources collected by multiple sound sensors, the second training sample includes sound signals of multiple historical non-vehicle media sound sources collected by multiple sound sensors, and the third training sample includes sound signals of at least one historical real sound source and at least one historical non-vehicle media sound source collected by multiple sound sensors; the quantity ratio of the first training sample, the second training sample and the third training sample in the target training set is a preset quantity ratio, and the quantity ratio of the third training sample is higher than the quantity ratio of the first training sample and the second training sample;

[0052] The initial end-to-end network model is trained using training samples in the target training set to update the network parameters of the initial end-to-end network model and obtain the target end-to-end network model.

[0053] The above-mentioned first training sample may include sound signals of multiple historical real sound sources collected by multiple sound sensors, wherein the sound signals of the historical real sound sources, illustratively, may be sound signals of 1 to M separate real sound sources distributed in N different sound zones in the car, M<=N.

[0054] The above-mentioned second training sample may include sound signals of multiple historical non-vehicle-end media sound sources collected by multiple sound sensors, wherein the sound signals of the historical non-vehicle-end media sound sources, for example, simulate non-vehicle-end media sounds by simulation means, thereby obtaining 0 to V sound signals of non-vehicle-end media sounds in the vehicle.

[0055] The third training sample may include at least one historical sound signal of a real person sound source and at least one historical sound signal of a non-vehicle media sound source collected by a plurality of sound sensors.

[0056] In this embodiment, the target training set includes a first training sample, a second training sample, and a third training sample, so that when the initial end-to-end network model is trained, the network parameters can be updated based on the first training sample, the second training sample, and the third training sample to obtain a target end-to-end network model with higher accuracy and robustness.

[0057] In the above-mentioned preset quantity ratio, the quantity ratio of the third training samples is higher than the quantity ratio of the first training samples and the second training samples. Exemplarily, it can be a 1:1:3 ratio. In this embodiment, the preset quantity ratio is not limited to this, and can also be other ratios, which are not specifically limited here.

[0058] In this embodiment, the ratio of the number of third training samples in the target training set is higher than the ratio of the number of first training samples and second training samples. The initial end-to-end network model can be trained using the training samples in the target training set to ensure the recognition performance of the target end-to-end network model for real human sound sources and non-vehicle media sound sources.

[0059] In some embodiments, the above-mentioned obtaining of the target training set may specifically include:

[0060] Acquire an initial sample, where the initial sample includes sound signals of a target historical sound source corresponding to a plurality of different sound source attributes collected by a plurality of sound sensors;

[0061] The sound signals generated by the target historical sound sources corresponding to the different sound source attributes are mixed and processed according to a preset ratio to obtain mixed speech in different scenarios; wherein the different scenarios include: real person mixed speech, non-vehicle side media mixed speech, and real person non-vehicle side media mixed speech;

[0062] Silence signal segments are configured for mixed speech in different scenarios to obtain training samples.

[0063] The above-mentioned multiple target historical sound sources may be multiple historical real-person sound sources, multiple historical non-vehicle media sound sources, or at least one historical real-person sound source and at least one historical non-vehicle media sound source.

[0064] The preset ratio of the above-mentioned real-person non-vehicle media mixed voice can be, for example, a ratio of 1:1:3. In this embodiment, the preset ratio is not limited to this, and can also be other ratios, which are not specifically limited here.

[0065] The above different scenarios may include: real person mixed voice, non-vehicle end media mixed voice, real person non-vehicle end media mixed voice. Exemplarily, real person mixed voice can be based on 1 to M single sound sources simulated in N (M <= N) different sound zones by simulation means, to construct the mixed voice collected by N microphones in the car; non-vehicle end media mixed voice can be based on 0 to V in-vehicle non-vehicle end media sounds simulated in-vehicle interference by simulation means, to construct the mixed voice collected by N microphones in the car; real person non-vehicle end media mixed voice can be based on 1 to M single sound sources simulated in N (M <= N) different sound zones by simulation means, and 0 to V in-vehicle non-vehicle end media sounds are added to simulate in-vehicle interference, to construct the mixed voice collected by N microphones in the car.

[0066] In this embodiment, the sound signals generated by the target historical sound sources corresponding to different sound source attributes are mixed and processed according to a preset ratio to obtain mixed speech in different scenarios, and silent signal segments are configured for the mixed speech in different scenarios to obtain training samples, thereby improving the accuracy of the target end-to-end network model.

[0067] In some embodiments, configuring silence signal segments for mixed speech in different scenarios to obtain training samples may specifically include:

[0068] A silent signal is configured between the sound signals of multiple target historical sound sources in the target mixed speech to obtain a target training sample corresponding to the initial sample. There is a silent signal between every two sound signals of the target historical sound sources in the target training sample. The duration ratio of the silent signal in the target training sample is within a preset ratio range. The target mixed speech is any one of a real person mixed speech, a non-vehicle media mixed speech and a real person non-vehicle media mixed speech.

[0069] The target training sample may be a first training sample, a second training sample or a third training sample.

[0070] The above-mentioned preset ratio range may be, for example, a ratio range of 40%-70%.

[0071] In this embodiment, by configuring a silent signal between the sound signals of multiple target historical sound sources, the proportion of silent segments can be increased while simulating a real vehicle scene, thereby reducing the impact of the sound source labeling deviation at the frame level.

[0072] In some embodiments, the above S103 may specifically include:

[0073] Performing attribute label processing on each sound source in the multi-channel audio signal to obtain a target multi-channel audio signal, where the target multi-channel audio signal includes a sound signal of at least one sound source and an attribute label of each sound source, where the attribute label is used to indicate that a sound source attribute corresponding to the sound source is a sound source attribute corresponding to a real sound source or a sound source attribute corresponding to a non-vehicle-side media sound source;

[0074] The above S104 may specifically include:

[0075] Based on the attribute labels of each sound source in the target multi-channel audio signal, the sound signal corresponding to the sound source whose attribute label is the sound source attribute corresponding to the real sound source is identified, and the target sound area corresponding to the sound source with the attribute corresponding to the real sound source is determined.

[0076] In this embodiment, a target multi-channel audio signal is obtained by performing attribute label processing on each sound source in a multi-channel audio signal. The target multi-channel audio signal includes a sound signal of at least one sound source and an attribute label of each sound source. The attribute label is used to indicate that the sound source attribute corresponding to the sound source is a real sound source or a non-vehicle-end media sound source. In this way, subsequent sound zone recognition processing only processes the sound signal whose attribute label indicates that the sound source attribute corresponds to a real sound source in the target multi-channel audio signal, thereby suppressing the sound signal of the non-vehicle-end media sound source.

[0077] In order to facilitate the understanding of the sound zone separation method of the in-car audio in the embodiment of the present application, the application process of the sound zone separation method of the in-car audio is described here, as follows:

[0078] Assume that there are N sound zones and microphones in the car, and the sound zones correspond to the microphones one by one. The speech signal collected by the microphone is X = [X1…X N ] T , including all the sound sources and noises in the car. Assume that there are M real sound sources in the car (M<=N), which are S=[S1…S M ] T , there are V non-car-end media audio in the car at the same time, which are I=[I1…I V ] T , then through online learning we can get Contains the output of N sound zones, and the estimated M real-person sound sources in the car are distributed in M ​​corresponding sound zones All media sounds played in the car that are not played by the car machine are suppressed.

[0079] Training phase:

[0080] Step 1: Use simulation methods to simulate 1 to M single sound sources distributed in N (M<=N) different sound zones, and add 0 to V non-car-end media sounds in the car to simulate in-car interference, construct mixed speech collected by N microphones in the car, and retain 1 to M single sound sources collected by the microphones as learning targets to obtain training set data;

[0081] Step 2: Input the mixed speech of pure human voice with noise in the car, the media sound with noise outside the car machine, and the mixed speech of noisy human voice + media sound outside the car machine in a ratio of 1:1:3. The mixed speech of different scenes contains 40%-70% of silent segments. While simulating the real car scene, increase the proportion of silent segments to reduce the influence of the deviation of the sound source labels at the frame level. Simulate the situation where the number of sound sources in the car is 1 to M and whether there is interference from the media sound outside the car machine. Train the end-to-end multi-sound zone separation network (that is, the initial end-to-end network model mentioned above), input the mixed speech, output the multi-sound zone separation speech in the car, and calculate the loss function of 1 to M single sound sources collected by N microphones and the output The loss function can be the commonly used L1 (mean absolute error, MAE) or L2 (mean square error, MSE). When the loss function loss meets the preset conditions, a trained end-to-end multi-zone separation network (i.e., the above-mentioned target end-to-end network model) is obtained. If there is no sound source in the current sound zone, then silence or pure noise is the learning target, which can ensure that each sound zone can achieve the purpose of separating and suppressing non-car-end media sound.

[0082] Testing phase:

[0083] Step 1: In an actual vehicle environment, N microphones collect N audio signals;

[0084] Step 2: Directly input the mixed audio into the end-to-end multi-zone separation network (i.e., the target end-to-end network model mentioned above) to obtain a separated voice signal that suppresses the interference of non-vehicle media sound in the vehicle.

[0085] In this embodiment, due to the hardware characteristics of the non-vehicle device itself, the played audio will have large differences in statistical characteristics such as the low frequency band [0-1kHz] compared to the sound emitted by a real person. This difference can be learned through an end-to-end multi-zone separation network, thereby achieving the purpose of end-to-end separation suppression.

[0086] Based on the in-vehicle audio sound zone separation method provided in the above embodiment, the present application also provides a specific implementation of the in-vehicle audio sound zone separation device. Please refer to the following embodiment.

[0087] like Figure 2 As shown, an in-vehicle audio sound zone separation device 200 provided in an embodiment of the present application may include the following modules: a first acquisition module 201 , a detection module 202 , a processing module 203 and an identification module 204 .

[0088] A first acquisition module 201 is used to acquire multi-channel audio signals collected by multiple sound sensors in the vehicle, where the multi-channel audio signals include sound signals generated by multiple sound sources with different sound source properties, and different sound sensors are set in different sound zones;

[0089] The detection module 202 is used to detect the sound source attributes of the sound signals of each sound source in the multi-channel audio signal based on the pre-trained target end-to-end network model, and obtain the sound source attributes corresponding to each sound signal in the multi-channel audio signal; the sound source attributes include the attributes of the real sound source and the attributes of the non-vehicle media sound source;

[0090] The processing module 203 is used to suppress the sound signal whose sound source attribute is not the sound source attribute of the vehicle-side media in the multi-channel audio signal, and obtain the sound signal corresponding to the sound source attribute of the real person;

[0091] The recognition module 204 is used to recognize the sound zone in the car of the sound signal corresponding to the real person's voice source attribute, and perform sound zone separation on the sound signal corresponding to the real person's voice source attribute according to the recognition result.

[0092] The sound zone separation device for in-car audio of the embodiment of the present application can detect the sound source attributes of the sound signals of each sound source in the multi-channel audio signals based on the pre-trained target end-to-end network model according to the multi-channel audio signals collected by multiple sound sensors in the car, and obtain the sound source attributes corresponding to each sound signal in the multi-channel audio signals; the sound source attributes include real person sound source attributes and non-car-end media sound source attributes; suppress the sound signals with sound source attributes of non-car-end media sound source attributes in the multi-channel audio signals, and obtain the sound signals corresponding to the real person sound source attributes; identify the sound zone in the car for the sound signal corresponding to the real person sound source attributes, and perform sound zone separation on the sound signal corresponding to the real person sound source attributes according to the identification result. In this way, the embodiment of the present application can determine whether the sound source attribute corresponding to each sound signal in the multi-channel audio signal is a real person sound source attribute or a non-vehicle media sound source attribute based on the multi-channel audio signals collected by multiple sound sensors in the car and based on the pre-trained target end-to-end network model, and suppress the sound signal with the non-vehicle media sound source attribute to reduce the interference of the sound signal corresponding to the non-vehicle media sound source attribute in the sound zone recognition in the car, thereby improving the accuracy of the sound zone recognition in the car, and further enhancing the user's voice interaction experience.

[0093] In some embodiments, the detection module 202 may include:

[0094] A detection unit, used to detect complete information of a sound signal in a multi-channel audio signal of each sound source in a preset frequency band, wherein the complete information is used to indicate whether the sound signal in the multi-channel audio signal of each sound source is complete in the preset frequency band;

[0095] A determination unit is used to determine whether each sound source is a real sound source or a non-vehicle-end media sound source based on the complete information of each sound source, wherein if the sound signal in the multi-channel audio signal is complete in a preset frequency band, the sound source attribute corresponding to the sound signal includes real sound source attributes; if the sound signal in the multi-channel audio signal is incomplete in the preset frequency band, the sound source attribute corresponding to the sound signal includes non-vehicle-end media sound source attributes.

[0096] As an implementation of the present application, in order to obtain a target end-to-end network model, the above-mentioned device 200 may further include:

[0097] A second acquisition module is used to acquire a target training set, the target training set includes multiple training samples, the multiple training samples include a first training sample, a second training sample and a third training sample, the first training sample includes sound signals of multiple historical real sound sources collected by multiple sound sensors, the second training sample includes sound signals of multiple historical non-vehicle media sound sources collected by multiple sound sensors, and the third training sample includes sound signals of at least one historical real sound source and at least one historical non-vehicle media sound source collected by multiple sound sensors; the quantity ratio of the first training sample, the second training sample and the third training sample in the target training set is a preset quantity ratio, and the quantity ratio of the third training sample is higher than the quantity ratio of the first training sample and the second training sample;

[0098] The training module is used to train the initial end-to-end network model through training samples in the target training set to update the network parameters of the initial end-to-end network model to obtain the target end-to-end network model.

[0099] In some embodiments, the second acquisition module may specifically include:

[0100] An acquisition unit, configured to acquire an initial sample, wherein the initial sample includes sound signals of a target historical sound source corresponding to a plurality of different sound source attributes collected by a plurality of sound sensors;

[0101] The first processing unit is used to mix and process the sound signals generated by the target historical sound sources corresponding to the different sound source attributes according to a preset ratio to obtain mixed speech in different scenarios; wherein the different scenarios include: real person mixed speech, non-vehicle side media mixed speech, and real person non-vehicle side media mixed speech;

[0102] The configuration unit is used to configure silence signal segments for mixed speech in different scenarios to obtain training samples.

[0103] In some embodiments, the above-mentioned configuration unit can be specifically used to configure silent signals between the sound signals of multiple target historical sound sources in the target mixed speech, so as to obtain a target training sample corresponding to the initial sample, wherein there is a silent signal between every two sound signals of the target historical sound sources in the target training sample, and the duration ratio of the silent signal in the target training sample is within a preset ratio range, and the target mixed speech is any one of a real-person mixed speech, a non-vehicle-side media mixed speech, and a real-person non-vehicle-side media mixed speech.

[0104] In some embodiments, the processing module 203 may specifically include:

[0105] A second processing unit is used to perform attribute label processing on each sound source in the multi-channel audio signal to obtain a target multi-channel audio signal, where the target multi-channel audio signal includes a sound signal of at least one sound source and an attribute label of each sound source, where the attribute label is used to indicate that a sound source attribute corresponding to the sound source is a sound source attribute corresponding to a real sound source or a sound source attribute corresponding to a non-vehicle-side media sound source;

[0106] The identification module 204 may specifically include:

[0107] The recognition unit is used to recognize the sound signal corresponding to the sound source whose attribute label is the sound source attribute corresponding to the real sound source based on the attribute label of each sound source in the target multi-channel audio signal, and determine the target sound area corresponding to the sound source with the attribute corresponding to the real sound source.

[0108] Based on the in-car audio sound zone separation method provided in the above embodiment, the present application also provides a specific implementation of the electronic device. Please refer to the following embodiment.

[0109] Figure 3 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.

[0110] The electronic device may include a processor 301 and a memory 302 storing computer program instructions.

[0111] Specifically, the processor 301 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0112] The memory 302 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 302 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 302 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 302 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 302 is a non-volatile solid-state memory.

[0113] In certain embodiments, the memory 302 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Thus, in general, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.

[0114] The processor 301 reads and executes the computer program instructions stored in the memory 302 to implement any one of the in-vehicle audio sound zone separation methods in the above embodiments.

[0115] In one example, the electronic device may further include a communication interface 303 and a bus 310. Figure 3 As shown, the processor 301, the memory 302, and the communication interface 303 are connected via a bus 310 and communicate with each other.

[0116] The communication interface 303 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.

[0117] Bus 310 includes hardware, software or both, and the parts of electronic equipment are coupled to each other. For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 310 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the application considers any suitable bus or interconnection.

[0118] The electronic device can execute the audio zone separation method of the in-car audio in the embodiment of the present application, thereby realizing the combination Figure 1 and Figure 2 A method and device for separating audio zones of in-car audio are described.

[0119] In addition, in combination with the in-vehicle audio sound zone separation method in the above embodiments, the present application embodiment can provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any of the in-vehicle audio sound zone separation methods in the above embodiments is implemented.

[0120] In combination with the in-vehicle audio sound zone separation method in the above embodiment, the present application embodiment can provide a vehicle for implementation. The vehicle includes: an electronic device for implementing any of the in-vehicle audio sound zone separation methods as described above.

[0121] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.

[0122] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0123] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.

[0124] Aspects of the present disclosure are described above with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs a specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0125] The above is only a specific implementation of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.

[0126] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprises" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

Claims

1. A method for separating audio zones in a car, characterized in that: The method comprises: Acquire multi-channel audio signals collected by multiple sound sensors in the vehicle, wherein the multi-channel audio signals include sound signals generated by multiple sound sources with different sound source properties, and different sound sensors are arranged in different sound zones; Based on the pre-trained target end-to-end network model, sound source attribute detection is performed on the sound signal of each sound source in the multi-channel audio signal to obtain the sound source attribute corresponding to each sound signal in the multi-channel audio signal; the sound source attribute includes the attribute of a real person sound source and the attribute of a non-vehicle-side media sound source; Suppressing the sound signal whose sound source attribute is not the sound source attribute of the vehicle-side media in the multi-channel audio signal, and obtaining the sound signal corresponding to the sound source attribute of the real person; The sound zone in the car of the sound signal corresponding to the real person's sound source attribute is identified, and the sound zone of the sound signal corresponding to the real person's sound source attribute is separated according to the identification result.

2. The method according to claim 1, characterized in that The performing sound source attribute detection on the sound signal of each sound source in the multi-channel audio signal to obtain the sound source attribute corresponding to each sound signal in the multi-channel audio signal includes: Detecting complete information of the sound signal in the multi-channel audio signals of each sound source in a preset frequency band, wherein the complete information is used to indicate whether the sound signal in the multi-channel audio signals of each sound source is complete in the preset frequency band; Based on the complete information of each of the sound sources, it is determined whether each of the sound sources is a real person sound source or a non-vehicle-end media sound source, wherein if the sound signal in the multi-channel audio signal is complete in the preset frequency band, the sound source attribute corresponding to the sound signal includes real person sound source attribute; if the sound signal in the multi-channel audio signal is incomplete in the preset frequency band, the sound source attribute corresponding to the sound signal includes non-vehicle-end media sound source attribute.

3. The method according to claim 2, characterized in that The training process of the target end-to-end network model includes: Acquire a target training set, the target training set comprising a plurality of training samples, the plurality of training samples comprising a first training sample, a second training sample and a third training sample, the first training sample comprising sound signals of a plurality of historical real sound sources collected by a plurality of sound sensors, the second training sample comprising sound signals of a plurality of historical non-vehicle media sound sources collected by the plurality of sound sensors, and the third training sample comprising sound signals of at least one historical real sound source and at least one historical non-vehicle media sound source collected by the plurality of sound sensors; the quantity ratio of the first training sample, the second training sample and the third training sample in the target training set is a preset quantity ratio, and the quantity ratio of the third training sample is higher than the quantity ratio of the first training sample and the second training sample; The initial end-to-end network model is trained using the training samples in the target training set to update the network parameters of the initial end-to-end network model and obtain the target end-to-end network model.

4. The method according to claim 3, characterized in that The step of obtaining a target training set includes: Acquire an initial sample, wherein the initial sample includes sound signals of a target historical sound source corresponding to a plurality of different sound source attributes collected by a plurality of sound sensors; The sound signals generated by the target historical sound sources corresponding to the different sound source attributes are mixed and processed according to a preset ratio to obtain mixed speech in different scenarios; wherein the different scenarios include: real person mixed speech, non-vehicle side media mixed speech, and real person non-vehicle side media mixed speech; Silence signal segments are configured for mixed speech in different scenarios to obtain training samples.

5. The method according to claim 4, characterized in that The step of respectively configuring silent signal segments for mixed speech in different scenarios to obtain training samples includes: A silent signal is respectively configured between the sound signals of multiple target historical sound sources in the target mixed speech to obtain a target training sample corresponding to the initial sample, wherein the silent signal is between every two sound signals of the target historical sound sources in the target training sample, and the duration ratio of the silent signal in the target training sample is within a preset ratio range, and the target mixed speech is any one of a real person mixed speech, a non-vehicle-side media mixed speech, and a real person non-vehicle-side media mixed speech.

6. The method according to claim 1, characterized in that The step of suppressing the sound signal whose sound source attribute is not the vehicle-side media sound source attribute in the multi-channel audio signal to obtain the sound signal corresponding to the real person sound source attribute includes: Performing attribute label processing on each of the sound sources in the multi-channel audio signal to obtain a target multi-channel audio signal, wherein the target multi-channel audio signal includes a sound signal of the at least one sound source and an attribute label of each of the sound sources, wherein the attribute label is used to indicate that the sound source attribute corresponding to the sound source is the sound source attribute corresponding to the real person sound source or the sound source attribute corresponding to the non-vehicle media sound source; The step of identifying the sound zone in the car for the sound signal corresponding to the real person sound source attribute, and performing sound zone separation on the sound signal corresponding to the real person sound source attribute according to the identification result, comprises: Based on the attribute labels of the sound sources in the target multi-channel audio signals, the sound signals corresponding to the sound sources whose attribute labels are the sound source attributes corresponding to the real voice sources are identified, and the target sound areas corresponding to the sound sources whose attributes are corresponding to the real voice sources are determined.

7. A sound zone separation device for in-car audio, characterized in that: include: A first acquisition module is used to acquire multi-channel audio signals collected by multiple sound sensors in the vehicle, wherein the multi-channel audio signals include sound signals generated by multiple sound sources with different sound source properties, and different sound sensors are arranged in different sound zones; A detection module, configured to perform sound source attribute detection on the sound signal of each sound source in the multi-channel audio signal based on a pre-trained target end-to-end network model, and obtain the sound source attribute corresponding to each sound signal in the multi-channel audio signal; the sound source attribute includes a real person sound source attribute and a non-vehicle media sound source attribute; A processing module, used for suppressing the sound signal whose sound source attribute is a non-vehicle-side media sound source attribute in the multi-channel audio signal, and obtaining the sound signal corresponding to the real person sound source attribute; The recognition module is used to recognize the sound zone in the car of the sound signal corresponding to the real person sound source attribute, and perform sound zone separation on the sound signal corresponding to the real person sound source attribute according to the recognition result.

8. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the method for separating sound zones of in-vehicle audio according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the method for separating sound zones of in-vehicle audio according to any one of claims 1 to 6 is implemented.

10. A vehicle, characterized in that: include: An electronic device, wherein the electronic device is used to implement the method for separating sound zones of in-vehicle audio according to any one of claims 1 to 6.

Citation Information

Cited By

  • Audio mark generation method and device

    CN120472868A

  • Mobile equipment voice interception method and system based on vehicle-mounted multi-voice-area scene

    CN121600922A