Sound area separation method, device and equipment for audio in vehicle, storage medium and vehicle
By using the target end-to-end network model to detect the sound source spatial position of the audio signal in the car and suppress the outside of the car interference, the problem of reducing the accuracy of sound separation in the car is solved and the voice interaction experience is improved.
Patent Information
- Application Number
- CN202311466847.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-06
AI Technical Summary
The separation of sound regions of the audio in the car has a problem of reduced accuracy in the voice interaction in the car. It is mainly because the microphone in the car is highly sensitive and easy to collect outside the car interference sound, which affects the accuracy of sound regions recognition.
By obtaining the multi-channel audio signals collected by multiple sound sensors in the car, using the pre-trained target end-to-end network model, the sound source spatial position detection is performed on each sound source, distinguishing the sound source inside and outside the car, and suppressing the sound source signal outside the car, thereby separating the sound region.
It improves the accuracy of sound area recognition in the car, reduces the impact of outside-car interference on sound area recognition, and improves the user's voice interaction experience.
Smart Images

Figure CN119943084A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of vehicle technology, and in particular, relates to a method, device, equipment and storage medium for separating sound zones of in-vehicle audio. Background Art
[0002] With the continuous development of vehicle technology, vehicles have more and more functions. In particular, passengers can interact with the vehicle terminal by inputting voice information. When passengers interact with the vehicle end by voice, in order to ensure that the vehicle end recognizes the voice of each passenger without being interfered by the voice of other people in the car, it is necessary to separate the received voice information according to the position in the car to realize the voice interaction between each passenger and the vehicle terminal. However, due to the small space inside the car and the high sensitivity of the microphone inside the car, the vehicle end can easily collect interference sounds outside the car (including human voices and noise outside the car), which interferes with the vehicle end's separation of the sound zone of voice information, resulting in reduced accuracy of sound zone recognition, which in turn affects the voice interaction experience of passengers in the car. Summary of the invention
[0003] The embodiments of the present application provide a method, device, equipment and computer storage medium for separating sound zones of in-vehicle audio, which can improve the accuracy of in-vehicle sound zone recognition and thus enhance the user's voice interaction experience.
[0004] In a first aspect, an embodiment of the present application provides a method for separating sound zones of in-vehicle audio, the method comprising:
[0005] Acquire multi-channel audio signals collected by multiple sound sensors in the vehicle, wherein the multi-channel audio signals include sound signals generated by sound sources corresponding to different spatial positions, and different sound sensors are arranged in different sound zones in the vehicle;
[0006] Based on the pre-trained target end-to-end network model, the sound source spatial position detection is performed on the sound signal of each sound source in the multi-channel audio signal to obtain the sound source spatial position corresponding to each sound signal in the multi-channel audio signal; the sound source spatial position includes the sound source position inside the vehicle and the sound source position outside the vehicle;
[0007] Suppressing the sound signal whose sound source spatial position is the sound source position outside the vehicle in the multi-channel audio signal, and obtaining the sound signal corresponding to the sound source position inside the vehicle;
[0008] The sound zone in the vehicle of the sound signal corresponding to the position of the sound source in the vehicle is identified, and the sound zone of the sound signal corresponding to the position of the sound source in the vehicle is separated according to the identification result.
[0009] In a second aspect, an embodiment of the present application provides a sound zone separation device for in-car audio, the device comprising:
[0010] A first acquisition module is used to acquire multi-channel audio signals collected by multiple sound sensors in the vehicle, wherein the multi-channel audio signals include sound signals generated by sound sources corresponding to different spatial positions, and different sound sensors are arranged in different sound zones in the vehicle;
[0011] A detection module, configured to detect the spatial position of each sound source in the multi-channel audio signal based on a pre-trained target end-to-end network model, and obtain the spatial position of the sound source corresponding to each sound signal in the multi-channel audio signal; the spatial position of the sound source includes the position of the sound source inside the vehicle and the position of the sound source outside the vehicle;
[0012] A suppression module, used to suppress the sound signal whose sound source spatial position is the sound source position outside the vehicle in the multi-channel audio signal, and obtain the sound signal corresponding to the sound source position inside the vehicle;
[0013] The recognition module is used to recognize the sound zone in the vehicle of the sound signal corresponding to the sound source position in the vehicle, and perform sound zone separation on the sound signal corresponding to the sound source position in the vehicle according to the recognition result.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the method for separating sound zones of in-vehicle speech as described in any one of the above items is implemented.
[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method for separating sound zones of in-vehicle speech as described in any one of the above is implemented.
[0016] In a fifth aspect, an embodiment of the present application provides a vehicle, comprising: an electronic device, wherein the electronic device is used to implement the sound zone separation method for in-vehicle audio as described in any one of the above.
[0017] The sound zone separation method, device, equipment, storage medium and vehicle for in-car audio of the embodiments of the present application can detect the sound source spatial position of the sound signal of each sound source in the multi-channel audio signal based on the pre-trained target end-to-end network model according to the multi-channel audio signals collected by multiple sound sensors in the car, and obtain the sound source spatial position corresponding to each sound signal in the multi-channel audio signal; the sound source spatial position includes the sound source position in the car and the sound source position outside the car; suppress the sound signal whose sound source spatial position in the multi-channel audio signal is the sound source position outside the car, and obtain the sound signal corresponding to the sound source position in the car; identify the sound zone in the car for the sound signal corresponding to the sound source position in the car, and perform sound zone separation on the sound signal corresponding to the sound source position in the car according to the identification result. In this way, the embodiment of the present application can determine, based on the multi-channel audio signals collected by multiple sound sensors in the car and the pre-trained target end-to-end network model, whether the sound source spatial position corresponding to each sound signal in the multi-channel audio signals is the sound source position inside the car or the sound source position outside the car, and suppress the sound signal at the sound source position outside the car to reduce the interference with the sound zone recognition in the car with the sound signal corresponding to the sound source position inside the car, thereby improving the accuracy of the sound zone recognition in the car and further enhancing the user's voice interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] Figure 1 It is a flowchart of a method for separating audio zones of in-vehicle audio provided by an embodiment of the present application;
[0020] Figure 2 is a structural schematic diagram of a sound zone separation device for in-vehicle audio provided by another embodiment of the present application;
[0021] Figure 3 It is a structural diagram of an electronic device provided by yet another embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0023] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0024] With the continuous development of vehicle technology, vehicles have more and more functions. In particular, passengers can interact with the vehicle terminal by inputting voice information. When passengers interact with the vehicle end by voice, in order to ensure that the vehicle end recognizes the voice of each passenger without being interfered by the voice of other people in the car, it is necessary to separate the received voice information according to the position in the car to realize the voice interaction between each passenger and the vehicle terminal. However, due to the small space inside the car and the high sensitivity of the microphone inside the car, the vehicle end can easily collect interference sounds outside the car (including human voices and noise outside the car), which interferes with the vehicle end's separation of the sound zone of voice information, resulting in reduced accuracy of sound zone recognition, which in turn affects the voice interaction experience of passengers in the car.
[0025] In order to solve the problems of the prior art, the embodiments of the present application provide a method, device, equipment, storage medium and vehicle for separating the sound zones of in-vehicle audio. The following first introduces the method for separating the sound zones of in-vehicle audio provided by the embodiments of the present application.
[0026] Figure 1 FIG. 1 is a flow chart of a method for separating audio zones of in-car audio provided by an embodiment of the present application. Figure 1 As shown, a method for separating audio zones of in-vehicle audio may include the following steps S101 to S104:
[0027] S101, acquiring multi-channel audio signals collected by multiple sound sensors in the vehicle, where the multi-channel audio signals include sound signals generated by sound sources corresponding to different spatial positions, and different sound sensors are arranged in different sound zones in the vehicle;
[0028] S102, based on the pre-trained target end-to-end network model, performing sound source spatial position detection on the sound signal of each sound source in the multi-channel audio signal to obtain the sound source spatial position corresponding to each sound signal in the multi-channel audio signal; the sound source spatial position includes the sound source position inside the vehicle and the sound source position outside the vehicle;
[0029] S103, suppressing the sound signal whose sound source spatial position is the position of the sound source outside the vehicle in the multi-channel audio signal, and obtaining the sound signal corresponding to the position of the sound source inside the vehicle;
[0030] S104: Identify the sound zone in the vehicle for the sound signal corresponding to the sound source position in the vehicle, and perform sound zone separation on the sound signal corresponding to the sound source position in the vehicle according to the identification result.
[0031] The sound zone separation method for in-car audio of an embodiment of the present application can detect the sound source spatial position of the sound signal of each sound source in the multi-channel audio signal based on the pre-trained target neural network model according to the multi-channel audio signals collected by multiple sound sensors in the car, and obtain the sound source spatial position corresponding to each sound signal in the multi-channel audio signal; the sound source spatial position includes the sound source position in the car and the sound source position outside the car; suppress the sound signal whose sound source spatial position in the multi-channel audio signal is the sound source position outside the car, and obtain the sound signal corresponding to the sound source position in the car; identify the sound zone in the car for the sound signal corresponding to the sound source position in the car, and perform sound zone separation on the sound signal corresponding to the sound source position in the car according to the identification result. In this way, the embodiment of the present application can determine, based on the multi-channel audio signals collected by multiple sound sensors in the car and the pre-trained target end-to-end network model, whether the sound source spatial position corresponding to each sound signal in the multi-channel audio signals is the sound source position inside the car or the sound source position outside the car, and suppress the sound signal at the sound source position outside the car to reduce the interference with the sound zone recognition in the car with the sound signal corresponding to the sound source position inside the car, thereby improving the accuracy of the sound zone recognition in the car and further enhancing the user's voice interaction experience.
[0032] The above-mentioned audio zone separation method for in-vehicle audio is mainly applied to vehicle terminals. The present application is not limited thereto, and can also be a terminal device with data processing capability such as a vehicle central control system.
[0033] The specific implementation methods of the above steps are introduced below.
[0034] In S101, the sound sensor may be an energy conversion device that converts a sound signal into an electrical signal, and may be, for example, a microphone.
[0035] The above-mentioned sound zones can be obtained by dividing the interior space of the vehicle into different spatial positions according to the seats, and each spatial position corresponds to a sound zone.
[0036] Since the multiple sound sensors correspond to the sound zones one by one, the multiple sound sensors can be arranged at different seats in the vehicle.
[0037] The above multi-channel audio signals are obtained by collecting the sound signals of at least one sound source by multiple sound sensors. Each audio signal corresponds to a sound sensor, so the multi-channel audio signals include the sound signals of at least one sound source in the vehicle. The object that is making sound is called the sound source.
[0038] The at least one sound source mentioned above may be all sound sources received by the microphone inside the vehicle, including sound sources inside the vehicle, sound sources outside the vehicle, noise, etc.
[0039] In S102, the in-vehicle sound source may represent a real person making a sound in the vehicle; and the out-vehicle sound source may represent a real person making a sound outside the vehicle.
[0040] The above target end-to-end network model is pre-trained and is an end-to-end neural network model. End-to-end means that the model directly outputs the target waveform after inputting the source waveform, without the need for Fourier transform to convert the time domain signal to the frequency domain. The above audio separation means separating two or more sound sources from a mixed speech.
[0041] The above-mentioned sound source spatial position detection is performed on the sound signal of each sound source in the multi-channel audio signal to obtain the sound source spatial position corresponding to each sound signal in the multi-channel audio signal. Exemplarily, it can be done by detecting the difference information of the sound signal in the multi-channel audio signal of each sound source, and the difference information is used to indicate whether the time difference or energy difference between the sound signals in the multi-channel audio signal of each sound source meets the preset condition; according to the difference information of each sound source, it is determined that each sound source is an in-vehicle sound source or an outside-vehicle sound source, wherein if the time difference or energy difference between the sound signals in the multi-channel audio signal meets the preset condition, the sound source spatial position corresponding to the sound signal includes the sound source position inside the vehicle; if the time difference or energy difference between the sound signals in the multi-channel audio signal does not meet the preset condition, the sound source spatial position corresponding to the sound signal includes the sound source position outside the vehicle.
[0042] In S103, the above-mentioned suppression of the sound signal whose sound source spatial position in the multi-channel audio signal is the position of the sound source outside the vehicle, and obtaining the sound signal corresponding to the position of the sound source inside the vehicle, exemplarily, can be to mute the sound signal of each sound source outside the vehicle in the multi-channel audio signal to obtain the target multi-channel audio signal, so that the target multi-channel audio signal does not include the sound signal of the sound source outside the vehicle, so that the sound signal of the sound source outside the vehicle is suppressed. Alternatively, it is also possible to perform spatial position label processing on each sound source in the multi-channel audio signal, the target multi-channel audio signal includes the sound signal of at least one sound source and the spatial position label of each sound source, the spatial position label is used to indicate that the spatial position corresponding to the sound source is the spatial position corresponding to the sound source inside the vehicle or the spatial position corresponding to the sound source outside the vehicle, so that the subsequent sound zone recognition processing only processes the sound signal indicated by the spatial position label as the sound source inside the vehicle in the target multi-channel audio signal, so that the sound signal of the sound source outside the vehicle is suppressed.
[0043] In S104, the sound zone in the vehicle corresponding to the sound source position in the vehicle is identified, and the sound zone of the sound signal corresponding to the sound source position in the vehicle is separated according to the identification result. Exemplarily, the sound zone of the sound source in the vehicle can be identified based on the difference between the sound signals of each sound source in the multi-channel audio signal, and the sound zone in which each sound source in the vehicle is located in at least one sound source is obtained. The sound zone in which each sound source in the vehicle is located can be the sound zone corresponding to the sound sensor of the optimal audio signal in the multi-channel audio signals of each sound source in the vehicle, wherein the optimal audio signal is the audio signal with the earliest sound signal reception time or the strongest sound signal reception energy in the multi-channel audio signals.
[0044] In some embodiments, the above S102 may specifically include:
[0045] Detecting difference information of sound signals in the multi-channel audio signals of each sound source, where the difference information is used to indicate whether a time difference or an energy difference between the sound signals in the multi-channel audio signals of each sound source meets a preset condition;
[0046] Based on the difference information of each sound source, it is determined that each sound source is an inside-vehicle sound source or an outside-vehicle sound source, wherein if the time difference or energy difference between the sound signals in the multi-channel audio signals meets the preset conditions, the sound source spatial position corresponding to the sound signal includes the inside-vehicle sound source position; if the time difference or energy difference between the sound signals in the multi-channel audio signals does not meet the preset conditions, the sound source spatial position corresponding to the sound signal includes the outside-vehicle sound source position.
[0047] Since there are multiple microphones in the car, the sound sources inside and outside the car can be distinguished based on the time difference, energy difference, etc. between the sound sources reaching each microphone. For example, the receiving time difference or receiving energy difference between the sound signals in the multi-channel audio signals of the sound source inside the car is smaller than the receiving time difference or receiving energy difference between the sound signals in the multi-channel audio signals of the sound source inside the car.
[0048] The above difference information is used to indicate whether the difference between the sound signals in the multi-channel audio signals of each sound source meets the preset condition. For example, it can indicate whether the reception time difference between the sound signals in the multi-channel audio signals of each sound source meets the preset time difference threshold, or it can also indicate whether the reception energy difference between the sound signals in the multi-channel audio signals of each sound source meets the preset energy difference threshold. In this embodiment, the preset time difference threshold and the preset energy difference threshold are set according to the actual situation of the user and are not specifically limited here.
[0049] In this embodiment, the target end-to-end network model obtained through pre-training can detect the spatial position of the sound source of each sound source in the multi-channel audio signal of at least one sound source, and accurately distinguish the spatial position of each sound source as the sound source position inside the vehicle or the sound source position outside the vehicle, thereby improving the accuracy of the sound zone recognition inside the vehicle, and further enhancing the user's voice interaction experience.
[0050] As an implementation of the present application, in order to obtain a target end-to-end network model, the training process of the target end-to-end network model may specifically include:
[0051] Obtain a target training set, the target training set includes multiple training samples, the multiple training samples include a first training sample, a second training sample and a third training sample, the first training sample includes sound signals of multiple historical in-vehicle sound sources collected by multiple sound sensors, the second training sample includes sound signals of multiple historical external sound sources collected by multiple sound sensors, and the third training sample includes at least one historical in-vehicle sound source and at least one historical external sound source collected by multiple sound sensors; the quantity ratio of the first training sample, the second training sample and the third training sample in the target training set is a preset quantity ratio, and the quantity ratio of the first training sample is higher than the quantity ratio of the third training sample; the quantity ratio of the third training sample is higher than the quantity ratio of the second training sample;
[0052] The initial end-to-end network model is trained using training samples in the target training set to update the network parameters of the initial end-to-end network model and obtain the target end-to-end network model.
[0053] The above-mentioned first training sample may include sound signals of multiple historical in-vehicle sound sources collected by multiple sound sensors, wherein the sound signals of the historical in-vehicle sound sources may, for example, be sound signals of 1 to M single sound sources distributed in N different sound zones in the vehicle, where M<=N.
[0054] The second training sample may include sound signals of multiple historical external sound sources collected by multiple sound sensors, wherein the sound signals of the historical external sound sources are, for example, simulated by simulation means to simulate a single external sound source, thereby obtaining 0 to V sound signals of a single external sound source.
[0055] The third training sample may include at least one historical sound signal from an in-vehicle sound source and at least one historical sound signal from an out-vehicle sound source collected by a plurality of sound sensors.
[0056] In this embodiment, the target training set includes a first training sample, a second training sample, and a third training sample, so that when the initial end-to-end network model is trained, the network parameters can be updated based on the first training sample, the second training sample, and the third training sample to obtain a target end-to-end network model with higher accuracy and robustness.
[0057] In the above-mentioned preset quantity ratio, the quantity ratio of the first training samples is higher than the quantity ratio of the second training samples and the third training samples. Exemplarily, it can be a 5:1:4 ratio. In this embodiment, the preset quantity ratio is not limited to this, and can also be other ratios, which are not specifically limited here.
[0058] In this embodiment, the ratio of the number of first training samples in the target training set is higher than the ratio of the number of second training samples and the number of third training samples. The ratio of the number of samples in the target training set is biased towards the scenario of the sound source inside the vehicle. When the initial end-to-end network model is trained by the training samples in the target training set, it can be ensured that the separation performance of the target end-to-end network model for the sound source inside the vehicle is not reduced.
[0059] In some embodiments, the above-mentioned obtaining of the target training set may specifically include:
[0060] Acquire an initial sample, where the initial sample includes sound signals of target historical sound sources corresponding to multiple different spatial positions collected by multiple sound sensors;
[0061] The sound signals generated by the target historical sound sources corresponding to different spatial positions are mixed and processed according to a preset ratio to obtain mixed speech in different scenarios; wherein the different scenarios include: mixed speech inside the vehicle, mixed speech outside the vehicle, and mixed speech inside and outside the vehicle;
[0062] Silence signal segments are configured for mixed speech in different scenarios to obtain training samples.
[0063] The above-mentioned multiple target historical sound sources may be multiple historical in-vehicle sound sources, multiple historical outside-vehicle sound sources, or at least one historical in-vehicle sound source and at least one historical outside-vehicle sound source.
[0064] The preset ratio of the mixed voice inside and outside the vehicle can be, for example, a ratio of 5:1:4. In this embodiment, the preset ratio is not limited to this, but can also be other ratios, which are not specifically limited here.
[0065] The above-mentioned different scenarios may include: mixed speech in the car, mixed speech outside the car, and mixed speech inside and outside the car. Exemplarily, the mixed speech inside the car can be a mixed speech collected by N microphones in the car by simulating 1 to M single sound sources distributed in N different sound zones by simulation means; the mixed speech outside the car can be a mixed speech collected by N microphones in the car by simulating 0 to V single sound sources outside the car by simulation means to simulate external interference; the mixed speech inside and outside the car can be a mixed speech collected by N microphones in the car by simulating 1 to M single sound sources distributed in N different sound zones by simulation means, and adding 0 to V single sound sources outside the car to simulate external interference.
[0066] In this embodiment, the sound signals generated by the target historical sound sources corresponding to different spatial positions are mixed and processed according to a preset ratio to obtain mixed speech in different scenarios, and silent signal segments are configured for the mixed speech in different scenarios to obtain training samples, thereby improving the accuracy of the target end-to-end network model.
[0067] In some embodiments, configuring silence signal segments for mixed speech in different scenarios to obtain training samples may specifically include:
[0068] A silent signal is configured between the sound signals of multiple target historical sound sources in the target mixed speech to obtain a target training sample corresponding to the initial sample. A silent signal is configured between every two sound signals of the target historical sound sources in the target training sample. The duration ratio of the silent signal in the target training sample is within a preset ratio range. The target mixed speech is any one of the mixed speech inside the vehicle, the mixed speech outside the vehicle, and the mixed speech inside and outside the vehicle.
[0069] The target training sample may be the first training sample, the second training sample or the third training sample.
[0070] The above-mentioned preset ratio range may be, for example, a ratio range of 40%-70%.
[0071] In this embodiment, by configuring a silent signal between the sound signals of multiple target historical sound sources, the proportion of silent segments can be increased while simulating a real vehicle scene, thereby reducing the impact of the sound source labeling deviation at the frame level.
[0072] In some embodiments, the above S103 may specifically include:
[0073] Performing spatial position label processing on each sound source in the multi-channel audio signal to obtain a target multi-channel audio signal, wherein the target multi-channel audio signal includes a sound signal of at least one sound source and a spatial position label of each sound source, wherein the spatial position label is used to indicate that the spatial position corresponding to the sound source is the spatial position corresponding to the sound source inside the vehicle or the spatial position corresponding to the sound source outside the vehicle;
[0074] The above S104 may specifically include:
[0075] Based on the spatial position label of each sound source in the target multi-channel audio signal, the sound signal corresponding to the sound source whose spatial position label is the spatial position corresponding to the in-vehicle sound source is identified, and the target sound area corresponding to the sound source at the spatial position corresponding to the in-vehicle sound source is determined.
[0076] In this embodiment, a target multi-channel audio signal is obtained by performing spatial position label processing on each sound source in a multi-channel audio signal. The target multi-channel audio signal includes a sound signal of at least one sound source and a spatial position label of each sound source. The spatial position label is used to indicate that the spatial position corresponding to the sound source is the spatial position corresponding to the sound source inside the vehicle or the spatial position corresponding to the sound source outside the vehicle. In this way, in subsequent sound zone recognition processing, only the sound signal whose spatial position label indicates that it corresponds to the spatial position of the sound source inside the vehicle in the target multi-channel audio signal is processed, thereby achieving the suppression of the sound signal of the sound source outside the vehicle.
[0077] In order to facilitate the understanding of the sound zone separation method of the in-car audio in the embodiment of the present application, the application process of the sound zone separation method of the in-car audio is described here, as follows:
[0078] Assume that there are N sound zones and microphones in the car, and the sound zones correspond to the microphones one by one. The speech signal collected by the microphone is X = [X1…X N ] T , including all the sound sources and noise inside and outside the car. Assume that there are M sound sources in the car (M<=N), which are S=[S1…S M ] T There are V sound sources outside the car, which are I=[I1…I V ] T , then through online learning we can get Contains the output of N sound zones, and the estimated M in-car sound sources are distributed in N corresponding sound zones All sound sources outside the vehicle are suppressed.
[0079] Training phase:
[0080] Step 1: Use simulation methods to simulate 1 to M single sound sources distributed in N different sound zones, and add 0 to V single sound sources outside the car to simulate external interference, construct mixed speech collected by N microphones in the car, and retain 1 to M single sound sources collected by the microphone as learning targets to obtain training set data;
[0081] Step 2: Input mixed speech inside the car, mixed speech outside the car, and mixed speech inside and outside the car in different proportions (considering the need to prioritize the performance of separation inside the car, the proportion can be biased towards the inside car scene, such as inside the car: outside the car: inside and outside the car = 5:1:4). The mixed speech of different scenes contains 40%-70% of silent segments. While simulating the real car scene, the proportion of silent segments is increased to reduce the impact of the deviation of the sound source labels at the frame level. The number of sound sources in the car is simulated to be 1 to M and whether there is human voice interference outside the car. The end-to-end multi-zone separation network (that is, the initial end-to-end network model mentioned above) is trained, mixed speech input, and multi-zone separation speech in the car is output. The loss function of 1-M single sound sources collected by N microphones and the output is calculated. The loss is the commonly used L1 (mean absolute error, MAE) or L2 (mean squared error, MSE). When the loss meets the preset conditions, a trained end-to-end multi-sound zone separation network (i.e., the above target end-to-end network model) is obtained. If the current speech does not have a sound source inside the car, then silence or pure noise is the learning target, which can ensure that each speech can achieve the purpose of sound zone separation and suppress the sound source outside the car.
[0082] Testing phase:
[0083] Step 1: In an actual vehicle environment, N microphones collect N mixed audio signals, where the N audio signals are signals of at least one sound source collected by the N microphones;
[0084] Step 2: Directly input the mixed audio into the end-to-end multi-zone separation network (i.e., the target end-to-end network model mentioned above) to obtain a separated speech signal with suppressed external interference.
[0085] In this embodiment, since there are multiple microphones in the car, the sound sources inside and outside the car can be distinguished based on the time difference, energy difference, etc. between the sound sources reaching each microphone. This difference is learned through an end-to-end multi-zone separation network, thereby achieving the purpose of end-to-end separation suppression.
[0086] Based on the in-vehicle audio sound zone separation method provided in the above embodiment, the present application also provides a specific implementation of the in-vehicle audio sound zone separation device. Please refer to the following embodiment.
[0087] like Figure 2 As shown, an in-vehicle audio sound zone separation device 200 provided in an embodiment of the present application may include the following modules: a first acquisition module 201 , a detection module 202 , a suppression module 203 and an identification module 204 .
[0088] A first acquisition module 201 is used to acquire multi-channel audio signals collected by multiple sound sensors in the vehicle, where the multi-channel audio signals include sound signals generated by sound sources corresponding to different spatial positions, and different sound sensors are arranged in different sound zones in the vehicle;
[0089] The detection module 202 is used to detect the sound source spatial position of each sound source in the multi-channel audio signal based on the pre-trained target end-to-end network model, and obtain the sound source spatial position corresponding to each sound signal in the multi-channel audio signal; the sound source spatial position includes the sound source position inside the vehicle and the sound source position outside the vehicle;
[0090] The suppression module 203 is used to suppress the sound signal whose sound source spatial position is the sound source position outside the vehicle in the multi-channel audio signal, and obtain the sound signal corresponding to the sound source position inside the vehicle;
[0091] The recognition module 204 is used to recognize the sound zone in the vehicle for the sound signal corresponding to the sound source position in the vehicle, and perform sound zone separation on the sound signal corresponding to the sound source position in the vehicle according to the recognition result.
[0092] The sound zone separation device for in-car audio of the embodiment of the present application can detect the sound source spatial position of the sound signal of each sound source in the multi-channel audio signal based on the pre-trained target end-to-end network model according to the multi-channel audio signals collected by multiple sound sensors in the car, and obtain the sound source spatial position corresponding to each sound signal in the multi-channel audio signal; the sound source spatial position includes the sound source position in the car and the sound source position outside the car; suppress the sound signal whose sound source spatial position in the multi-channel audio signal is the sound source position outside the car, and obtain the sound signal corresponding to the sound source position in the car; identify the sound zone in the car for the sound signal corresponding to the sound source position in the car, and perform sound zone separation on the sound signal corresponding to the sound source position in the car according to the identification result. In this way, the embodiment of the present application can determine, based on the multi-channel audio signals collected by multiple sound sensors in the car and the pre-trained target end-to-end network model, whether the sound source spatial position corresponding to each sound signal in the multi-channel audio signals is the sound source position inside the car or the sound source position outside the car, and suppress the sound signal at the sound source position outside the car to reduce the interference with the sound zone recognition in the car with the sound signal corresponding to the sound source position inside the car, thereby improving the accuracy of the sound zone recognition in the car and further enhancing the user's voice interaction experience.
[0093] In some embodiments, the detection module 202 may include:
[0094] a detection unit, configured to detect difference information of sound signals in the multi-channel audio signals of each sound source, wherein the difference information is used to indicate whether a time difference or an energy difference between the sound signals in the multi-channel audio signals of each sound source meets a preset condition;
[0095] A determination unit is used to determine whether each sound source is an in-vehicle sound source or an outside-vehicle sound source based on difference information of each sound source, wherein if a time difference or an energy difference between sound signals in a multi-channel audio signal meets a preset condition, the sound source spatial position corresponding to the sound signal includes the in-vehicle sound source position; if a time difference or an energy difference between sound signals in a multi-channel audio signal does not meet the preset condition, the sound source spatial position corresponding to the sound signal includes the outside-vehicle sound source position.
[0096] As an implementation of the present application, in order to obtain a target end-to-end network model, the above-mentioned device 200 may further include:
[0097] A second acquisition module is used to acquire a target training set, the target training set includes multiple training samples, the multiple training samples include a first training sample, a second training sample and a third training sample, the first training sample includes sound signals of multiple historical in-vehicle sound sources collected by multiple sound sensors, the second training sample includes sound signals of multiple historical external sound sources collected by multiple sound sensors, and the third training sample includes sound signals of at least one historical in-vehicle sound source and at least one historical external sound source collected by multiple sound sensors; the quantity ratio of the first training sample, the second training sample and the third training sample in the target training set is a preset quantity ratio, and the quantity ratio of the first training sample is higher than the quantity ratio of the third training sample; the quantity ratio of the third training sample is higher than the quantity ratio of the second training sample;
[0098] The training module is used to train the initial end-to-end network model through training samples in the target training set to update the network parameters of the initial end-to-end network model to obtain the target end-to-end network model.
[0099] In some embodiments, the second acquisition module may specifically include:
[0100] An acquisition unit, used to acquire an initial sample, wherein the initial sample includes sound signals of target historical sound sources corresponding to multiple different spatial positions collected by multiple sound sensors;
[0101] The first processing unit is used to mix and process the sound signals generated by the target historical sound sources corresponding to different spatial positions according to a preset ratio to obtain mixed speech in different scenarios; wherein the different scenarios include: mixed speech inside the vehicle, mixed speech outside the vehicle, and mixed speech inside and outside the vehicle;
[0102] The configuration unit is used to configure silence signal segments for mixed speech in different scenarios to obtain training samples.
[0103] In some embodiments, the above-mentioned configuration unit can be specifically used to configure silent signals between the sound signals of multiple target historical sound sources in the target mixed speech, so as to obtain a target training sample corresponding to the initial sample, wherein there is a silent signal between every two sound signals of the target historical sound sources in the target training sample, and the duration ratio of the silent signal in the target training sample is within a preset ratio range, and the target mixed speech is any one of the mixed speech inside the vehicle, the mixed speech outside the vehicle, and the mixed speech inside and outside the vehicle.
[0104] In some embodiments, the suppression module 203 may specifically include:
[0105] a second processing unit, configured to perform spatial position label processing on each sound source in the multi-channel audio signal to obtain a target multi-channel audio signal, wherein the target multi-channel audio signal includes a sound signal of at least one sound source and a spatial position label of each sound source, wherein the spatial position label is used to indicate that the spatial position corresponding to the sound source is the spatial position corresponding to the sound source inside the vehicle or the spatial position corresponding to the sound source outside the vehicle;
[0106] The identification module 204 may specifically include:
[0107] The recognition unit is used to recognize the sound signal corresponding to the sound source whose spatial position label is the spatial position corresponding to the in-vehicle sound source based on the spatial position label of each sound source in the target multi-channel audio signal, and determine the target sound area corresponding to the sound source at the spatial position corresponding to the in-vehicle sound source.
[0108] Based on the in-car audio sound zone separation method provided in the above embodiment, the present application also provides a specific implementation of the electronic device. Please refer to the following embodiment.
[0109] Figure 3 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.
[0110] The electronic device may include a processor 301 and a memory 302 storing computer program instructions.
[0111] Specifically, the processor 301 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0112] The memory 302 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 302 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 302 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 302 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 302 is a non-volatile solid-state memory.
[0113] In certain embodiments, the memory 302 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Thus, in general, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.
[0114] The processor 301 reads and executes the computer program instructions stored in the memory 302 to implement any one of the in-vehicle audio sound zone separation methods in the above embodiments.
[0115] In one example, the electronic device may further include a communication interface 303 and a bus 310. Figure 3 As shown, the processor 301, the memory 302, and the communication interface 303 are connected via a bus 310 and communicate with each other.
[0116] The communication interface 303 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0117] Bus 310 includes hardware, software or both, and the parts of electronic equipment are coupled to each other. For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 310 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the application considers any suitable bus or interconnection.
[0118] The electronic device can execute the audio zone separation method of the in-car audio in the embodiment of the present application, thereby realizing the combination Figure 1 and Figure 2 A method and device for separating audio zones of in-car audio are described.
[0119] In addition, in combination with the in-vehicle audio sound zone separation method in the above embodiments, the present application embodiment can provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any of the in-vehicle audio sound zone separation methods in the above embodiments is implemented.
[0120] In combination with the in-vehicle audio sound zone separation method in the above embodiment, the present application embodiment can provide a vehicle for implementation. The vehicle includes: an electronic device for implementing any of the in-vehicle audio sound zone separation methods as described above.
[0121] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.
[0122] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0123] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.
[0124] Aspects of the present disclosure are described above with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs a specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0125] The above is only a specific implementation of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application.
[0126] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprises" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
Claims
1. A method for separating audio zones in a car, characterized in that: The method comprises: Acquire multi-channel audio signals collected by multiple sound sensors in the vehicle, wherein the multi-channel audio signals include sound signals generated by sound sources corresponding to different spatial positions, and different sound sensors are arranged in different sound zones in the vehicle; Based on the pre-trained target end-to-end network model, the sound source spatial position detection is performed on the sound signal of each sound source in the multi-channel audio signal to obtain the sound source spatial position corresponding to each sound signal in the multi-channel audio signal; the sound source spatial position includes the sound source position inside the vehicle and the sound source position outside the vehicle; Suppressing the sound signal whose sound source spatial position is the sound source position outside the vehicle in the multi-channel audio signal, and obtaining the sound signal corresponding to the sound source position inside the vehicle; The sound zone in the vehicle of the sound signal corresponding to the position of the sound source in the vehicle is identified, and the sound zone of the sound signal corresponding to the position of the sound source in the vehicle is separated according to the identification result.
2. The method according to claim 1, characterized in that The detecting of the spatial position of each sound source of the sound signal in the multi-channel audio signal to obtain the spatial position of the sound source corresponding to each sound signal in the multi-channel audio signal includes: Detecting difference information of sound signals in the multi-channel audio signals of each sound source, wherein the difference information is used to indicate whether a time difference or an energy difference between the sound signals in the multi-channel audio signals of each sound source meets a preset condition; According to the difference information of each of the sound sources, it is determined that each of the sound sources is an in-vehicle sound source or an outside-vehicle sound source, wherein if the time difference or energy difference between the sound signals in the multi-channel audio signals meets the preset condition, the sound source spatial position corresponding to the sound signal includes the sound source position inside the vehicle; if the time difference or energy difference between the sound signals in the multi-channel audio signals does not meet the preset condition, the sound source spatial position corresponding to the sound signal includes the sound source position outside the vehicle.
3. The method according to claim 2, characterized in that The training process of the target end-to-end network model includes: Acquire a target training set, the target training set includes multiple training samples, the multiple training samples include a first training sample, a second training sample and a third training sample, the first training sample includes sound signals of multiple historical in-vehicle sound sources collected by multiple sound sensors, the second training sample includes sound signals of multiple historical outside-vehicle sound sources collected by the multiple sound sensors, and the third training sample includes at least one historical in-vehicle sound source and at least one historical outside-vehicle sound source collected by the multiple sound sensors; the quantity ratio of the first training sample, the second training sample and the third training sample in the target training set is a preset quantity ratio, and the quantity ratio of the first training sample is higher than the quantity ratio of the third training sample; the quantity ratio of the third training sample is higher than the quantity ratio of the second training sample; The initial end-to-end network model is trained using the training samples in the target training set to update the network parameters of the initial end-to-end network model and obtain the target end-to-end network model.
4. The method according to claim 3, characterized in that The obtaining of the target training set comprises: Acquire an initial sample, wherein the initial sample includes sound signals of target historical sound sources corresponding to multiple different spatial positions collected by multiple sound sensors; Mixing and processing the sound signals generated by the target historical sound sources corresponding to different spatial positions according to a preset ratio to obtain mixed speech in different scenarios; wherein the different scenarios include: mixed speech inside the vehicle, mixed speech outside the vehicle, and mixed speech inside and outside the vehicle; Silence signal segments are configured for mixed speech in different scenarios to obtain training samples.
5. The method according to claim 4, characterized in that The step of respectively configuring silent signal segments for mixed speech in different scenarios to obtain training samples includes: A silent signal is respectively configured between the sound signals of multiple target historical sound sources in the target mixed speech to obtain a target training sample corresponding to the initial sample, wherein the silent signal is between every two sound signals of the target historical sound sources in the target training sample, and the duration ratio of the silent signal in the target training sample is within a preset ratio range, and the target mixed speech is any one of the mixed speech inside the vehicle, the mixed speech outside the vehicle, and the mixed speech inside and outside the vehicle.
6. The method according to claim 1, characterized in that The step of suppressing the sound signal whose sound source spatial position is the sound source position outside the vehicle in the multi-channel audio signal to obtain the sound signal corresponding to the sound source position inside the vehicle includes: Performing spatial position label processing on each of the sound sources in the multi-channel audio signal to obtain a target multi-channel audio signal, wherein the target multi-channel audio signal includes a sound signal of the at least one sound source and a spatial position label of each of the sound sources, wherein the spatial position label is used to indicate that the spatial position corresponding to the sound source is the spatial position corresponding to the in-vehicle sound source or the spatial position corresponding to the out-vehicle sound source; The step of identifying the sound zone in the vehicle for the sound signal corresponding to the position of the sound source in the vehicle, and performing sound zone separation on the sound signal corresponding to the position of the sound source in the vehicle according to the identification result, comprises: Based on the spatial position labels of each sound source in the target multi-channel audio signal, the sound signal corresponding to the sound source whose spatial position label is the spatial position corresponding to the in-vehicle sound source is identified, and the target sound area corresponding to the sound source at the spatial position corresponding to the in-vehicle sound source is determined.
7. A sound zone separation device for in-car audio, characterized in that: The method comprises: A first acquisition module is used to acquire multi-channel audio signals collected by multiple sound sensors in the vehicle, wherein the multi-channel audio signals include sound signals generated by sound sources corresponding to different spatial positions, and different sound sensors are arranged in different sound zones in the vehicle; A detection module, configured to detect the spatial position of each sound source in the multi-channel audio signal based on a pre-trained target end-to-end network model, and obtain the spatial position of the sound source corresponding to each sound signal in the multi-channel audio signal; the spatial position of the sound source includes the position of the sound source inside the vehicle and the position of the sound source outside the vehicle; A suppression module, used to suppress the sound signal whose sound source spatial position is the sound source position outside the vehicle in the multi-channel audio signal, and obtain the sound signal corresponding to the sound source position inside the vehicle; The recognition module is used to recognize the sound zone in the vehicle of the sound signal corresponding to the sound source position in the vehicle, and perform sound zone separation on the sound signal corresponding to the sound source position in the vehicle according to the recognition result.
8. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the method for separating the sound zones of in-vehicle speech as described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by the processor, the method for separating the sound zones of in-vehicle speech as described in any one of claims 1-6 is implemented.
10. A vehicle, characterized in that: include: An electronic device, wherein the electronic device is used to implement the method for separating the sound zones of in-vehicle speech as described in any one of claims 1 to 6.