In-vehicle communication method, device, equipment and medium

By encoding and time-frequency analysis of the in-car call signal, determining the mask estimate, and eliminating various interference sound components in the in-car call signal, the sound quality problem caused by error accumulation in the existing technology is solved, and efficient interference sound removal and sound quality assurance are achieved.

CN120835110APending Publication Date: 2025-10-24BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410455716.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-16
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

In existing in-car communication systems, the interference sound processing step has error accumulation, which affects the call quality and cannot effectively guarantee the call quality of the in-car communication system.

Method used

By encoding the in-car call signal, obtaining the frequency characteristics and performing time-frequency analysis, the mask values ​​of various interference sound components are determined, and masking processing is performed using mask estimation to uniformly eliminate various interference sound components and avoid error accumulation caused by serial processing.

Benefits of technology

The interference noise removal effect of the in-car call signal is improved, the sound quality is ensured, the error accumulation is avoided, and the sound quality of the in-car call is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120835110A_ABST
    Figure CN120835110A_ABST
Patent Text Reader

Abstract

The invention discloses an in-vehicle call method, device, equipment and medium, and the method comprises the steps: carrying out the coding of an in-vehicle call signal under the condition of obtaining the in-vehicle call signal, and obtaining the frequency characteristics of the in-vehicle call signal, the in-vehicle call signal comprising various interference sound components and an in-vehicle clean signal; performing time-frequency analysis on the frequency characteristics to obtain time-frequency characteristics of the in-vehicle call signal, the time-frequency characteristics being used for representing distribution of various interference sound components in the in-vehicle call signal; performing decoding processing on the time-frequency characteristics, determining mask values of various interference sound components in the in-vehicle call signal on each time-frequency point, and obtaining mask estimation of the in-vehicle call signal; and performing masking processing on the in-vehicle call signal based on mask estimation, and eliminating various interference sound components in the in-vehicle call signal to obtain an in-vehicle clean signal. According to the embodiment of the invention, the interference sound removal effect of the in-vehicle call signal can be improved, and the sound quality of the in-vehicle call is effectively ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of in-vehicle communication, and particularly relates to an in-vehicle communication method, device, equipment and medium. BACKGROUND

[0002] The in-vehicle communication system refers to a communication system for facilitating communication between people at different positions in a vehicle in a noise environment during driving, using a microphone and a loudspeaker in the vehicle to realize normal communication between different seats in the vehicle.

[0003] In the related art, the original signal generated by the in-vehicle personnel during in-vehicle communication usually contains various interference sounds. In order to remove the various interference sound components, the current in-vehicle communication scheme usually performs serial processing on the various interference sound components carried in the original signal. For example, the original signal is first subjected to media echo cancellation, then the signal after the media echo cancellation is subjected to feedback echo cancellation, and then the signal after the feedback echo cancellation is subjected to noise cancellation to obtain a final clean signal. Since there are usually errors in the interference sound processing steps, the serial processing will cause the errors of each interference sound processing step to accumulate, thereby affecting the interference sound removal effect of the original signal and failing to guarantee the sound quality of the in-vehicle communication system. SUMMARY

[0004] The purpose of the embodiments of the present application is to provide an in-vehicle communication method and device, which can improve the interference sound removal effect of the in-vehicle communication signal and effectively guarantee the sound quality of the in-vehicle communication.

[0005] In a first aspect, the embodiments of the present application provide an in-vehicle communication method, which comprises: in the case of obtaining an in-vehicle communication signal, performing encoding processing on the in-vehicle communication signal to obtain a frequency feature of the in-vehicle communication signal, wherein the in-vehicle communication signal comprises various interference sound components and an in-vehicle clean signal; performing time-frequency analysis on the frequency feature to obtain a time-frequency feature of the in-vehicle communication signal, wherein the time-frequency feature is used to represent the distribution of the various interference sound components in the in-vehicle communication signal; performing decoding processing on the time-frequency feature to determine the mask value of the various interference sound components in the in-vehicle communication signal at each time-frequency point, and obtain a mask estimation of the in-vehicle communication signal; and performing masking processing on the in-vehicle communication signal based on the mask estimation to eliminate the various interference sound components in the in-vehicle communication signal, and obtain the in-vehicle clean signal.

[0006] In some implementable manners of the first aspect, in the case of obtaining the in-vehicle communication signal, the encoding processing on the in-vehicle communication signal to obtain the frequency feature of the in-vehicle communication signal comprises: performing feature extraction on the in-vehicle communication signal by a target encoder to obtain a first sub-frequency feature and M second sub-frequency features corresponding to M interference sound components; and the frequency feature comprises the first sub-frequency feature and the M second sub-frequency features.

[0007] In some possible implementation manners of the first aspect, the target encoder comprises a first encoder and M second encoders, and the feature extraction on the in-vehicle conversation signal by the target encoder to obtain the first sub-frequency feature and the M second sub-frequency features corresponding to the M types of interference sound components comprises: performing feature extraction on the in-vehicle conversation signal by the first encoder to obtain the first sub-frequency feature; and performing feature extraction on the in-vehicle conversation signal by the M second encoders corresponding to the M types of interference sound components to obtain the M second sub-frequency features corresponding to the M types of interference sound components, wherein the second encoders are preset autoencoders for the interference sound components.

[0008] In some possible implementation manners of the first aspect, the in-vehicle clean signal is obtained by performing the masking processing on the in-vehicle conversation signal based on the mask estimation and eliminating the types of interference sound components in the in-vehicle conversation signal, and the masking processing comprises: multiplying the in-vehicle conversation signal by the mask estimation to obtain the in-vehicle clean signal.

[0009] In some possible implementation manners of the first aspect, the time-frequency analysis is performed on the frequency features to obtain the time-frequency features of the in-vehicle conversation signal, and the time-frequency analysis comprises: determining the distribution of the types of interference sound components at each time-frequency point of the in-vehicle conversation signal by analyzing the relationship between the types of interference sound components and the in-vehicle conversation signal in the feature domain to obtain the time-frequency features of the in-vehicle conversation signal.

[0010] In some possible implementation manners of the first aspect, before the encoding processing is performed on the in-vehicle conversation signal, the method further comprises: obtaining media echo data, feedback echo data, noise data and clean speech data; mixing the media echo data, the feedback echo data, the noise data and the clean speech data by using preset signal-to-noise ratios, preset signal-to-feedback ratios and preset signal-to-interference ratios to obtain mixed data; determining a training mask estimation based on the mixed data and the clean speech data; training a preset neural network model by taking the mixed data as input training data and taking the training mask estimation as output training data to obtain an interference sound removal model; and wherein the input of the interference sound removal model is the in-vehicle conversation signal, and the output is the mask estimation of the in-vehicle conversation signal.

[0011] In some possible implementation manners of the first aspect, after the in-vehicle clean signal is obtained by performing the masking processing on the in-vehicle conversation signal based on the mask estimation and eliminating the types of interference sound components in the in-vehicle conversation signal, the method further comprises: repairing a distorted part of the in-vehicle clean signal by using a post-filtering algorithm to obtain a target clean signal.

[0012] In some possible implementation manners of the first aspect, the in-vehicle conversation signal is a multi-channel microphone signal, and after the in-vehicle clean signal is obtained by performing the masking processing on the in-vehicle conversation signal based on the mask estimation and eliminating the types of interference sound components in the in-vehicle conversation signal, the method further comprises: performing channel separation on the in-vehicle clean signal to obtain a single-channel in-vehicle clean signal under each channel.

[0013] In a second aspect, an in-vehicle intercom device is provided. The device includes an encoding module configured to encode an in-vehicle intercom signal to obtain a frequency feature of the in-vehicle intercom signal, the in-vehicle intercom signal including various interference sound components and a clean in-vehicle signal; a time-frequency analysis module configured to perform time-frequency analysis on the frequency feature to obtain a time-frequency feature of the in-vehicle intercom signal, the time-frequency feature being indicative of a distribution of the various interference sound components in the in-vehicle intercom signal; a decoding module configured to decode the time-frequency feature to determine a mask value of each of the various interference sound components at each time-frequency point in the in-vehicle intercom signal, and obtain a mask estimation of the in-vehicle intercom signal; and a masking module configured to mask the in-vehicle intercom signal based on the mask estimation to eliminate the various interference sound components in the in-vehicle intercom signal, and obtain the clean in-vehicle signal.

[0014] In some implementations of the second aspect, the encoding module is specifically configured to: extract features of the in-vehicle intercom signal by a target encoder to obtain a first sub-frequency feature and M second sub-frequency features corresponding to the M interference sound components; and wherein the frequency feature includes the first sub-frequency feature and the M second sub-frequency features.

[0015] In some implementations of the second aspect, the target encoder includes a first encoder and M second encoders, and the encoding module is specifically configured to: extract features of the in-vehicle intercom signal by the first encoder to obtain the first sub-frequency feature; and extract features of the in-vehicle intercom signal by the M second encoders corresponding to the M interference sound components to obtain the M second sub-frequency features corresponding to the M interference sound components, wherein the second encoders are preset autoencoders for the interference sound components.

[0016] In some implementations of the second aspect, the masking module is specifically configured to: multiply the in-vehicle intercom signal by the mask estimation to obtain the clean in-vehicle signal.

[0017] In some implementations of the second aspect, the time-frequency analysis module is specifically configured to: analyze a relationship between the various interference sound components and the in-vehicle intercom signal in a feature domain to determine a distribution of the various interference sound components at each time-frequency point in the in-vehicle intercom signal, and obtain the time-frequency feature of the in-vehicle intercom signal.

[0018] In some implementable manners of the second aspect, the apparatus further comprises: an acquisition module, configured to acquire the media echo data, the feedback echo data, the noise data and the clean speech data before the in-vehicle conversation signal is processed by encoding; a data mixing module, configured to mix the media echo data, the feedback echo data, the noise data and the clean speech data to obtain mixed data by using a preset signal-to-noise ratio, a preset signal-to-feedback ratio and a preset signal-to-interference ratio; a determination module, configured to determine a trained mask estimation based on the mixed data and the clean speech data; and a training module, configured to train the preset neural network model by taking the mixed data as input training data and the trained mask estimation as output training data, to obtain the interference sound removal model, wherein the input of the interference sound removal model is the in-vehicle conversation signal, and the output is the mask estimation of the in-vehicle conversation signal.

[0019] In some implementable manners of the second aspect, the apparatus further comprises a post-filtering module, configured to perform post-filtering on a distorted part of the in-vehicle clean signal by using a post-filtering algorithm to obtain a target clean signal after the in-vehicle clean signal is obtained by performing the masking processing on the in-vehicle conversation signal based on the mask estimation to eliminate various interference sound components in the in-vehicle conversation signal.

[0020] In some implementable manners of the second aspect, the in-vehicle conversation signal is a multi-channel microphone signal, and the apparatus further comprises a channel separation module, configured to perform channel separation on the in-vehicle clean signal to obtain a single-channel in-vehicle clean signal under each channel after the in-vehicle clean signal is obtained by performing the masking processing on the in-vehicle conversation signal based on the mask estimation to eliminate various interference sound components in the in-vehicle conversation signal.

[0021] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory storing computer program instructions; the processor implements the steps of the in-vehicle conversation method of the first aspect when executing the computer program instructions.

[0022] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, the computer readable storage medium stores computer program instructions, and the computer program instructions are executed by a processor to implement the steps of the in-vehicle conversation method of the first aspect.

[0023] In a fifth aspect, an embodiment of the present application provides a computer program product stored in a non-volatile storage medium, and the computer program product is executed by at least one processor to implement the steps of the in-vehicle conversation method of the first aspect.

[0024] In a sixth aspect, an embodiment of the present application provides a chip, which comprises a processor and a communication interface, the communication interface and the processor are coupled, and the processor is configured to run a program or an instruction to implement the steps of the in-vehicle conversation method of the first aspect.

[0025] The present application provides an in-car call method, apparatus, device, and medium. The present application uniformly and simultaneously removes various interference sound components, including media echo, feedback echo, and noise. Specifically, by encoding the in-car call signal and performing time-frequency analysis on the encoded frequency characteristics, the time-frequency characteristics of the in-car call signal can be obtained. The time-frequency characteristics can reflect the timing information and frequency information of the in-car call signal. Therefore, based on the time-frequency characteristics, the distribution of various interference sound components in the in-car call signal can be accurately expressed. That is, the relationship between the various interference sound components in the in-car call signal and the clean signal in the car is clarified in the feature domain. This relationship is the key to eliminating the various interference sound components in the in-car call signal and retaining the clean signal in the car. Therefore, by decoding the time-frequency features and determining the mask values ​​of various interference sound components in the in-car call signal at various time-frequency points, a mask estimate can be obtained. The mask estimate is used to eliminate various interference sound components in the in-car call signal, rather than a specific interference sound component. Therefore, by performing masking processing on the in-car call signal based on the mask estimate, all interference sound components can be eliminated from the clean signal in the car at one time. There is no need to execute a corresponding interference sound elimination step for each type of interference sound component. Therefore, it can avoid the error accumulation caused by serial execution of multiple interference sound processing steps, thereby improving the interference sound removal effect of the in-car call signal and effectively ensuring the sound quality of the in-car call. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present application, a brief introduction to the drawings required for use in the embodiments of the present application is given below.

[0027] Figure 1 This is a flow chart of an in-car call method provided by an embodiment of the present application;

[0028] Figure 2 is a flowchart of an in-car communication method provided by another embodiment of the present application;

[0029] Figure 3 This is a schematic structural diagram of an in-car communication device provided by an embodiment of the present application;

[0030] Figure 4 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] The features and exemplary embodiments of various aspects of the present application will be described in detail below with reference to the drawings. In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, but not to limit the present application. The present application can be implemented without some of the specific details for those skilled in the art. The following description of the embodiments is only to provide a better understanding of the present application by showing examples of the present application.

[0032] At present, the original signal generated by the in-vehicle personnel when talking in the vehicle usually contains various interference sounds. In order to remove various interference sounds, the current in-vehicle communication scheme usually serially processes various interference sound components carried in the original signal, for example, first performing media echo cancellation on the original signal, then performing feedback echo cancellation on the signal after media echo cancellation, and then performing noise cancellation on the signal after feedback echo cancellation to obtain the final clean signal. Since there are usually errors in the interference sound processing steps, the serial processing will cause the errors of each interference sound processing step to accumulate, thereby affecting the interference sound removal effect of the original signal and unable to guarantee the call sound quality of the in-vehicle communication system.

[0033] The in-vehicle communication method provided by the embodiments of the present application will be described in detail below with reference to the drawings, specific embodiments and application scenarios.

[0034] Figure 1 is a flowchart of the in-vehicle communication method provided by an embodiment of the present application. The execution subject of the in-vehicle communication method can be a vehicle.

[0035] The in-vehicle communication method of the present application will be described below taking the execution subject of the in-vehicle communication method as a vehicle. It should be noted that the above execution subject and application scenario do not constitute a limitation on the present application.

[0036] As shown in Figure 1 The in-vehicle communication method provided by the embodiments of the present application can include steps 110-140.

[0037] Step 110, in the case of obtaining an in-vehicle communication signal, performing encoding processing on the in-vehicle communication signal to obtain the frequency characteristics of the in-vehicle communication signal, wherein the in-vehicle communication signal includes various interference sound components and an in-vehicle clean signal;

[0038] Step 120, performing time-frequency analysis on the frequency characteristics to obtain the time-frequency characteristics of the in-vehicle communication signal, wherein the time-frequency characteristics are used to represent the distribution of various interference sound components in the in-vehicle communication signal;

[0039] Step 130, the time-frequency feature is decoded to determine the mask value of each type of interference sound component in the in-vehicle communication signal at each time-frequency point, and the mask estimation of the in-vehicle communication signal is obtained.

[0040] Step 140, based on the mask estimation, the in-vehicle communication signal is masked to eliminate each type of interference sound component in the in-vehicle communication signal, and the clean in-vehicle signal is obtained.

[0041] The in-vehicle communication method provided by the embodiment of the application uniformly and simultaneously removes each type of interference sound component including media echo, feedback echo and noise. Specifically, by performing encoding processing on the in-vehicle communication signal and time-frequency analysis on the encoded frequency feature, the time-frequency feature of the in-vehicle communication signal can be obtained. The time-frequency feature can reflect the time sequence information and frequency information of the in-vehicle communication signal. Therefore, based on the time-frequency feature, the distribution of each type of interference sound component in the in-vehicle communication signal can be accurately expressed, that is, the relationship between each type of interference sound component in the in-vehicle communication signal and the clean in-vehicle signal in the feature domain is determined. The relationship is the key to eliminating each type of interference sound component in the in-vehicle communication signal and retaining the clean in-vehicle signal. Therefore, by decoding the time-frequency feature to determine the mask value of each type of interference sound component in the in-vehicle communication signal at each time-frequency point, the mask estimation can be obtained. The mask estimation is used to eliminate each type of interference sound component in the in-vehicle communication signal, rather than a certain type of interference sound component. Therefore, based on the mask estimation, the masking processing on the in-vehicle communication signal can eliminate all interference sound components from the clean in-vehicle signal at one time, without performing a corresponding interference sound elimination step for each type of interference sound component. Thus, the error accumulation caused by serially performing multiple interference sound processing steps can be avoided, the interference sound removal effect of the in-vehicle communication signal is improved, and the sound quality of the in-vehicle communication is effectively ensured.

[0042] The specific implementation of the above steps will be described in detail below in conjunction with specific embodiments.

[0043] Step 110 is related to, after obtaining the in-vehicle communication signal, the in-vehicle communication signal is encoded to obtain the frequency feature of the in-vehicle communication signal.

[0044] In step 110, the in-vehicle communication signal is a single frame of original signal without signal processing. Therefore, the in-vehicle communication signal contains each type of interference sound component in addition to the clean in-vehicle signal, specifically including echo component and noise component. The echo component can be divided into media echo and feedback echo. The media echo is, for example, the music sound played in the vehicle, and the feedback echo is, for example, the amplified voice of the in-vehicle communication. The noise component is, for example, environmental background noise, air conditioner working sound, friction and collision sound, etc. Therefore, each type of interference sound component can include media echo, feedback echo and noise. The frequency feature is the characteristic information of the interference sound component in the frequency.

[0045] The application can encode the in-vehicle conversation signal through a target encoder to obtain frequency characteristics of the in-vehicle conversation signal. The target encoder is a device that compiles and converts a channel compressed signal into a signal form available for communication, transmission and storage. The target encoder can extract high-dimensional representation information of the in-vehicle conversation signal to obtain the frequency characteristics.

[0046] In some embodiments of the application, when the in-vehicle conversation signal is obtained, the step 110 can specifically include: extracting features of the in-vehicle conversation signal through a target encoder to obtain a first sub-frequency characteristic and M second sub-frequency characteristics corresponding to M interference sound components.

[0047] wherein M is a positive integer, and the frequency characteristics include the first sub-frequency characteristic and the M second sub-frequency characteristics.

[0048] In some embodiments of the application, the target encoder can include a first encoder and M second encoders. The step of extracting features of the in-vehicle conversation signal through the target encoder to obtain the first sub-frequency characteristic and the M second sub-frequency characteristics corresponding to the M interference sound components can specifically include the following steps:

[0049] extracting features of the in-vehicle conversation signal through the first encoder to obtain the first sub-frequency characteristic;

[0050] extracting features of the in-vehicle conversation signal through the M second encoders corresponding to the M interference sound components to obtain the M second sub-frequency characteristics corresponding to the M interference sound components.

[0051] wherein the second encoder is a preset autoencoder for the interference sound component, and each type of interference sound component has an associated matched second encoder. The second encoder can identify the corresponding interference sound component from the in-vehicle conversation signal and extract the frequency characteristics of the corresponding interference sound component.

[0052] In the embodiments of the application, considering that different types of interference sound have different characteristics, for example, the characteristic of media echo is relatively stable, and the characteristic of feedback echo is unstable and short-time signal intensity is high, therefore, the M second encoders are matched for the M interference sound components. Extracting features of the in-vehicle conversation signal through the M second encoders related to the M interference sound components can accurately identify the M interference sound components from the in-vehicle conversation signal, realize signal separation of the M interference sound components in the in-vehicle conversation signal, accurately extract the frequency characteristics of each type of interference sound component, and ensure the correlation degree of each frequency characteristic and each type of interference sound component. In this way, based on the frequency characteristics, time-frequency analysis is performed to obtain time-frequency characteristics that can accurately represent each type of interference sound component, which can more easily remove each type of interference sound component and improve the elimination ability of the mask estimation for each type of interference sound component.

[0053] Traditional filtering methods are based on the assumption of steady-state random process, and can only better handle steady-state noise, and cannot well model and eliminate non-steady-state noise.

[0054] In some embodiments of the present application, in order to effectively identify the non-steady-state noise component, Figure 2 is a flowchart of a car-to-car communication method provided by another embodiment of the present application, and the method can further include Figure 2 The steps 210-240 are shown.

[0055] Step 210, obtaining media echo data, feedback echo data, noise data and clean speech data;

[0056] Step 220, mixing the media echo data, the feedback echo data, the noise data and the clean speech data using a preset signal-to-noise ratio, a preset signal-to-feedback ratio and a preset signal-to-interference ratio to obtain mixed data;

[0057] Step 230, determining a training mask estimate based on the mixed data and the clean speech data;

[0058] Step 240, training a preset neural network model using the mixed data as input training data and the training mask estimate as output training data to obtain an interference sound removal model;

[0059] The input of the interference sound removal model is the car-to-car communication signal, and the output is the mask estimate of the car-to-car communication signal; the training mask estimate multiplied by the mixed data can obtain the clean speech data, so the training mask estimate can be determined in the case that the mixed data and the clean speech data are known.

[0060] In the embodiments of the present application, by training the preset neural network model using the mixed data as input training data and the training mask estimate as output training data, the preset neural network model can learn the separation and extraction ability of various interference sound components, and obtain the interference sound removal model. The interference sound removal model can use the time sequence and frequency feature learning ability of the neural network, can effectively identify the non-steady-state noise component, and eliminate it through time-frequency masking, effectively eliminate the non-steady-state noise component in the car-to-car communication signal, and improve the signal quality of the clean signal in the car. Using the method of deep neural network, the traditional adaptive filter is avoided in the state of difficult convergence in transient noise and nonlinear echo scene, poor echo cancellation ability and poor feedback elimination ability, which leads to howling of the car-to-car communication system, and the car-to-car communication system based on neural network has better robustness and can ensure the sound quality of the communication in different scenes.

[0061] In some embodiments of the present application, the interference sound removal model comprises a first encoder, that is, the first encoder can be synchronously trained and optimized while training the interference sound removal model.

[0062] In some embodiments of the present application, before step 110, the method can further comprise the following steps: obtaining a far-end reference signal and an in-vehicle microphone signal; splicing and compressing the far-end reference signal and the in-vehicle microphone signal in the channel dimension to obtain an in-vehicle talk signal.

[0063] Specifically, the far-end reference signal is a signal played from a receiver or a loudspeaker spk and received by a microphone after spatial propagation, and the near-end speech signal also enters the microphone to obtain an in-vehicle microphone signal, so that the microphone receives the superposition of the far-end reference signal and the in-vehicle microphone signal.

[0064] Regarding step 120, time-frequency analysis is performed on the frequency characteristics to obtain time-frequency characteristics of the in-vehicle talk signal.

[0065] In step 120, time-frequency analysis is used to analyze the relationship between various interference sound components and the in-vehicle talk signal in the feature domain, so the time-frequency characteristics can represent the distribution of various interference sound components in the in-vehicle talk signal in the feature domain dimension.

[0066] In some embodiments of the present application, the above step 120 of performing time-frequency analysis on the frequency characteristics to obtain time-frequency characteristics of the in-vehicle talk signal can specifically comprise:

[0067] By analyzing the relationship between various interference sound components and the in-vehicle talk signal in the feature domain, the distribution of various interference sound components at each time-frequency point of the in-vehicle talk signal is determined, and the time-frequency characteristics of the in-vehicle talk signal are obtained.

[0068] Specifically, time-frequency analysis is time-frequency joint domain analysis, which is used to analyze time-varying non-stationary signals. Time-frequency analysis methods provide joint distribution information of the time domain and the frequency domain, and clearly describe the relationship between the signal frequency and time. The basic idea of time-frequency analysis is to design a joint function of time and frequency to describe the energy density or intensity of the signal at different times and frequencies. This joint function of time and frequency is simply referred to as time-frequency distribution. Using time-frequency distribution to analyze signals can give the instantaneous frequency and its amplitude at each time, and can perform time-frequency filtering and time-varying signal research. Through time-frequency analysis, time-series information and frequency information of the frequency characteristics can be associated, and time-frequency characteristics are obtained.

[0069] In some embodiments of the present application, the frequency features can be subjected to time-frequency analysis by using a time modeling module in a recurrent neural network (RNN) or a temporal convolutional network (TCN) to obtain time-frequency features of the in-vehicle conversation signal.

[0070] In the embodiments of the present application, the recurrent neural network has memory, parameter sharing and Turing completeness, so it has certain advantages in learning the nonlinear characteristics of sequences. Due to the existence of recurrent connections, RNN can well capture the long-term dependencies in sequences. It can also pass context information by keeping past information in hidden states, and can process variable-length sequence data to adapt to changes in input sequence length. Due to the cyclical nature of RNN, it is naturally suitable for processing time series data and can naturally handle time.

[0071] Regarding step 130, the time-frequency features are decoded to determine the mask values of various interference sound components in the in-vehicle conversation signal at each time-frequency point, and the mask estimation of the in-vehicle conversation signal is obtained.

[0072] In step 130, the time-frequency features can be converted from the feature domain to the time-frequency domain by the decoder, so that the distribution of various interference sound components in the in-vehicle conversation signal at each time-frequency point can be obtained, and then the mask estimation can be obtained based on the distribution.

[0073] Regarding step 140, the in-vehicle conversation signal is masked based on the mask estimation to eliminate various interference sound components in the in-vehicle conversation signal, and the in-vehicle clean signal is obtained.

[0074] In some embodiments of the present application, the mask estimation can include mask values of various interference sound components in the in-vehicle conversation signal at each time-frequency point, and the above step 140 can mask the in-vehicle conversation signal based on the mask estimation to eliminate various interference sound components in the in-vehicle conversation signal, and obtain the in-vehicle clean signal, which can specifically include:

[0075] The in-vehicle conversation signal is multiplied by the mask estimation to obtain the in-vehicle clean signal.

[0076] In some embodiments of the present application, after the above step 140 masks the in-vehicle conversation signal based on the mask estimation to eliminate various interference sound components in the in-vehicle conversation signal, and obtains the in-vehicle clean signal, the method can further include the following steps:

[0077] A post-filtering algorithm is used to repair the distorted part of the in-vehicle clean signal to obtain the target clean signal.

[0078] Specifically, any one of the post-filtering algorithms in the Wiener filtering method, the spectral subtraction method, the recurrent neural network algorithm and the convolutional neural network algorithm can be adopted to repair the distorted part of the clean in-vehicle signal.

[0079] In the embodiments of the present application, after the echo and feedback elimination of the in-vehicle conversation signal in steps 110-140, part of the clean speech signal is eliminated, causing distortion of the clean in-vehicle signal. Therefore, through post-filtering processing, the distorted signal can be corrected, and the residual echo and feedback in the clean in-vehicle signal can be filtered out, thereby improving the signal quality of the target clean signal.

[0080] In some embodiments of the present application, the in-vehicle conversation signal is a multi-channel microphone signal. After the masking processing of the in-vehicle conversation signal based on the mask estimation in step 140, the method can further include the following steps:

[0081] The clean in-vehicle signal is separated into channels to obtain a single-channel clean in-vehicle signal under each channel.

[0082] The multi-channel microphone signal is a mixed microphone signal collected by multiple microphones in the vehicle under multiple channels.

[0083] For example, the interior of the vehicle includes multiple rows of seats. In the scenario of in-vehicle communication between the passengers in the first row and the third row, the front-row microphone signal of the passengers in the first row is collected by the microphones in the first row, and the rear-row microphone signal of the passengers in the third row is collected by the microphones in the third row. The front-row microphone signal and the rear-row microphone signal are mixed to obtain a multi-channel microphone signal. After the media echo, feedback echo and noise elimination of the multi-channel microphone signal in steps 110-130, a clean in-vehicle signal can be obtained. The clean in-vehicle signal contains two channels (the first-row microphone and the third-row microphone). Based on this, the clean in-vehicle signal is separated into a clean microphone signal under a single channel to obtain the clean microphone signals of the first row and the third row. The clean microphone signal of the third row is played through the loudspeaker of the first row, and the clean microphone signal of the first row is played through the loudspeaker of the third row.

[0084] It can be understood that the in-vehicle conversation method provided in the embodiments of the present application can be executed by a terminal or a control module in the terminal for executing the in-vehicle conversation method. The in-vehicle conversation device will be described in detail below.

[0085] Figure 3 is a structural schematic diagram of an in-vehicle conversation device provided by the embodiments of the present application. As shown in Figure 3As shown, the in-vehicle intercom device 300 can include an encoding module 310, a time-frequency analysis module 320, a decoding module 330, and a masking module 340.

[0086] The encoding module 410 is configured to, in a case where an in-vehicle intercom signal is acquired, perform encoding processing on the in-vehicle intercom signal to obtain a frequency feature of the in-vehicle intercom signal, wherein the in-vehicle intercom signal includes various interference sound components and a clean in-vehicle signal.

[0087] The in-vehicle intercom device provided in the present application uniformly and simultaneously removes various interference sound components including media echo, feedback echo, and noise, specifically, by performing encoding processing on the in-vehicle intercom signal and performing time-frequency analysis on the encoded frequency feature, the time-frequency feature of the in-vehicle intercom signal can be obtained, which can reflect the time sequence information and frequency information of the in-vehicle intercom signal, and thus the distribution of various interference sound components in the in-vehicle intercom signal can be accurately expressed based on the time-frequency feature, that is, the relationship between the various interference sound components and the clean in-vehicle signal in the feature domain in the in-vehicle intercom signal is clear, which is the key to eliminating the various interference sound components in the in-vehicle intercom signal and retaining the clean in-vehicle signal. Therefore, by performing decoding processing on the time-frequency feature to determine the mask value of each type of interference sound component at each time-frequency point, the mask estimate can be obtained, which is used to eliminate the various interference sound components in the in-vehicle intercom signal, rather than a certain interference sound component, and thus based on the mask estimate, all interference sound components in the clean in-vehicle signal can be eliminated at one time, without performing a corresponding interference sound elimination step for each type of interference sound component, thereby avoiding the error accumulation caused by serially performing multiple interference sound processing steps, improving the interference sound removal effect of the in-vehicle intercom signal, and effectively ensuring the sound quality of the in-vehicle intercom.

[0088] In some embodiments of the present application, the encoding module 310 is specifically configured to perform feature extraction on the in-vehicle intercom signal by a target encoder to obtain a first sub-frequency feature and M second sub-frequency features corresponding to M types of interference sound components, and the frequency feature includes the first sub-frequency feature and the M second sub-frequency features.

[0089] In some implementable manners of the second aspect, the target encoder comprises a first encoder and M second encoders, and the encoding module 310 is specifically configured to: perform feature extraction on the in-vehicle conversation signal by the first encoder to obtain a first sub-frequency feature; and perform feature extraction on the in-vehicle conversation signal by the M second encoders corresponding to the M types of interference sound components to obtain M second sub-frequency features corresponding to the M types of interference sound components, wherein the second encoders are preset autoencoders for the interference sound components.

[0090] In some implementable manners of the second aspect, the masking module 340 is specifically configured to: multiply the in-vehicle conversation signal by the mask estimate to obtain a clean in-vehicle signal.

[0091] In some implementable manners of the second aspect, the time-frequency analysis module 320 is specifically configured to: determine the distribution of the M types of interference sound components on each time-frequency point of the in-vehicle conversation signal by analyzing the relationship between the M types of interference sound components and the in-vehicle conversation signal in the feature domain to obtain a time-frequency feature of the in-vehicle conversation signal.

[0092] In some implementable manners of the second aspect, the apparatus further comprises: an acquisition module configured to acquire media echo data, feedback echo data, noise data and clean speech data before encoding processing of the in-vehicle conversation signal; a data mixing module configured to mix the media echo data, the feedback echo data, the noise data and the clean speech data by using a preset signal-to-noise ratio, a preset signal-to-feedback ratio and a preset signal-to-interference ratio to obtain mixed data; a determination module configured to determine a training mask estimate based on the mixed data and the clean speech data; and a training module configured to train a preset neural network model by taking the mixed data as input training data and taking the training mask estimate as output training data to obtain an interference sound removal model, wherein the input of the interference sound removal model is the in-vehicle conversation signal, and the output is a mask estimate of the in-vehicle conversation signal.

[0093] In some implementable manners of the second aspect, the apparatus further comprises: a post-filtering module configured to, after the masking processing of the in-vehicle conversation signal based on the mask estimate to eliminate the M types of interference sound components in the in-vehicle conversation signal to obtain a clean in-vehicle signal, repair a distorted part of the clean in-vehicle signal by using a post-filtering algorithm to obtain a target clean signal.

[0094] In some implementable manners of the second aspect, the in-vehicle conversation signal is a multi-channel microphone signal, and the apparatus further comprises: a channel separation module configured to, after the masking processing of the in-vehicle conversation signal based on the mask estimate to eliminate the M types of interference sound components in the in-vehicle conversation signal to obtain a clean in-vehicle signal, perform channel separation on the clean in-vehicle signal to obtain a single-channel clean in-vehicle signal under each channel.

[0095] The in-vehicle conversation apparatus provided by the embodiments of the present application canFigures 1-2 The method embodiments of the method are implemented by the electronic device, and the same technical effects can be achieved. To avoid repetition, details are not described here.

[0096] Figure 4 is a hardware structure schematic diagram of an electronic device provided by an embodiment of the present application.

[0097] As Figure 4 shown, the electronic device 400 includes a memory 401, a processor 402, and a computer program stored on the memory 401 and executable on the processor 402.

[0098] In one example, the processor 402 described above can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or can be configured to implement one or more integrated circuits of the embodiments of the present application.

[0099] The memory 401 can include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical, or other physical / tangible memory storage devices. Therefore, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions and, when the software is executed (e.g., by one or more processors), is operable to perform the operations described with reference to the card opening method in the embodiments according to the first aspect of the present application.

[0100] The processor 402 runs the computer program corresponding to the executable program code stored in the memory 401 by reading the executable program code, for implementing the card opening method in the embodiments of the first aspect described above.

[0101] In some examples, the electronic device 400 can also include a communication interface 403 and a bus 404. As Figure 4 shown, the memory 401, the processor 402, and the communication interface 403 are connected through the bus 404 and complete communication with each other.

[0102] The communication interface 403 is mainly used to realize the communication between the modules, devices, units, and / or devices in the embodiments of the present application. Input devices and / or output devices can also be accessed through the communication interface 403.

[0103] Bus 404 includes a hardware, software, or both that couples components of electronic device 400 to each other. As an example and not by way of limitation, bus 404 can include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand (IB) interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-E) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or some other bus or combination of buses. Bus 404 can include one or more buses, where appropriate. Although this application describes and illustrates a particular bus, this application contemplates any suitable bus or interconnect.

[0104] The electronic device provided by the embodiments of the present application can realize Figures 1-2 The method embodiments of the present application can realize the various processes realized by the electronic device, and can realize the same technical effects. To avoid repetition, the details are not described here.

[0105] In combination with the in-vehicle intercom method in the above embodiments, the embodiments of the present application can provide a computer storage medium to realize. The computer storage medium has computer program instructions stored thereon; the computer program instructions are executed by a processor to realize the steps of any one of the in-vehicle intercom methods in the above embodiments.

[0106] In combination with the in-vehicle intercom method in the above embodiments, the embodiments of the present application can provide a computer program product to realize. The (computer) program product is stored in a non-volatile storage medium, and the program product is executed by at least one processor to realize the steps of any one of the in-vehicle intercom methods in the above embodiments.

[0107] The embodiment of the present application further provides a chip, which comprises a processor and a communication interface, the communication interface is coupled with the processor, the processor is used for running programs or instructions, realizes each process of the in-vehicle communication method embodiment, and can achieve the same technical effects. To avoid repetition, details are not described here.

[0108] It should be understood that the chip mentioned in the embodiment of the present application can also be referred to as a system chip, a system chip, a chip system or a system on chip, etc.

[0109] It should be noted that the present application is not limited to the specific configurations and processes described above and shown in the drawings. For the sake of brevity, detailed descriptions of well-known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method processes of the present application are not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between steps, after understanding the spirit of the present application.

[0110] The functional blocks shown in the structural block diagram described above can be implemented as hardware, software, firmware or their combination. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are program or code segment used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted on a transmission medium or communication link through a data signal carried in a carrier wave. The "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable medium include electronic circuit, semiconductor memory device, ROM, flash memory, erasable ROM (EROM), floppy disk, CD-ROM, optical disk, hard disk, optical fiber medium, radio frequency (RF) link, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0111] It should be further noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or in an order different from the embodiments, or several steps can be executed simultaneously.

[0112] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other processing device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other processing device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0113] The above solely describes specific implementations of the present application. For the purpose of description and brevity, the specific working process of the system, module and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described herein. It should be understood that the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed in the present application, and these modifications or replacements shall be covered within the protection scope of the present application.

Claims

1. An intercom method characterized by comprising: The method comprises: When an in-car call signal is obtained, encoding the in-car call signal to obtain a frequency characteristic of the in-car call signal, wherein the in-car call signal includes various interference sound components and a clean in-car signal; Performing time-frequency analysis on the frequency features to obtain time-frequency features of the in-car call signal, wherein the time-frequency features are used to characterize the distribution of the various types of interference sound components in the in-car call signal; Decoding the time-frequency features to determine mask values ​​of various types of interfering sound components in the in-car call signal at various time-frequency points, thereby obtaining a mask estimate of the in-car call signal; The in-car call signal is masked based on the mask estimation to eliminate the various interference sound components in the in-car call signal to obtain a clean in-car signal.

2. The method according to claim 1, characterized in that When the in-car call signal is acquired, encoding the in-car call signal to obtain a frequency characteristic of the in-car call signal includes: Performing feature extraction on the in-car call signal using a target encoder to obtain a first sub-frequency feature and M second sub-frequency features corresponding to M types of interference sound components; The frequency feature includes the first sub-frequency feature and the M second sub-frequency features.

3. The method of claim 2, wherein, The target encoder includes a first encoder and M second encoders. The target encoder is used to extract features from the in-car call signal to obtain a first sub-frequency feature and M second sub-frequency features corresponding to M types of interference sound components, including: performing feature extraction on the in-car call signal by the first encoder to obtain the first sub-frequency feature; The in-car call signal is feature extracted by using M second encoders corresponding to the M types of interference sound components to obtain M second sub-frequency features corresponding to the M types of interference sound components, wherein the second encoder is a preset autoencoder for the interference sound components.

4. The method of claim 1, wherein, The masking process is performed on the in-car call signal based on the mask estimation to eliminate the various interference sound components in the in-car call signal to obtain a clean in-car signal, including: The in-car call signal is multiplied by the mask estimate to obtain the in-car clean signal.

5. The method of claim 1, wherein, The performing time-frequency analysis on the frequency characteristics to obtain the time-frequency characteristics of the in-car call signal includes: By analyzing the relationship between the various interference sound components and the in-car call signal in the feature domain, the distribution of the various interference sound components at each time-frequency point of the in-car call signal is determined, and the time-frequency characteristics of the in-car call signal are obtained.

6. The method of claim 1, wherein, Before encoding the in-car call signal, the method further includes: Acquire media echo data, feedback echo data, noise data, and clean voice data; Mixing the media echo data, the feedback echo data, the noise data, and the clean voice data using a preset signal-to-noise ratio, a preset signal-to-return ratio, and a preset signal-to-interference ratio to obtain mixed data; determining a training mask estimate based on the mixed data and the clean speech data; The mixed data is taken as input training data, and the training mask estimation is taken as output training data, and a preset neural network model is trained to obtain an interference sound removal model. The input of the interference sound removal model is the in-vehicle talk signal, and the output is a mask estimation of the in-vehicle talk signal.

7. The method of claim 1, wherein, After the in-vehicle clean signal is obtained by performing the masking processing on the in-vehicle talk signal based on the mask estimation and eliminating the various interference sound components in the in-vehicle talk signal, the method further comprises: A post-filtering algorithm is used to repair the distorted part of the in-vehicle clean signal to obtain a target clean signal.

8. The method of claim 1, wherein, The in-vehicle talk signal is a multi-channel microphone signal, and after the in-vehicle clean signal is obtained by performing the masking processing on the in-vehicle talk signal based on the mask estimation and eliminating the various interference sound components in the in-vehicle talk signal, the method further comprises: The in-vehicle clean signal is channel-separated to obtain a single-channel in-vehicle clean signal under each channel.

9. An intercom device, characterized by comprising: The device comprises: An encoding module configured to, in a case where an in-vehicle talk signal is acquired, perform encoding processing on the in-vehicle talk signal to obtain a frequency feature of the in-vehicle talk signal, wherein the in-vehicle talk signal comprises various interference sound components and an in-vehicle clean signal; A time-frequency analysis module configured to perform time-frequency analysis on the frequency feature to obtain a time-frequency feature of the in-vehicle talk signal, wherein the time-frequency feature is used to represent a distribution of the various interference sound components in the in-vehicle talk signal; A decoding module configured to perform decoding processing on the time-frequency feature to determine a mask value of each type of interference sound component in the in-vehicle talk signal at each time-frequency point, and obtain a mask estimation of the in-vehicle talk signal; A masking module configured to perform masking processing on the in-vehicle talk signal based on the mask estimation to eliminate the various interference sound components in the in-vehicle talk signal, and obtain an in-vehicle clean signal.

10. An electronic device, comprising: The device comprises a processor and a memory having computer program instructions stored therein; The processor, when executing the computer program instructions, implements the steps of the in-vehicle talk method according to any one of claims 1-8.

11. A computer readable storage medium, characterized in that, The computer program instructions are stored on the computer readable storage medium, and when executed by the processor, implement the steps of the in-vehicle talk method according to any one of claims 1-8.