Echo cancellation method, device, equipment, medium and vehicle

By using a deep learning model to eliminate echoes in the vehicle cabin, the problems of incomplete echo cancellation and high computational resource consumption are solved, achieving efficient and low-cost echo cancellation.

CN120877751APending Publication Date: 2025-10-31BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410544205.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing echo cancellation methods cannot completely eliminate echoes in vehicle cabins, and they consume high computational resources and have high development costs.

Method used

By acquiring the mixed audio signal and reference signal received by the microphone in the vehicle cabin, echo cancellation is performed using deep learning models such as recurrent neural networks or convolutional neural networks, reducing the computational resource requirements.

Benefits of technology

It improves the accuracy of echo cancellation, reduces echo residue, and saves computing resources and development costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877751A_ABST
    Figure CN120877751A_ABST
Patent Text Reader

Abstract

The invention relates to an echo cancellation method, device and equipment, a medium and a vehicle. According to the method, the mixed audio signal received by the microphone in the vehicle cabin and the reference signal corresponding to the mixed audio signal are acquired, the mixed audio signal comprises the echo signal and the human voice signal sent by the speaker, and the reference signal is the audio signal played by the loudspeaker in the cabin controlled by the vehicle equipment; the echo signal is a reference signal received by the microphone; the mixed audio signal and the reference signal are input into the preset target audio processing model, the echo signal in the mixed audio signal is eliminated based on the target audio processing model, the target audio signal corresponding to the mixed audio signal is obtained, and the target audio processing model can run only by consuming few computing resources; therefore, computing resources can be saved, the development cost can be reduced, the target audio processing model can accurately identify echoes, the accuracy of echo cancellation is improved, echo residues are reduced, and the effect of echo cancellation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to an echo cancellation method, apparatus, device, medium, and vehicle. Background Technology

[0002] Inside the vehicle cabin, users can interact with the in-vehicle infotainment system and control it to execute corresponding commands. In in-vehicle voice interaction scenarios, the microphone receives audio signals emitted by the speakers, i.e., echoes.

[0003] In related technologies, a common solution for echo removal is to use algorithms such as adaptive filtering to eliminate echoes. However, these methods are only based on signal estimation and cannot eliminate echoes to the greatest extent. There are still many echoes left. Moreover, the signal processing-based algorithms require a lot of computing resources, have a low performance ceiling, and have high development costs. Summary of the Invention

[0004] To address the aforementioned technical problems, this disclosure provides an echo cancellation method, apparatus, device, medium, and vehicle.

[0005] The first aspect of this disclosure provides an echo cancellation method, comprising:

[0006] Acquire the mixed audio signal received by the microphone in the vehicle cabin and acquire the reference signal corresponding to the mixed audio signal. The mixed audio signal includes an echo signal and a human voice signal emitted by the speaker. The reference signal is the audio signal played by the speakers in the cabin controlled by the vehicle's infotainment system. The echo signal is the reference signal received by the microphone.

[0007] The mixed audio signal and the reference signal are input into a preset target audio processing model. Based on the target audio processing model, the echo signal in the mixed audio signal is eliminated to obtain the target audio signal corresponding to the mixed audio signal.

[0008] A second aspect of this disclosure provides an echo cancellation device, comprising:

[0009] The first acquisition module is used to acquire the mixed audio signal received by the microphone in the vehicle cabin and to acquire the reference signal corresponding to the mixed audio signal. The mixed audio signal includes an echo signal and a human voice signal emitted by the speaker. The reference signal is the audio signal played by the speaker in the cabin controlled by the vehicle equipment. The echo signal is the reference signal received by the microphone.

[0010] The echo cancellation module is used to input the mixed audio signal and the reference signal into a preset target audio processing model. Based on the target audio processing model, the echo signal in the mixed audio signal is cancelled to obtain the target audio signal corresponding to the mixed audio signal.

[0011] A third aspect of this disclosure provides an in-vehicle infotainment device, including a memory and a processor, wherein the memory stores a computer program that, when executed by the processor, can implement the echo cancellation method of the first aspect described above.

[0012] A fourth aspect of this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the echo cancellation method of the first aspect described above.

[0013] The fifth aspect of this disclosure provides a vehicle including the echo cancellation device of the second aspect and / or the vehicle-mounted equipment of the third aspect and / or the computer-readable storage medium of the fourth aspect, which can implement the echo cancellation method of the first aspect.

[0014] The technical solution provided in this disclosure has the following advantages compared with the prior art:

[0015] This disclosure involves acquiring a mixed audio signal received by a microphone in a vehicle cabin and acquiring a corresponding reference signal for the mixed audio signal. The mixed audio signal includes an echo signal and a speaker's voice signal. The reference signal is the audio signal played by the speakers in the cabin controlled by the vehicle's infotainment system, and the echo signal is the reference signal received by the microphone. The mixed audio signal and the reference signal are input into a preset target audio processing model. Based on the target audio processing model, the echo signal in the mixed audio signal is eliminated to obtain the target audio signal corresponding to the mixed audio signal. Echo cancellation can be performed on the mixed audio signal received by the microphone based on the target audio processing model. The target audio processing model only requires minimal computing resources to run, which can save computing resources and reduce development costs. Compared with traditional echo cancellation algorithms, the target audio processing model can accurately identify echoes, improve the accuracy of echo cancellation, greatly reduce echo residue, and improve the effect of echo cancellation. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0017] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of an echo cancellation method provided in an embodiment of this disclosure;

[0019] Figure 2 This is a flowchart of another echo cancellation method provided in this embodiment of the disclosure;

[0020] Figure 3 This is a flowchart of a target audio processing model training method provided in an embodiment of this disclosure;

[0021] Figure 4 This is a schematic diagram of the structure of an echo cancellation device provided in an embodiment of this disclosure;

[0022] Figure 5 This is a schematic diagram of the structure of a vehicle-mounted device provided in an embodiment of this disclosure. Detailed Implementation

[0023] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0024] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0025] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0026] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0027] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0028] Figure 1 This is a flowchart of an echo cancellation method provided in an embodiment of the present disclosure, which can be executed by an in-vehicle infotainment system installed in a vehicle cabin. This system can be understood as any electronic device with processing and computing capabilities.

[0029] like Figure 1 As shown, the echo cancellation method provided in this embodiment includes the following steps:

[0030] Step 110: Obtain the mixed audio signal received by the microphone in the vehicle cabin and obtain the reference signal corresponding to the mixed audio signal. The mixed audio signal includes an echo signal and a human voice signal emitted by the speaker. The reference signal is the audio signal played by the speakers in the cabin controlled by the vehicle's infotainment system. The echo signal is the reference signal received by the microphone.

[0031] In this embodiment of the disclosure, at least one microphone and at least one speaker are installed in the vehicle cabin. The vehicle infotainment system installed in the vehicle cabin can acquire the audio signals received by all the microphones in the vehicle cabin to obtain a mixed audio signal, and acquire the reference signal corresponding to the mixed audio signal.

[0032] Reference signals can be understood as audio signals played by speakers in the cabin controlled by the vehicle's infotainment system, such as media signals and / or response signals emitted by the speakers controlled by the vehicle's infotainment system.

[0033] The mixed audio signal can include echo signals and the speaker's voice signal. The echo signal can be understood as the reference signal received by the microphone, that is, the audio signal played by the speakers in the cabin controlled by the vehicle's infotainment system.

[0034] Human voice signal can be understood as the audio signal emitted by the speaker.

[0035] Step 120: Input the mixed audio signal and the reference signal into the preset target audio processing model. Based on the target audio processing model, eliminate the echo signal in the mixed audio signal to obtain the target audio signal corresponding to the mixed audio signal.

[0036] In this embodiment of the disclosure, after obtaining the mixed audio signal and the reference signal, the vehicle-mounted device can input the mixed audio signal and the reference signal into a preset target audio processing model. Based on the target audio processing model, the echo signal in the mixed audio signal is eliminated to obtain the target audio signal corresponding to the mixed audio signal. That is, the target audio signal does not include the echo signal and only includes the human voice signal.

[0037] The target audio processing model can be understood as a deep learning model, such as a recurrent neural network (RNN) or a convolutional neural network (CNN).

[0038] For example, when people are speaking simultaneously in the driver's area and the second-row left area of ​​the vehicle cabin, the vehicle's infotainment system controls the speakers in the cabin to emit media signals and response signals, i.e., reference signals. The mixed audio signal received by the microphone includes the voice signals from the driver's area and the second-row left area, as well as the echo signal formed by the reference signal received by the microphone. The mixed audio signal and the reference signal are input into a preset target audio processing model. Based on the target audio processing model, the echo signal in the mixed audio signal is eliminated to obtain the target audio signal corresponding to the mixed audio signal.

[0039] In this embodiment, a mixed audio signal received by a microphone in the vehicle cabin and a corresponding reference signal are obtained. The mixed audio signal includes an echo signal and a speaker's voice signal. The reference signal is the audio signal played by the speakers in the cabin controlled by the vehicle's infotainment system, and the echo signal is the reference signal received by the microphone. The mixed audio signal and the reference signal are input into a preset target audio processing model. Based on the target audio processing model, the echo signal in the mixed audio signal is eliminated to obtain the target audio signal corresponding to the mixed audio signal. Echo cancellation can be performed on the mixed audio signal received by the microphone based on the target audio processing model. The target audio processing model only requires a small amount of computing resources to run, which can save computing resources and reduce development costs. Compared with traditional echo cancellation algorithms, the target audio processing model can accurately identify echoes, improve the accuracy of echo cancellation, greatly reduce echo residue, and improve the effect of echo cancellation.

[0040] Figure 2 This is a flowchart of an echo cancellation method provided in an embodiment of the present disclosure, which can be executed by an in-vehicle infotainment system installed in a vehicle cabin. This system can be understood as any electronic device with processing and computing capabilities.

[0041] like Figure 2 As shown, the echo cancellation method provided in this embodiment includes the following steps:

[0042] Step 210: Obtain the mixed audio signal received by the microphone in the vehicle cabin and obtain the reference signal corresponding to the mixed audio signal. The mixed audio signal includes an echo signal and a human voice signal emitted by the speaker. The reference signal is the audio signal played by the speakers in the cabin controlled by the vehicle's infotainment system. The echo signal is the reference signal received by the microphone.

[0043] Step 220: Obtain the initial reference signals of the first number of channels corresponding to the mixed audio signal. The initial reference signals of the first number of channels are the audio signals played by the first number of speakers in the cockpit controlled by the vehicle's infotainment system.

[0044] In this embodiment of the disclosure, the vehicle-mounted device can acquire initial reference signals for a first number of channels corresponding to the mixed audio signal. The initial reference signals for the first number of channels can be understood as the audio signals played by the first number of speakers in the cabin controlled by the vehicle-mounted device.

[0045] Step 230: Downmix the initial reference signals of the first number of channels to obtain reference signals of the second number of channels, where the second number is less than the first number.

[0046] In this embodiment of the present disclosure, the vehicle-mounted device can downmix the initial reference signals of a first number of channels to obtain reference signals of a second number of channels, wherein the second number is less than the first number.

[0047] For example, the first quantity is K, the second quantity is K1, and K1 <K。

[0048] This reduces the amount of data in the reference signal, decreases the amount of data processing required for subsequent models, and improves the efficiency of data processing for subsequent models.

[0049] In some embodiments, the vehicle-mounted device can identify the signal quality of the initial reference signal of each of the first number of channels; delete the initial reference signals of the channels whose signal quality is less than a preset quality threshold, and obtain the reference signals of the second number of channels.

[0050] Signal quality can be understood as a parameter for evaluating the strength of an audio signal. The stronger the audio signal, the higher the signal quality; the weaker the audio signal, the lower the signal quality.

[0051] The preset quality threshold can be set as needed; there are no restrictions here.

[0052] In other embodiments, the vehicle-mounted device can identify the signal quality of the initial reference signal of each of the first number of channels; sort the signal quality of each channel in descending order of signal quality to obtain a quality sequence; starting from the initial signal quality of the quality sequence, select the initial reference signal corresponding to the second number of signal qualities as the reference signal corresponding to the mixed audio signal.

[0053] In some other embodiments, the vehicle-mounted equipment can merge the initial reference signals of a first number of channels according to a preset channel merging rule to obtain reference signals of a second number of channels.

[0054] The preset channel merging rules can be set as needed, such as merging the initial reference signals of any two channels, etc., which are not limited here.

[0055] For example, if the number of speakers in the vehicle cabin is 6, that is, the initial number is 6, the initial reference signals of the 6 channels are combined in pairs to obtain reference signals of 3 channels.

[0056] Step 240: Input the mixed audio signal and the reference signal into the preset target audio processing model. Based on the target audio processing model, eliminate the echo signal in the mixed audio signal to obtain the target audio signal corresponding to the mixed audio signal.

[0057] Therefore, the amount of data in the reference signal can be reduced, the data processing load of the target audio processing model can be decreased, and the data processing efficiency of the target audio processing model can be improved. Echo cancellation can be performed on the mixed audio signal received by the microphone based on the target audio processing model. The target audio processing model only needs to consume very few computing resources to run, which can save computing resources and reduce development costs. Compared with traditional echo cancellation algorithms, the target audio processing model can accurately identify echoes, improve the accuracy of echo cancellation, greatly reduce echo residue, and improve the effect of echo cancellation.

[0058] In some embodiments of this disclosure, before acquiring the mixed audio signal received by the microphone in the vehicle cabin and acquiring the reference signal corresponding to the mixed audio signal, the vehicle-mounted device may perform... Figure 3 A flowchart of a target audio processing model training method is provided, such as... Figure 3 As shown, the target audio processing model training method provided in this embodiment includes the following steps:

[0059] Step 310: Obtain a preset number of training sample data and the target data corresponding to each training sample data. Each training sample data includes a target mixed audio signal received by the microphone in the vehicle cabin and a target reference signal corresponding to the target mixed audio signal. The target mixed audio signal includes a target echo signal and a target human voice signal emitted by at least one speaker. The target reference signal is an audio signal played by the speakers in the cabin controlled by the vehicle's infotainment system. The target echo signal is the target reference signal received by the microphone. The target data corresponding to the training sample data consists of the target human voice signal.

[0060] In this embodiment of the disclosure, the vehicle-mounted device can acquire a preset number of training sample data and the target data corresponding to each training sample data.

[0061] Each training sample data set includes a target mixed audio signal received by a microphone in the vehicle cabin and a target reference signal corresponding to the target mixed audio signal. The target mixed audio signal may include a target echo signal and a target human voice signal emitted by at least one speaker. The target reference signal can be understood as the audio signal played by the speakers in the cabin controlled by the vehicle's infotainment system. The target echo signal can be understood as the target reference signal received by the microphone. The target data corresponding to the training sample data consists of the target human voice signal.

[0062] The preset quantity can be set as needed; there is no limit here.

[0063] In some embodiments, obtaining a preset number of training sample data and the target data corresponding to each training sample data may include steps 3101-3108:

[0064] Step 3101: Control multiple speakers in the cockpit to play multiple clean audio signals with different audio attributes.

[0065] In this embodiment of the disclosure, the vehicle-mounted equipment can control multiple speakers in the cabin to play multiple clean audio signals with different audio attributes.

[0066] Audio properties can include sound effects, volume, and other attributes.

[0067] Step 3102: Obtain multiple clean audio signals received by the microphone in the cockpit from multiple speakers, and determine each clean audio signal received by the microphone as the initial echo signal.

[0068] Step 3103: Obtain the original audio signals emitted by multiple speakers received by the microphone in the cockpit, and determine each original audio signal received by the microphone as the initial human voice signal.

[0069] The initial voice signal may include at least one of the following: voice signals of different volumes, voice signals from speakers in different vocal ranges within the cabin, and voice signals from different numbers of speakers.

[0070] Step 3104: Combine multiple initial echo signals into a preset number of echo sets to obtain a preset number of target echo signals; combine multiple initial human voice signals into a preset number of human voice sets to obtain a preset number of target human voice signals.

[0071] Step 3105: Mix a preset number of target echo signals and a preset number of target human voice signals to obtain a preset number of target mixed audio signals.

[0072] Step 3106: For the initial echo signal contained in the target echo signal in each target mixed audio signal, determine the clean audio signal played by the vehicle-mounted device controlling the speaker corresponding to the initial echo signal as the target reference signal corresponding to the target mixed audio signal.

[0073] Step 3107: Construct a preset number of training sample data based on a preset number of target mixed audio signals and the target reference signals corresponding to the target mixed audio signals.

[0074] Step 3108: Based on the target human voice signal in the training sample data, construct the target data corresponding to the training sample data.

[0075] Step 320: Using the target data as the training target, train the audio processing model based on the training sample data and the target data to obtain the trained target audio processing model. The target audio processing model is a model that eliminates the target echo signal in the target mixed audio signal based on the target reference signal corresponding to the target mixed audio signal in the training sample data.

[0076] In this embodiment of the disclosure, the vehicle-mounted device can train the audio processing model based on each training sample data and the target data corresponding to the training sample data, using the target data as the training target. This allows the audio processing model to eliminate the target echo signal in the target mixed audio signal based on the target reference signal corresponding to the target mixed audio signal in the training sample data. After multiple rounds of training, the target loss function converges, resulting in a well-trained target audio processing model.

[0077] In some embodiments, using target data as the training target, the audio processing model is trained based on training sample data and target data to obtain a trained target audio processing model, which may include steps 3201-3203:

[0078] Step 3201 inputs the training sample data and the target data corresponding to the training sample data into the audio processing model. Based on the audio processing model, the target echo signal in the target mixed audio signal in the training sample data is eliminated to obtain the initial audio signal corresponding to the target mixed audio signal.

[0079] Step 3202: Calculate the difference between the audio signals of each microphone channel in the initial audio signal and the target data corresponding to the training sample data, and determine the maximum difference as the function value of the target loss function.

[0080] In this embodiment of the disclosure, the vehicle-mounted device can determine the maximum difference among the differences between the audio signals of each microphone channel in the initial audio signal corresponding to the target mixed audio signal as the function value of the target loss function. That is, the desired effect of the model is that the echo of each microphone channel can be completely eliminated.

[0081] Each microphone corresponds to one microphone channel. The initial audio signal includes the audio signals from each microphone channel.

[0082] Step 3203: Based on the target loss function, iteratively train the audio processing model on a preset number of training sample data and the target data corresponding to each training sample data. When the function value of the target loss function is less than the preset loss threshold, the trained target audio processing model is obtained.

[0083] The preset loss threshold can be set as needed, and is not limited here.

[0084] Therefore, a target audio processing model can be trained. Based on the target audio processing model, echo cancellation is performed on the mixed audio signal received by the microphone. The target audio processing model only requires very little computing resources to run, which can save computing resources and reduce development costs. Compared with traditional echo cancellation algorithms, the target audio processing model can accurately identify echoes, improve the accuracy of echo cancellation, greatly reduce echo residue, and improve the effect of echo cancellation.

[0085] Figure 4 This is a schematic diagram of an echo cancellation device provided in an embodiment of this disclosure. This device can be understood as the aforementioned vehicle infotainment system or a functional module within the aforementioned vehicle infotainment system. For example... Figure 4 As shown, the echo cancellation device 400 includes:

[0086] The first acquisition module 410 is used to acquire the mixed audio signal received by the microphone in the vehicle cabin and to acquire the reference signal corresponding to the mixed audio signal. The mixed audio signal includes an echo signal and a human voice signal emitted by the speaker. The reference signal is the audio signal played by the speaker in the cabin controlled by the vehicle equipment. The echo signal is the reference signal received by the microphone.

[0087] The echo cancellation module 420 is used to input the mixed audio signal and the reference signal into a preset target audio processing model, and based on the target audio processing model, cancel the echo signal in the mixed audio signal to obtain the target audio signal corresponding to the mixed audio signal.

[0088] Optionally, the first acquisition module mentioned above includes:

[0089] The acquisition submodule is used to acquire the initial reference signals of the first number of channels corresponding to the mixed audio signal. The initial reference signals are the audio signals played by the first number of speakers in the cockpit controlled by the vehicle equipment.

[0090] The downmixing submodule is used to downmix the initial reference signals of a first number of channels to obtain reference signals of a second number of channels, where the second number is less than the first number.

[0091] Optionally, the above-mentioned downmixing submodule includes:

[0092] The identification unit is used to identify the signal quality of the reference signal in each of the first number of channels;

[0093] The deletion unit is used to delete the initial reference signals of channels whose signal quality is less than a preset quality threshold, so as to obtain the target reference signals of a second number of channels.

[0094] Alternatively, a sorting unit is used to sort the signal quality of each channel in descending order of signal quality to obtain a quality sequence;

[0095] The selection unit is used to select an initial reference signal corresponding to a second number of signal qualities, starting from the initial signal quality of the quality sequence, as the reference signal corresponding to the mixed audio signal;

[0096] The merging unit is used to merge the initial reference signals of a first number of channels according to a preset channel merging rule to obtain reference signals of a second number of channels.

[0097] Alternatively, the first calculation unit is used to calculate and sum the initial reference signals of a first number of channels to obtain the total reference signal;

[0098] The second calculation unit is used to calculate the ratio of the total reference signal to the second number of channels to obtain the reference signal for the second number of channels.

[0099] Optionally, the above echo cancellation device includes:

[0100] The second acquisition module is used to acquire a preset number of training sample data and target data corresponding to each training sample data. Each training sample data includes a target mixed audio signal received by a microphone in the vehicle cabin and a target reference signal corresponding to the target mixed audio signal. The target mixed audio signal includes a target echo signal and a target human voice signal emitted by at least one speaker. The target reference signal is an audio signal played by a speaker in the cabin controlled by the vehicle's infotainment system. The target echo signal is the target reference signal received by the microphone. The target data corresponding to the training sample data consists of the target human voice signal.

[0101] The training module is used to train the audio processing model based on the training sample data and the target data, using the target data as the training target, to obtain the trained target audio processing model. The target audio processing model is a model that eliminates the target echo signal in the target mixed audio signal based on the target reference signal corresponding to the target mixed audio signal in the training sample data.

[0102] Optionally, the second acquisition module mentioned above includes:

[0103] The control submodule is used to control multiple speakers in the cockpit to play clean audio signals with different audio properties;

[0104] The first acquisition submodule is used to acquire multiple clean audio signals received by the microphone in the cockpit from multiple speakers, and to determine each clean audio signal received by the microphone as the initial echo signal.

[0105] The second acquisition submodule is used to acquire the original audio signals emitted by multiple speakers received by the microphone in the cockpit, and to determine each original audio signal received by the microphone as the initial human voice signal.

[0106] The combination submodule is used to combine multiple initial echo signals into a preset number of echo sets to obtain a preset number of target echo signals, and to combine multiple initial human voice signals into a preset number of human voice sets to obtain a preset number of target human voice signals.

[0107] The mixing submodule is used to mix a preset number of target echo signals and a preset number of target human voice signals to obtain a preset number of target mixed audio signals;

[0108] The determination submodule is used to determine the clean audio signal played by the vehicle-mounted device controlled speaker corresponding to the initial echo signal contained in the target echo signal in each target mixed audio signal as the target reference signal corresponding to the target mixed audio signal.

[0109] The first construction submodule is used to construct a preset number of training sample data based on a preset number of target mixed audio signals and the target reference signals corresponding to the target mixed audio signals.

[0110] The second construction submodule is used to construct the target data corresponding to the training sample data based on the target human voice signal in the training sample data.

[0111] Optionally, the above training module includes:

[0112] The echo cancellation submodule is used to input the training sample data and the target data corresponding to the training sample data into the audio processing model. Based on the audio processing model, the target echo signal in the target mixed audio signal in the training sample data is cancelled to obtain the initial audio signal corresponding to the target mixed audio signal.

[0113] The calculation submodule is used to calculate the difference between the audio signals of each microphone channel in the initial audio signal and the target data corresponding to the training sample data, and to determine the maximum difference as the function value of the target loss function.

[0114] The training submodule is used to iteratively train the audio processing model on a preset number of training sample data and the target data corresponding to each training sample data based on the target loss function. When the function value of the target loss function is less than the preset loss threshold, the trained target audio processing model is obtained.

[0115] The echo cancellation device provided in this disclosure can implement the method of any of the above embodiments, and its execution method and beneficial effects are similar, so they will not be described again here.

[0116] This disclosure also provides an in-vehicle infotainment device, which includes a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, it can implement the methods of any of the above embodiments. The execution method and beneficial effects are similar and will not be described again here.

[0117] Figure 5 This is a schematic diagram of the structure of a vehicle infotainment device provided in an embodiment of this disclosure, as shown below. Figure 5 As shown, the vehicle-mounted device 500 may include a processor 510 and a memory 520. The memory 520 stores a computer program 521. When the computer program 521 is executed by the processor 510, it can implement the method provided in any of the above embodiments. The execution method and beneficial effects are similar and will not be described again here.

[0118] Of course, for the sake of simplicity, Figure 5 Only some of the components of the in-vehicle infotainment system 500 relevant to the present invention are shown in this illustration; components such as buses, input / output interfaces, input devices, and output devices are omitted. In addition, the in-vehicle infotainment system 500 may include any other suitable components depending on the specific application.

[0119] This disclosure provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the methods of any of the above embodiments. The execution method and beneficial effects are similar, and will not be described again here.

[0120] The aforementioned computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0121] The computer program described above can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer device, partially on the user's device, as a standalone software package, partially on the user's computer device and partially on a remote computer device, or entirely on a remote computer device or server.

[0122] This disclosure provides a vehicle that may include the echo cancellation device described above and / or the vehicle infotainment system described above and / or the computer-readable storage medium described above, which can implement the methods of any of the above embodiments. The execution mode and beneficial effects are similar, and will not be described again here.

[0123] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0124] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0125] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An echo cancellation method, characterized in that, include: Acquire a mixed audio signal received by a microphone in the vehicle cabin and acquire a reference signal corresponding to the mixed audio signal. The mixed audio signal includes an echo signal and a human voice signal emitted by a speaker. The reference signal is an audio signal played by a speaker in the cabin controlled by a vehicle-mounted device. The echo signal is the reference signal received by the microphone. The mixed audio signal and the reference signal are input into a preset target audio processing model. Based on the target audio processing model, the echo signal in the mixed audio signal is eliminated to obtain the target audio signal corresponding to the mixed audio signal.

2. The method according to claim 1, characterized in that, The step of obtaining the reference signal corresponding to the mixed audio signal includes: Obtain the initial reference signals of a first number of channels corresponding to the mixed audio signal, wherein the initial reference signals of the first number of channels are audio signals played by the vehicle-mounted equipment controlling the first number of speakers in the cabin; The initial reference signals of the first number of channels are downmixed to obtain reference signals of the second number of channels, where the second number is less than the first number.

3. The method according to claim 2, characterized in that, The step of downmixing the initial reference signals of the first number of channels to obtain reference signals of the second number of channels includes: Identify the signal quality of the initial reference signal for each of the first number of channels; The initial reference signals of the channels whose signal quality is less than a preset quality threshold are deleted to obtain a second number of reference signals for the channels. Alternatively, the signal quality of each channel can be sorted in descending order of signal quality to obtain a quality sequence; Starting from the initial signal quality of the quality sequence, select the initial reference signal corresponding to the second number of signal qualities as the reference signal corresponding to the mixed audio signal; Alternatively, the initial reference signals of the first number of channels can be merged according to a preset channel merging rule to obtain reference signals of the second number of channels; Alternatively, the initial reference signals of the first number of channels can be calculated and summed to obtain the total reference signal; Calculate the ratio of the total reference signal to the second number to obtain the reference signals for the second number of channels.

4. The method according to claim 1, characterized in that, Before acquiring the mixed audio signal received by the microphone in the vehicle cabin and acquiring the reference signal corresponding to the mixed audio signal, the method includes: Acquire a preset number of training sample data and target data corresponding to each training sample data. Each training sample data includes a target mixed audio signal received by a microphone in the vehicle cabin and a target reference signal corresponding to the target mixed audio signal. The target mixed audio signal includes a target echo signal and a target human voice signal emitted by at least one speaker. The target reference signal is an audio signal played by a speaker in the cabin controlled by the vehicle's infotainment system. The target echo signal is the target reference signal received by the microphone. The target data corresponding to the training sample data is composed of the target human voice signal. Using the target data as the training target, the audio processing model is trained based on the training sample data and the target data to obtain a trained target audio processing model. The target audio processing model is a model that eliminates the target echo signal in the target mixed audio signal based on the target reference signal corresponding to the target mixed audio signal in the training sample data.

5. The method according to claim 4, characterized in that, The step of obtaining a preset number of training sample data and the target data corresponding to each training sample data includes: Control multiple speakers within the cockpit to play multiple clean audio signals with different audio properties; The system acquires multiple clean audio signals received by the microphone in the cockpit and played by multiple speakers, and determines each clean audio signal received by the microphone as an initial echo signal. The system acquires the original audio signals emitted by multiple speakers received by the microphone in the cockpit, and determines each original audio signal received by the microphone as the initial human voice signal. Multiple initial echo signals are combined into a preset number of echo sets to obtain a preset number of target echo signals; multiple initial human voice signals are combined into a preset number of human voice sets to obtain a preset number of target human voice signals. The preset number of target echo signals and the preset number of target human voice signals are mixed to obtain the preset number of target mixed audio signals; For each of the target mixed audio signals, the initial echo signal contained in the target echo signal is used to determine the clean audio signal played by the vehicle-mounted device controlled speaker corresponding to the initial echo signal as the target reference signal corresponding to the target mixed audio signal. Based on the preset number of target mixed audio signals and the target reference signals corresponding to the target mixed audio signals, a preset number of training sample data are constructed; Based on the target human voice signal in the training sample data, construct the target data corresponding to the training sample data.

6. The method according to claim 4, characterized in that, The step of training an audio processing model using the target data as the training target, and training the audio processing model based on the training sample data and the target data to obtain a trained target audio processing model includes: The training sample data and the target data corresponding to the training sample data are input into the audio processing model. Based on the audio processing model, the target echo signal in the target mixed audio signal in the training sample data is eliminated to obtain the initial audio signal corresponding to the target mixed audio signal. Calculate the difference between the audio signals of each microphone channel contained in the initial audio signal and the target data corresponding to the training sample data, and determine the maximum difference among the differences as the function value of the target loss function; Based on the target loss function, the audio processing model is iteratively trained on the preset number of training sample data and the target data corresponding to each training sample data. When the function value of the target loss function is less than the preset loss threshold, the trained target audio processing model is obtained.

7. An echo cancellation device, characterized in that, include: The first acquisition module is used to acquire a mixed audio signal received by a microphone in the vehicle cabin and to acquire a reference signal corresponding to the mixed audio signal. The mixed audio signal includes an echo signal and a human voice signal emitted by a speaker. The reference signal is an audio signal played by a speaker in the cabin controlled by a vehicle-mounted device. The echo signal is the reference signal received by the microphone. The echo cancellation module is used to input the mixed audio signal and the reference signal into a preset target audio processing model, and based on the target audio processing model, cancel the echo signal in the mixed audio signal to obtain the target audio signal corresponding to the mixed audio signal.

8. A vehicle-mounted infotainment system, characterized in that, include: A memory and a processor, wherein the memory stores a computer program that, when executed by the processor, implements the echo cancellation method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the echo cancellation method as described in any one of claims 1-6.

10. A vehicle, characterized in that, The echo cancellation method as described in any one of claims 1-6 is implemented by including the echo cancellation device as described in claim 7 and / or the vehicle infotainment system as described in claim 8 and / or the computer-readable storage medium as described in claim 9.