Voice interaction method, device and equipment in cabin environment, medium and product

By acquiring and processing multiple audio signals and reference signals, and using dynamic filters to remove echoes, the problems of howling and echo interference in the cockpit environment were solved, achieving higher accuracy in voice and audio playback.

CN121963733APending Publication Date: 2026-05-01GUANGZHOU KUGOU COMP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU KUGOU COMP TECH CO LTD
Filing Date
2026-03-20
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In a cockpit environment, the sound from the speaker amplifier causes feedback and echo, which severely interferes with voice quality and makes karaoke impossible. Existing technologies such as handheld microphones and single-channel acoustic echo cancellation solutions have problems with low convenience or insufficient echo cancellation capability.

Method used

Multiple audio signals and reference signals are acquired, echoes are removed through dynamically updated filters, a speech estimation signal is generated, and playback is performed based on the speech estimation signal. The multiple audio signals and reference signals are used to finely remove echo interference, thereby improving the playback accuracy of speech and audio.

Benefits of technology

It effectively suppresses howling and echo, ensuring smooth karaoke performances and improving the accuracy of voice and audio playback in the cabin environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963733A_ABST
    Figure CN121963733A_ABST
Patent Text Reader

Abstract

The invention discloses a voice interaction method and device in a cabin environment, equipment, a medium and a product, and the method comprises the steps: collecting at least two paths of audio signals in the cabin environment, and enabling the at least two paths of audio signals to be signals which are collected under the condition of a power amplifier in the cabin environment and comprise voice audios; acquiring at least two paths of reference signals; removing the i-th reference signal from the i-th audio signal to obtain an i-th voice estimation signal; and playing the acquired voice audio based on the at least two paths of voice estimation signals. According to the invention, the audio collected under different microphones and the influence of the audio played by the loudspeaker on the different microphones are expressed through the multiple paths of audio signals and the multiple paths of collection signals, echo interference is removed from the audio signals in a finer and more accurate manner, and the playing accuracy of the collected voice audio in the cabin environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Voice interaction methods, devices, equipment, media and products in cockpit environments Technical Field

[0001] This application relates to the field of human-computer interaction, and in particular to a voice interaction method, device, equipment, medium and product in a cockpit environment. Background Technology

[0002] With the development of new energy vehicles, the ways of entertainment in the cabin environment have also undergone significant changes. Karaoke in the cabin environment refers to singing karaoke using the multi-zone microphones and sound system built into the cabin.

[0003] In related technologies, the user's voice is collected by the multi-zone microphone built into the cockpit and then played through the speaker. In other words, the user can sing karaoke in the cockpit environment by simply singing along with the accompaniment.

[0004] However, due to the high audio complexity in the cabin environment, the presence of speaker amplifiers, and a large amount of howling and echo, the voice quality is severely interfered with, making it impossible to sing karaoke. Summary of the Invention

[0005] This application provides a voice interaction method, device, equipment, medium, and product in a cockpit environment. The technical solution provided by this application includes the following aspects.

[0006] According to one aspect of the embodiments of this application, a voice interaction method in a cockpit environment is provided. The method includes: acquiring at least two audio signals in the cockpit environment, wherein the at least two audio signals are signals containing voice audio acquired when the power amplifier is used in the cockpit environment, wherein the i-th audio signal is acquired by the i-th microphone in the cockpit environment, and i is a positive integer; acquiring at least two reference signals, wherein the i-th reference signal is used to represent the signal of the speaker amplifier corresponding to the i-th microphone; removing the i-th reference signal from the i-th audio signal to obtain the i-th voice estimation signal; and playing the acquired voice audio based on the at least two voice estimation signals.

[0007] According to one aspect of the embodiments of this application, a voice interaction device in a cockpit environment is provided. The device includes: a acquisition module for acquiring at least two audio signals in the cockpit environment, wherein the at least two audio signals are signals containing voice audio acquired when the power amplifier is activated in the cockpit environment, wherein the i-th audio signal is acquired by the i-th microphone in the cockpit environment, and i is a positive integer; an acquisition module for acquiring at least two reference signals, wherein the i-th reference signal is used to represent the signal of the speaker amplifier corresponding to the i-th microphone; a processing module for removing the i-th reference signal from the i-th audio signal to obtain an i-th voice estimation signal; and a playback module for playing the acquired voice audio based on the at least two voice estimation signals.

[0008] In an optional embodiment, the processing module includes: a filtering unit for filtering the at least two reference signals based on a dynamically updated filter to obtain at least two echo estimation signals; a calculation unit for removing the i-th echo estimation signal from the i-th audio signal to obtain the i-th speech estimation signal; the processing module further includes: an updating unit for dynamically updating the filter based on the at least two speech estimation signals.

[0009] In an optional embodiment, the computing unit is configured to perform a first preprocessing on the i-th audio signal to obtain the i-th desired signal; and to subtract the i-th echo estimation signal from the i-th desired signal to obtain a residual signal as the i-th speech estimation signal; wherein the first preprocessing includes at least one of DC removal processing and pre-emphasis processing, the DC removal processing being used to remove the DC component in the audio signal, and the pre-emphasis processing being used to compensate for the propagation loss of the speech audio.

[0010] In an optional embodiment, the computing unit is further configured to perform a second preprocessing on the at least two reference signals to obtain at least two preprocessed signals; the filtering unit is further configured to filter the at least two preprocessed signals based on a dynamically updated filter to obtain the at least two echo estimation signals; wherein the second preprocessing includes pre-emphasis processing and time-frequency domain conversion processing, the pre-emphasis processing is used to compensate for the propagation loss of the loudspeaker power amplifier signal, and the time-frequency domain conversion processing is used to convert the reference signal in the time domain into a frequency domain signal.

[0011] In an optional embodiment, the filtering unit is further configured to filter the at least two preprocessed signals based on the dynamically updated filter to obtain at least two filtered signals; the calculation unit is further configured to perform time-frequency domain inverse processing on the at least two filtered signals to obtain the at least two echo estimation signals.

[0012] In an optional embodiment, the filtering unit is further configured to filter at least two reference signals in the (n+1)th round based on the (n+1)th filter dynamically updated in the (n)th round, to obtain at least two echo estimation signals in the (n+1)th round.

[0013] In an optional embodiment, the updating unit is further configured to update the filter coefficients of the filter based on the at least two speech estimation signals, the at least two reference signals, and the step size factor.

[0014] In an optional embodiment, the playback module is further configured to predict the sounding location of the speech audio in the cockpit environment based on at least two speech estimation signals, the at least two speech estimation signals including the i-th speech estimation signal; fuse the at least two speech estimation signals based on the sounding location to obtain a fused speech signal; and play the fused speech signal.

[0015] In an optional embodiment, the playback module is further configured to acquire the pitch energy distribution of the at least two speech estimation signals; and acquire the speech activity probability of the at least two speech estimation signals; and predict the sounding location of the speech audio in the cockpit environment based on the pitch energy distribution and the speech activity probability.

[0016] In an optional embodiment, the playback module is further configured to perform energy calculation on the i-th speech estimation signal in multiple sound zones to obtain energy values ​​corresponding to the multiple sound zones respectively; and sort the multiple sound zones according to the energy values ​​from largest to smallest to obtain the sound zone energy distribution corresponding to the i-th speech estimation signal.

[0017] In an optional embodiment, the playback module is further configured to obtain the fusion weights corresponding to the at least two speech estimation signals based on the distance between the sound source and the microphone; and to perform weighted fusion of the at least two speech estimation signals based on the fusion weights corresponding to the at least two speech estimation signals to obtain the fused speech signal.

[0018] According to one aspect of the embodiments of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the above-described voice interaction method in a cockpit environment.

[0019] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, the computer program being loaded and executed by a processor to implement the above-described voice interaction method in a cockpit environment.

[0020] According to one aspect of the embodiments of this application, a computer program product is provided, the computer program product including a computer program stored in a computer-readable storage medium, and a processor reading from the computer-readable storage medium and executing the computer program to implement the above-described voice interaction method in a cockpit environment.

[0021] The technical solution provided in this application can bring at least the following beneficial effects: acquiring multiple audio signals and obtaining multiple reference signals, thereby generating a speech estimation signal based on the multiple audio signals and multiple reference signals, and playing speech audio based on the speech estimation signal. By expressing the audio acquired under different microphones and the influence of different microphones on the audio played by the speaker through multiple audio signals and multiple acquisition signals, echo interference is removed from the audio signal in a more refined and accurate manner, improving the playback accuracy of the acquired speech audio in the cockpit environment, and significantly improving the echo situation when acquiring and playing speech audio in the power amplifier scenario. Attached Figure Description

[0022] Figure 1 is a schematic diagram of a voice interaction process in a cockpit environment provided by an exemplary embodiment of this application; Figure 2 is a schematic diagram of a voice interaction system in a cockpit environment provided by an embodiment of this application; Figure 3 is a flowchart of a voice interaction method in a cockpit environment provided by an exemplary embodiment of this application; Figure 4 is a schematic diagram of acquiring at least two audio signals and at least two reference signals provided by an exemplary embodiment of this application; Figure 5 is a schematic diagram of performing a first preprocessing on at least two audio signals provided by an exemplary embodiment of this application; Figure 6 is a flowchart of a voice interaction method in a cockpit environment provided by another exemplary embodiment of this application; Figure 7 is a schematic diagram of the process of processing audio signals and reference signals provided by an exemplary embodiment of this application; Figure 8 is a flowchart of a voice interaction method in a cockpit environment provided by another exemplary embodiment of this application; Figure 9 is a schematic diagram of the process of predicting the sound location in a cockpit environment provided by an exemplary embodiment of this application; Figure 10 is a structural block diagram of a voice interaction device in a cockpit environment provided by another exemplary embodiment of this application; Figure 11 is a structural block diagram of a voice interaction device in a cockpit environment provided by another exemplary embodiment of this application; Figure 12 is a structural block diagram of a computer device provided by an embodiment of this application. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0024] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0025] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0026] It should be understood that although the terms first, second, etc., may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, a first parameter may also be referred to as a second parameter, and similarly, a second parameter may also be referred to as a first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0027] It should be noted that this application may display prompt interfaces, pop-ups, or output voice prompts before and during the collection of user data. These prompt interfaces, pop-ups, or voice prompts are used to inform the user that their data is being collected. This ensures that the application only begins the steps for collecting user data after receiving confirmation from the user regarding the prompt interface or pop-up; otherwise (i.e., without user confirmation), the steps for collecting user data end, meaning no user data is collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of related user data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0028] With the development of new energy vehicles, the entertainment options in car cabins will also undergo significant changes, with karaoke in the cabin environment becoming a popular form of entertainment.

[0029] Among them, "cockpit karaoke without microphone" refers to singing karaoke using the multi-zone microphone and sound system built into the cockpit. In the enclosed cockpit space, due to the small space environment of the cockpit, plus multiple high-sensitivity microphones and more than 10 speakers, if a voice enhancement preprocessing algorithm is not performed, there will be a lot of howling, echo, etc., which will seriously interfere with the voice quality and make it impossible to sing karaoke.

[0030] Among the related technologies, there is at least one karaoke solution in the cabin environment: (1) singing karaoke in the cabin environment using a handheld microphone; (2) a single-channel acoustic echo cancellation (AEC) solution.

[0031] Among them, the solution for karaoke with a handheld microphone requires users to hold the microphone, which is not very convenient.

[0032] For single-channel AEC solutions, the advantages of multi-microphone recording and multiple reference signals are not fully utilized, which often results in insufficient echo cancellation capabilities and significant speech damage.

[0033] This application provides a speech enhancement preprocessing algorithm that effectively suppresses howling and echo, thereby ensuring the smooth progress of karaoke.

[0034] Schematic, Figure 1 is a diagram illustrating a voice interaction process in a cockpit environment according to an exemplary embodiment of this application. As shown in Figure 1, the cockpit environment includes a speaker 110 and multiple microphones 120. The multiple microphones 120 are used to collect at least two audio signals. Since the cockpit environment also includes a speaker 110, the audio signals collected by the multiple microphones 120 may include both voice audio obtained from the user's speech and audio played by the speaker 110. The microphones 120 in the cockpit environment refer to sound acquisition components that are configured in the cockpit and are installed in the cockpit at the time of vehicle manufacture, with relatively fixed positions. For example, sound acquisition components configured on the top of the driver's seat, the top of the passenger seat, or on the steering wheel or A-pillar. This embodiment of the application does not limit the configuration position of the microphones 120. The microphones 120 are implemented as fixed / relatively fixed components, rather than movable devices for the user to hold.

[0035] The i-th audio signal is acquired by the i-th microphone in the cockpit environment, where i is a positive integer.

[0036] In addition, at least two reference signals are acquired. The i-th reference signal is used to express the signal of the speaker 110 power amplifier corresponding to the i-th microphone. That is, the i-th reference signal is used to express the signal of the speaker 110 power amplifier that may be acquired by the i-th microphone during the process of acquiring voice audio by the i-th microphone.

[0037] After filtering at least two reference signals using a dynamically updated filter, at least two echo estimation signals are obtained. Then, the i-th echo estimation signal is removed from the i-th audio signal to obtain the i-th speech estimation signal.

[0038] Optionally, in this embodiment, the dynamic update of the filter is performed based on the reference signal and the speech estimation signal predicted in each round. Optionally, the filter for the current round is updated based on the reference signal, the speech estimation signal and the update step size, so as to obtain the filter to be used in the next round.

[0039] The acquired speech audio is played based on at least two speech estimation signals, wherein the acquired speech audio is played through speaker 110 based on at least two speech estimation signals. Additionally, the filter is dynamically updated based on at least two speech estimation signals.

[0040] The voice interaction method in the cockpit environment provided in this application embodiment is implemented by a computer device. The voice interaction method in the cockpit environment can be implemented by the vehicle terminal alone, or it can be implemented by the vehicle terminal and the server working together.

[0041] In this scenario, when the voice interaction method in the cockpit environment is implemented solely by the in-vehicle terminal, i.e., a multimedia program running on the in-vehicle terminal side that provides a karaoke function, at least two audio signals from the cockpit environment are acquired on the in-vehicle terminal side, along with reference signals corresponding to multiple microphones (i.e., at least two reference signals are acquired). Based on these reference signals, noise and echo are removed from the at least two audio signals to obtain the voice audio, which is then played through the speakers in the cockpit environment.

[0042] When the voice interaction method in the cockpit environment is jointly implemented by the vehicle terminal and the server, for example, the vehicle terminal collects at least two audio signals in the cockpit environment and obtains reference signals corresponding to multiple microphones, that is, it obtains at least two reference signals, and sends at least two audio signals and at least two reference signals to the server. The server then denoises and removes echoes from the at least two audio signals based on the at least two reference signals to obtain voice audio, which is then sent to the vehicle terminal and played through the speakers in the cockpit environment.

[0043] Please refer to Figure 2, which shows a schematic diagram of a voice interaction system in a cockpit environment according to an embodiment of this application. The computer system 200 includes a terminal 220, or a terminal 220 and a server 240. Specifically, when the voice interaction method in the cockpit environment is implemented by the terminal and server configured and executed, the computer system 200 includes a terminal 220 and a server 240.

[0044] The device type of terminal 220 includes at least one of the following: in-vehicle terminal, smartphone, laptop, smartwatch, desktop computer, tablet, intelligent robot, augmented reality (AR) device, and virtual reality (VR) device. The aforementioned terminal 220 is used within a cockpit environment. For example, in a cockpit environment, the in-vehicle terminal collects at least two audio signals and at least two reference signals, sends them to the smartphone for filtering and other processing, and then plays the filtered audio / voice through the cockpit speakers.

[0045] In this embodiment, terminal 220 is implemented as an in-vehicle terminal, which is connected to server 240 via a wireless network or wired network. Terminal 220 runs a multimedia program. The multimedia program has multimedia content playback functionality, and in this embodiment, it also has a karaoke function. This multimedia program can also be implemented as at least one of various programs such as a game program, a multimedia playback program, a financial payment program, or an instant messaging program. In this embodiment, a multimedia playback program is used as an example, meaning that the multimedia playback program has a karaoke function. By triggering the karaoke function, a karaoke song can be selected, and microphone-free karaoke can be performed within the cabin environment.

[0046] In one example, the voice interaction method in this cockpit environment is implemented by the cooperation of terminal 220 and server 240. Illustratively, server 240 provides multimedia program execution data to terminal 220, such as song data that can be used for karaoke. Terminal 220 receives the song data sent by server 240, thereby being able to display the lyrics and play the background audio of the song.

[0047] During the microphone-free karaoke process, terminal 220 acquires at least two audio signals from the cabin environment; and acquires at least two reference signals; filters the at least two reference signals based on dynamically updated filters to obtain at least two echo estimation signals; removes the i-th echo estimation signal from the i-th audio signal to obtain the i-th speech estimation signal; plays the acquired speech audio based on the at least two speech estimation signals; and dynamically updates the filters based on the at least two speech estimation signals.

[0048] Those skilled in the art will understand that the number of the aforementioned devices can be more or less. For example, there may be only one device, or there may be dozens or hundreds of devices, or even more. This application does not limit the number or type of devices.

[0049] Server 240 includes at least one of one server or multiple servers. In some embodiments, server 240 is implemented as a private server connected to terminal 220, and the model provided by server 240 is a private model deployed on the private server. Optionally, server 240 undertakes the main computing work, and terminal 220 undertakes the secondary computing work; or, server 240 undertakes the secondary computing work, and terminal 220 undertakes the main computing work; or, server 240 and terminal 220 adopt a distributed computing architecture for collaborative computing.

[0050] It is worth noting that the aforementioned server 240 can be implemented as a physical server or as a cloud server in the cloud. In some embodiments, the aforementioned server 240 can also be implemented as a node in a blockchain system.

[0051] Please refer to Figure 3, which shows a flowchart of a voice interaction method in a cockpit environment provided by an exemplary embodiment of this application. The method can be executed by a computer device (which can be implemented as a terminal 220 or a server 240 as shown in Figure 2), or the method can be jointly executed by the terminal and the server. In this embodiment, the method is executed by a terminal / vehicle terminal as an example. As shown in Figure 3, the method includes at least one of the following steps 310 to 340.

[0052] Step 310: Collect at least two audio signals from the cockpit environment.

[0053] At least two audio signals are signals containing voice and audio collected in the cockpit environment with the power amplifier. Among them, the i-th audio signal is collected by the i-th microphone in the cockpit environment, where i is a positive integer. Therefore, at least two audio signals are collected by at least two microphones in the cockpit environment, and there is a one-to-one correspondence between the audio signals and the microphones.

[0054] In other words, the number of at least two audio signals corresponds to the number of microphones in the cockpit environment; or, the number of at least two audio signals corresponds to the number of microphones in the cockpit environment that are in operation. A microphone in operation refers to a microphone that performs audio signal acquisition, and the acquired audio signal is subsequently processed as one of at least two audio signals.

[0055] In some embodiments, the microphones in working state are all the microphones in the cabin environment, or the microphones in working state are the microphones enabled by an activation operation, such as selecting to enable multiple microphones in the front half of the cabin in the control interface of the vehicle terminal, or selecting to enable multiple microphones in the rear half of the cabin by triggering a physical button in the cabin environment, or selecting to enable multiple specified microphones in the cabin environment by an application bound to the vehicle terminal on a mobile terminal. This embodiment does not limit the activation method of multiple microphones.

[0056] In some embodiments, the cabin environment includes multiple microphones, such as: four microphones distributed on the left, right, left, and right sides of the front cabin; or six microphones distributed on the left, right, middle, left, right, and middle sides of the front cabin; or, for example, in the cabin environment of a multi-purpose vehicle (MPV), they may be distributed on the left, right, left, right, and right sides of the front cabin, middle cabin, and rear cabin; or, for example, in the cabin environment of a bus, one or more microphones may be distributed around each seat, resulting in dozens of microphones in the bus cabin environment. This application does not limit the number or distribution of microphones. Different vehicle types may have the same or different number of microphones; and different vehicle types may have the same or different microphone distribution locations.

[0057] In this embodiment, m microphones are used as an example for illustration, where m ≥ 2 and m is an integer. In this embodiment, m is used as an example with a value of 4. For example, in a cabin environment, the left side of the front cabin is where the driver's position is distributed with microphones; the right side of the front cabin is where the co-pilot's position is distributed with microphones; and the left side and right side of the rear cabin are where the passengers are distributed with microphones.

[0058] At least two audio signals were acquired in the cockpit environment with the amplifier in operation. Among them, at least two audio signals also include voice audio signals, which are acquired in the cockpit environment with the amplifier in operation. In other words, voice audio is acquired and amplified by the speaker in the cockpit environment.

[0059] Speech audio refers to audio obtained by a speaker, which can be a human being, an animal, or the like. This embodiment will use a human being as an example for explanation.

[0060] The audio acquisition scenarios include at least one of the following: 1. Microphone-free karaoke scenario; that is, a scenario where karaoke is performed using the microphone built into the cabin environment without a handheld microphone. In some embodiments, for the microphone-free karaoke scenario, the background music, such as the accompaniment, is first amplified by the speakers in the cabin environment. The user can then sing along with the accompaniment, and the microphone in the cabin environment captures the user's singing audio (i.e., voice audio), which is then amplified by the speakers, thus achieving the microphone-free karaoke effect by amplifying both the accompaniment and the singing audio. However, in this scenario, the speakers are used to amplify both the accompaniment and the user's singing audio. Therefore, when the microphone acquires the audio signal, it can capture both the singing signal from the user and the signal from the speaker's own amplifier. Therefore, to avoid echo effects, it is necessary to remove the signal from the speaker's own amplifier from the audio signal acquired by the microphone.

[0061] 2. In-vehicle notification scenario; for example, the cabin environment is implemented as a bus cabin environment. In scenarios such as tour groups, the tour guide or driver needs to make notifications within the cabin environment, such as reminding passengers to fasten their seat belts. While playing background noise, the driver / tour guide can turn on microphones in multiple designated locations and issue the notification aloud. The microphones can capture both the speaker's voice audio and background noise. Therefore, to avoid echo effects, the signal from the speaker's amplifier needs to be removed from the audio signal captured by the microphone.

[0062] It is worth noting that the above-mentioned voice and audio acquisition scenarios are merely illustrative examples, and the embodiments of this application do not limit the specific voice and audio acquisition scenarios.

[0063] Optionally, in this embodiment, at least two audio signals refer to signals directly acquired by a microphone without preprocessing or other processing. The at least two audio signals are either analog signals acquired by a microphone or digital signals acquired by a microphone.

[0064] In this embodiment of the application, at least two audio signals are implemented as digital signals for illustration.

[0065] The i-th audio signal is acquired by the i-th microphone in the cockpit environment. That is, the audio signal acquired by the i-th microphone in the cockpit environment is the i-th audio signal. For illustration, the i-th microphone is the microphone distributed on the left side of the front half of the cockpit. The i-th audio signal acquired by the i-th microphone is the audio signal that can be acquired on the left side of the front half of the cockpit. For example, if the sound in the rear half of the cockpit is far away from the microphone, the audio signal will be weaker, while if the sound in the front half of the cockpit is close to the microphone, the audio signal will be stronger.

[0066] Step 320: Obtain at least two reference signals.

[0067] The i-th reference signal is used to represent the signal of the speaker amplifier corresponding to the i-th microphone.

[0068] In some embodiments, when acquiring at least two reference signals, the acquisition method includes at least one of the following: 1. Determine the speaker corresponding to each microphone, wherein each microphone corresponds to one or more speakers, and receive the signal returned by the speaker. Then, the signal fed back by the speaker corresponding to the i-th microphone is the reference signal corresponding to the i-th microphone. Wherein, if the i-th microphone corresponds to multiple speakers, the fused value of the signals fed back by the multiple speakers, such as the average signal, is the reference signal corresponding to the i-th microphone.

[0069] Among them, the speaker corresponding to the microphone is the speaker closest to the microphone; or, the speaker corresponding to the microphone is a speaker within a preset distance range from the microphone; or, the speaker corresponding to the microphone is a speaker that conforms to a preset positional relationship with the microphone.

[0070] For example, for the microphone on the left side of the front cabin, the audio returned by the speaker on the left side of the front cabin is used as the reference signal for the microphone on the left side of the front cabin.

[0071] The signal fed back by the speaker is the signal played by the speaker.

[0072] 2. Determine the signal fusion function corresponding to each microphone. The signal fusion function is used to express the fusion method of the signals returned by multiple speakers. After obtaining the signals returned by multiple speakers, fuse the signals returned by each speaker according to the signal fusion function corresponding to the i-th microphone to obtain the reference signal corresponding to the i-th microphone.

[0073] The signal fusion function is used to weight and fuse the signals returned by each speaker, with the signal from the speaker closer to the microphone having a higher weight. For example, for the microphone on the left side of the front cabin, the signal returned by the speaker on the left side of the front cabin has a higher weight in the signal fusion function, while the signal returned by the speaker on the right side of the front cabin or the rear cabin has a lower weight.

[0074] The signal fusion function expresses the influence of each speaker on the signal acquisition of the microphone, and the reference signal of each microphone is obtained based on the influence amplitude. The reference signal corresponding to different microphones is predicted in a more accurate way, and the reference signal is removed from the audio signal acquired by the microphone, which improves the accuracy of removing the reference signal from the audio signal.

[0075] It is worth noting that the above-mentioned method of obtaining the reference signal is only an illustrative example, and this embodiment does not limit it.

[0076] In some embodiments, different microphones may correspond to the same reference signal, or different microphones may correspond to different reference signals. This application embodiment uses the example of different microphones corresponding to different reference signals for illustration.

[0077] In some embodiments, steps 310 and 320 are two parallel steps. Step 310 can be executed first, followed by step 320; or step 320 can be executed first, followed by step 310; or steps 310 and 320 can be executed simultaneously. Specifically, steps 310 and 320 are executed simultaneously for the k-th time period. That is, at least two audio signals in the cockpit environment during the k-th time period are acquired; and at least two reference signals are acquired within the k-th time period, where k is a positive integer. For example, at least two audio signals within one second of the current time period in the cockpit environment are acquired; and at least two reference signals are acquired within the current time period.

[0078] At least two audio signals and at least two reference signals are signals acquired for the same / similar time period, so that the microphone acquisition results expressed by the at least two audio signals and the amplifier effects of the speaker expressed by the at least two reference signals correspond to the same / similar time period.

[0079] Schematic, Figure 4 is a schematic diagram of the acquisition of at least two audio signals and at least two reference signals provided in an exemplary embodiment of this application. As shown in Figure 4, on the time axis, four audio signals 410 are acquired through four microphones, and four reference signals 420 are acquired, wherein the four audio signals 410 and the four reference signals 420 correspond to the same / similar time periods.

[0080] Among them, the 4 microphones collect 4 audio signals 410 can be a cockpit environment including 4 microphones, each microphone collecting one audio signal, or the cockpit environment including more than 4 microphones, such as 6 microphones, with 4 of them enabled to be in working state, and each of the 4 microphones in working state collecting one audio signal.

[0081] Step 330: Remove the i-th reference signal from the i-th audio signal to obtain the i-th speech estimation signal.

[0082] Optionally, at least two reference signals are filtered based on the dynamically updated filter to obtain at least two echo estimation signals. The i-th echo estimation signal is removed from the i-th audio signal to obtain the i-th speech estimation signal.

[0083] Among them, the dynamically updated filter refers to the filter that is continuously updated, that is, the filter that is constantly updated during the speech audio processing.

[0084] By using dynamically updated filters to filter the reference signal, the filters are maintained to perform filtering on the reference signal in an appropriate manner. Since the filter updates are based on the previously acquired reference signal and the predicted speech estimation signal, the updated filters are better able to adapt to the current filtering requirements of the reference signal, thus improving the processing efficiency and accuracy of the reference signal.

[0085] In some embodiments, based on the processing of the reference signal and the audio signal in the (n-1)th time period, the filter applied in the (n-1)th time period is updated to obtain the filter applied in the nth time period, where n≥1.

[0086] Filtering a reference signal is a signal processing operation used to enhance, weaken, or remove specific frequency components from the reference signal. By altering the frequency response characteristics of the reference signal, filtering can achieve various audio processing effects, such as noise removal and enhancement of sound within a specific frequency range. In other words, filtering processes a reference signal to extract, enhance, or suppress specific frequency components. A filter is a circuit or algorithm capable of frequency-selectively processing a signal; in this embodiment, a filter implementation as an algorithm is used for illustration.

[0087] Low-pass filtering refers to a filtering method that allows low-frequency signals to pass through while attenuating high-frequency signals; high-pass filtering refers to a filtering method that allows high-frequency signals to pass through while attenuating low-frequency signals; band-pass filtering refers to a filtering method that attenuates signals within a specific frequency range while allowing signals outside that range to pass through. In this embodiment, the filtering method is not limited.

[0088] In some embodiments, filtering of the reference signal is achieved through convolutional network layers.

[0089] In one scheme where a filter is implemented using a convolutional network layer, the convolutional network layer convolves the filter coefficients (i.e., filter weights) with the reference signal to filter the reference signal. In some embodiments, at least two reference signals are filtered based on dynamically updated filter weights to obtain at least two echo estimation signals.

[0090] Optionally, the residual signal is obtained by subtracting the i-th echo estimation signal from the i-th audio signal and using it as the i-th speech estimation signal.

[0091] Specifically, when subtracting the i-th echo estimation signal from the i-th audio signal, the audio signal and the echo estimation signal are first aligned on the time axis. Then, the signal value of the sampling point on the audio signal is subtracted from the signal value of the sampling point on the echo estimation signal to obtain a new signal value, which is used as the signal value of the sampling point in the speech estimation signal.

[0092] In some embodiments, since the audio signal is a signal collected by the microphone after propagation through the air, a first preprocessing is performed on the i-th audio signal to obtain the i-th expected signal. The i-th expected signal is then subtracted from the i-th echo estimation signal to obtain the residual signal as the i-th speech estimation signal. The first preprocessing includes at least one of DC removal processing and pre-emphasis processing. The DC removal processing is used to remove the DC component in the audio signal, and the pre-emphasis processing is used to compensate for the propagation loss of the speech audio.

[0093] In some embodiments, a notch filter is used to remove the DC component from the i-th audio signal. The DC component refers to a signal with a frequency of zero in the audio signal, meaning its value is constant or changes very slowly. In audio signals, DC signals are typically caused by the bias voltage of electronic devices or the DC offset of sensors.

[0094] Pre-emphasis processing of audio signals is an audio signal processing technique that compensates for the attenuation of high-frequency components during transmission or recording by increasing the amplitude of these components. The core idea of ​​pre-emphasis processing is to compensate for the attenuation of high-frequency components during signal transmission. Optionally, pre-emphasis processing is typically implemented using a high-pass filter or a similar circuit. A high-pass filter amplifies the high-frequency components of the signal while having less impact on the low-frequency components. In this way, when the signal passes through the transmission medium, the attenuation of high-frequency components is compensated, resulting in a flatter signal spectrum and reduced distortion caused by high-frequency attenuation.

[0095] By performing a first preprocessing on the audio signal, misleading DC components in the audio signal can be removed, and / or signal loss caused by the spatial propagation of speech audio in the cabin environment can be compensated, thereby improving the accuracy of the speech estimation signal obtained after subsequent filtering of the audio signal and improving the accuracy of human voice processing in microphone-free karaoke scenarios.

[0096] Schematic, Figure 5 is a schematic diagram of performing a first preprocessing on at least two audio signals according to an exemplary embodiment of this application. As shown in Figure 5, after at least two audio signals are acquired, the DC component in the at least two audio signals is first removed by a notch filter 510. The output data of the notch filter 510 is processed by a pre-emphasis processing 520 to simulate the loss of the audio signal during air propagation. The pre-emphasis processing 520 outputs the desired signal. The i-th audio signal is obtained after passing through the notch filter 510 and the pre-emphasis processing 520.

[0097] After acquiring at least two expected signals, at least two echo estimation signals are subtracted from the at least two expected signals to obtain the residual signal, which is the aforementioned speech estimation signal. Specifically, the i-th expected signal is subtracted from the i-th echo estimation signal to obtain the i-th residual signal, which is also the i-th speech estimation signal.

[0098] Step 340: Play the acquired speech audio based on at least two speech estimation signals.

[0099] Optionally, the filter is dynamically updated based on at least two speech estimation signals.

[0100] In other words, after obtaining at least two voice estimation signals, on the one hand, the voice audio is generated based on the voice estimation signals and played through the speaker, thereby enabling users to speak without a microphone in the cabin environment, such as singing karaoke without a microphone, making notifications without a microphone, etc.

[0101] On the other hand, after determining the speech estimation signal, the filter is dynamically updated based on the speech estimation signal. Optionally, based on the (n+1)th filter obtained by dynamically updating in the nth round, at least two reference signals in the (n+1)th round are filtered to obtain at least two echo estimation signals in the (n+1)th round.

[0102] In other words, the use and updating of filters is an iterative process.

[0103] In summary, the method provided in this application acquires multiple audio signals and multiple reference signals, thereby generating a speech estimation signal based on the multiple audio signals and multiple reference signals, and playing speech audio based on the speech estimation signal. By expressing the audio acquired under different microphones and the influence of different microphones on the audio played by the speaker through multiple audio signals and multiple acquisition signals, echo interference is removed from the audio signal in a more refined and accurate manner, improving the playback accuracy of the acquired speech audio in the cockpit environment, and significantly improving the echo situation when acquiring and playing speech audio in the power amplifier scenario.

[0104] In an optional embodiment, when removing the reference signal from the audio signal, the reference signal is first filtered to obtain the echo estimation signal, and the filter is dynamically updated.

[0105] Figure 6 is a flowchart of a voice interaction method in a cockpit environment provided by another exemplary embodiment of this application. The method can be executed by a computer device (which can be implemented as a terminal 220 or a server 240 as shown in Figure 2), or the method can be jointly executed by the terminal and the server. In this embodiment, taking the method as being executed by a terminal / vehicle terminal as an example, as shown in Figure 6, before or during step 330, the above step 330 can also be implemented as at least one of the following steps.

[0106] Step 610: Filter at least two reference signals based on the dynamically updated filter to obtain at least two echo estimation signals.

[0107] In some embodiments, a second preprocessing is performed on at least two reference signals to obtain at least two preprocessed signals. The at least two preprocessed signals are then filtered based on dynamically updated filters to obtain at least two echo estimation signals. The second preprocessing includes pre-emphasis processing and time-frequency domain transformation processing. The pre-emphasis processing is used to compensate for the propagation loss of the loudspeaker amplifier signal, and the time-frequency domain transformation processing is used to convert the time-domain reference signal into a frequency-domain signal.

[0108] By performing a second preprocessing on the reference signal, the signal loss caused by the speaker amplifier signal propagating in the cabin environment is compensated, and the time domain dimension is converted into the frequency domain dimension to better participate in subsequent signal analysis and processing. As a result, the accuracy of the speech estimation signal obtained by removing the reference signal from the audio signal is higher, which improves the accuracy of human voice processing in the microphoneless karaoke scenario.

[0109] Optionally, at least two preprocessed signals are filtered based on the dynamically updated filters to obtain at least two filtered signals; time-frequency domain inverse processing is performed on the at least two filtered signals to obtain the at least two echo estimation signals, wherein the time-frequency domain inverse processing is used to convert the frequency-domain filtered signals into time-domain echo estimation signals.

[0110] Optionally, based on the (n+1)th filter obtained by dynamic update in the nth round, at least two reference signals in the (n+1)th round are filtered to obtain at least two echo estimation signals in the (n+1)th round.

[0111] Step 620: Remove the i-th echo estimation signal from the i-th audio signal to obtain the i-th speech estimation signal.

[0112] Optionally, the filter is dynamically updated based on at least two speech estimation signals.

[0113] In some embodiments, the filter coefficients of the filter are updated based on at least two speech estimation signals, at least two reference signals, and a step size factor.

[0114] In this method, at least two speech estimation signals and at least two reference signals are used together to calculate the filter coefficients; or, each set of speech estimation signals and reference signals is used to calculate the filter coefficients separately, and the calculated filter coefficients are integrated to obtain the filter coefficients used to update the filter.

[0115] In some embodiments, the example of using at least two speech estimation signals and at least two reference signals to calculate the filter coefficients is illustrated. Schematic, at least two speech estimation signals and at least two reference signals form a vector for calculating the filter coefficients.

[0116] As shown in Formula 1 and Formula 2 below.

[0117] Formula 1:

[0118] in, This represents the step size factor used in the nth round. Let represent the reference signal acquired in the nth round. This represents the transpose of the vector formed by the reference signals acquired in the nth round.

[0119] Formula 2:

[0120] in, This represents the filter coefficients used in the (n+1)th round of prediction. This indicates the filtering system used in the nth round. This indicates at least two voice estimation signals.

[0121] Updating the filter coefficients using the speech estimation signal, reference signal, and compensation factor improves the accuracy of filter updates and avoids problems caused by excessively large or small updates, which can lead to poor filter performance.

[0122] Figure 7 is a schematic diagram of the process of processing audio signals and reference signals provided in an exemplary embodiment of this application. Referring to Figure 7, the processing of at least two audio signals and at least two reference signals will be described.

[0123] 1. The multiple audio signals are processed by notch filter 710 to remove the DC component.

[0124] The multiple audio signals are acquired through multiple microphones in the cockpit environment. The notch filter 710 is used to remove the DC component from the audio signals by filtering.

[0125] The notch filter 710 removes the DC component from multiple audio signals. These audio signals can be input to the notch filter 710 simultaneously or sequentially. Taking sequential input as an example, after the i-th audio signal is input to the notch filter 710, the notch filter 710 removes the DC component from the i-th audio signal.

[0126] 2. The output of the notch filter 710 is pre-emphasized 720.

[0127] By pre-emphasis processing 720, the loss of the simulated signal during spatial propagation is obtained, thus obtaining the desired signal d. The desired signal d is the signal obtained by predicting and compensating for the loss of the signal during spatial propagation.

[0128] 3. Multiple reference signals undergo pre-emphasis processing 730.

[0129] The pre-emphasis processing 730 for the reference signal is used to simulate the loss caused by the signal played by the loudspeaker propagating in space.

[0130] 4. The signal after pre-emphasis processing 730 is processed by Fast Fourier Transform (FFT) 740 to obtain the frequency domain signal of the reference signal.

[0131] Among them, the frequency domain signal (or preprocessed signal) is the signal obtained by time-frequency domain transformation of the reference signal expressed in the frequency domain dimension.

[0132] 5. Convolve the frequency domain signal and the filter coefficients by 750 to obtain the echo estimation signal in the frequency domain dimension.

[0133] Specifically, the i-th frequency domain signal is convolved with the current filter coefficients to obtain the echo estimation signal of the i-th frequency domain dimension.

[0134] 6. The echo estimation signal in the frequency domain is processed by inverse time-frequency domain processing, namely inverse fast Fourier transform (IFFT) 760, to obtain the echo estimation signal in the time domain.

[0135] In other words, IFFT760 converts the echo estimation signal expressed in the frequency domain into an echo estimation signal expressed in the time domain. Expressing the echo estimation signal in the frequency domain involves establishing the correspondence between frequency ranges and the echo estimation signals, such as the distribution of the echo estimation signals within a first frequency range. Expressing the echo estimation signal in the time domain involves establishing the correspondence between timestamps and the echo estimation signals, such as the intensity of the echo estimation signal at a first timestamp and the intensity of the echo estimation signal at a second timestamp.

[0136] 7. Subtract the time-domain echo estimation signal from the expected signal d to obtain the residual signal.

[0137] 8. Perform de-emphasis processing 770 on the residual signal to obtain the residual output signal, also known as the speech estimation signal.

[0138] 9. Additionally, filter coefficients are updated 780 based on the frequency domain signals of the estimated speech signal and the reference signal.

[0139] In summary, the method provided in this application acquires multiple audio signals and multiple reference signals, thereby generating a speech estimation signal based on the multiple audio signals and multiple reference signals, and playing speech audio based on the speech estimation signal. By expressing the audio acquired under different microphones and the influence of different microphones on the audio played by the speaker through multiple audio signals and multiple acquisition signals, echo interference is removed from the audio signal in a more refined and accurate manner, improving the playback accuracy of the acquired speech audio in the cockpit environment, and significantly improving the echo situation when acquiring and playing speech audio in the power amplifier scenario.

[0140] By iteratively updating the filter, the (n+1)th filter is obtained by dynamically updating it in the nth round. This allows the use of prior knowledge from previous rounds in the (n+1)th round, improving the accuracy of filtering the reference signal in the (n+1)th round and thus improving the accuracy of audio signal processing.

[0141] In an optional embodiment, after acquiring the speech estimation signal, the location of the sound source within the cockpit can also be predicted, thereby playing speech audio based on the sound source location.

[0142] Figure 8 is a flowchart of a voice interaction method in a cockpit environment provided by another exemplary embodiment of this application. The method can be executed by a computer device (which can be implemented as a terminal 220 or a server 240 as shown in Figure 2), or the method can be jointly executed by the terminal and the server. In this embodiment, taking the method as being executed by a terminal / vehicle terminal as an example, as shown in Figure 8, the above step 340 can also be implemented to include at least one of the following steps.

[0143] Step 810: Predict the location of the speech audio in the cockpit environment based on at least two speech estimation signals.

[0144] At least two speech estimation signals include the i-th speech estimation signal, which is the estimated signal obtained by removing the i-th reference signal corresponding to the i-th microphone from the audio signal collected by the i-th microphone. In other words, the signal obtained by the speaker that the i-th microphone may collect is removed from the audio signal collected by the i-th microphone, so as to restore the sound signal of the sound-producing subject collected by the i-th microphone.

[0145] Optionally, the region energy distribution of at least two speech estimation signals is obtained; and the speech activity probability of at least two speech estimation signals is obtained; based on the region energy distribution and the speech activity probability, the location of the speech audio in the cockpit environment is predicted.

[0146] Specifically, when obtaining the energy distribution of the speech regions, energy calculation is performed on the i-th speech estimation signal in multiple speech regions to obtain the energy values ​​corresponding to each speech region. The multiple speech regions are sorted from largest to smallest according to the energy values ​​to obtain the energy distribution of the speech regions corresponding to the i-th speech estimation signal.

[0147] In some embodiments, multiple frames of speech estimation signals are buffered for each channel, and the energy distribution of the speech region is identified based on the buffered multiple frames of speech estimation signals.

[0148] A frequency range refers to a sub-range of frequencies within which a musical instrument, human voice, or any other sound-producing entity can emit sound. Frequency range energy distribution refers to the energy intensity of an audio signal within these different frequency ranges, reflecting the energy distribution of the signal across the frequency spectrum.

[0149] Optionally, energy calculation is performed on at least two speech estimation signals. For each speech estimation signal, 10 frames of speech estimation signals are divided and buffered in 80ms increments, and the voice region energy of each signal is predicted based on the buffered 10 frames. The calculation is performed at a single-channel granularity, with each channel corresponding to one microphone; therefore, the voice region energy distribution of the speech estimation signal in the audio signal acquired by each microphone is calculated.

[0150] The multi-channel energy of at least two speech estimation signals is sorted from largest to smallest. That is, for the i-th speech estimation signal, energy calculation is performed on the i-th speech estimation signal in multiple voice regions to obtain the energy values ​​corresponding to the multiple voice regions respectively. Then, the multiple voice regions of the i-th speech estimation signal are sorted in descending order of energy values ​​to obtain the voice region energy distribution of the i-th speech estimation signal.

[0151] In addition, at least two speech estimation signals are mixed, and the mixed signal is subjected to Voice Activity Detection (VAD) to obtain the speech activity probability. VAD is a speech processing technique used to detect the presence of a speech signal; that is, in this embodiment, the probability of a subject uttering speech in at least two speech estimation signals is expressed based on the speech activity probability.

[0152] Therefore, based on the energy distribution of the sound region and the probability of speech activity, it is possible to predict the location of speech audio in the cockpit environment.

[0153] When the probability of speech activity reaches a preset probability threshold, it indicates that there is a speaker making a sound. Then, the location of the speaker in the cockpit environment is predicted based on the energy distribution of the sound regions of each estimated speech signal.

[0154] Conversely, if the probability of speech activity is less than the preset probability threshold, it means that no subject is speaking, and the prediction of the speech location is canceled.

[0155] For example, if the preset probability threshold is 60%, then when the probability of voice activity reaches 60%, the location of the sound source in the cockpit environment is predicted based on the energy distribution of the sound regions in each estimated voice signal. Illustratively, the higher the energy of a specified sound region in the estimated voice signal, the closer the sound source is to the microphone corresponding to that estimated voice signal; conversely, the lower the energy of a specified sound region in the estimated voice signal, the farther the sound source is from the microphone corresponding to that estimated voice signal.

[0156] The designated pitch range is either a pre-defined pitch range corresponding to the vocal subject, or a pitch range determined based on the timestamp on the track's timeline and the vocal subject.

[0157] To illustrate, based on the timestamp of the track playback and the speaker, the frequency range is determined to be 340 Hz – 540 Hz. The higher the energy in the estimated speech signal within this frequency range, the closer the speaker's location is to the microphone corresponding to that estimated speech signal.

[0158] Step 820: Based on the location of the sound, fuse at least two speech estimation signals to obtain a fused speech signal.

[0159] In some embodiments, based on the distance between the sound source and the microphone, fusion weights corresponding to at least two speech estimation signals are obtained respectively, and the at least two speech estimation signals are weighted and fused based on the fusion weights corresponding to the at least two speech estimation signals respectively to obtain a fused speech signal.

[0160] In some embodiments, the closer the sound source is to the microphone, the greater the fusion weight of the speech estimation signal corresponding to the microphone. In other words, the distance between the sound source and the microphone is negatively correlated with the fusion weight of the speech estimation signal.

[0161] The system loads a pre-designed microphone weighting factor based on the sound source location, and then mixes the multiple speech estimation signals based on the weighting factor to output the mixing result.

[0162] The vocal position is obtained based on the information obtained from the above-mentioned sound region energy distribution and speech activity probability prediction. The correspondence between the pre-designed vocal position and the microphone weight factor is obtained, so as to find the fusion weight corresponding to each microphone based on the vocal position, and fuse the speech estimation signals corresponding to each microphone based on the fusion weight to obtain the fused speech signal.

[0163] Step 830: Play the fused audio signal.

[0164] The fused voice signal is played through speakers in the cabin environment.

[0165] Schematic, Figure 9 is a schematic diagram of the process of predicting the sound location in the cockpit environment provided by an exemplary embodiment of this application. As shown in Figure 9, firstly, energy calculation 910 is performed on at least two speech estimation signals. Energy calculation 910 is used to determine the energy distribution in multiple sound regions of each of the at least two speech estimation signals. That is, the energy distribution of the speech estimation signal in multiple sound regions is calculated. After performing energy calculation 910, energy sorting 920 is performed on the multiple sound regions in the speech estimation signal. That is, the multiple sound regions are sorted according to the energy value from largest to smallest to obtain the sound region energy distribution corresponding to each speech estimation signal.

[0166] In addition, at least two speech estimation signals are mixed, and VAD930 is performed based on the mixed signals to determine the speech activity probability in the at least two speech estimation signals. The higher the probability, the more likely the at least two speech estimation signals include the subject's voice.

[0167] Based on the probability of speech activity and the energy distribution of sound regions, position recognition is performed 940 to predict the speech location of the subject within the cockpit environment. Weighting is then performed 950 based on the speech location to obtain the fusion weights for each microphone at that location, with microphones closer to the speech location receiving higher fusion weights. At least two estimated speech signals are then weighted and fused based on these weights 960 to obtain a fused speech signal, which is then output.

[0168] In summary, the method provided in this application acquires multiple audio signals and multiple reference signals, thereby generating a speech estimation signal based on the multiple audio signals and multiple reference signals, and playing speech audio based on the speech estimation signal. By expressing the audio acquired under different microphones and the influence of different microphones on the audio played by the speaker through multiple audio signals and multiple acquisition signals, echo interference is removed from the audio signal in a more refined and accurate manner, improving the playback accuracy of the acquired speech audio in the cockpit environment, and significantly improving the echo situation when acquiring and playing speech audio in the power amplifier scenario.

[0169] By predicting the location of the speech audio in the cockpit environment, the acquisition capabilities of different microphones can be determined. The microphone closer to the sound source obviously has a stronger ability to acquire speech audio. Therefore, a higher fusion weight is applied to the speech estimation signal on the corresponding path of the microphone, which improves the playback accuracy of speech audio.

[0170] Figure 10 is a structural block diagram of a voice interaction device in a cockpit environment provided by another exemplary embodiment of this application. As shown in Figure 10, the device includes: a acquisition module 1010, used to acquire at least two audio signals in the cockpit environment, wherein the at least two audio signals are signals containing voice audio acquired when the power amplifier is activated in the cockpit environment, wherein the i-th audio signal is acquired by the i-th microphone in the cockpit environment, and i is a positive integer; an acquisition module 1020, used to acquire at least two reference signals, wherein the i-th reference signal is used to represent the signal of the speaker amplifier corresponding to the i-th microphone; a processing module 1030, used to remove the i-th reference signal from the i-th audio signal to obtain the i-th voice estimation signal; and a playback module 1040, used to play the acquired voice audio based on the at least two voice estimation signals.

[0171] In an optional embodiment, as shown in FIG11, the processing module 1030 includes: a filtering unit 1031, used to filter the at least two reference signals based on a dynamically updated filter to obtain at least two echo estimation signals; a calculation unit 1032, used to remove the i-th echo estimation signal from the i-th audio signal to obtain the i-th speech estimation signal; the processing module 1030 further includes: an updating unit 1033, used to dynamically update the filter based on the at least two speech estimation signals.

[0172] In an optional embodiment, the computing unit 1032 is configured to perform a first preprocessing on the i-th audio signal to obtain the i-th desired signal; and to subtract the i-th echo estimation signal from the i-th desired signal to obtain a residual signal as the i-th speech estimation signal; wherein the first preprocessing includes at least one of DC removal processing and pre-emphasis processing, the DC removal processing being used to remove the DC component in the audio signal, and the pre-emphasis processing being used to compensate for the propagation loss of the speech audio.

[0173] In an optional embodiment, the computing unit 1032 is further configured to perform a second preprocessing on the at least two reference signals to obtain at least two preprocessed signals; the filtering unit 1031 is further configured to filter the at least two preprocessed signals based on a dynamically updated filter to obtain the at least two echo estimation signals; wherein, the second preprocessing includes pre-emphasis processing and time-frequency domain conversion processing, the pre-emphasis processing is used to compensate for the propagation loss of the loudspeaker power amplifier signal, and the time-frequency domain conversion processing is used to convert the reference signal in the time domain into a frequency domain signal.

[0174] In an optional embodiment, the filtering unit 1031 is further configured to filter the at least two preprocessed signals based on the dynamically updated filter to obtain at least two filtered signals; the calculation unit 1032 is further configured to perform time-frequency domain inverse processing on the at least two filtered signals to obtain the at least two echo estimation signals.

[0175] In an optional embodiment, the filtering unit 1031 is further configured to filter at least two reference signals in the (n+1)th round based on the (n+1)th filter dynamically updated in the (n)th round, to obtain at least two echo estimation signals in the (n+1)th round.

[0176] In an optional embodiment, the updating unit 1033 is further configured to update the filter coefficients of the filter based on the at least two speech estimation signals, the at least two reference signals, and the step size factor.

[0177] In an optional embodiment, the playback module 1040 is further configured to predict the sounding location of the speech audio in the cockpit environment based on at least two speech estimation signals, the at least two speech estimation signals including the i-th speech estimation signal; fuse the at least two speech estimation signals based on the sounding location to obtain a fused speech signal; and play the fused speech signal.

[0178] In an optional embodiment, the playback module 1040 is further configured to acquire the pitch energy distribution of the at least two speech estimation signals; and acquire the speech activity probability of the at least two speech estimation signals; and predict the sounding location of the speech audio in the cockpit environment based on the pitch energy distribution and the speech activity probability.

[0179] In an optional embodiment, the playback module 1040 is further configured to perform energy calculation on the i-th speech estimation signal in multiple sound zones to obtain energy values ​​corresponding to the multiple sound zones respectively; and sort the multiple sound zones according to the energy values ​​from largest to smallest to obtain the sound zone energy distribution corresponding to the i-th speech estimation signal.

[0180] In an optional embodiment, the playback module 1040 is further configured to obtain the fusion weights corresponding to the at least two speech estimation signals based on the distance between the sound source and the microphone; and to perform weighted fusion of the at least two speech estimation signals based on the fusion weights corresponding to the at least two speech estimation signals to obtain the fused speech signal.

[0181] In summary, the apparatus provided in this application acquires multiple audio signals and multiple reference signals, thereby generating a speech estimation signal based on the multiple audio signals and multiple reference signals, and playing speech audio based on the speech estimation signal. By expressing the audio acquired under different microphones and the influence of different microphones on the audio played by the speaker through multiple audio signals and multiple acquisition signals, echo interference is removed from the audio signal in a more refined and accurate manner, improving the playback accuracy of the acquired speech audio in the cockpit environment, and significantly improving the echo situation when acquiring and playing speech audio in the power amplifier scenario.

[0182] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the content structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0183] Please refer to Figure 12, which shows a structural block diagram of a computer device 1200 provided in one embodiment of this application. The computer device 1200 can be a terminal or server in the computer system shown in Figure 2. In this embodiment, the computer device 1200 is implemented as a vehicle-mounted terminal to implement the voice interaction method in the cockpit environment provided in the above embodiments. Specifically, the computer device 1200 typically includes a processor 1201 and a memory 1202.

[0184] Processor 1201 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1201 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). Processor 1201 may also include a main processor and a coprocessor. The main processor, also known as the CPU, is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1201 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1201 may also include an Artificial Intelligence (AI) processor, which is used to handle computational operations related to machine learning.

[0185] The memory 1202 may include one or more computer-readable storage media, which may be non-transitory. The memory 1202 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1202 are used to store a computer program configured to be executed by one or more processors to implement the above-described voice interaction method in a cockpit environment.

[0186] In some embodiments, the computer device 1200 may also optionally include other components 1203: a peripheral device interface and at least one peripheral device. The processor 1201, memory 1202, and peripheral device interface can be connected via a bus or signal lines. Each peripheral device can be connected to the peripheral device interface via a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of: radio frequency circuitry, a display screen, audio circuitry, and a power supply.

[0187] Those skilled in the art will understand that the structure shown in FIG12 does not constitute a limitation on the computer device 1200, and may include more or fewer components than shown, or combine certain components, or employ different component arrangements.

[0188] In an exemplary embodiment, a computer-readable storage medium is also provided, wherein a computer program is stored in the storage medium, and the computer program, when executed by a processor, implements the voice interaction method in the cockpit environment described above. Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. The random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM).

[0189] In an exemplary embodiment, a computer program product is also provided, the computer program product including a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the computer device to perform the above-described voice interaction method in a cockpit environment.

[0190] It should be noted that the collection and processing of relevant data (such as account information, list data, etc.) in this application should strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0191] It should be understood that "multiple" as used in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0192] Furthermore, the step numbers described herein are merely illustrative of one possible execution order between steps. In some other embodiments, the steps may not be executed in the order of their numbers, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.

[0193] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A voice interaction method in a cockpit environment, characterized in that, The method includes: acquiring at least two audio signals in the cockpit environment, wherein the at least two audio signals are signals containing speech audio acquired under the condition of power amplification in the cockpit environment, wherein the i-th audio signal is acquired by the i-th microphone in the cockpit environment, and i is a positive integer; acquiring at least two reference signals, wherein the i-th reference signal is used to represent the signal of the speaker power amplification corresponding to the i-th microphone; removing the i-th reference signal from the i-th audio signal to obtain the i-th speech estimation signal; and playing the acquired speech audio based on the at least two speech estimation signals.

2. The method according to claim 1, characterized in that, The step of removing the i-th reference signal from the i-th audio signal to obtain the i-th speech estimation signal includes: filtering the at least two reference signals based on a dynamically updated filter to obtain at least two echo estimation signals; removing the i-th echo estimation signal from the i-th audio signal to obtain the i-th speech estimation signal; the method further includes: dynamically updating the filter based on the at least two speech estimation signals.

3. The method according to claim 2, characterized in that, The step of removing the i-th echo estimation signal from the i-th audio signal to obtain the i-th speech estimation signal includes: performing a first preprocessing on the i-th audio signal to obtain an i-th desired signal; subtracting the i-th echo estimation signal from the i-th desired signal to obtain a residual signal as the i-th speech estimation signal; wherein the first preprocessing includes at least one of DC removal processing and pre-emphasis processing, the DC removal processing being used to remove the DC component in the audio signal, and the pre-emphasis processing being used to compensate for the propagation loss of the speech audio.

4. The method according to claim 2, characterized in that, The filtering of the at least two reference signals by the dynamically updated filter to obtain at least two echo estimation signals includes: performing a second preprocessing on the at least two reference signals to obtain at least two preprocessed signals; filtering the at least two preprocessed signals by the dynamically updated filter to obtain the at least two echo estimation signals; wherein the second preprocessing includes pre-emphasis processing and time-frequency domain conversion processing, the pre-emphasis processing is used to compensate for the propagation loss of the loudspeaker power amplifier signal, and the time-frequency domain conversion processing is used to convert the reference signal in the time domain into a frequency domain signal.

5. The method according to claim 4, characterized in that, The step of filtering the at least two preprocessed signals using the dynamically updated filter to obtain the at least two echo estimation signals includes: filtering the at least two preprocessed signals using the dynamically updated filter to obtain at least two filtered signals; and performing time-frequency domain inverse processing on the at least two filtered signals to obtain the at least two echo estimation signals.

6. The method according to claim 2, characterized in that, The filter obtained based on dynamic updates filters the at least two reference signals to obtain at least two echo estimation signals, including: the filter obtained based on the (n+1)th round dynamically updated to filter at least two reference signals in the (n+1)th round to obtain at least two echo estimation signals in the (n+1)th round.

7. The method according to claim 2, characterized in that, The step of dynamically updating the filter based on at least two speech estimation signals includes updating the filter coefficients based on the at least two speech estimation signals, the at least two reference signals, and a step size factor.

8. The method according to any one of claims 1 to 3, characterized in that, The step of acquiring at least two reference signals includes: determining the signal fusion function corresponding to the microphone, wherein the signal fusion function is used to express the fusion method of multiple speaker return signals; acquiring the signals returned by multiple speakers, and fusing the signals returned by multiple speakers according to the signal fusion function corresponding to the i-th microphone to obtain the i-th reference signal corresponding to the i-th microphone.

9. The method according to any one of claims 1 to 3, characterized in that, Playing the acquired speech audio based on at least two speech estimation signals includes: predicting the sounding location of the speech audio in the cockpit environment based on at least two speech estimation signals, wherein the at least two speech estimation signals include the i-th speech estimation signal; fusing the at least two speech estimation signals based on the sounding location to obtain a fused speech signal; and playing the fused speech signal.

10. The method according to claim 9, characterized in that, The method of predicting the sound location of the speech audio in the cockpit environment based on at least two speech estimation signals includes: obtaining the pitch energy distribution of the at least two speech estimation signals; and obtaining the speech activity probability of the at least two speech estimation signals; and predicting the sound location of the speech audio in the cockpit environment based on the pitch energy distribution and the speech activity probability.

11. The method according to claim 10, characterized in that, The step of obtaining the region energy distribution of the at least two speech estimation signals includes: performing energy calculation on the i-th speech estimation signal in multiple regions to obtain energy values ​​corresponding to the multiple regions respectively; sorting the multiple regions according to the energy values ​​from largest to smallest to obtain the region energy distribution corresponding to the i-th speech estimation signal.

12. The method according to claim 9, characterized in that, The step of fusing the at least two speech estimation signals based on the sound source location to obtain a fused speech signal includes: obtaining the fusion weights corresponding to the at least two speech estimation signals based on the distance between the sound source location and the microphone; and weighting and fusing the at least two speech estimation signals based on the fusion weights corresponding to the at least two speech estimation signals to obtain the fused speech signal.

13. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the voice interaction method in a cockpit environment as described in any one of claims 1 to 12.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the voice interaction method in a cockpit environment as described in any one of claims 1 to 12.

15. A computer program product, characterized in that, The computer program product includes a computer program stored in a computer-readable storage medium, and a processor reads from and executes the computer program to implement the voice interaction method in a cockpit environment as described in any one of claims 1 to 12.