Display device and echo cancellation method

By employing audio mixing technology in the multi-channel system of the display device, audio signals are directly acquired from the output of the power amplifier unit and echo cancellation is performed, solving the high cost problem caused by external voice chips and improving the accuracy of echo cancellation and voice interaction performance.

CN121728307APending Publication Date: 2026-03-24HISENSE VISUAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-23
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing display devices require an external voice chip to eliminate echo, resulting in higher costs.

Method used

Display devices using multi-channel systems perform audio mixing through a power amplifier unit, directly acquiring audio signals from the output end, and combining this with a controller for echo cancellation, thus avoiding the need for an external voice chip.

Benefits of technology

It reduces the cost of echo cancellation, improves the accuracy of signal acquisition and echo cancellation, ensures voice interaction performance, and achieves a balance between cost and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121728307A_ABST
    Figure CN121728307A_ABST
Patent Text Reader

Abstract

The invention relates to a display device and an echo cancellation method. The equipment comprises a power amplification unit which is configured to perform power amplification on original audios of a plurality of system sound channels of the display equipment and output respective audio signals of each system sound channel; the plurality of system sound channels are at least three system sound channels; the loudspeaker is configured to play sound signals corresponding to the audio signals; a microphone configured to receive an echo signal of the sound signal; the controller is configured to perform sound mixing processing on the audio signals of the system sound channels based on a preset number of sound mixing sound channels to obtain sound mixing signals; the preset number of sound mixing channels is determined based on the equipment performance of the display equipment; and performing echo cancellation on the echo signal according to the sound mixing signal. According to the invention, the echo cancellation cost can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of display device technology, and more particularly to a display device and an echo cancellation method. Background Technology

[0002] With the rapid development of speech recognition technology, far-field voice interaction is becoming increasingly common. For example, smart TVs and smart speakers utilize microphone array technology to increase the distance of human-computer interaction and achieve far-field voice interaction. In far-field speech recognition scenarios, while the microphone picks up the user's voice commands, it will inevitably also pick up the echoes caused by the audio played by the display device.

[0003] In the implementation of microphone echo cancellation, the display device usually uses one or two external voice chips to collect microphone signals and echo signals, thereby completing the corresponding echo cancellation and speech recognition. This undoubtedly brings cost issues.

[0004] Therefore, the current method of using an external voice chip on the display device for echo cancellation has the disadvantage of high cost. Summary of the Invention

[0005] This application provides a display device and an echo cancellation method to solve the problem of high cost of traditional technologies.

[0006] In a first aspect, some embodiments provide a display device, including: a power amplifier unit configured to amplify the power of raw audio from a plurality of system channels of the display device and output an audio signal for each of the system channels; the plurality of system channels are at least three system channels;

[0007] A speaker is configured to play sound signals corresponding to each of the said audio signals;

[0008] A microphone is configured to receive the echo signal of the sound signal;

[0009] The controller is configured as follows:

[0010] Based on a preset number of mixing channels, the audio signals of each system channel are mixed to obtain a mixed signal; the preset number of mixing channels is determined based on the device performance of the display device.

[0011] Based on the mixed signal, echo cancellation is performed on the echo signal.

[0012] The aforementioned display device acquires audio signals from each system channel at the output of the power amplifier unit. This eliminates the need for a separate external voice chip on the display device for echo cancellation, effectively reducing audio signal acquisition costs and consequently echo cancellation costs. Furthermore, the audio signal output by the power amplifier unit is closest to the sound signal played by the speakers, ensuring accurate signal acquisition and consequently accurate echo cancellation. It's understandable that as the number of system channels in the display device increases, the number of audio signal acquisition channels also increases, which would undoubtedly put some pressure on the display device's resources, affecting the echo cancellation effect and consequently the voice interaction performance. Therefore, the aforementioned display device employs a mixing processing technique. Based on a preset number of mixing channels, the audio signals from each system channel are mixed to obtain a mixed signal. The preset number of mixing channels is determined in advance based on the device's performance. This ensures that even in a multi-channel system, the display device can still effectively cancel echo signals, improving the accuracy of echo cancellation and thus improving the voice interaction experience, achieving a balance between echo cancellation costs and voice interaction performance.

[0013] Secondly, some embodiments also provide an echo cancellation method applied to the display device provided in the first aspect, the method comprising:

[0014] The original audio from multiple system channels of the display device is amplified by power, and each system channel is output as its own audio signal; the multiple system channels are at least three system channels.

[0015] Play the sound signals corresponding to each of the aforementioned audio signals;

[0016] The echo signal received from the sound signal;

[0017] Based on a preset number of mixing channels, the audio signals of each system channel are mixed to obtain a mixed signal; the preset number of mixing channels is determined based on the device performance of the display device.

[0018] Based on the mixed signal, echo cancellation is performed on the echo signal.

[0019] The aforementioned echo cancellation method acquires audio signals from each system channel at the output of the power amplifier unit. This eliminates the need for a separate external voice chip on the display device, effectively reducing audio signal acquisition costs and consequently echo cancellation costs. Furthermore, the audio signal output by the power amplifier unit is closest to the sound signal played by the speakers, ensuring accurate signal acquisition and consequently accurate echo cancellation. It's understandable that as the number of system channels in the display device increases, the number of audio signal acquisition channels also increases, which undoubtedly puts pressure on the display device's resources, affecting the echo cancellation effect and consequently the voice interaction performance. Therefore, the aforementioned display device employs a mixing processing technique. Based on a preset number of mixing channels, the audio signals from each system channel are mixed to obtain a mixed signal. The preset number of mixing channels is determined in advance based on the display device's performance. This ensures that even in a multi-channel system, the display device can still effectively cancel echo signals, improving the accuracy of echo cancellation and thus enhancing the voice interaction experience, achieving a balance between echo cancellation costs and voice interaction performance.

[0020] Thirdly, some embodiments also provide an echo cancellation device, the device comprising:

[0021] An audio signal output module is used to amplify the power of the original audio from multiple system channels of the display device and output the audio signal for each of the multiple system channels; the multiple system channels are at least three system channels.

[0022] A sound signal playback module is used to play the sound signals corresponding to each of the aforementioned audio signals;

[0023] An echo signal receiving module is used to receive the echo signal of the sound signal;

[0024] The audio mixing module is used to mix the audio signals of each system channel based on a preset number of mixing channels to obtain a mixed signal; the preset number of mixing channels is determined based on the device performance of the display device.

[0025] An echo cancellation module is used to cancel the echo signal based on the mixed signal.

[0026] The aforementioned echo cancellation device acquires audio signals from each system channel at the output of the power amplifier unit. This eliminates the need for a separate external voice chip on the display device, effectively reducing audio signal acquisition costs and consequently echo cancellation costs. Furthermore, the audio signal output by the power amplifier unit is closest to the sound signal played by the speakers, ensuring accurate signal acquisition and consequently accurate echo cancellation. It's understandable that as the number of system channels in the display device increases, the number of audio signal acquisition channels also increases, which undoubtedly puts pressure on the display device's resources, affecting the echo cancellation effect and consequently the voice interaction performance. Therefore, the aforementioned display device employs a mixing processing technique. Based on a preset number of mixing channels, the audio signals from each system channel are mixed to obtain a mixed signal. The preset number of mixing channels is determined in advance based on the display device's performance. This ensures that even in a multi-channel system, the display device can still effectively cancel echo signals, improving the accuracy of echo cancellation and thus enhancing the voice interaction experience, achieving a balance between echo cancellation costs and voice interaction performance.

[0027] Fourthly, some embodiments also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the echo cancellation method provided in the second aspect.

[0028] The aforementioned computer-readable storage medium, by acquiring audio signals from each system channel at the output of the power amplifier unit, eliminates the need for a separate external voice chip on the display device to complete the echo signal acquisition, thereby effectively reducing the cost of audio signal acquisition and echo cancellation. Furthermore, the audio signal output by the power amplifier unit is closest to the sound signal played by the speakers, ensuring the accuracy of signal acquisition and consequently the accuracy of subsequent echo cancellation. It is understandable that as the number of system channels in the display device increases, the number of audio signal acquisition channels also increases, which undoubtedly puts pressure on the display device's resources, affecting the echo cancellation effect and consequently the voice interaction performance. Therefore, the aforementioned display device employs a mixing processing technique, that is, based on a preset mixing channel number, mixing the audio signals of each system channel to obtain a mixed signal. The preset mixing channel number is determined in advance based on the device's performance, ensuring that even in a multi-channel system, the display device can still effectively cancel echo signals, improving the accuracy of echo cancellation and thus enhancing the voice interaction experience, achieving a balance between echo cancellation cost and voice interaction performance.

[0029] Fifthly, some embodiments also provide a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the echo cancellation method provided in the second aspect.

[0030] The aforementioned computer program product, by acquiring audio signals from each system channel at the output of the power amplifier unit, eliminates the need for a separate external voice chip on the display device to complete the echo signal acquisition, thereby effectively reducing the cost of audio signal acquisition and echo cancellation. Furthermore, the audio signal output by the power amplifier unit is closest to the sound signal played by the speakers, ensuring the accuracy of signal acquisition and consequently the accuracy of subsequent echo cancellation. It is understandable that as the number of system channels in the display device increases, the number of audio signal acquisition channels also increases, which undoubtedly puts pressure on the display device's resources, affecting the echo cancellation effect and consequently the voice interaction performance. Therefore, the aforementioned display device employs a mixing processing technique, that is, based on a preset mixing channel number, mixing the audio signals of each system channel to obtain a mixed signal. The preset mixing channel number is determined in advance based on the device's performance, ensuring that even in a multi-channel system, the display device can still effectively cancel echo signals, improving the accuracy of echo cancellation and thus enhancing the voice interaction experience, achieving a balance between echo cancellation cost and voice interaction performance. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in some embodiments of this application;

[0033] Figure 2 This is a schematic diagram of the hardware configuration of a display device provided in some embodiments of this application;

[0034] Figure 3 This is a schematic diagram of the hardware configuration of the control device provided in some embodiments of this application;

[0035] Figure 4 This is a schematic diagram of the software configuration of a display device provided in some embodiments of this application;

[0036] Figure 5 This is a schematic diagram illustrating the process of a display device performing far-field voice interaction in one embodiment of this application;

[0037] Figure 6 This is a schematic diagram illustrating the principle of a display device for eliminating microphone echo signals in one embodiment of this application;

[0038] Figure 7 This is a schematic diagram illustrating the principle of signal acquisition using an external voice chip in one embodiment of this application;

[0039] Figure 8 This is a schematic diagram illustrating the principle of signal acquisition using two external voice chips in one embodiment of this application.

[0040] Figure 9 This is a schematic diagram of the functional module architecture of the display device for echo cancellation in one embodiment of this application;

[0041] Figure 10 This is a schematic diagram of the power amplifier unit in one embodiment of this application;

[0042] Figure 11 This is a flowchart illustrating the voice interaction process of the display device in a specific embodiment of this application;

[0043] Figure 12 This is a flowchart illustrating an echo cancellation method in one embodiment of this application;

[0044] Figure 13 This is a flowchart illustrating the module interaction of an echo cancellation device in one embodiment of this application. Detailed Implementation

[0045] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.

[0046] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0047] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0048] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0049] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0050] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in some embodiments of this application. For example... Figure 1 As shown, users can operate the display device 200 via touch operation, mobile terminal 300, and control device 100. For example, control device 100 can be a remote control, stylus, gamepad, etc.

[0051] In this embodiment, the display device 200 generally refers to a device with screen display and data processing capabilities. For example, the display device 200 includes, but is not limited to, smart TVs, computers, tablets, wearable devices, virtual reality devices, augmented reality devices, etc.

[0052] The mobile terminal 300 can function as a control device for human-computer interaction between the user and the display device 200. It can also function as a communication device for establishing a communication connection with the display device 200 and exchanging data. In some embodiments, the mobile terminal 300 can have software applications installed on it and communicate with the display device 200 via network communication protocols to achieve one-to-one control and data communication. Furthermore, it can transmit audio and video content displayed on the mobile terminal 300 to the display device 200 for synchronized display.

[0053] like Figure 1 The diagram also shows that the display device 200 communicates with the server 400 via various communication methods. The display device 200 can communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks. The server 400 can provide various content and interactive features to the display device 200. The server 400 can be a cluster or multiple clusters, and may include one or more types of servers.

[0054] Display device 200 can provide broadcast television reception function, and can also be equipped with intelligent network television function that provides computer support function, including but not limited to network television, smart television, Internet Protocol television (IPTV), etc.

[0055] Figure 2 Provided for some embodiments of this application Figure 1 Hardware configuration block diagram of display device 200.

[0056] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.

[0057] In some embodiments, detector 230 is used to acquire signals from the external environment or to interact with the outside world. For example, detector 230 includes a light receiver, a sensor for acquiring ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to acquire external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.

[0058] In some embodiments, the display 260 includes display function components for presenting images and driving components for driving image display. The display 260 is used to receive and display image signals output from the controller 250. For example, the display 260 can be used to display video content, image content, menu control interface components, and user control UI interfaces, etc.

[0059] In some embodiments, the communication device 220 is a component used to communicate with external devices or the server 400 according to various communication protocol types. The display device 200 may have multiple communication devices 220 depending on the supported communication methods. For example, when the display device 200 supports wireless network communication, it may have a communication device 220 with WiFi functionality. When the display device 200 supports Bluetooth connectivity, it needs to have a communication device 220 with Bluetooth functionality.

[0060] The communication device 220 enables the display device 200 to communicate with external devices or the server 400 via wireless or wired connections. Wired connections utilize data cables, interfaces, or other components to connect the display device 200 to external devices. Wireless connections utilize wireless signals or wireless networks. The display device 200 can directly establish a connection with external devices or indirectly through gateways, routers, or other connection devices.

[0061] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first to an nth interface for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory. The controller 250 controls the overall operation of the display device 200.

[0062] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0063] In some embodiments, a user can input user commands through a graphical user interface (GUI) displayed on a display 260, and the user input interface receives user input commands through the graphical user interface (GUI).

[0064] In some embodiments, the audio output device 270 can be a built-in speaker of the display device 200 or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may also be provided with an external audio output terminal, through which the audio output device can be connected to the display device 200 to output sound from the display device 200.

[0065] In some embodiments, the user input interface 280 can be used to receive instructions from user input.

[0066] Figure 3 Provided for some embodiments of this application Figure 1 Hardware configuration block diagram of the central control device 100. (See diagram below.) Figure 3 As shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.

[0067] The control device 100 is configured to control the display device 200, and to receive user input operation commands and convert the operation commands into commands that the display device 200 can recognize and respond to, thus acting as an intermediary for interaction between the user and the display device 200.

[0068] In some embodiments, the control device 100 may be an intelligent device. For example, the control device 100 may be equipped with various applications for controlling the display device 200 according to user needs.

[0069] In some embodiments, such as Figure 1As shown, the mobile terminal 300 or other smart electronic devices can perform similar functions to the control device 100 after installing the application of the control display device 200.

[0070] The controller 110 includes a processor 112, RAM 113 and ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation of the control device 100, as well as the communication and cooperation between internal components and the external and internal data processing functions.

[0071] Under the control of the controller 110, the communication interface 130 enables communication of control signals and data signals with the display device 200. The communication interface 130 may include at least one of other near-field communication modules such as WiFi chip 131, Bluetooth module 132, and NFC module 133.

[0072] User input / output interface 140, wherein the input interface includes at least one of other input interfaces such as microphone 141, touchpad 142, sensor 143, and button 144.

[0073] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The control device 100 is configured with the communication interface 130, such as a WiFi, Bluetooth, or NFC module, which can encode user input commands via WiFi, Bluetooth, or NFC protocols and send them to the display device 200.

[0074] The memory 190 is used to store various operating programs, data, and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can also store various control signal instructions input by the user.

[0075] The power supply 180 is used to provide operating power support for the various components of the control device 100 under the control of the controller.

[0076] In order to perform user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program used to manage and control the hardware and software resources in the display device 200. The operating system can (control the display device) provide a user interface, allowing users to interact with the display device 200 and supporting the running of various applications.

[0077] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system that is deeply customized based on a specific operating platform, or an independent operating system specifically developed for display devices.

[0078] An operating system can be divided into different modules or levels based on the functions it implements, for example... Figure 4 As shown, in some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the System Library layer, and the Kernel layer.

[0079] In some embodiments, the application layer provides services and interfaces for applications, enabling the display device 200 to run applications and interact with the user based on the applications. The application layer may contain at least one application, which may be a built-in Windows program, system settings program, or clock program of the operating system; or it may be an application developed by a third-party developer. In specific implementations, the application packages in the application layer are not limited to the examples above.

[0080] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.

[0081] like Figure 4 As shown, the application framework layer in this embodiment includes a view system, managers, and content providers. The view system designs and implements the application's interface and interactions, and includes lists, grids, text boxes, and buttons. The managers include at least one of the following modules: an activity manager for interacting with all running activities in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0082] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling changes to the display window, such as shrinking the display window, shaking the display, or distorting the display.

[0083] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library contained in the system runtime library layer, such as the C / C++ instruction library, to implement the functions to be performed by the framework layer.

[0084] In some embodiments, the kernel layer is a functional layer situated between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, ... Figure 4 As shown, hardware drivers can be configured in the kernel layer. The kernel layer can contain at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0085] It should be noted that the above examples are merely a simple division of operating system functions and do not limit the specific form of the operating system of the display device 200 in this application embodiment. Depending on the function of the display device, the type of operating system, and other factors, the number of levels and the specific level type of the operating system may be expressed in other forms.

[0086] Figure 5 A schematic diagram of a display device performing far-field voice interaction is shown in one embodiment of this application.

[0087] The display device uses a built-in microphone array and audio codec to acquire microphone and speaker signals. First, high-pass filtering removes low-frequency interference from the circuitry. Then, the signal processing section includes echo cancellation, beamforming, dereverberation, speech activity detection, and noise suppression to obtain the processed signal. Next, dynamic gain adjustment is applied to the processed signal. When active speech is detected, it is transmitted to the speech engine for speech recognition. Finally, semantic understanding and voice playback technologies are combined to achieve far-field voice interaction.

[0088] Figure 6 A schematic diagram illustrating the principle of a display device for eliminating microphone echo signals is shown in one embodiment of this application.

[0089] In far-field voice interaction scenarios, echo is often one of the important factors affecting the voice interaction effect. For example, when a display device plays a song or video, it generates signal B. The microphone will simultaneously pick up the sound emitted by the user, i.e., signal A and signal B. The purpose of echo cancellation is to accurately distinguish the difference between these two sounds. For example, by using an AEC (Acoustic Echo Cancellation) echo cancellation component, signal B is eliminated, and only signal A is retained, thereby completing effective voice interaction.

[0090] Most display devices are equipped with multi-channel systems to enhance sound effects. Because the signals in a multi-channel system are not absolutely correlated—that is, the audio signals of each channel are independent and do not affect each other—it is necessary to acquire the audio signal of each channel for echo cancellation. Currently, it is common practice to use one or two external voice chips, such as DSP (Digital Signal Processing) chips, to complete the echo acquisition of the multi-channel system. The acquired signals are transmitted to the voice chip to perform echo cancellation, beamforming, dereverberation, voice activity detection, noise suppression, and other processing. Figure 7 As shown, Figure 7 This illustration shows a schematic diagram of the principle of signal acquisition using an external voice chip connected to a display device in one embodiment of this application. A DSP chip, namely a CSK4002 single chip, is connected to the display device. The CSK4002 voice chip, in conjunction with two ADCs (analog-to-digital converters), acquires audio signals from the multi-channel display device as echo cancellation feedback signals. The ADCs can be ES7210 four-channel ADCs, and two ADCs can achieve up to eight channels of feedback signal acquisition. A microphone array can be directly connected to the CSK4002 chip via a PDM (Pulse Density Modulation) interface to send the acquired microphone signals to the CSK4002 chip. Thus, the microphone signals and multi-channel feedback signals undergo echo cancellation, beamforming, dereverberation, noise reduction, and far-field voice wake-up within the CSK4002 chip, completing the far-field voice interaction process for the display device in a multi-channel system.

[0091] Further references are available. Figure 8 , Figure 8This illustration shows a schematic diagram of the principle of signal acquisition using two external voice chips in a display device according to an embodiment of this application. Two RK3308G chips are connected in series, labeled DSP1 and DSP2 respectively. DSP1 is responsible for acquiring microphone signals and audio signals from the multi-channel system, as well as performing echo cancellation, while DSP2 is responsible for signal noise reduction and far-field speech recognition.

[0092] It can be seen that whether using one or two external DSP chips, the hardware system design is relatively complex. In addition, the required auxiliary components such as ADCs naturally increase the hardware cost of the solution. Therefore, the current method of using external voice chips to collect the echo signal from the display device in a multi-channel system for echo cancellation has the disadvantage of high cost.

[0093] To solve the above problems, Figure 9 This illustration shows a schematic diagram of the functional module architecture for echo cancellation in a display device according to an embodiment of this application. The display device provided in this application includes a power amplifier unit 910, a speaker 920, a microphone 930, and a controller 940. The controller 940 may house the main chip (System on Chip) of the display device.

[0094] The power amplifier unit 910 is configured to amplify the raw audio from multiple system channels of the display device and output an audio signal for each system channel. The speaker 920 is configured to play the corresponding sound signal for each audio signal. The microphone 930 is configured to receive the echo signal of the sound signal. The controller 940 is configured to mix the audio signals of each system channel based on a preset number of mixing channels to obtain a mixed signal, and to perform echo cancellation on the echo signal according to the mixed signal.

[0095] In this embodiment, the power amplifier unit can amplify the raw audio from multiple system channels of the display device, thereby outputting the audio signal for each system channel. Here, a system channel refers to the channels configured in the display device; multiple system channels refer to at least three system channels, such as 2.1 channels, 2.1.2 channels, or 3.1.2 channels. It is understood that acquiring the audio signal of a 2.0 channel (stereo channel) typically does not place a significant resource burden on the display device; therefore, this embodiment addresses display devices configured with at least three system channels. Raw audio can refer to audio data without any audio processing. The audio signal refers to the signal after the raw audio has been processed by the power amplifier unit.

[0096] Optionally, the power amplifier unit may include multiple power amplifiers, each capable of amplifying power for different channels, such as... Figure 10As shown, taking a 2.1.2 channel as an example, a 2.1.2 channel includes two main channels, namely the left and right channels, one bass channel and two sky channels. The power amplifier unit can include power amplifier 1, power amplifier 2 and power amplifier 3. Power amplifier 1 is responsible for amplifying the original audio of the two main channels, power amplifier 2 is responsible for amplifying the original audio of the bass channel, and power amplifier 3 is responsible for amplifying the original audio of the two sky channels.

[0097] Specifically, the controller sends the raw audio from each system channel to the corresponding power amplifier via the I2S OUT (Integrated Interchip Sound OUT) interface. Each power amplifier amplifies the raw audio and outputs the audio signal for each system channel. Simultaneously, the controller acquires the audio signal output by each power amplifier via the I2S IN interface.

[0098] Understandably, directly acquiring audio signals from the output of the power amplifier eliminates the need for an external voice chip on the display device, reducing the cost of signal acquisition and thus the cost of echo cancellation.

[0099] A loudspeaker can play the audio signals corresponding to the audio signals of each system channel. Here, the audio signal can refer to the audio signal of the sound played by the loudspeaker.

[0100] Specifically, after each power amplifier completes the power amplification process of each original audio signal and obtains the audio signals of each system channel, it can directly drive the speakers to play the sound signals corresponding to each audio signal.

[0101] A microphone can receive echo signals of sound signals, as well as voice commands initiated by the user. An echo is an acoustic phenomenon; when a sound signal travels and encounters a large reflective surface such as a wall, it is reflected, forming a reflected sound wave, which is the echo signal. Voice commands are instructions issued by the user to wake up the display device and perform corresponding interactive operations.

[0102] Optionally, a four-microphone linear array, consisting of four digital microphones arranged at equal intervals in a straight line, can be deployed to receive the echo signal of the sound signal and the voice commands initiated by the user, thereby improving the quality of the voice commands and thus improving the accuracy of subsequent speech recognition.

[0103] Specifically, in far-field voice interaction scenarios, when a user initiates a voice command, the speaker of the display device is also playing a sound signal. At this time, the microphone will simultaneously collect the voice command and the echo signal of the sound signal, and send both the voice command and the echo signal to the controller. The controller will then perform echo cancellation on the echo signal based on the audio signals of each system channel collected, retaining only the voice command, thereby executing the interactive operation corresponding to the voice command.

[0104] Optionally, the microphone can directly send echo signals and voice commands to the controller via the PDM interface.

[0105] Understandably, while acquiring voice commands and echo signals through the PDM interface and acquiring audio signals from each system channel through the output of the power amplifier can effectively reduce the cost of echo cancellation, considering that the number of system channels on display devices will increase in practical applications, this will undoubtedly put some pressure on the display devices' resources. As shown in Table 1, Table 1 shows the average CPU (Central Processing Unit) of the display device when acquiring signals under different system channels. Taking three system channels, namely 2.1 channel, 2.0.2 channel, and 2.1.2 channel, as examples, under the same display device's main chip (MT9653 chip) and the same test scenario (display device boot-up home page, device volume 20), signal acquisition was carried out under different configurations, and the corresponding average CPU values ​​of the display devices are given. As can be seen, in 2.1 channel audio, using a configuration of 4 microphones + 3 feedback signals (i.e., acquiring 4 microphone signals and 3 feedback signals), the average CPU usage of the display device's main chip reached 39.6% during signal acquisition; in 2.0.2 channel audio, it reached 45.2%; and in 2.1.2 channel audio, it reached 51.2%. This means that as the number of channels increases, the average CPU usage also gradually increases. Since the CPU resources of the display device are limited, this can lead to stuttering or crashes on the display's homepage, affecting the user's voice interaction experience.

[0106] Table 1. Average CPU usage of display devices during signal acquisition under different system channels.

[0107]

[0108] Therefore, this embodiment introduces a mixing process. The controller mixes the audio signals acquired from the power amplifier's output to obtain a mixed signal. This mixed signal is then used to cancel the echo signal from the speaker, preserving the user-initiated voice commands. The mixing process involves combining multi-channel audio signals into a signal with a preset number of mixing channels. The preset number of mixing channels is determined based on the display device's performance and can be understood as the number of channels in the mixed signal. Device performance refers to the display device's ability to perform echo cancellation.

[0109] Specifically, the controller can perform performance analysis on the display device in advance, and then configure an appropriate number of mixing channels for the display device based on its performance. After acquiring each audio signal from the output of each power amplifier, the audio signals can be mixed based on the specified number of mixing channels, that is, the multi-channel audio signals are combined into an audio signal with the specified number of mixing channels. Based on the obtained mixed signal, echo cancellation is performed on the echo signal.

[0110] It should be noted that the mixing process can be performed after the controller receives the echo signal sent by the microphone, or after each power amplifier outputs its audio signal.

[0111] The above technical solution has the following advantages or beneficial effects: By acquiring the audio signals of each system channel from the output of the power amplifier unit, on the one hand, there is no need to separately add a voice chip to the display device to complete the echo signal acquisition, thereby effectively reducing the cost of audio signal acquisition and thus reducing the cost of echo cancellation. On the other hand, the audio signal output by the power amplifier unit is closest to the sound signal played by the speaker, which can ensure the accuracy of signal acquisition and thus ensure the accuracy of subsequent echo cancellation. It is understandable that as the number of system channels of the display device increases, the number of audio signal acquisition channels will also increase, which will undoubtedly put a certain resource pressure on the display device, thereby affecting the effect of echo cancellation and thus affecting the voice interaction performance. Therefore, the above display device adopts a mixing processing technology, that is, based on a preset mixing channel number, the audio signals of each system channel are mixed to obtain a mixed signal. The preset mixing channel number is determined in advance based on the device performance of the display device. This ensures that even in a multi-channel system, the display device can still effectively cancel the echo signal, improve the accuracy of echo cancellation, thereby improving the voice interaction experience and achieving a balance between echo cancellation cost and voice interaction performance.

[0112] In one embodiment, the controller is further configured to acquire the signal type of the audio signal.

[0113] The signal types include multi-channel signals and non-multi-channel signals. Multi-channel signals are sound signals that include at least three audio channels, while non-multi-channel signals are sound signals that include at most two audio channels. An audio channel is a channel of a sound signal.

[0114] It's important to note that a sound signal is not the same as an audio signal. In other words, the audio signal output by the power amplifier unit is not the same as the sound signal played by the speakers. This is because, in practical applications, the audio signal output by the power amplifier unit may undergo further audio processing to form a new audio signal before being transmitted to the speakers for playback. Taking a 2.1.2 channel system as an example, the power amplifier unit outputs a 2.1.2 channel audio signal, which includes the audio signals of two main channels, one subwoofer channel, and two sky channels. In practical applications, these audio signals can be merged to form a 2.0 channel signal, which only includes the two main channels, and then transmitted to the speakers for playback. In this case, the speakers will play the 2.0 channel audio signal. Therefore, this embodiment needs to determine the type of sound signal to select different mixing processing methods based on different types of sound signals.

[0115] Specifically, after receiving the echo signal of the sound signal sent by the microphone, the controller can first identify the signal type of the sound signal, so that different mixing processing methods can be used for different types of sound signals.

[0116] Therefore, in the process of mixing the audio signals of each system channel based on the preset number of mixing channels to obtain the mixed signal, the controller is further configured to: mix the audio signals of each system channel based on the preset number of mixing channels and signal type to obtain the mixed signal.

[0117] Specifically, when the audio signal type is a multi-channel signal, it means that the audio signals of each system channel of the display device are directly output as the corresponding multi-channel audio signals without undergoing audio processing other than the power amplifier unit. In this case, there is no correlation between the audio signals of each system channel, and mixing can be performed directly. However, when the audio signal type is a non-multi-channel signal, such as stereo or mono, it means that the audio signals of each system channel of the display device have undergone audio processing other than the power amplifier unit, such as audio blending, thus forming a non-multi-channel audio signal. In this case, it is necessary to further analyze whether there is a correlation between the audio signals of each system channel, and then perform mixing based on the results of the correlation analysis.

[0118] The above technical solution has the following advantages or beneficial effects: by acquiring the signal type of the sound signal, and performing different mixing processes based on the signal type, the accuracy of the mixing process can be improved, thereby improving the accuracy of subsequent echo cancellation based on the mixed signal, and thus improving the performance of voice interaction.

[0119] In one embodiment, the controller, in the process of mixing the audio signals of each system channel based on a preset number of mixing channels and signal type to obtain a mixed signal, is further configured to: determine the target channel for mixing and other system channels besides the target channel among multiple system channels based on the preset number of mixing channels; acquire the first audio signal of the target channel and the second audio signals of the other system channels; and mix the first audio signal and the second audio signal based on the signal type to obtain a mixed signal.

[0120] The target channel can refer to the system channel used for mixing. For example, taking system channel 2.1.2 (including 2 main channels, 1 subwoofer channel, and 2 sky channels) as an example, assuming the preset number of mixing channels is 2, then the target channels can be the 2 main channels. Assuming the preset number of mixing channels is 3, then the target channels can be the 2 main channels and 1 subwoofer channel. The first audio signal refers to the audio signal from the target channel, and the second audio signal refers to the audio signal from other system channels.

[0121] Of course, in practical applications, the target channel can be selected according to the actual situation, and there are no restrictions on the specific selection method for the target channel.

[0122] Specifically, the controller, based on a preset number of mixing channels, determines the target channel for mixing from multiple system channels, as well as the other system channels. It then acquires the first audio signal from the target channel and the second audio signals from the other system channels, and mixes the first and second audio signals according to their signal types to obtain a mixed signal. The mixing process may involve adding the second audio signal to the first audio signal.

[0123] The above technical solution has the following advantages or beneficial effects: By determining the target channel and other system channels for mixing, and then mixing the first audio signal of the target channel with the second audio signals of other system channels based on the signal type, a mixed signal is obtained. This improves the efficiency of the mixing process, thereby improving the efficiency of subsequent echo cancellation and ultimately enhancing voice interaction performance.

[0124] In one embodiment, the controller is further configured to, during the process of mixing the first audio signal and the second audio signal based on the signal type to obtain a mixed signal, add the second audio signal to the first audio signal when the audio signal is a multi-channel signal to obtain a mixed signal.

[0125] Specifically, when the audio signal is a multi-channel signal, it means that there is no correlation between the audio signals. Therefore, the controller can add the second audio signal of other system channels to the first audio signal of the target channel to complete the mixing process and obtain the mixed signal.

[0126] In practical applications, the second audio signal from other system channels can be added to the first audio signal of the target channel by adjusting the data format of the audio signal. For example, taking a display device with a 2.1.2 system channel configuration and a preset mixing channel count of 2 (i.e., the target channels are two main channels (left and right)), assuming the audio signal is a 2.1.2 channel signal, which is a multi-channel signal, this means that the audio signal output by the power amplifier unit directly forms the corresponding number of channel audio signals without undergoing any other audio processing. Therefore, there is no correlation between the audio signals. By adjusting the data format of the two main channel audio signals, the final two-channel mixed signal is obtained, namely channel 1 and channel 2. The data format of channel 1 is: channel number identifier + left channel + bass + sky tone left, and the data format of channel 2 is: channel number identifier + right channel + bass + sky tone right. The channel number flag is used to record the number of channels present in the mixed channel. In this example, the channel number flags for both channel 1 and channel 2 are 3.

[0127] Understandably, bass frequencies are not differentiated by left or right, so in order to ensure data format balance, both mix channels contain the audio signal of the bass channel.

[0128] The above technical solution has the following advantages or beneficial effects: When the sound signal is a multi-channel signal, a mixed signal is obtained by adding the second audio signal of other system channels to the first audio signal of the target channel. This ensures that the mixed signal contains audio data from all system channels, thereby guaranteeing the integrity and accuracy of the mixed signal, and consequently ensuring the accuracy of subsequent echo cancellation.

[0129] In one embodiment, during the process of mixing the first audio signal and the second audio signal based on the signal type to obtain a mixed signal, the controller is further configured to: when the audio signal is a non-multichannel signal, perform correlation analysis on the first audio signal and the second audio signal to obtain correlation analysis results; when the correlation analysis results indicate that the first audio signal does not contain the second audio signal, add the second audio signal to the first audio signal to obtain the mixed signal.

[0130] Correlation analysis refers to the process of analyzing the correlation between a first audio signal and a second audio signal. Specifically, the correlation refers to the inclusion relationship between the first and second audio signals, that is, whether the first audio signal contains the second audio signal. The results of the correlation analysis can include whether the first audio signal does not contain the second audio signal, or whether the first audio signal contains the second audio signal.

[0131] It is understandable that when the audio signal is not a multi-channel signal, it means that the audio signals output by the power amplifier unit have undergone other audio processing, such as audio fusion, to form a non-multi-channel audio signal. In other words, the target channel may contain audio signals from other system channels. Therefore, it is necessary to further analyze the correlation between the first audio signal of the target channel and the second audio signals of other system channels.

[0132] Specifically, when the audio signal is not a multi-channel signal, the controller needs to further analyze whether the first audio signal of the target channel contains the second audio signal of other system channels. If the first audio signal does not contain the second audio signal, it means that there is no correlation between the first audio signal and the second audio signal, that is, they are independent of each other. To ensure the integrity of the mixed signal, the second audio signal can be directly added to the first audio signal to obtain the mixed signal.

[0133] The above technical solution has the following advantages or beneficial effects: when the sound signal is not a multi-channel signal, by further analyzing the correlation between the audio signals of the system channels to perform corresponding mixing processing, the accuracy and efficiency of mixing processing can be improved, thereby improving the accuracy and efficiency of subsequent echo cancellation.

[0134] In one embodiment, the controller is further configured to: when the correlation analysis result indicates that the first audio signal contains the second audio signal, use the first audio signal as a mix signal and delete the second audio signal.

[0135] It is understandable that when the correlation analysis results indicate that the first audio signal contains the second audio signal, it means that the target channel already contains the audio signals of all system channels. Therefore, the first audio signal of the target channel can be directly used as the mixing signal, and the second audio signals of other system channels can be deleted.

[0136] The above technical solution has the following advantages or beneficial effects: when the first audio signal contains the second audio signal, it means that the target channel already contains the audio signals of all system channels. Directly retaining the first audio signal of the target channel and deleting the second audio signals of other system channels can improve the efficiency of mixing and avoid interference from repeated audio signals.

[0137] In one embodiment, during the process of performing correlation analysis on the first audio signal and the second audio signal to obtain the correlation analysis result, the controller is further configured to: execute a voice interaction task of voice command based on the first audio signal to obtain a first voice interaction result; the first voice interaction result includes a first voice wake-up rate; execute a voice interaction task of voice command based on each audio signal to obtain a second voice interaction result; the second voice interaction result includes a second voice wake-up rate; and perform correlation analysis on the first audio signal and the second audio signal based on the comparison result between the first voice wake-up rate and the second voice wake-up rate to obtain the correlation analysis result.

[0138] Here, the voice interaction task can refer to the interactive task indicated by a voice command. The first voice interaction result refers to the voice interaction result obtained by executing the voice interaction task based on the first audio signal of the target channel. The first voice wake-up rate refers to the probability that the display device successfully recognizes the voice command and responds based on the first audio signal. The second voice interaction result refers to the voice interaction result obtained by executing the voice interaction task based on the audio signals of all system channels. The second voice wake-up rate refers to the probability that the display device successfully recognizes the voice command and responds based on all audio signals. The comparison result can refer to the difference between the first voice wake-up rate and the second voice wake-up rate.

[0139] Optionally, the voice interaction results may also include, but are not limited to, voice recognition accuracy, response rate, voice false wake-up rate, and voice wake-up method.

[0140] Understandably, the purpose of correlation analysis is to analyze whether the target channel contains audio signals from other system channels. Therefore, in this embodiment, the first audio signal of the target channel is selected for echo cancellation to perform the corresponding voice interaction task and obtain the first voice wake-up rate. Simultaneously, echo cancellation is also performed using audio signals from all system channels to perform the corresponding voice interaction task and obtain the second voice wake-up rate. The correlation between the first and second audio signals is then analyzed based on the difference between the first and second voice wake-up rates.

[0141] Specifically, if the difference between the first voice wake-up rate and the second voice wake-up rate is less than a preset threshold, it indicates that the effect of echo cancellation using the audio signal of the target channel is consistent with the effect of echo cancellation using the audio signals of all system channels, and it can be considered that the target channel contains audio signals from other system channels. If the difference between the first voice wake-up rate and the second voice wake-up rate is not less than the preset threshold, it indicates that the effect of echo cancellation using the audio signal of the target channel is inconsistent with the effect of echo cancellation using the audio signals of all system channels, and it can be considered that the target channel does not contain audio signals from other system channels.

[0142] The preset threshold can be a standard value pre-defined to account for the difference between the first and second voice wake-up rates. In this embodiment, the preset threshold can be determined based on the channel frequency, relevant parameters of the EQ equalizer, and the feedback coefficient of the AEC echo cancellation component. Channel frequency refers to the frequency range of the audio signal. The EQ equalizer is used to gain or attenuate one or more frequency bands in the audio signal to adjust the timbre. Relevant parameters of the EQ equalizer can include the adjustment frequency point, gain parameters (used to adjust the gain or attenuation at the adjustment frequency point), and quantization parameters (used to set the width of the frequency band to be gained or attenuated). The feedback coefficient is a coefficient fed back by the AEC echo cancellation component after echo cancellation based on the first audio signal and the entire audio signal, respectively. Specifically, the preset threshold can be the product of the channel frequency, relevant parameters of the EQ equalizer, and the feedback coefficient of the AEC echo cancellation component. In this embodiment, the preset threshold can be set to 90%.

[0143] The above technical solution has the following advantages or beneficial effects: Echo cancellation and voice interaction are performed on the first audio signal and each of the other audio signals respectively, thereby comparing the differences in the two voice interaction results (i.e., the voice wake-up rate) to analyze the correlation between the first and second audio signals. This improves the accuracy of the correlation analysis results, thus improving the accuracy of the audio mixing process.

[0144] In one embodiment, the controller is further configured to: decode the audio signal to obtain the number of audio channels of the audio signal; and determine the signal type of the audio signal based on the number of audio channels.

[0145] The number of audio channels refers to the number of audio channels in a sound signal.

[0146] Specifically, after receiving a sound signal, the controller can use audio decoding tools or audio decoding technology to decode the sound signal, obtain the number of audio channels of the sound signal, and then determine the signal type of the sound signal based on the number of audio channels.

[0147] For example, when the number of audio channels in the decoded sound signal is greater than or equal to 3, the sound signal is considered to be a multi-channel signal; when the number of audio channels in the decoded sound signal is less than or equal to 2, the sound signal is considered to be a non-multi-channel signal.

[0148] The above technical solution has the following advantages or beneficial effects: by obtaining the number of audio channels of the sound signal through audio decoding technology, the signal type of the sound signal can be determined, thereby improving the accuracy of the sound signal type and thus improving the accuracy of the mixing process.

[0149] In one embodiment, the controller is further configured to: acquire device configuration parameters of the display device; perform performance analysis on the display device based on the device configuration parameters to obtain the performance analysis results of the display device; and determine the preset number of mixing channels that matches the performance analysis results.

[0150] The device configuration parameters can refer to the basic configuration parameters of the display device, including but not limited to processor parameters, memory, sound mode, sound enhancement parameters, and audio system parameters. The performance analysis results are used to characterize the display device's ability to perform echo cancellation. In this embodiment, the performance analysis results can be the CPU limit value for the display device to perform echo cancellation.

[0151] Specifically, the controller can obtain the basic configuration parameters of the display device in advance, and then perform performance analysis on the display device based on the basic configuration parameters to obtain the CPU limit value for echo cancellation of the display device, and then determine the number of mixing channels that match the CPU limit value.

[0152] For example, assuming that after performance analysis of the display device, the CPU limit for echo cancellation is 40%, and according to Table 1, the display device can only achieve signal acquisition for 2.1 system channels, so the number of mixing channels can be set to 2.

[0153] The above technical solution has the following advantages or beneficial effects: By analyzing the performance of the display device, the limit value for echo cancellation can be obtained, thereby determining the number of mixing channels that match this limit value for subsequent mixing processing. In this way, even in a multi-channel system, the display device can complete echo cancellation within its capabilities, ensuring sufficient resources and avoiding stuttering or crashes due to insufficient resources, thus improving the voice interaction experience.

[0154] In one specific embodiment, the display device is configured with at least three system channels and includes a power amplifier unit, speakers, a microphone, and a controller. The power amplifier unit includes multiple power amplifiers, and the controller houses the display device's SOC chip.

[0155] Figure 11 A flowchart illustrating the voice interaction process of a display device in a specific embodiment of this application is shown.

[0156] Specifically, the controller pre-acquires the device configuration parameters of the display device, performs performance analysis on the display device based on the device configuration parameters, obtains the performance analysis results of the display device, and then determines the preset number of mixing channels that matches the performance analysis results.

[0157] The controller sends the raw audio for each system channel to each power amplifier via the I2S OUT interface. Each power amplifier amplifies the raw audio and outputs its own audio signal for each system channel, which is then played by the speakers. Simultaneously, the controller acquires the audio signals output by each power amplifier via the I2S IN interface. Based on a preset number of mixing channels, the controller determines the target channel for mixing and the other system channels from among the multiple system channels. It then acquires the first audio signal for the target channel and the second audio signals for the other system channels.

[0158] The microphone array, composed of microphones, can first determine whether the display device is playing an audio signal when it receives a voice command from the user. If not, it can directly perform beamforming, dereverberation, noise suppression, and voice wake-up operations on the voice command. If the audio signal is playing, it can simultaneously acquire the echo signal of the audio signal and send the voice command and echo signal to the controller through the PDM interface.

[0159] The controller decodes the sound signal to obtain the number of audio channels, thereby determining the signal type of the sound signal, i.e., whether the sound signal belongs to a multi-channel signal. If it belongs to a multi-channel signal, it performs mixing processing, that is, adding the second audio signal to the first audio signal to obtain the mixed signal.

[0160] When the audio signal is not a multi-channel signal, the audio signal correlation judgment module is entered. Based on the first audio signal, a voice interaction task using voice commands is executed to obtain a first voice wake-up rate. Based on each audio signal, a voice interaction task using voice commands is executed to obtain a second voice wake-up rate. Based on the comparison between the first and second voice wake-up rates, it is determined whether the target channel contains audio signals from other system channels.

[0161] If the difference between the first voice wake-up rate and the second voice wake-up rate is less than a preset threshold, it indicates that using the audio signal of the target channel for echo cancellation has the same effect as using the audio signal of all system channels for echo cancellation. Therefore, it can be assumed that the target channel contains audio signals from other system channels. In this case, the first audio signal is used as the mixing signal, and the second audio signal is deleted.

[0162] If the difference between the first voice wake-up rate and the second voice wake-up rate is not less than a preset threshold, it is considered that the target channel does not contain audio signals from other system channels. Then, mixing processing is performed, that is, the second audio signal is added to the first audio signal to obtain a mixed signal.

[0163] The mixed signal is sent to the AEC echo cancellation component to cancel the echo signal. After the echo cancellation is completed, a series of operations such as beamforming, dereverberation, noise reduction, and voice wake-up are performed to complete the entire voice interaction process.

[0164] The above technical solution has the following advantages or beneficial effects: By acquiring the audio signals of each system channel from the output of the power amplifier unit, on the one hand, there is no need to separately add a voice chip to the display device to complete the echo signal acquisition, thereby effectively reducing the cost of audio signal acquisition and thus reducing the cost of echo cancellation. On the other hand, the audio signal output by the power amplifier unit is closest to the sound signal played by the speaker, which can ensure the accuracy of signal acquisition and thus ensure the accuracy of subsequent echo cancellation. It is understandable that as the number of system channels of the display device increases, the number of audio signal acquisition channels will also increase, which will undoubtedly put a certain resource pressure on the display device, thereby affecting the effect of echo cancellation and thus affecting the voice interaction performance. Therefore, the above display device adopts a mixing processing technology, that is, based on a preset mixing channel number, the audio signals of each system channel are mixed to obtain a mixed signal. The preset mixing channel number is determined in advance based on the device performance of the display device. This ensures that even in a multi-channel system, the display device can still effectively cancel the echo signal, improve the accuracy of echo cancellation, thereby improving the voice interaction experience and achieving a balance between echo cancellation cost and voice interaction performance.

[0165] Based on the same inventive concept, this application also provides an echo cancellation method applied to the aforementioned display device. The solution provided by this method is similar to the implementation described in the above-described display device embodiments. Therefore, the specific limitations of one or more echo cancellation embodiments provided below can be found in the limitations of the display device described above, and will not be repeated here.

[0166] In some embodiments, this application also provides an echo cancellation method applied to the aforementioned display device. In this embodiment, such as Figure 12 As shown, the echo cancellation method includes the following steps:

[0167] Step S1202: Amplify the power of the original audio of the multiple system channels of the display device and output the audio signal of each system channel; the multiple system channels are at least three system channels;

[0168] Step S1204: Play the sound signals corresponding to each audio signal;

[0169] Step S1206: Receive the echo signal of the sound signal;

[0170] Step S1208: Based on the preset number of mixing channels, the audio signals of each system channel are mixed to obtain a mixed signal; the preset number of mixing channels is determined based on the device performance of the display device.

[0171] Step S1210: Based on the mixed signal, perform echo cancellation on the echo signal.

[0172] In some embodiments, the echo cancellation method further includes: acquiring the signal type of the sound signal; the signal type includes multi-channel signals and non-multi-channel signals; a multi-channel signal refers to a sound signal that includes at least three audio channels, and a non-multi-channel signal refers to a sound signal that includes at most two audio channels. Step S1208 includes: mixing the audio signals of each system channel based on a preset number of mixing channels and signal type to obtain a mixed signal.

[0173] In some embodiments, the audio signals of each system channel are mixed based on a preset number of mixing channels and signal type to obtain a mixed signal, including: determining the target channel for mixing and other system channels besides the target channel among multiple system channels based on the preset number of mixing channels; acquiring the first audio signal of the target channel and the second audio signals of the other system channels; and mixing the first audio signal and the second audio signal based on the signal type to obtain a mixed signal.

[0174] In some embodiments, based on the signal type, the first audio signal and the second audio signal are mixed to obtain a mixed signal, including: when the audio signal is a multi-channel signal, adding the second audio signal to the first audio signal to obtain a mixed signal.

[0175] In some embodiments, based on the signal type, the first audio signal and the second audio signal are mixed to obtain a mixed signal, including: when the sound signal is a non-multichannel signal, performing correlation analysis on the first audio signal and the second audio signal to obtain correlation analysis results; when the correlation analysis results indicate that the first audio signal does not contain the second audio signal, adding the second audio signal to the first audio signal to obtain the mixed signal.

[0176] In some embodiments, the echo cancellation method further includes: when the correlation analysis results indicate that the first audio signal contains the second audio signal, using the first audio signal as a mixing signal and deleting the second audio signal.

[0177] In some embodiments, the echo cancellation method further includes: receiving a voice command. A correlation analysis is performed on the first audio signal and the second audio signal to obtain a correlation analysis result, including: executing a voice interaction task based on the first audio signal to obtain a first voice interaction result; the first voice interaction result includes a first voice wake-up rate; executing a voice interaction task based on each audio signal to obtain a second voice interaction result; the second voice interaction result includes a second voice wake-up rate; and performing a correlation analysis on the first audio signal and the second audio signal based on a comparison between the first and second voice wake-up rates to obtain a correlation analysis result.

[0178] In some embodiments, obtaining the signal type of the sound signal includes: decoding the sound signal to obtain the number of audio channels of the sound signal; and determining the signal type of the sound signal based on the number of audio channels.

[0179] In some embodiments, the echo cancellation method further includes: obtaining device configuration parameters of the display device; performing performance analysis on the display device based on the device configuration parameters to obtain performance analysis results of the display device; and determining a preset number of mixing channels that matches the performance analysis results.

[0180] Based on the same inventive concept, this application also provides an echo cancellation device for implementing the echo cancellation method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more echo cancellation device embodiments provided below can be found in the limitations of the echo cancellation method described above, and will not be repeated here.

[0181] In some embodiments, this application also provides an echo cancellation device applied to the aforementioned display device. In this embodiment, as... Figure 13 As shown, Figure 13 A flowchart illustrating the module interaction of an echo cancellation device is shown, wherein the echo cancellation device includes:

[0182] The audio signal output module is used to amplify the power of the raw audio from multiple system channels of the display device and output the audio signal for each system channel; the multiple system channels are at least three system channels.

[0183] The playback module is used to play the sound signals corresponding to each audio signal.

[0184] The echo acquisition module is used to receive the echo signal of the sound signal;

[0185] The audio mixing module is used to mix the audio signals of each system channel based on a preset number of mixing channels to obtain a mixed signal; the preset number of mixing channels is determined based on the device performance of the display device.

[0186] The echo cancellation module is used to cancel the echo signal based on the mixed signal.

[0187] In some embodiments, the echo cancellation device is further configured to: acquire the signal type of the sound signal; the signal type includes multi-channel signals and non-multi-channel signals; a multi-channel signal refers to a sound signal that includes at least three audio channels, and a non-multi-channel signal refers to a sound signal that includes at most two audio channels. The mixing module is further configured to: perform mixing processing on the audio signals of each system channel based on a preset number of mixing channels and signal type to obtain a mixed signal.

[0188] In some embodiments, the mixing processing module further includes: a channel determination unit, used to determine, based on a preset number of mixing channels, the target channel for mixing processing and other system channels besides the target channel among multiple system channels; a signal acquisition unit, used to acquire a first audio signal of the target channel and a second audio signal of the other system channels; and a mixing processing unit, used to mix the first audio signal and the second audio signal based on the signal type to obtain a mixed signal.

[0189] In some embodiments, the mixing processing unit is further configured to: when the audio signal is a multi-channel signal, add a second audio signal to the first audio signal to obtain a mixed signal.

[0190] In some embodiments, the mixing processing unit further includes: a correlation analysis subunit, used to perform correlation analysis on the first audio signal and the second audio signal when the audio signal is a non-multichannel signal, and obtain a correlation analysis result; and a mixing processing subunit, used to add the second audio signal to the first audio signal when the correlation analysis result indicates that the first audio signal does not contain the second audio signal, to obtain a mixed signal.

[0191] In some embodiments, the echo cancellation device is further configured to: when the correlation analysis results indicate that the first audio signal contains the second audio signal, use the first audio signal as a mixing signal and delete the second audio signal.

[0192] In some embodiments, the echo cancellation device is further configured to: receive voice commands. The correlation analysis subunit is further configured to: execute a voice interaction task based on a first audio signal to obtain a first voice interaction result; the first voice interaction result includes a first voice wake-up rate; execute a voice interaction task based on each audio signal to obtain a second voice interaction result; the second voice interaction result includes a second voice wake-up rate; and perform correlation analysis on the first and second audio signals based on a comparison between the first and second voice wake-up rates to obtain a correlation analysis result.

[0193] In some embodiments, the echo cancellation device is also used to: decode the sound signal to obtain the number of audio channels of the sound signal; and determine the signal type of the sound signal based on the number of audio channels.

[0194] In some embodiments, the echo cancellation device is further configured to: acquire device configuration parameters of the display device; perform performance analysis on the display device based on the device configuration parameters to obtain performance analysis results of the display device; and determine a preset number of mixing channels that matches the performance analysis results.

[0195] Each module in the aforementioned projection image processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0196] In some embodiments, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described method steps.

[0197] In some embodiments, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described method steps.

[0198] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0199] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0200] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A display device, characterized in that, include: The power amplifier unit is configured to amplify the raw audio of multiple system channels of the display device and output an audio signal for each of the system channels. The plurality of system channels comprises at least three system channels; A speaker is configured to play sound signals corresponding to each of the said audio signals; A microphone is configured to receive the echo signal of the sound signal; The controller is configured as follows: Based on a preset number of mixing channels, the audio signals of each system channel are mixed to obtain a mixed signal; the preset number of mixing channels is determined based on the device performance of the display device. Based on the mixed signal, echo cancellation is performed on the echo signal.

2. The display device according to claim 1, characterized in that, The controller is further configured to: The signal type of the sound signal is obtained; the signal type includes multi-channel signal and non-multi-channel signal; the multi-channel signal refers to a sound signal that includes at least three audio channels, and the non-multi-channel signal refers to a sound signal that includes at most two audio channels. The controller, in the process of mixing the audio signals of each system channel based on a preset number of mixing channels to obtain a mixed signal, is further configured to: Based on the preset number of mixing channels and the signal type, the audio signals of each system channel are mixed to obtain a mixed signal.

3. The display device according to claim 2, characterized in that, The controller, in the process of mixing the audio signals of each system channel based on a preset number of mixing channels and the signal type to obtain a mixed signal, is further configured to: Based on the preset number of mixing channels, the target channel for mixing and other system channels besides the target channel are determined among the plurality of system channels. Acquire the first audio signal of the target channel and the second audio signals of the other system channels; Based on the signal type, the first audio signal and the second audio signal are mixed to obtain a mixed signal.

4. The display device according to claim 3, characterized in that, The controller, during the process of mixing the first audio signal and the second audio signal based on the signal type to obtain the mixed signal, is further configured to: When the sound signal is a multi-channel signal, the second audio signal is added to the first audio signal to obtain a mixed signal.

5. The display device according to claim 3, characterized in that, The controller, during the process of mixing the first audio signal and the second audio signal based on the signal type to obtain the mixed signal, is further configured to: When the sound signal is not a multi-channel signal, a correlation analysis is performed on the first audio signal and the second audio signal to obtain the correlation analysis results; When the correlation analysis result indicates that the first audio signal does not contain the second audio signal, the second audio signal is added to the first audio signal to obtain a mixed signal.

6. The display device according to claim 5, characterized in that, The controller is further configured to: When the correlation analysis results indicate that the first audio signal contains the second audio signal, the first audio signal is used as the mixed signal, and the second audio signal is deleted.

7. The display device according to claim 3, characterized in that, The microphone is further configured to receive voice commands; the controller, in the process of performing correlation analysis on the first audio signal and the second audio signal to obtain the correlation analysis result, is further configured to: Based on the first audio signal, the voice interaction task of the voice command is executed to obtain a first voice interaction result; the first voice interaction result includes a first voice wake-up rate; Based on each of the audio signals, the voice interaction task of the voice command is executed to obtain a second voice interaction result; the second voice interaction result includes a second voice wake-up rate; Based on the comparison between the first voice wake-up rate and the second voice wake-up rate, a correlation analysis is performed on the first audio signal and the second audio signal to obtain the correlation analysis results.

8. The display device according to claim 1, characterized in that, The controller is further configured to: during the process of acquiring the signal type of the sound signal. The sound signal is decoded to obtain the number of audio channels of the sound signal; Based on the number of audio channels, the signal type of the sound signal is determined.

9. The display device according to claim 1, characterized in that, The controller is further configured to: determine the preset number of mixing channels during the process of determining the number of mixing channels. Obtain the device configuration parameters of the display device; Based on the device configuration parameters, a performance analysis is performed on the display device to obtain the performance analysis results of the display device; Determine the preset number of mix channels that matches the performance analysis results.

10. An echo cancellation method, characterized in that, Applied to a display device as described in any one of claims 1 to 9, the method comprises: The original audio from multiple system channels of the display device is amplified by power, and each system channel is output as its own audio signal; the multiple system channels are at least three system channels. Play the sound signals corresponding to each of the aforementioned audio signals; The echo signal received from the sound signal; Based on a preset number of mixing channels, the audio signals of each system channel are mixed to obtain a mixed signal; the preset number of mixing channels is determined based on the device performance of the display device. Based on the mixed signal, echo cancellation is performed on the echo signal.