Display device and echo cancellation method
By using the filter coefficient compensation method of microphone array and speaker array in display devices, the problem of poor echo cancellation effect in far-field interaction is solved, and the accuracy of speech recognition and user experience are improved.
Patent Information
- Application Number
- CN202010819984.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-14
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2040-08-14
AI Technical Summary
In far-field interaction, the prior art echo cancellation effect is poor, resulting in poor speech recognition accuracy and user experience, especially in the complex layout of multi-speakers and microphones in smart TVs.
The microphone array and speaker array of the display device are used to eliminate echo by calculating the coefficients of the filter, and the filter coefficient compensation is used to compensate the positional relationship between the microphone array and the speaker array, respectively, and the echo generated by the speakers in the first and second areas is processed.
It improves the wake-up rate and user experience of far-field interaction, improves the effect of echo cancellation, and ensures the wake-up response speed.
Smart Images

Figure CN114078480B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of display devices, and in particular, to a display device and an echo cancellation method. Background Art
[0002] With the rapid development of speech recognition technology, speech interaction has become an important human-computer interaction method. According to the distance of human-computer interaction, speech interaction can be divided into near-field interaction and far-field interaction. Among them, in some application scenarios, such as when performing speech interaction with a smart TV, far-field interaction is more convenient for users.
[0003] During far-field interaction, in order to avoid people deliberately increasing the volume, microphones can be set at multiple different positions of the smart TV to improve the ability to collect user voices. However, when the smart TV is in the playback state, the more microphones there are, the more complex the echo is; in addition, for a smart TV with multiple channels, multiple speakers are distributed at different positions, which will further increase the echo complexity. To improve the accuracy of speech recognition, echo cancellation is required.
[0004] In the related art, by mixing all channels according to the left and right channels, that is, mixing the signals of all channels on the left side of the smart TV into one path, and mixing the signals of all channels on the right side into one path, so as to form two reference signals similar to the traditional 2.0 channels, and using the reference signals to cancel part of the sound actually received by the microphone to achieve echo cancellation. However, due to the overlapping parts of different channels in different frequency bands, for the sound in the physical space, after circuit mixing, some frequency band sounds will be lost, and at the same time, the delay information in the space is also lost, etc., which easily leads to poor echo cancellation effect, and then leads to a poor experience of far-field interaction. Summary of the Invention
[0005] To solve the above technical problems, this application provides a display device and an echo cancellation method.
[0006] In a first aspect, this application provides a display device, which includes:
[0007] A display;
[0008] A microphone array, including microphones distributed in a first area and a second area;
[0009] A speaker array, including speakers distributed in the first area and the second area;
[0010] A controller, respectively connected to the display, the microphone array and the speaker array, and the controller is configured to:
[0011] Obtain the system reference signal of the speaker and the audio signal received by the microphone respectively;
[0012] Calculate the coefficients of the first filter according to the system reference signal, where the first filter is used to filter the echo generated by the speakers in the first area;
[0013] Compensate the coefficients of the first filter according to the positional relationship between the microphone array and the speaker array to obtain the coefficients of the second filter, where the second filter is used to filter the echo generated by the speakers in the second area;
[0014] Perform echo cancellation on the audio signal according to the system reference signal, the coefficients of the first filter, and the coefficients of the second filter to obtain a clean signal.
[0015] In some embodiments, the first area and the second area are symmetric about the central axis of the display, the microphones in the first area and the microphones in the second area are symmetric about the central axis, and the speakers in the first area and the speakers in the second area are symmetric about the central axis.
[0016] In some embodiments, the compensating the coefficients of the first filter according to the positional relationship between the microphone array and the speaker array to obtain the coefficients of the second filter includes:
[0017] Calculate the sound attenuation ratio between the speakers at symmetric positions in the first area and the second area to each microphone;
[0018] Calculate the power amplifier ratio between the speakers at symmetric positions in the first area and the second area;
[0019] Compensate the coefficients of the first filter according to the product of the sound attenuation ratio and the power amplifier ratio to obtain the coefficients of the second filter.
[0020] Coefficients of the first filter Coefficients of the first filter 4. The display device according to claim 1, wherein calculating the coefficients of the first filter according to the system reference signal Coefficients of the first filter includes:
[0021] Calculate the coefficients of the adaptive filter for the first speaker according to the system reference signal to obtain the coefficients of the first filter.
[0022] In a second aspect, an embodiment of the present application provides an echo cancellation method for the display device described in the first aspect. The method includes:
[0023] Obtain the system reference signal of the speaker and the audio signal received by the microphone respectively;
[0024] Calculate the coefficients of a first filter based on the system reference signal, where the first filter is used to filter the echo generated by the speakers in the first region;
[0025] Compensate the coefficients of the first filter according to the positional relationship between the microphone array and the speaker array to obtain the coefficients of a second filter, where the second filter is used to filter the echo generated by the speakers in the second region;
[0026] Perform echo cancellation on the audio signal according to the system reference signal, the coefficients of the first filter, and the coefficients of the second filter to obtain the clean signal corresponding to the audio signal.
[0027] In some embodiments, the first region and the second region are symmetric about the central axis of the display, the microphones in the first region and the microphones in the second region are symmetric about the central axis, and the speakers in the first region and the speakers in the second region are symmetric about the central axis.
[0028] The beneficial effects of the display device and the echo cancellation method provided in this application include:
[0029] In the embodiments of this application, the display area of the display device is divided into symmetric first and second regions, the coefficients of a first filter for filtering the echo between the speakers and microphones in the first region are calculated, and then, according to the positional relationship between the speakers and microphones in the first and second regions, the coefficients of the first filter are compensated to obtain the coefficients of a second filter. Echo cancellation is performed on the audio signal according to the system reference signal, the coefficients of the first filter, and the coefficients of the second filter to obtain a clean signal. In the embodiments of this application, by filtering the echo between each speaker and microphone respectively using the corresponding coefficients of the first filter or the second filter, the interruption wake-up rate of far-field interaction can be improved. Among them, the coefficients of the second filter are obtained by compensating the coefficients of the first filter, which has less computational workload compared with calculating according to the system reference signal, and can ensure the wake-up response speed of far-field interaction. The embodiments of this application can improve the user experience of far-field interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] To more clearly illustrate the technical solutions of this application, the accompanying drawings required for the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0031] Figure 1 Exemplarily shown is a schematic diagram of an operation scenario between a display device and a control device according to some embodiments;
[0032] Figure 2 An exemplary hardware configuration block diagram of the display device 200 according to some embodiments is shown;
[0033] Figure 3 An exemplary hardware configuration block diagram of the control device 100 according to some embodiments is shown;
[0034] Figure 4 An exemplary software configuration schematic diagram of the display device 200 according to some embodiments is shown;
[0035] Figure 5 An exemplary display schematic diagram of the icon control interface of the application program in the display device 200 according to some embodiments is shown;
[0036] Figure 6 An exemplary distribution schematic diagram of the microphone array and the speaker array according to some embodiments is shown;
[0037] Figure 7 An exemplary audio transmission schematic diagram according to some embodiments is shown;
[0038] Figure 8 An exemplary flowchart of the echo cancellation method according to some embodiments is shown;
[0039] Figure 9 An exemplary echo cancellation schematic diagram according to some embodiments is shown;
[0040] Figure 10 An exemplary schematic diagram of the calculation method of the coefficients of the second filter according to some embodiments is shown. Detailed implementation manners
[0041] To make the objectives, implementation manners and advantages of the present application clearer, the following will clearly and completely describe the exemplary implementation manners of the present application with reference to the accompanying drawings in the exemplary embodiments of the present application. Apparently, the described exemplary embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0042] Based on the exemplary embodiments described in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope protected by the appended claims of the present application. In addition, although the disclosed content in the present application is introduced according to one or several exemplary examples, it should be understood that each aspect of these disclosed contents can also constitute a complete implementation manner alone.
[0043] It should be noted that the brief description of terms in this application is only for facilitating the understanding of the following described embodiments, rather than intending to limit the embodiments of this application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.
[0044] In this application, terms such as "first", "second", "third", etc. in the specification, claims and the above-mentioned drawings are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise indicated. It should be understood that such terms can be interchanged under appropriate circumstances, for example, it is possible to implement in an order other than those given in the illustration or description of the embodiments of this application.
[0045] In addition, the terms "comprising" and "having" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device comprising a series of components does not necessarily have to be limited to those components clearly listed, but may include other components not clearly listed or inherent to these products or devices.
[0046] The term "module" used in this application refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or a combination of hardware or / and software code that can perform functions related to that element.
[0047] The term "remote control" used in this application refers to a component of an electronic device (such as the display device disclosed in this application), which can generally wirelessly control the electronic device within a relatively short distance range. Generally, it is connected to the electronic device using infrared and / or radio frequency (RF) signals and / or Bluetooth, and may also include functional modules such as WiFi, wireless USB, Bluetooth, motion sensors, etc. For example: a handheld touch remote control replaces most of the physical built-in hard keys in a general remote control device with a user interface on the touch screen.
[0048] The term "gesture" used in this application refers to a user behavior where the user makes a change in hand shape or a hand movement and other actions to express an intended idea, action, purpose / or result.
[0049] Figure 1 Exemplarily shows a schematic diagram of an operation scenario between a display device and a control device according to an embodiment. As Figure 1 shown, the user can operate the display device 200 through the mobile terminal 300 and the control device 100.
[0050] In some embodiments, the control device 100 may be a remote control. The communication between the remote control and the display device includes infrared protocol communication, Bluetooth protocol communication, and other short - distance communication methods, etc., and controls the display device 200 through wireless or other wired methods. The user can input user instructions through the buttons on the remote control, voice input, control panel input, etc. to control the display device 200. For example, the user can input corresponding control instructions through the volume increase / decrease buttons, channel control buttons, up / down / left / right movement buttons, voice input button, menu button, power - on / off button, etc. on the remote control to achieve the function of controlling the display device 200.
[0051] In some embodiments, a mobile terminal, a tablet computer, a computer, a laptop computer, and other intelligent devices can also be used to control the display device 200. For example, an application program running on the intelligent device is used to control the display device 200. Through configuration, the application program can provide various controls for the user in an intuitive user interface (UI) on the screen associated with the intelligent device.
[0052] In some embodiments, the mobile terminal 300 and the display device 200 can install software applications and achieve connection communication through network communication protocols to achieve the purpose of one - to - one control operation and data communication. For example, a control instruction protocol can be established between the mobile terminal 300 and the display device 200, and the remote control keyboard can be synchronized to the mobile terminal 300. By controlling the user interface on the mobile terminal 300, the function of controlling the display device 200 can be achieved. It is also possible to transmit the audio - video content displayed on the mobile terminal 300 to the display device 200 to achieve the synchronous display function.
[0053] As Figure 1 It is also shown that the display device 200 also conducts data communication with the server 400 through various communication methods. The display device 200 is allowed to communicate and connect through a local area network (LAN), a wireless local area network (WLAN), and other networks. The server 400 can provide various contents and interactions for the display device 200. By way of example, the display device 200 receives software program updates or accesses a remotely stored digital media library by sending and receiving information, as well as through electronic program guide (EPG) interaction. The server 400 can be a cluster or multiple clusters, and can include one or more types of servers. Other network service contents such as video - on - demand and advertising services are provided through the server 400.
[0054] The display device 200 can be a liquid crystal display, an OLED display, or a projection display device. The specific type, size, and resolution of the display device are not limited. Those skilled in the art can understand that the display device 200 can make some changes in performance and configuration according to needs.
[0055] In addition to providing the function of receiving broadcast television, the display device 200 can additionally provide the function of a smart network television that supports computer functions, including but not limited to, network television, smart television, Internet Protocol Television (IPTV), etc.
[0056] Figure 2 An exemplary hardware configuration block diagram of the display device 200 according to an exemplary embodiment is shown.
[0057] In some embodiments, the display device 200 includes at least one of a controller 250, a tuner demodulator 210, a communicator 220, a detector 230, an input / output interface 255, a display 275, an audio output interface 285, a memory 260, a power supply 290, a user interface 265, and an external device interface 240.
[0058] In some embodiments, the display 275 is a component for receiving an image signal output from a first processor and displaying video content, images, and a menu control interface.
[0059] In some embodiments, the display 275 includes a display screen component for presenting a picture and a driving component for driving image display.
[0060] In some embodiments, the displayed video content can come from broadcast television content, that is, various broadcast signals that can be received through wired or wireless communication protocols. Or, various image contents sent from a network server end received through a network communication protocol can be displayed.
[0061] In some embodiments, the display 275 is used to present a user control UI interface generated in the display device 200 and used to control the display device 200.
[0062] In some embodiments, depending on the type of the display 275, a driving component for driving the display is further included.
[0063] In some embodiments, the display 275 is a projection display, and may further include a projection device and a projection screen.
[0064] In some embodiments, the communicator 220 is a component for communicating with an external device or an external server according to various communication protocol types. For example: the communicator may include at least one of a Wifi chip, a Bluetooth communication protocol chip, a wired Ethernet communication protocol chip, other network communication protocol chips such as a near-field communication protocol chip, and an infrared receiver.
[0065] In some embodiments, the display device 200 can establish the sending and receiving of control signals and data signals with an external control device 100 or a content providing device through the communicator 220.
[0066] In some embodiments, the user interface 265 can be used to receive infrared control signals from a control device 100 (such as an infrared remote control).
[0067] In some embodiments, the detector 230 is a signal for the display device 200 to collect the external environment or interact with the outside.
[0068] In some embodiments, the detector 230 includes a light receiver, a sensor for collecting the intensity of ambient light, and can adaptively change display parameters by collecting ambient light.
[0069] In some embodiments, the detector 230 may further include an image collector, such as a camera, which can be used to collect external environmental scenes, as well as to collect user attributes or user interaction gestures, and can adaptively change display parameters, and can also recognize user gestures to achieve the function of interacting with users.
[0070] In some embodiments, the detector 230 may further include a temperature sensor, etc., such as by sensing the ambient temperature.
[0071] In some embodiments, the display device 200 can adaptively adjust the display color temperature of the image. For example, in an environment with a relatively high temperature, the display device 200 can be adjusted to display an image with a cooler color temperature, or in an environment with a relatively low temperature, the display device 200 can be adjusted to display an image with a warmer color temperature.
[0072] In some embodiments, the detector 230 may further include a sound collector, such as a microphone, which can be used to receive the user's voice. Exemplarily, it includes a voice signal for the user to control the display device 200, or collects ambient sound to identify the type of ambient scene, so that the display device 200 can adapt to ambient noise.
[0073] In some embodiments, such as Figure 2 As shown, the input / output interface 255 is configured to enable data transmission between the controller 250 and other external devices or other controllers 250. Such as receiving video signal data, audio signal data, or command instruction data from an external device.
[0074] In some embodiments, the external device interface 240 may include, but is not limited to: a high-definition multimedia interface HDMI interface, an analog or data high-definition component input interface, a composite video input interface, a USB input interface, an RGB port, or any one or more of these interfaces. It can also be a composite input / output interface formed by the above multiple interfaces.
[0075] In some embodiments, such as Figure 2As shown, the tuner demodulator 210 is configured to receive broadcast television signals through wired or wireless reception, and can perform modulation and demodulation processes such as amplification, mixing, and resonance, and demodulate audio and video signals from multiple wireless or wired broadcast television signals. The audio and video signals can include the television audio and video signals carried in the television channel frequencies selected by the user, as well as EPG data signals.
[0076] In some embodiments, the frequency points demodulated by the tuner demodulator 210 are controlled by the controller 250. The controller 250 can issue control signals according to user selections, so that the modem responds to the television signal frequency selected by the user and modulates and demodulates the television signal carried by that frequency.
[0077] In some embodiments, broadcast television signals can be classified into terrestrial broadcast signals, cable broadcast signals, satellite broadcast signals, or Internet broadcast signals, etc., according to different television signal broadcast formats. Or they can be classified into digital modulation signals, analog modulation signals, etc., according to different modulation types. Or they can be classified into digital signals, analog signals, etc., according to different signal types.
[0078] In some embodiments, the controller 250 and the tuner demodulator 210 can be located in different separate devices. That is, the tuner demodulator 210 can also be in an external device of the main device where the controller 250 is located, such as an external set-top box. In this way, the set-top box outputs the television audio and video signals obtained by modulating and demodulating the received broadcast television signals to the main device, and the main device receives the audio and video signals through the first input / output interface.
[0079] In some embodiments, the controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 can control the overall operation of the display device 200. For example: in response to receiving a user command for selecting a UI object to be displayed on the display 275, the controller 250 can perform operations related to the object selected by the user command.
[0080] In some embodiments, the object can be any one of the selectable objects, such as a hyperlink or an icon. Operations related to the selected object, such as: operations for displaying a page, document, image, etc. connected to the hyperlink, or operations for executing a program corresponding to the icon. The user command for selecting a UI object can be a command input through various input devices (such as a mouse, keyboard, touchpad, etc.) connected to the display device 200 or a voice command corresponding to the user's spoken words.
[0081] Such as Figure 2As shown, the controller 250 includes at least one of a random access memory 251 (RAM), a read-only memory 252 (ROM), a video processor 270, an audio processor 280, other processors 253 (such as a graphics processing unit (GPU), a central processing unit 254 (CPU)), a communication interface, and a communication bus 256. Among them, the communication bus connects each component.
[0082] In some embodiments, the RAM 251 is used to store temporary data of the operating system or other running programs.
[0083] In some embodiments, the ROM 252 is used to store instructions for various system startups.
[0084] In some embodiments, the ROM 252 is used to store a basic input / output system, called the Basic Input Output System (BIOS). It is used to complete the power-on self-test of the system, initialize each functional module in the system, drive programs for the basic input / output of the system, and boot the operating system.
[0085] In some embodiments, when a power-on signal is received, the power of the display device 200 starts to boot up. The CPU runs the system startup instructions in the ROM 252 and copies the temporary data of the operating system stored in the memory to the RAM 251 to facilitate the startup or running of the operating system. After the operating system starts up, the CPU then copies the temporary data of various application programs in the memory to the RAM 251, and then, to facilitate the startup or running of various application programs.
[0086] In some embodiments, the CPU processor 254 is used to execute the instructions of the operating system and application programs stored in the memory, and execute various application programs, data, and content according to various interaction instructions received externally, so as to finally display and play various audio and video contents.
[0087] In some exemplary embodiments, the CPU processor 254 may include multiple processors. The multiple processors may include a main processor and one or more sub-processors. The main processor is used to execute some operations of the display device 200 in the pre-power-on mode and / or the operation of displaying a picture in the normal mode. One or more sub-processors are used for an operation in a standby mode or other states.
[0088] In some embodiments, the graphics processor 253 is configured to generate various graphic objects, such as icons, operation menus, and graphic displays of user input instructions. It includes an arithmetic unit that performs operations by receiving various interactive instructions input by the user and displays various objects according to display attributes. It also includes a renderer that renders various objects obtained from the arithmetic unit, and the rendered objects are used for display on the display.
[0089] In some embodiments, the video processor 270 is configured to receive an external video signal and perform video processing such as decompression, decoding, scaling, noise reduction, frame rate conversion, resolution conversion, image synthesis, etc. according to the standard codec protocol of the input signal, so as to obtain a signal that can be directly displayed or played on the display device 200.
[0090] In some embodiments, the video processor 270 includes a demultiplexing module, a video decoding module, an image synthesis module, a frame rate conversion module, a display formatting module, etc.
[0091] Among them, the demultiplexing module is used to perform demultiplexing processing on the input audio-visual data stream. For example, if the input is MPEG-2, the demultiplexing module demultiplexes it into a video signal and an audio signal, etc.
[0092] The video decoding module is used to process the demultiplexed video signal, including decoding and scaling processing, etc.
[0093] The image synthesis module, such as an image synthesizer, is used to superimpose and mix the GUI signal generated by the graphics generator according to user input or self-generation with the video image after scaling processing to generate an image signal for display.
[0094] The frame rate conversion module is used to convert the input video frame rate, such as converting the 60Hz frame rate to 120Hz frame rate or 240Hz frame rate, and usually the format is implemented by means of frame interpolation.
[0095] The display formatting module is used to receive the video output signal after frame rate conversion and change the signal to conform to the display format signal, such as outputting an RGB data signal.
[0096] In some embodiments, the graphics processor 253 and the video processor can be integrated or separately provided. When integrated, it can perform processing on the graphic signal output to the display. When separately provided, they can perform different functions respectively, such as the GPU+FRC (Frame Rate Conversion) architecture.
[0097] In some embodiments, the audio processor 280 is configured to receive an external audio signal, decompress and decode it according to the standard codec protocol of the input signal, and perform processing such as noise reduction, digital-to-analog conversion, and amplification to obtain a sound signal that can be played on the speaker.
[0098] In some embodiments, the video processor 270 may be composed of one or more chips. The audio processor may also be composed of one or more chips.
[0099] In some embodiments, the video processor 270 and the audio processor 280 may be separate chips or may be integrated with the controller in one or more chips.
[0100] In some embodiments, the audio output receives the sound signal output by the audio processor 280 under the control of the controller 250, such as the speaker 286, and in addition to the speaker carried by the display device 200 itself, it can be output to the external audio output terminal of the sound generating device of the external device, such as the external audio interface or the earphone interface, etc. It may also include a short-range communication module in the communication interface, for example, a Bluetooth module for Bluetooth speaker sound output.
[0101] The power supply 290, under the control of the controller 250, provides power supply support for the display device 200 with the power input from the external power supply. The power supply 290 may include a built-in power circuit installed inside the display device 200 or may be an external power supply installed outside the display device 200, with a power interface for providing external power in the display device 200.
[0102] The user interface 265 is configured to receive the user's input signal and then send the received user input signal to the controller 250. The user input signal may be a remote control signal received through an infrared receiver or various user control signals received through the network communication module.
[0103] In some embodiments, the user inputs a user command through the control device 100 or the mobile terminal 300, and the user input interface then, based on the user's input, the display device 200 responds to the user's input through the controller 250.
[0104] In some embodiments, the user can input a user command in the graphical user interface (GUI) displayed on the display 275, and then the user input interface receives the user input command through the graphical user interface (GUI). Alternatively, the user can input a user command by inputting a specific sound or gesture, and then the user input interface recognizes the sound or gesture through the sensor to receive the user input command.
[0105] In some embodiments, a "user interface" is a media interface for interaction and information exchange between an application or an operating system and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. The common manifestation form of a user interface is a graphical user interface (GUI), which refers to a user interface related to computer operations displayed in a graphical manner. It can be an interface element such as an icon, a window, a control, etc. displayed on the display screen of an electronic device, where the control can include visible interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, Widgets, etc.
[0106] The memory 260 includes various software modules stored for driving the display device 200. For example, various software modules stored in the first memory include at least one of a basic module, a detection module, a communication module, a display control module, a browser module, and various service modules, etc.
[0107] The basic module is a low-level software module for signal communication between various hardware in the display device 200 and for sending processing and control signals to upper-level modules. The detection module is a management module for collecting various information from various sensors or user input interfaces, performing digital-to-analog conversion, and analysis and management.
[0108] For example, the voice recognition module includes a voice parsing module and a voice instruction database module. The display control module is a module for controlling the display to display image content, and can be used to play multimedia image content and information such as UI interfaces, etc. The communication module is a module for controlling and data communication with external devices. The browser module is a module for performing data communication between browsing servers. The service module is a module for providing various services and various application programs. At the same time, the memory 260 is also used to store received external data and user data, images of various items in various user interfaces, and visual effect diagrams of focus objects, etc.
[0109] Figure 3 An exemplary configuration block diagram of the control device 100 according to an exemplary embodiment is shown. As Figure 3 shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.
[0110] The control device 100 is configured to control the display device 200, and can receive input operation instructions from a user, and convert the operation instructions into instructions recognizable and responsive by the display device 200, acting as an interaction intermediary between the user and the display device 200. For example, when the user operates the channel plus and minus keys on the control device 100, the display device 200 responds to the channel plus and minus operations.
[0111] In some embodiments, the control device 100 can be an intelligent device. For example, the control device 100 can install various applications for controlling the display device 200 according to user requirements.
[0112] In some embodiments, such as Figure 1 as shown, after installing the application for controlling the display device 200, the mobile terminal 300 or other intelligent electronic devices can perform functions similar to those of the control device 100. For example, the user can install the application and use various function keys or virtual buttons on the graphical user interface provided on the mobile terminal 300 or other intelligent electronic devices to implement the functions of the physical buttons of the control device 100.
[0113] The controller 110 includes a processor 112, a RAM 113, a ROM 114, a communication interface 130, and a communication bus. The controller is used to control the operation and operation of the control device 100, as well as the communication and cooperation between internal components and the data processing functions between the external and internal.
[0114] Under the control of the controller 110, the communication interface 130 realizes the communication of control signals and data signals with the display device 200. For example, the received user input signal is sent to the display device 200. The communication interface 130 can include at least one of other near-field communication modules such as a WiFi chip 131, a Bluetooth module 132, and an NFC module 133.
[0115] The user input / output interface 140, where the input interface includes at least one of a microphone 141, a touchpad 142, a sensor 143, a button 144, and other input interfaces. For example, the user can implement the user instruction input function through actions such as voice, touch, gesture, and pressing. The input interface converts the received analog signal into a digital signal and then converts the digital signal into a corresponding instruction signal and sends it to the display device 200.
[0116] The output interface includes an interface for sending the received user instruction to the display device 200. In some embodiments, it can be an infrared interface or a radio frequency interface. For example, when it is an infrared signal interface, the user input instruction needs to be converted into an infrared control signal according to the infrared control protocol and sent to the display device 200 through the infrared sending module. Another example is that when it is a radio frequency signal interface, the user input instruction needs to be converted into a digital signal, then modulated according to the radio frequency control signal modulation protocol, and sent to the display device 200 by the radio frequency sending terminal.
[0117] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The communication interface 130 is configured in the control device 100, such as modules like WiFi, Bluetooth, NFC, etc., which can encode the user input instructions through the WiFi protocol, or the Bluetooth protocol, or the NFC protocol and send them to the display device 200.
[0118] The memory 190 is used to store various operating programs, data, and applications for driving and controlling the control device 200 under the control of the controller. The memory 190 can store various control signal instructions input by the user.
[0119] The power supply 180 is used to provide operating power support for each component of the control device 100 under the control of the controller. It can be a battery and related control circuits.
[0120] In some embodiments, the system may include a Kernel, a command parser (shell), a file system, and application programs. The Kernel, shell, and file system together constitute the basic operating system structure, which allows users to manage files, run programs, and use the system. After power-on, the Kernel starts, activates the Kernel space, abstracts the hardware, initializes the hardware parameters, etc., and runs and maintains the virtual memory, scheduler, signals, and inter-process communication (IPC). After the Kernel starts, the Shell and user application programs are loaded. The application programs are compiled into machine code after startup to form a process.
[0121] See Figure 4 , in some embodiments, the system is divided into four layers, from top to bottom are the Application layer (referred to as the "application layer" for short), the Application Framework layer (referred to as the "framework layer" for short), the Android runtime and the System Library layer (referred to as the "system runtime library layer" for short), and the Kernel layer.
[0122] In some embodiments, at least one application program runs in the application layer. These application programs can be window (Window) programs, system setting programs, clock programs, camera applications, etc. that come with the operating system; or they can be application programs developed by third-party developers, such as the HiSee program, the KTV program, the Magic Mirror program, etc. In specific implementations, the application program packages in the application layer are not limited to the above examples. Actually, they can also include other application program packages, and the embodiments of the present application do not limit this.
[0123] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer includes some predefined functions. The application framework layer is equivalent to a processing center that determines the actions of the applications in the application layer. Applications can access resources in the system and obtain system services during execution through the API interface.
[0124] As Figure 4 shown, in the embodiments of this application, the application framework layer includes Managers, Content Provider, etc. Among them, the Managers include at least one of the following modules: The ActivityManager is used to interact with all the activities running in the system; the Location Manager is used to provide access to the system location service for system services or applications; the Package Manager is used to retrieve various information related to the application packages currently installed on the device; the NotificationManager is used to control the display and clearing of notification messages; the Window Manager is used to manage icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.
[0125] In some embodiments, the ActivityManager is used to: manage the life cycles of various applications and the usual navigation back functions, such as controlling the exit of applications (including switching the user interface currently displayed in the display window to the system desktop), opening, and backing (including switching the user interface currently displayed in the display window to the previous-level user interface of the currently displayed user interface), etc.
[0126] In some embodiments, the Window Manager is used to manage all window programs, such as obtaining the screen size of the display, determining whether there is a status bar, locking the screen, taking screenshots, and controlling the changes in the display window (such as shrinking the display window, jittering the display, distorting the display, etc.).
[0127] In some embodiments, the system runtime layer provides support for the upper layer, i.e., the framework layer. When the framework layer is used, the Android operating system will run the C / C++ libraries included in the system runtime layer to implement the functions that the framework layer needs to achieve.
[0128] In some embodiments, the kernel layer is the layer between hardware and software. As Figure 4As shown, the kernel layer includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, touch sensor, pressure sensor, etc.).
[0129] In some embodiments, the kernel layer further includes a power driver module for power management.
[0130] In some embodiments, Figure 4 the software program and / or module corresponding to the software architecture in Figure 2 or Figure 3 is stored in the first memory or the second memory shown.
[0131] In some embodiments, for a display device with touch function, taking split-screen operation as an example, the display device receives an input operation (such as split-screen operation) applied by the user on the display screen. The kernel layer can generate a corresponding input event according to the input operation and report the event to the application framework layer. The activity manager of the application framework layer sets the window mode (such as multi-window mode) corresponding to the input operation, as well as the window position and size, etc. The window management of the application framework layer draws the window according to the settings of the activity manager, and then sends the drawn window data to the display driver of the kernel layer, and the display driver displays the corresponding application interface in different display areas of the display screen.
[0132] In some embodiments, as Figure 5 shown in, the application layer includes at least one application that can display corresponding icon controls on the display, such as: live TV application icon control, video-on-demand application icon control, media center application icon control, application center icon control, game application icon control, etc.
[0133] In some embodiments, the live TV application can provide live TV through different signal sources. For example, the live TV application can use the input from cable TV, wireless broadcast, satellite service or other types of live TV services to provide TV signals. Also, the live TV application can display the video of the live TV signal on the display device 200.
[0134] In some embodiments, the video-on-demand application can provide videos from different storage sources. Different from the live TV application, the video-on-demand provides the display of videos from certain storage sources. For example, the video-on-demand can come from the server side of cloud storage, from a local hard disk storage containing stored video programs.
[0135] In some embodiments, a media center application can provide applications for playing various multimedia contents. For example, the media center can provide services different from live TV or video on demand, and users can access various images or audio through the media center application.
[0136] In some embodiments, an application center can provide storage for various applications. The applications can be games, applications, or other applications related to a computer system or other devices but can run on a smart TV. The application center can obtain these applications from different sources, store them in a local storage, and then they can be run on the display device 200.
[0137] The hardware or software architecture in some embodiments can be based on the introduction in the above embodiments, and in some embodiments, it can be based on other similar hardware or software architectures, as long as the technical solutions of this application can be implemented.
[0138] In some embodiments, the application center can be provided with a voice assistant application to implement intelligent voice services, such as searching for media assets, adjusting the volume, etc. Users can wake up the voice assistant application by sending a voice signal to the display device. The voice signal can be some preset wake-up words. After the voice assistant application is woken up, users can interact with the voice assistant application to perform voice control on the display device.
[0139] In some embodiments, when a user inputs a voice signal, such as a wake-up word, to the display device and the display device is not playing an audio or video, the audio signal received by the microphone on the display device includes the user's voice signal and the noise signal of the environment where the display device is located. Since the noise signal is usually of a lower volume and is quite different from the user's voice signal, in this case, the probability that the controller of the display device receives and accurately recognizes the voice signal is relatively high, and it is relatively easy to wake up the voice assistant application. The probability of receiving and accurately recognizing the voice signal can be referred to as the interruption wake-up rate, and the speed at which the voice assistant application responds to the voice signal can be referred to as the wake-up response speed.
[0140] In some embodiments, when a user inputs a voice signal, such as a wake-up word, to the display device and the display device is playing an audio or video, the audio signal received by the microphone on the display device, in addition to the user's voice signal and the noise signal, also includes the audio signal being played by the display device. The audio signal being played by the display device can be referred to as echo. Since the volume of the echo may be large and the echo may also contain human voices, it is easy to cover the user's voice signal, which will interfere with the voice signal and result in a lower interruption wake-up rate of the display device, making it difficult to wake up the voice assistant application. By eliminating the echo, the interruption wake-up rate can be increased. However, correspondingly, if the elimination algorithm is too complex, it will cause a significant reduction in the wake-up response speed, affecting the user experience.
[0141] In some embodiments, to facilitate users to perform voice control on the display device even when they are far away from the display device, the display device may be provided with a microphone array including multiple microphones at different positions to improve the ability to collect user voice signals.
[0142] In some embodiments, to improve the sound effect of the display device, the display device may be provided with a speaker array including multiple speakers distributed at different positions.
[0143] See Figure 6 , which is a schematic diagram of the distribution of the microphone array and the speaker array according to some embodiments. As Figure 6 shown, the microphone array 500 may be distributed at the upper border of the display device, and the speaker array may be distributed at the upper border and the lower border of the display device. The speaker array may include speaker 601, speaker 602, speaker 603, speaker 604, speaker 605, and speaker 606.
[0144] See Figure 7 , which is a schematic diagram of audio transmission according to some embodiments. As Figure 7 shown, the microphone array may include 6 microphones, namely microphone 1, microphone 2, microphone 3, microphone 4, microphone 5, and microphone 6. The speaker array may include 6 speakers, namely speaker a, speaker b, speaker c, speaker d, speaker e, and speaker f. Among them, speaker a may be Figure 6 the speaker 601 in Figure 6 the speaker 602 in Figure 6 the speaker 603 in Figure 6 the speaker 604 in Figure 6 the speaker 605 in Figure 6 the speaker 606 in
[0145] Figure 7 In, the lines between the microphones and the speakers indicate that the sound emitted by the speakers can be transmitted to the microphones, and the dashed line indicates the central axis of the display device. The microphone array and the speaker array may be symmetrically distributed on both sides of the central axis of the display device. The area on the left side of the central axis of the display device is the left channel area, and the area on the right side of the central axis of the display device is the right channel area. The left channel area may be referred to as the first area, and the right channel area may be referred to as the second area, or the left channel area may be the second area, and the right channel area is the second area.
[0146] It can be seen that the microphones 1 - 6 can respectively receive the echoes of the speakers a - f. Among them, the echo paths of the microphone 1 and the microphone 6 are symmetrical, the echo paths of the microphone 2 and the microphone 5 are symmetrical, and the echo paths of the microphone 3 and the microphone 4 are symmetrical.
[0147] It should be noted that in actual implementation, affected by the internal space structure of the display device, the respective position errors of the microphone and the speaker, etc., the echo paths may be difficult to achieve complete symmetry.
[0148] Based on Figure 7 the shown audio transmission schematic diagram, in order to improve the interruption wake-up rate while ensuring a relatively fast wake-up response speed, the embodiment of the present application provides an echo cancellation method. Refer to Figure 8 , and this echo cancellation method may include the following steps:
[0149] Step S110: Obtain the system reference signal of the speaker and the audio signal received by the microphone respectively.
[0150] In some embodiments, when the user inputs a voice signal to the display device and the display device is playing media resources, the controller of the display device can obtain the media resource data currently being played by the system and the audio signal received by each microphone, analyze the media resource data, and obtain the media resource audio signal being played by each of the speakers. This media resource audio signal can be referred to as the system reference signal. The audio signal received by the microphone includes the above-mentioned voice signal and the media resource audio signal.
[0151] In some embodiments, after the user inputs a voice signal to the display device, the audio signal of the microphone will include this voice signal.
[0152] For example, the audio signals of the microphones 1 to 6 can be: M1 - M6.
[0153] Step S120: Calculate the coefficients of the first filter according to the system reference signal, and the first filter is used to filter the echo generated by the speakers in the first area.
[0154] In some embodiments, an echo filter can be configured for the audio signal transmitted between each speaker and the microphone to filter the echo. The total number of filters is the product of the number of speakers and the number of microphones. The echo filter between the speakers and the microphones in the first area can be referred to as the first filter, and the filtering coefficient of the first filter is referred to as the coefficient of the first filter; the echo filter between the speakers and the microphones in the second area can be referred to as the second filter, and the filtering coefficient of the second filter is referred to as the coefficient of the second filter.
[0155] For example, in Figure 7In the audio transmission schematic diagram shown, there are a total of 36 echo filters. The coefficients of the first filter may include: W1a to W6a, W1b to W6b, W1c to W6c. The coefficients of the second filter may include: W1d to W6d, W1e to W6e, W1f to W6f.
[0156] In some embodiments, the echo filter may adopt an adaptive filter. According to the system reference signal, the coefficients of the adaptive filter can be calculated for the first speaker to obtain the coefficients of the first filter.
[0157] See Figure 9 , for the echo cancellation schematic diagram according to some embodiments, as Figure 9 shown, for a remote input signal x(n), after passing through an unknown echo path w(n), the signal y(n) is obtained, y(n) = x(n) * w(n). After adding the observation noise v(n), the desired signal d(n) is obtained, d(n) = y(n) + v(n). The estimated echo signal obtained after x(n) passes through the adaptive filter w^(n) is w^T(n)x(n). Subtracting the estimated echo signal from the desired signal d(n) gives the error signal e(n), that is, e(n) = d(n) - w^T(n)x(n). The smaller the value of the error signal, the closer the echo path estimated by the adaptive filter is to the actual echo path.
[0158] An adaptive algorithm is used to adjust the weight vector of the adaptive filter so that the estimated echo path w^(n) gradually approaches the true echo path w(n). Obviously, in the AEC (Acoustic Echo Cancelation) problem, the selection of the adaptive filter plays a very crucial role in the performance of echo cancellation. Here, taking the adaptive algorithm as the LMS (Least Mean Square) algorithm as an example, the adjustment process of the adaptive filter will be described.
[0159] Using the least mean square error criterion, by taking the derivative of the estimated echo signal and setting it equal to 0, the w(n) that makes the error |e(n)| 2 the smallest is obtained. Among them, since |e(n)| is not differentiable at the minimum point, |e(n)| 2 is used to calculate w(n). In the LMS algorithm, the coefficient iteration formula of its filter is:
[0160]
[0161] The iterative process of the coefficients of the filter can include three aspects:
[0162] 1. Output y(n) through the FIR filter:
[0163]
[0164] In formula (2), i represents the index of the sampling point, N represents the maximum number of sampling points, for example, N = 3200.
[0165] 2. Calculate the error according to formula (2):
[0166] e(n) = d(n) - y(n) (3)
[0167] 3. Update the weights of the FIR vector to prepare for the next iteration until steady-state convergence is obtained according to the least mean square error criterion. At this time, the weight vector is the coefficient of the desired filter. Among them, the FIR vector is the weight vector 2μe(n), and the weight update formula is as follows:
[0168] w(n + 1) = w(n) + 2μe(n)x(n) (4)
[0169] Step S130: According to the positional relationship between the microphone array and the speaker array, compensate the coefficients of the first filter to obtain the coefficients of the second filter, and the second filter is used to filter the echo generated by the speakers in the second area.
[0170] In some embodiments, the positional relationship between the speakers in the first area and the speakers in the second area may be a symmetric relationship. For example, speaker a and speaker f are symmetric about the central axis of the display device, speaker b and speaker d are symmetric about the central axis of the display device, and speaker c and speaker d are symmetric about the central axis of the display device.
[0171] In some embodiments, the positional relationship between the microphones in the first area and the microphones in the second area may be a symmetric relationship. For example, speaker a and speaker f are symmetric about the central axis of the display device, speaker b and speaker d are symmetric about the central axis of the display device, and speaker c and speaker d are symmetric about the central axis of the display device.
[0172] According to the above symmetric relationship, the coefficients of the first filter can be compensated to obtain the coefficients of the second filter. Refer to Figure 10 , which is a schematic diagram of the calculation method of the coefficients of the second filter according to some embodiments. For example, Figure 10 as shown, this calculation method may include steps S310 - S330.
[0173] Step S310: Calculate the sound attenuation ratio between the speakers at the symmetric positions in the first area and the second area to each microphone.
[0174] In some embodiments, the filtering difference of the filter is affected by the sound attenuation difference. According to the sound attenuation formula: AdiV = 10lg[1 / (4πr^2)], the sound attenuation ratio x1 between the speakers at the symmetric positions in the first region and the second region and each microphone can be calculated. Here, r is the sound path length, which can be specifically obtained by pre-measuring the distance between the microphone and the speaker. The calculation formula is as follows:
[0175] x1 = 10lg(L1a / L6f) (5)
[0176] In formula (5), L1a is the distance from speaker a to microphone 1, L1a is the distance from speaker a to microphone 1, and L6f is the distance from speaker f to microphone 6.
[0177] Step S320: Calculate the power amplifier ratio between the speakers at the symmetric positions in the first region and the second region.
[0178] In some embodiments, the filtering difference of the filter is affected by the power amplifier ratio between the speakers. Calculate the power amplifier ratio x2 between the speakers at the symmetric positions in the first region and the second region. The calculation formula is as follows:
[0179] x2 = Pa / Pf (6)
[0180] In formula (6), Pa is the power of speaker a, and Pf is the power of speaker f.
[0181] Step S330: Compensate the coefficients of the first filter according to the product of the sound attenuation ratio and the power amplifier ratio to obtain the coefficients of the second filter.
[0182] According to the product of the sound attenuation ratio x1 and the power amplifier ratio x2, the filtering difference between the first filter and the second filter can be obtained as follows:
[0183] C1 = 10lg(L1a / L6f)*Pa / Pf; C2 = 10lg(L2a / L5f)*Pb / Pe; C3 = 10lg(L3a / L4f)*Pc / Pd;
[0184] D1 = 10lg(L1b / L6e)*Pb / Pe; D2 = 10lg(L2b / L5e)*Pb / Pe; D3 = 10lg(L3b / L4e)*Pb / Pe;
[0185] E1 = 10lg(L1c / L6d)*Pc / Pd; E2 = 10lg(L2c / L5d)*Pc / Pd; E3 = 10lg(L3c / L4d)*Pc / Pd (7)
[0186] In formula (7), C1 is the filtering difference coefficient between the coefficient W1a of the first filter and the coefficient W6f of the second filter. The filtering difference coefficient is used to characterize the filtering performance gap between the two filters. C2 is the filtering difference coefficient between the coefficient W2a of the first filter and the coefficient W5f of the second filter. C3 is the filtering difference coefficient between the coefficient W3a of the first filter and the coefficient W4f of the second filter. D1 is the filtering difference coefficient between the coefficient W1b of the first filter and the coefficient W6e of the second filter. D2 is the filtering difference coefficient between the coefficient W2b of the first filter and the coefficient W5e of the second filter. D3 is the filtering difference coefficient between the coefficient W3b of the first filter and the coefficient W4e of the second filter. E1 is the filtering difference coefficient between the coefficient W1c of the first filter and the coefficient W6d of the second filter. E2 is the filtering difference coefficient between the coefficient W2c of the first filter and the coefficient W5d of the second filter. E3 is the filtering difference coefficient between the coefficient W3c of the first filter and the coefficient W4d of the second filter.
[0187] In some embodiments, each system reference signal is spatially symmetric with respect to the microphone in terms of physical position, and the data of the reference signal is also symmetric. That is, without considering other spatial reflection factors, such as reverberation factors, the information from the first speaker reaching the 6 microphones and the information from the second speaker reaching the 6 microphones at the corresponding positions only have a time delay difference. Therefore, after calculating the coefficients of the first filter, the coefficients of the first filter can be compensated for spatial position according to the filtering difference coefficient to obtain the coefficients of the second filter.
[0188] For example, the coefficients of the second filter can be calculated according to the following formula:
[0189] W1a = W6f * C1; W2a = W5f * C2; W3a = W4f * C3;
[0190] W1b = W6e * D1; W2b = W5e * D2; W3b = W4e * D3;
[0191] W1c = W6d * E1; W2c = W5d * E2; W3c = W4d * E3 (8)
[0192] Step S140: Perform echo cancellation on the audio signal according to the system reference signal, the coefficients of the first filter, and the coefficients of the second filter to obtain a clean signal.
[0193] In some embodiments, after calculating the coefficients of the first filter and the coefficients of the second filter, echo cancellation can be performed on the audio signal, and the calculation formula is as follows:
[0194] A1 = M1 – Ra * W1a – Rb * W1b – Rc * W1c – Rd * W1d – Re * W1e – Rf * W1f
[0195] A2 = M2 – Ra * W2a – Rb * W2b – Rc * W2c – Rd * W2d – Re * W2e – Rf * W2f
[0196] A3 = M3 – Ra * W3a – Rb * W3b – Rc * W3c – Rd * W3d – Re * W3e – Rf * W3f
[0197] A4 = M4 – Ra * W4a – Rb * W4b – Rc * W4c – Rd * W4d – Re * W4e – Rf * W4f
[0198] A5 = M5 – Ra * W5a – Rb * W5b – Rc * W5c – Rd * W5d – Re * W5e – Rf * W5f
[0199] A6 = M6 – Ra * W6a – Rb * W6b – Rc * W6c – Rd * W6d – Re * W6f – Rf * W6f (9)
[0200] (9) In the formula, M1 is the audio signal received by microphone 1, M2 is the audio signal received by microphone 2, M3 is the audio signal received by microphone 3, M4 is the audio signal received by microphone 4, M5 is the audio signal received by microphone 5, and M6 is the audio signal received by microphone 6.
[0201] Ra is the reference signal of speaker a, Rb is the reference signal of speaker b, Rc is the reference signal of speaker c, Rd is the reference signal of speaker d, Re is the reference signal of speaker e, and Rf is the reference signal of speaker f.
[0202] Ra * W1a, Ra * W1a, Rc * W1c, Ra * W2a, Rb * W2b, Rc * W2c, Ra * W3a, Rb * W3b, Rc * W3c, Ra * W4a, Rb * W4b, Rc * W4c, Ra * W5a, Rb * W5b, Rc * W5c, Ra * W6a, Rb * W6b, Rc * W6c are the first product signals, and Rd * W1d, Re * W1e, Rf * W1f, Rd * W2d, Re * W2e, Rf * W2f, Rd * W3d, Re * W3e, Rf * W3f, Rd * W4d, Re * W4e, Rf * W4f, Rd * W5d, Re * W5e, Rf * W5f, Rd * W6d, Re * W6f, Rf * W6f are the second product signals.
[0203] A1 is the pure signal component corresponding to the audio signal received by microphone 1, A2 is the pure signal component corresponding to the audio signal received by microphone 2, A3 is the pure signal component corresponding to the audio signal received by microphone 3, A4 is the pure signal component corresponding to the audio signal received by microphone 4, A5 is the pure signal component corresponding to the audio signal received by microphone 5, and A6 is the pure signal component corresponding to the audio signal received by microphone 6.
[0204] Superimposing all the pure signal components can obtain the pure signal corresponding to the audio signal.
[0205] Furthermore, the controller of the display device can perform speech recognition on the pure signal to obtain a user request, and control the display device according to the user request. For example, after performing speech recognition on the pure signal, if the user request is "reduce the volume", then reduce the volume of the currently playing audio or video.
[0206] As can be seen from the above embodiments, in the embodiments of the present application, the display area of the display device is divided into symmetric first and second regions, the coefficients of the first filter for filtering the echo between the speaker and the microphone in the first region are calculated, and then, according to the positional relationship between the speakers and microphones between the first region and the second region, the coefficients of the first filter are compensated to obtain the coefficients of the second filter. The audio signal is subjected to echo cancellation according to the system reference signal, the coefficients of the first filter, and the coefficients of the second filter to obtain a pure signal. In the embodiments of the present application, by filtering the echo between each speaker and microphone respectively using the corresponding coefficients of the first filter or the second filter, the interruption wake-up rate of far-field interaction can be improved. Among them, the coefficients of the second filter are obtained by compensating the coefficients of the first filter, and the computational workload is smaller compared to calculating based on the system reference signal, which can ensure the wake-up response speed of far-field interaction. The embodiments of the present application can improve the user experience of far-field interaction.
[0207] Since the above embodiments are all described by reference and combination on the basis of other methods, and there are identical parts among different embodiments, the same or similar parts among the various embodiments in this specification can be referred to each other. Details are not elaborated herein again.
[0208] It should be noted that in this specification, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a circuit structure, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such circuit structure, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the circuit structure, article or device comprising the element.
[0209] After considering the specification and practicing the disclosure of the present invention, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present application and include well-known knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the content of the claims. The above embodiments of the present application do not constitute a limitation on the protection scope of the present application.
Claims
1. A display device, characterized in that: include: monitor; a microphone array, comprising microphones distributed in a first area and a second area, wherein the first area and the second area are symmetrical about a central axis of the display, and the microphones in the first area and the second area are symmetrical about the central axis; A speaker array includes speakers distributed in the first area and speakers in the second area, wherein the speakers in the first area and the speakers in the second area are symmetrical about the central axis; A controller is connected to the display, the microphone array, and the speaker array, respectively, and is configured to: respectively acquiring a system reference signal of the loudspeaker and an audio signal received by the microphone; Calculating coefficients of a first filter according to the system reference signal, where the first filter is used to filter echoes generated by the speakers in the first area, and the second filter is used to filter echoes generated by the speakers in the second area; Calculating a sound attenuation ratio between a loudspeaker at a symmetrical position between the first area and the second area and each of the microphones; Calculating a power amplifier ratio between loudspeakers at symmetrical positions in the first area and the second area; Compensating the coefficients of the first filter according to the product of the sound attenuation ratio and the power amplification ratio to obtain coefficients of the second filter, where the second filter is used to filter the echo generated by the speaker in the second area; Echo cancellation is performed on the audio signal according to the system reference signal, the coefficients of the first filter, and the coefficients of the second filter to obtain a pure signal.
2. The display device according to claim 1, wherein The performing echo cancellation on the audio signal according to the system reference signal, the coefficient of the first filter, and the coefficient of the second filter to obtain a pure signal corresponding to the audio signal includes: performing echo cancellation on the microphone according to the system reference signal, the coefficient of the first filter, and the coefficient of the second filter to obtain a pure signal component; The sum of the pure signal components of all the microphones is calculated to obtain a pure signal corresponding to the audio signal.
3. The display device according to claim 2, wherein The method of performing echo cancellation on the microphone according to the system reference signal, the coefficient of the first filter, and the coefficient of the second filter to obtain a pure signal component includes: Calculating a first product signal of the system reference signal and a coefficient of a first filter; Calculating a second product signal of the system reference signal and the coefficient of a second filter; The first product signal and the second product signal are respectively subtracted from the audio signal to obtain a pure signal component.
4. An echo cancellation method for a display device, characterized in that: The display device includes a display, a microphone array, and a speaker array, the microphone array includes microphones distributed in a first area and a second area, the microphone array includes speakers distributed in the first area and speakers in the second area, the first area and the second area are symmetrical about a central axis of the display, the microphones in the first area and the microphones in the second area are symmetrical about the central axis, and the speakers in the first area and the speakers in the second area are symmetrical about the central axis; The echo cancellation method comprises: respectively acquiring a system reference signal of the loudspeaker and an audio signal received by the microphone; Calculating coefficients of a first filter according to the system reference signal, where the first filter is used to filter echoes generated by the speakers in the first area, and the second filter is used to filter echoes generated by the speakers in the second area; Calculating a sound attenuation ratio between a loudspeaker at a symmetrical position between the first area and the second area and each of the microphones; Calculating a power amplifier ratio between loudspeakers at symmetrical positions in the first area and the second area; Compensating the coefficients of the first filter according to the product of the sound attenuation ratio and the power amplification ratio to obtain coefficients of the second filter, where the second filter is used to filter the echo generated by the speaker in the second area; Echo cancellation is performed on the audio signal according to the system reference signal, the coefficients of the first filter, and the coefficients of the second filter to obtain a pure signal corresponding to the audio signal.