Display device and far-field voice echo cancellation method
By using the power amplifier unit and controller in the display device to extract the target reference audio data, combined with nonlinear optimization and sound effect processing, the problem of far-field voice interaction function being interfered with by background audio is solved, achieving efficient voice recognition and reducing hardware costs.
Patent Information
- Application Number
- CN202510828921.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-09
AI Technical Summary
The far-field voice interaction function of the display device is interfered by background audio, resulting in poor voice command recognition. The additional setting of voice processing chips increases design costs, which is not conducive to the popularization of functions.
By configuring a power amplifier unit and a controller in the display device, the target reference audio data is extracted using the sampling point identifier in the power amplifier unit register. Combined with nonlinear optimization and sound processing units, the interference audio in the voice commands is filtered out and the voice recognition effect is improved.
It improves the recognition accuracy and stability of voice commands, reduces hardware costs, simplifies design complexity, and enhances the reliability of far-field voice interaction.
Smart Images

Figure CN120612951A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of display devices, and in particular to a display device and a far-field voice echo cancellation method. Background Art
[0002] Display devices can support voice interaction, allowing users to interact with the display device through voice interaction. If hardware resources support it, the display device can provide users with far-field voice interaction, allowing users to interact with the display device through voice interaction even when they are far away from the device.
[0003] While far-field voice interaction can provide convenience and flexibility for users using display devices, it is significantly affected by background audio, particularly when the display device is outputting audio. This audio can interfere with the display device's ability to recognize voice commands in far-field voice interaction scenarios.
[0004] To address the issue of the display device's output audio interfering with far-field voice command recognition, a dedicated voice processing chip can be installed in the display device to address audio interference. However, installing an additional voice processing chip increases the design cost of the display device and hinders the widespread adoption of far-field voice interaction functionality. Summary of the Invention
[0005] The present application provides a display device and a far-field voice echo cancellation method to solve the problem of high design cost of far-field voice interaction systems.
[0006] In a first aspect, the present application provides a display device, comprising:
[0007] monitor.
[0008] The audio output device is configured to: output media audio;
[0009] The user input interface is configured to: receive a voice command; the user input interface is connected to the controller; the voice command includes interference audio; the interference audio is background noise corresponding to the media audio;
[0010] A power amplifier unit, which is connected to the controller and the audio output device respectively;
[0011] The controller is configured as:
[0012] When the audio output device outputs the media audio, the recollection point identifier stored in the power amplifier unit register is read; the power amplifier unit register is used to store the operation configuration information of the power amplifier unit; the operation configuration information of the power amplifier unit includes the recollection point identifier;
[0013] Extracting target reference audio data from the power amplifier unit based on the recollection point represented by the recollection point identifier; wherein the target reference audio data is at least media resource audio data that has undergone nonlinear optimization;
[0014] In response to the voice instruction, filtering out interference audio in the voice instruction based on target reference audio data to obtain a target voice instruction;
[0015] Execute a target task determined based on the target voice instruction.
[0016] In this way, when the controller outputs media audio through the audio output device, it can read the sampling point identifier in the power amplifier unit register to determine the sampling point for reading the target reference audio data, and then extract the target reference audio data based on the sampling point. When the controller receives a voice command, it can filter the target reference audio data containing the voice command based on the target reference audio data to extract a clearer voice command, thereby improving the recognition effect of the voice command.
[0017] In some feasible embodiments, the operation configuration information further includes a sampling channel identifier. The sampling channel identifier is used to indicate the sampling data channel used by the power amplifier unit to transmit the target reference audio data. The controller extracts the target reference audio data from the power amplifier unit based on the nonlinear sampling point represented by the nonlinear sampling point identifier, and is specifically configured to:
[0018] Read the acquisition channel ID in the power amplifier unit register.
[0019] Determine the acquisition data channel of the power amplifier unit according to the acquisition channel identifier.
[0020] Extract target reference audio data from the retrieved data channel.
[0021] In this way, by reading the sampling channel identifier in the power amplifier unit's register, the controller can determine the sampling data channel used by the power amplifier unit to transmit the target reference audio data, and then receive / extract the target reference audio data from the sampling data channel. The sampling channel identifier setting can distinguish the data channel used by the power amplifier unit to transmit the target reference audio data and the data channel used to output audio data to the speaker, preventing data transmission conflicts.
[0022] In some feasible embodiments, the system further includes: a sound effect processing unit, which is connected to the controller and the power amplifier unit respectively.
[0023] When the audio output device outputs media audio, the controller is specifically configured as follows:
[0024] Input the media audio data into the audio processing unit.
[0025] The sound effect processing unit is controlled to input the media audio data processed by the sound effect into the power amplifier unit.
[0026] The power amplifier unit is controlled to input the modulated media audio data into the audio output device.
[0027] In this way, based on the audio data processing capability provided by the sound processing unit and the modulation and optimization capability of the power amplifier unit for audio data, the audio format output by the display device can be enriched, thereby improving the audio output effect of the display device.
[0028] In some feasible embodiments, the power amplifier unit includes a first power amplifier unit and a second power amplifier unit. The sound effect processing unit is connected to the first power amplifier unit, and the controller is connected to the second power amplifier unit.
[0029] The controller controls the sound effect processing unit to input the media audio data processed by the sound effect into the power amplifier unit, and is specifically configured as follows:
[0030] Input the media audio data into the audio processing unit.
[0031] The audio effect processing unit is controlled to input the media audio data into the first power amplifier unit.
[0032] The controller executes inputting the media asset audio data into the sound effect processing unit, and is further configured to:
[0033] The media audio data is input to the second power amplifier unit.
[0034] In this way, when the display device outputs audio data, it is connected to the first power amplifier unit through the sound effect processing chip and to the second power amplifier unit through the controller. Two signal input sources are arranged in the display device so that the audio output device can output media audio by combining the audio data processed by the sound effect processing chip and the native audio data directly transmitted by the controller. Among them, based on the sound effect processing chip, the multi-channel audio characteristics can be adjusted to enrich the output effect of the media audio, and based on the native audio data, the authenticity of the audio data can be guaranteed. In addition, based on the layout of the two signal input sources, the audio output effect of the display device is improved.
[0035] In some feasible embodiments, the controller extracts target reference audio data from the power amplifier unit based on the nonlinear recapture point represented by the nonlinear recapture point identifier, and is specifically configured to:
[0036] Extracting first reference audio data from a first power amplifier unit according to a first re-collection point represented by a first re-collection point identifier, wherein the first re-collection point identifier is stored in a first power amplifier unit register corresponding to the first power amplifier unit.
[0037] The second reference audio data is extracted from the second power amplifier unit according to the second recapture point represented by the second recapture point identifier. The first reference audio data and the second reference audio data correspond to different sound effects. The first recapture point and the second recapture point represent different audio data processing stages.
[0038] The first reference audio data and the second reference audio data are fused to obtain target reference audio data.
[0039] In this way, the controller can retrieve reference audio data from the first power amplifier unit and the second power amplifier unit respectively. Based on the characteristics retained in each of the first reference audio data and the second reference audio data, the target reference audio data can be fused. This allows the target reference audio data to be consistent with the interference audio generated by the media audio output by the audio output device, thereby facilitating filtering of interference audio in voice commands based on the target reference audio data.
[0040] In some feasible embodiments, an analog-to-digital conversion unit is further included. The analog-to-digital conversion unit is connected to the first power amplifier unit and the controller. The controller extracts target reference audio data from the power amplifier unit based on the nonlinear recapture point represented by the nonlinear recapture point identifier, and is specifically configured as follows:
[0041] The first power amplifier unit is controlled to input the modulated media audio data into the analog-to-digital conversion unit to obtain third reference audio data. The second reference audio data and the third reference audio data have the same phase. The second reference audio data and the third reference audio data have different phases.
[0042] The third reference audio data output by the analog-to-digital conversion unit is received.
[0043] In this way, by configuring the analog-to-digital conversion unit, the first reference audio data extracted from the first power amplifier unit can be phase-modulated to generate third reference audio data, so that the phase of the third reference audio data is consistent with the phase of the second audio reference data. If the phases of the third reference audio data and the second reference audio data are consistent, they can be further fused to generate target reference audio data, thereby facilitating the subsequent filtering of interfering audio in voice commands.
[0044] In some feasible embodiments, the controller performs filtering of interfering audio in the voice command based on the target reference audio data, and is specifically configured to:
[0045] The signal-to-noise ratio of the voice command is adjusted based on the audio characteristics of the target reference audio data and the audio data corresponding to the interference audio to obtain the target voice command; the signal-to-noise ratio is used to characterize the signal strength relationship between the main audio data of the voice command and the audio data of the interference audio; the signal-to-noise ratio of the target voice command is greater than the signal-to-noise ratio of the voice command.
[0046] In this way, the controller can adjust the signal-to-noise ratio of the voice command based on the audio similarity between the target reference audio data and the interfering audio. The valid signal is the voice command, and the invalid signal is the interfering audio. By adjusting the signal-to-noise ratio of the voice command, the signal-to-noise ratio of the voice command and the interfering audio is adjusted, thereby filtering out the interfering audio while retaining the filtering effect of the voice command, thereby facilitating the subsequent accurate recognition of the voice command.
[0047] In some feasible embodiments, before the controller executes the target task indicated by the voice instruction, it is further configured to:
[0048] Parsing the target voice command to obtain semantic information of the voice command;
[0049] The target task is determined based on the mapping relationship between semantic information and the target task.
[0050] In this way, the clarity of voice commands can be improved, which is beneficial to improving the accuracy of voice recognition. Then, corresponding semantic information can be obtained based on the voice commands, and the target task can be determined based on the semantic information.
[0051] In some feasible embodiments, a digital silicon microphone is further included. The digital silicon microphone is arranged in a linear array in the preset sound collection area of the display and is connected to the user input interface.
[0052] The digital silicon microphone is configured as:
[0053] Collect voice commands.
[0054] Input the voice command into the user input interface.
[0055] This allows digital silicon microphones to directly convert voice commands into digital output signals, which helps improve the signal-to-noise ratio of voice commands. Furthermore, the array arrangement effectively enhances the display device's ability to receive far-field voice signals.
[0056] In a second aspect, the present application provides a far-field voice echo cancellation method, the steps of which include:
[0057] When the audio output device outputs the media audio, the recollection point identifier stored in the power amplifier unit register is read; the power amplifier unit register is used to store the operation configuration information of the power amplifier unit; the operation configuration information of the power amplifier unit includes the recollection point identifier;
[0058] If a nonlinear re-collection point identifier is read, reference audio data is extracted from the power amplifier unit based on the nonlinear re-collection point represented by the nonlinear re-collection point identifier; wherein the reference audio data is the audio data of the media resource to be optimized that has undergone nonlinear optimization;
[0059] In response to the voice instruction, filtering out interference audio in the voice instruction based on reference audio data to obtain a target voice instruction;
[0060] Execute a target task determined based on the target voice instruction.
[0061] In this way, when the display device outputs media audio through the audio output device, it can read the sampling point identifier in the power amplifier unit register to determine the sampling point for reading the target reference audio data, and then extract the target reference audio data based on the sampling point. When the display device receives a voice command, it can filter the voice command based on the target reference audio data to extract a clearer voice command, thereby improving the recognition effect of the voice command. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0063] Figure 1 A schematic diagram of a display device operation scenario provided in some embodiments of the present application;
[0064] Figure 2 A schematic diagram of the hardware configuration of a display device provided in some embodiments of the present application;
[0065] Figure 3 A schematic diagram of software configuration of a display device provided in some embodiments of the present application;
[0066] Figure 4 An audio processing architecture diagram of a display device provided in some embodiments of the present application;
[0067] Figure 5 An internal interaction diagram of a display device in a far-field voice interaction scenario provided by some embodiments of the present application;
[0068] Figure 6 A flowchart of task execution of a display device in a far-field voice interaction scenario provided by some embodiments of the present application;
[0069] Figure 7 A schematic diagram of a recovery point provided for some embodiments of the present application;
[0070] Figure 8 A schematic diagram of register identifier allocation provided for some embodiments of the present application;
[0071] Figure 9 A diagram of an audio processing architecture of a display device including an audio processing unit provided in some embodiments of the present application;
[0072] Figure 10 The mapping diagrams provided for some embodiments of the present application are schematic. DETAILED DESCRIPTION
[0073] The following embodiments are described in detail, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numbers in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following embodiments are not intended to represent all possible implementations consistent with the present application. They are merely examples of systems and methods consistent with certain aspects of the present application, as detailed in the claims.
[0074] The following embodiments are described in detail, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numbers in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following embodiments are not intended to represent all possible implementations consistent with the present application. They are merely examples of systems and methods consistent with certain aspects of the present application, as detailed in the claims.
[0075] It should be noted that the brief descriptions of terms in this application are only for the purpose of facilitating the understanding of the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.
[0076] In the specification and claims of this application and the accompanying drawings, the terms "first," "second," "third," etc. are used to distinguish similar or similar objects or entities, and are not necessarily intended to limit a particular order or sequence, unless otherwise noted. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances.
[0077] The terms "comprise," "include," and "have," and any variations thereof, are intended to cover but not exclude inclusion; for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.
[0078] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functionality associated with that element.
[0079] In the embodiments of the present application, the display device 200 generally refers to a device capable of displaying images and processing data. For example, the display device 200 includes but is not limited to a smart TV, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, etc.
[0080] Figure 1This is a schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application. Figure 1 As shown in FIG, a user can operate the display device 200 through touch operation, the mobile terminal 300 and the control device 100. For example, the control device 100 can be a remote controller, a stylus pen, a handle, etc.
[0081] The mobile terminal 300 can function as a control device for performing human-computer interaction between a user and the display device 200. The mobile terminal 300 can also function as a communication device for establishing a communication connection with the display device 200 and exchanging data. In some embodiments, the mobile terminal 300 can install software applications with the display device 200, enabling connection and communication via a network communication protocol, enabling one-to-one control operations and data communication. Audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 for synchronized display.
[0082] like Figure 1 As shown in FIG, the display device 200 also communicates data with the server 400 through various communication methods. The display device 200 may be allowed to communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks.
[0083] The display device 200 may provide a broadcast receiving television function, and may also additionally provide an intelligent network television function with a computer support function, including but not limited to network television, smart TV, Internet Protocol television (IPTV), etc.
[0084] Figure 2 Some embodiments of this application provide Figure 1 2 is a block diagram of the hardware configuration of the display device 200.
[0085] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.
[0086] In some embodiments, detector 230 is used to collect signals from the external environment or external interactions. For example, detector 230 may include a millimeter-wave radar, which can be used to detect whether a user is within a preset range. Detector 230 may also include a sound collector to collect user input voice commands.
[0087] In some embodiments, the display 260 includes a display component for presenting images and a driver component for driving image display. The display 260 is configured to receive image signals output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, menu control interface components, and user control UI interfaces.
[0088] In some embodiments, the communication device 220 is a component used to communicate with an external device or server 400 according to various communication protocol types. The display device 200 can be provided with multiple communication devices 220 depending on the supported communication methods. For example, if the display device 200 supports wireless network communication, the display device 200 can be provided with a communication device 220 including WiFi functionality. If the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including Bluetooth functionality.
[0089] The communication device 220 can establish a communication connection between the display device 200 and an external device or server 400 via a wireless or wired connection. A wired connection can connect the display device 200 to an external device via a data cable, an interface, or other components. A wireless connection can connect the display device 200 to an external device via a wireless signal or wireless network. The display device 200 can establish a connection with an external device directly or indirectly through a gateway, router, or connection device.
[0090] In some embodiments, the controller 250 may include at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processor, and a power processor, and first to nth interfaces for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in a memory. The controller 250 controls the overall operation of the display device 200.
[0091] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.
[0092] In some embodiments, the user may input a user command through a graphical user interface (GUI) displayed on the display 260 , and the user input interface receives the user input command through the graphical user interface (GUI).
[0093] In some embodiments, the audio output device 270 may be a local speaker of the display device 200, or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may further be provided with an external audio output terminal, through which the audio output device may be connected to the display device 200 to output the sound of the display device 200.
[0094] In some embodiments, user input interface 280 may be configured to receive user input commands. User input interface 280 may include at least one of a microphone, a touchpad, a sensor, a remote control, and the like. Furthermore, display device 200 may receive user input commands based on user input interface 280 to perform interactive functions with the user.
[0095] To facilitate user interaction, in some embodiments, the display device 200 may run an operating system. An operating system is a computer program used to manage and control the hardware and software resources of the display device 200. The operating system can control the display device to provide a user interface. For example, the operating system can directly control the display device to provide a user interface, or it can provide a user interface by running an application program. The operating system also allows the user to interact with the display device 200.
[0096] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for display devices.
[0097] The operating system can be divided into different modules or layers according to the functions implemented, e.g. Figure 3 As shown, in some embodiments, the system is divided into four layers, from top to bottom, namely, the application layer (referred to as "application layer"), the application framework layer (referred to as "framework layer"), the system library layer and the kernel layer.
[0098] In some embodiments, the application layer provides services and interfaces for applications, enabling the display device 200 to run applications and interact with the user based on the applications. The application layer can host at least one application, which can include built-in window programs, system settings programs, clock programs, and the like, or applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the examples above.
[0099] The framework layer provides applications with an application programming interface (API) and programming framework. The application framework layer includes predefined functions. The application framework layer acts as a processing center, determining the actions taken by applications in the application layer. Through the API, applications can access system resources and services during execution.
[0100] like Figure 3 As shown, in the embodiment of the present application, the application framework layer includes a view system, managers, content providers, etc., wherein the view system can design and implement the interface and interaction of the application, and the view system includes lists, grids, text boxes, buttons, etc. The manager includes at least one of the following modules: an activity manager for interacting with all activities running in the system; a location manager for providing system services or applications with access to the system location service; a package manager for retrieving various information related to the application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.
[0101] In some embodiments, the activity manager is used to manage the lifecycle of each application and common navigation back functions, such as controlling application exit, opening, and back. The window manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, taking screenshots, and controlling changes in display windows, such as shrinking, shaking, or distorting the display window.
[0102] In some embodiments, the system runtime layer can provide support for the framework layer. When the framework layer is used, the operating system will run the instruction library contained in the system runtime layer, such as the C / C++ instruction library, to implement the functions to be implemented by the framework layer.
[0103] In some embodiments, the kernel layer is a functional layer between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. Figure 3As shown, the kernel layer can be configured with hardware drivers, and the drivers included in the kernel layer can be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.
[0104] based on Figure 1-Figure 3 The software architecture, hardware architecture and the coordination mechanism between various levels mentioned in the text enable the system to handle multiple tasks efficiently. For example, the display device 200 can output media audio through the audio output device, and then collect external audio data through a sound collection device such as a microphone connected to the user input interface to realize the voice interaction function. Among them, the voice interaction function can include a near-field voice interaction function and a far-field voice interaction function. The near-field voice interaction function can be combined with a remote control to realize voice input, while the far-field voice interaction function refers to the user initiating a voice command at a distance from the display device 200, and the display device 200 can still respond to the voice command, that is, the display device 200 based on the far-field voice interaction function can provide users with a more natural and convenient interaction method.
[0105] It is understandable that in a scenario where the display device 200 outputs media audio via an audio output device, when a user inputs a voice command, the media audio will become the background audio for the voice command, i.e., interference audio. That is, the microphone array connected to the user input interface of the display device 200 will receive both the voice command input by the user and the interference audio corresponding to the media audio output by the audio output device. Therefore, when a user inputs a voice command, the display device 200 needs to accurately extract the user's voice command from the external audio data containing the voice command and the interference audio, and process it using voice recognition technology to ensure accurate execution of the command, while suppressing the impact of the interference audio and improving the stability and reliability of voice interaction.
[0106] In some embodiments, a dedicated audio processing chip, such as a DSP (Digital Signal Processing), may be configured in the display device 200. The display device 200 equipped with the audio processing chip can filter the echo signal corresponding to the interfering audio in the external audio data through the audio processing chip, and perform noise reduction processing on the external audio data.
[0107] It is understandable that adding an audio processing chip to the display device 200 will increase the development cost of the display device. From a hardware perspective, it will increase the complexity and power consumption of the circuit board, thereby increasing the difficulty of hardware design. From a software perspective, adding an audio processing chip requires the controller 250 to provide additional drivers and algorithm support to achieve real-time processing and optimization of audio data. Therefore, by adding an audio processing chip specifically for extracting voice commands from external audio data to support the far-field voice interaction function, although the accuracy and stability of voice recognition are improved, it also brings about an increase in cost and power consumption. As a result, the far-field voice function is difficult to popularize in the display device 200.
[0108] In order to solve the above problems, Figure 4 As shown, an embodiment of the present application provides a display device 200. The display device 200 includes a display 260, which can display a media playback screen and an interaction screen with a user.
[0109] The audio output device 270 can be used to play media audio.
[0110] The user input interface can be connected to a microphone / microphone array to receive external audio data through the microphone array. In this embodiment of the present application, the external audio data includes voice commands and / or interference audio corresponding to the media audio. That is, when the user does not issue a voice command, the external audio data includes background noise such as interference audio. When the user issues a voice command, the external audio data includes voice commands and interference audio.
[0111] In some embodiments, the microphones may be four digital silicon microphones arranged linearly in a preset sound collection area of the display 260. The microphone array formed by the digital silicon microphones inputs the collected external audio data into the controller 250 via a user input interface (Pulse Density Modulation interface, PDM).
[0112] The power amplifier unit modulates and optimizes the media data before outputting it to the audio output device 270. For example, the power amplifier unit can perform nonlinear optimization processing steps such as three-band dynamic range control (3-band DRC) and parametric equalization, as well as PWM (Pulse Width Modulation) modulation on the media audio data to ensure high efficiency of the media audio signal.
[0113] Controller 250, in the embodiment of the present application, the controller 250 may be a SOC (System On Chip). Figure 4 As shown, based on the digital signal transmission characteristics of digital silicon microphones, digital signals can be directly transmitted between the controller 250 and the microphone array via the PDM interface. Direct transmission via the PDM interface reduces signal conversion losses, which helps reduce the loss of audio data corresponding to voice commands, thereby ensuring voice recognition quality.
[0114] The controller 250 and the power amplifier unit can also be directly connected to transmit digital signals based on the I2S (Integrated Interchip Sound) interface. Through the direct connection of the I2S interface, the controller 250 can accurately control the audio output of the power amplifier unit to ensure the simultaneous optimization of sound quality and far-field speech recognition. In addition, multiple data channels can be pre-configured between the controller 250 and the power amplifier unit, and the controller 250 can also retrieve audio data from the power amplifier unit during the audio output process to filter out interference audio in external audio data.
[0115] like Figure 5 and Figure 6 As shown, the controller 250 is configured to:
[0116] S100: When the audio output device outputs media audio, read the recollection point identifier stored in the power amplifier unit register.
[0117] In some embodiments, the power amplifier unit register is used to store the operation configuration information of the power amplifier unit. The operation configuration information of the power amplifier unit includes but is not limited to the sampling point identifier of the power amplifier unit. The sampling point identifier can be used to confirm the specific location where the controller 250 samples the audio data from the power amplifier unit. The specific location here refers to the specific stage where the power amplifier unit processes the media audio data, such as the buffer after DRC processing or the data segment before PWM modulation. When processing media audio data, the power amplifier unit will process it in a certain order, and the data processed at each stage can be cached in the buffer. For example, the sampling point is after the DRC stage, and the controller 250 can sample the media audio data processed by DRC based on the sampling point.
[0118] In some embodiments, the sampling point identifier of the power amplifier unit can be divided into "before" and "after", where "before" means sampling data before the nonlinear optimization stage, and "after" means sampling data after the nonlinear optimization stage.
[0119] In some embodiments, both the recovery point identifier and the recovery data channel identifier can be configured using multiple bit fields in a register. For example, bits 0-1 of the power amplifier unit register are used to store the recovery point identifier. A combination of "01" can correspond to the "before" identifier, and a combination of "10" can correspond to the "after" identifier. This allows the controller to precisely select the recovery timing and data channel based on the recovery point identifier and recovery data channel identifier read from the power amplifier.
[0120] It is understood that bit segments can be pre-allocated based on the entity object corresponding to the identifier. For example, when the retrieved data channel identifier can correspond to different pin combinations of a power amplifier unit, since the power amplifier unit has a large number of pins and a large number of pin combinations, more bit segments need to be allocated to support sufficient identifier combinations.
[0121] By properly allocating bit segments, each identifier uniquely corresponds to a specific sampling location and data channel, thereby improving the flexibility and configurability of the system. It should be noted that the configuration of the power amplifier unit registers is based on parameters pre-set by engineers, so the display device 200 will strictly operate according to these parameters during operation, and the controller 250 will also read the pre-set parameter values when reading the identifiers. This preset mechanism ensures the stability and consistency of the system and avoids runtime errors.
[0122] S200: If a nonlinear re-sampling point identifier is read, target reference audio data is extracted from the power amplifier unit based on the nonlinear re-sampling point represented by the nonlinear re-sampling point identifier.
[0123] In some embodiments, during the media audio output process, the controller 250 may read a re-acquisition point identifier from a power amplifier unit register at a preset period and extract target reference audio data from the power amplifier unit based on the re-acquisition point indicated by the re-acquisition point identifier. When the controller 250 reads a nonlinear re-acquisition point identifier, it may extract target reference audio data from the power amplifier unit based on the re-acquisition point indicated by the nonlinear re-acquisition point identifier. The target reference audio data is data that has undergone nonlinear optimization processing, such as PEQ and 3-band DRC.
[0124] It should be noted that the audio characteristics of the data after nonlinear optimization are highly similar to the audio data output by the final audio output device 270. Therefore, the data extracted by the controller 250 at the nonlinear sampling point can be used as the optimization basis, that is, the target reference audio data.
[0125] like Figure 7As shown, by setting the recapture point at the nonlinear recapture point (after), the recaptured signal, after PEQ and 3-band DRC processing, more accurately reflects the speaker's sound quality characteristics. This allows controller 250 to accurately capture the audio characteristics after nonlinear processing when extracting target reference audio data. Based on these captured audio characteristics, controller 250 can filter out interfering audio from the external audio data, ultimately extracting the more effective audio data corresponding to the voice command.
[0126] S300 : In response to a voice instruction, filter out interference audio in external audio data based on target reference audio data to extract the voice instruction from the external audio data.
[0127] In some embodiments, when no voice interaction is involved, the controller 250 will still periodically read the acquisition point identifier in the power amplifier unit register and extract the corresponding audio data. If no wake-up word is detected, the controller 250 can discard the target reference audio data obtained at this time. In the case of voice interaction, the controller 250 can filter the interference audio in the voice command based on the acquired target reference audio data.
[0128] It should be noted that the voice command can wake up the far-field voice interaction function of the display device 200. At this time, the controller 250 can filter the interference audio in the received voice command according to the retrieved target reference audio data. For example, it can be determined whether to start the far-field voice interaction function based on the monitoring of the wake-up word. Since the wake-up word is a pre-set sensitive input signal, the wake-up word can be recognized based on a simple acoustic recognition module, and then the far-field voice interaction function can be woken up when the user inputs a voice command. Among them, the wake-up design of the far-field voice interaction function is not the focus of the description of the embodiment of this application, and a general design can be adopted. Its design scheme is not described here.
[0129] The embodiment of the present application aims to improve the accuracy of the voice recognition engine in recognizing voice commands by filtering out the interference audio in the voice commands based on the collected target reference audio data and combining the target reference audio data. After the interference audio in the voice commands is filtered out, it can be input into the voice recognition engine. Based on the filtering process of the interference audio, the recognition accuracy of the voice recognition engine for voice commands can be effectively improved.
[0130] This approach allows the system to accurately capture user voices in complex environments, reducing misrecognition and improving the user experience. Especially in far-field voice interaction scenarios where background audio noise is high, using the sampled signal as the target reference audio to filter out interfering audio from external audio data effectively improves the accuracy and stability of voice recognition.
[0131] S400: Execute a target task determined based on the target voice instruction.
[0132] In some embodiments, a mapping relationship between voice commands and target tasks may be pre-configured in the storage space of the display device 200. Thus, after determining the target voice command, the controller 250 may determine the target task based on the target voice command and the pre-stored mapping relationship.
[0133] For example, when a user issues a voice command of "start my favorite app", the controller 250 can determine and execute the target task of "start my favorite app" according to the voice command and the pre-stored mapping relationship after filtering out the interfering audio and determining the target command.
[0134] like Figure 8 As shown, the operation configuration information stored in the power amplifier unit register also includes a sampling channel identifier. The sampling channel identifier can be used to control the pin in the power amplifier unit for transmitting the target reference audio data to the controller 250. In other words, the controller 250 can determine the sampling data channel for receiving the target reference audio data based on the sampling channel identifier. Furthermore, when executing the nonlinear sampling point characterized by the nonlinear sampling point identifier to extract the target reference audio data from the power amplifier unit, the controller 250 is specifically configured as follows:
[0135] Read the sampling channel identifier in the power amplifier unit register.
[0136] The acquisition data channel of the power amplifier unit is determined according to the acquisition channel identifier.
[0137] The target reference audio data is extracted from the retrieved data channel.
[0138] In some embodiments, the retrieved data channel identifier can be used to determine the pin or channel in the power amplifier unit used to transmit the target reference audio data back to the controller 250. By storing the retrieved point identifier and the retrieved data channel identifier of the power amplifier unit in a register, the controller 250 can accurately locate the retrieved location, ensuring accurate acquisition of the target reference audio data.
[0139] like Figure 8 As shown, the recovery channel identifier and the recovery point identifier can be stored together in the same power amplifier unit register and located in different bit segments in the register. In this way, the controller 250 can obtain the recovery point identifier and the recovery channel identifier respectively by reading the information of different bit segments.
[0140] It should be noted that the sampling channel identifier can also be pre-set. This pre-determined sampling channel identifier can be used to designate a specific pin or channel within the power amplifier unit for transmitting the target reference audio data, thereby simplifying the process by which the controller 250 extracts the target reference audio data from the power amplifier unit. The controller 250 only needs to extract the target reference audio data from the sampling data channel indicated by the sampling channel identifier.
[0141] It is understood that there is an association between the nonlinear recapture point indicated by the recapture point identifier and the pin / channel of the power amplifier unit specified by the recapture channel identifier. That is, the pin / channel specified by the recapture channel identifier is used only to transmit audio data at the data processing stage indicated by the nonlinear recapture point, i.e., the target reference audio data, to the controller 250.
[0142] In this way, during operation, the power amplifier unit can also output media audio data that has undergone a complete modulation process, including nonlinear optimization and PWM modulation, to the audio output device 270 via other pins, enabling the audio output device 270 to output high-quality audio. Therefore, based on the design of the sampling point identifier and the sampling channel identifier, it is possible to distinguish between the sampling data channel and the audio output channel, thereby avoiding confusion between the sampling data and the output data, ensuring that the controller 250 accurately obtains the target reference audio data and ensures the normal output of the audio data.
[0143] like Figure 9 As shown, in some embodiments, the display device 200 is further configured with an audio processing unit. The audio processing unit can be a digital signal processing chip (DSP) specifically configured to process audio effects. It should be noted that the digital signal processing chip is only used to process the audio effects of the audio data to form a rich audio output architecture, thereby enabling the audio output device 270 to output complex audio effects such as Dolby and Devialet.
[0144] In the case where the display device 200 is provided with a sound effect processing unit, the controller 250 is configured as follows:
[0145] The media audio data is input into the audio effect processing unit.
[0146] The sound effect processing unit is controlled to input the media audio data processed by the sound effect into the power amplifier unit.
[0147] The power amplifier unit is controlled to input the modulated media audio data into the audio output device.
[0148] In some embodiments, the audio processing unit is disposed between the controller 250 and the power amplifier unit. That is, the controller 250 first inputs the media audio to the audio processing unit, which then optimizes the media audio before transmitting the optimized media audio to the power amplifier unit.
[0149] In this way, by optimizing the media audio through the sound processing unit, it is possible to achieve effects such as reducing distortion, enhancing bass, optimizing mid- and high-frequency sounds, and thus improving the quality of the output audio.
[0150] It should be noted that the audio processing unit can optimize audio data to achieve rich audio output effects and audio switching effects. However, for the stability and real-time performance of audio output, the audio output architecture needs to be configured based on various factors such as the number of amplifiers and the number of channels.
[0151] like Figure 9 As shown, taking the display device 200 supporting a 5.1.2 multi-channel system as an example, the display device 200 may include five amplifier units, three of which may correspond to the main left channel, main right channel, subwoofer, and center channel of the multi-channel audio architecture, and the remaining two amplifier units may correspond to auxiliary channels. For the sake of audio data processing efficiency, the sound effect processing unit can be connected to the amplifier units corresponding to important channels such as the main channel, thereby centrally processing key audio signals and ensuring audio signal quality.
[0152] The remaining two amplifier units, acting as auxiliary channels, can provide auxiliary sound effects for the main channel. Furthermore, these two amplifier units can directly process the media audio data output by the controller 250 to ensure its authenticity. This ensures that the audio ultimately output by the audio output device 270 retains the authenticity of the media audio and also features the rich sound effects added by the sound effects processing chip.
[0153] It is understood that, based on the description in the above embodiment, the power amplifier unit can be divided into a first power amplifier unit and a second power amplifier unit. For example, the power amplifier unit that receives media audio from the sound processing chip is the first power amplifier unit, and the power amplifier unit that receives media audio from the controller 250 is the second power amplifier unit. The first power amplifier unit is responsible for processing the key audio channel, while the second power amplifier unit processes the auxiliary audio channel. The two work together to ensure the layering and three-dimensionality of the sound effects.
[0154] The controller 250 is configured to:
[0155] Inputting the media audio data into the sound effect processing unit;
[0156] Controlling the audio effect processing unit to input the media audio data into the first power amplifier unit;
[0157] The controller executes inputting the media asset audio data into the sound effect processing unit, and is further configured to:
[0158] The media audio data is input into the second power amplifier unit.
[0159] It is understood that the controller 250 is physically connected to the second power amplifier unit (via physical wiring preset in the circuit board), and the controller 250 is physically connected to the sound processing unit, and the sound processing unit is physically connected to the first power amplifier unit (via physical wiring preset in the circuit board). Therefore, the controller 250 can directly input the media audio to the second power amplifier unit, and input the media audio to the sound processing unit. After processing the media audio, the sound processing unit can input the processed media audio to the first power amplifier unit.
[0160] This hierarchical audio processing architecture enriches the sound effects of media assets while ensuring their authenticity. Furthermore, concentrating audio processing on the primary channel facilitates requirements such as audio-visual synchronization.
[0161] It is understood that the media audio data received by the first amplifier unit is processed by the sound effects chip, resulting in certain differences between the media audio data received by the first amplifier unit and the media audio data received by the second amplifier unit. During the media audio output process, upon receiving the media audio data transmitted by the two amplifier units, the audio output device 270 fuses the two sets of media audio data to output the fused audio. Therefore, the controller 250 needs to fuse the retrieved reference audio data to obtain target reference audio data that has the highest similarity to the final output audio.
[0162] The controller 250 is configured as follows:
[0163] First reference audio data is extracted from the first power amplifier unit according to the first recollection point represented by the first recollection point identifier.
[0164] The second reference audio data is extracted from the second power amplifier unit according to the second recollection point represented by the second recollection point identifier. The first reference audio data and the second reference audio data have different phases; the first reference audio data and the second reference audio data correspond to different sound effects; the first recollection point and the second recollection point represent different audio data processing stages;
[0165] The first reference audio data and the second reference audio data are fused to obtain target reference audio data.
[0166] In some embodiments, different registers can be pre-configured for different power amplifier units, namely, power amplifier unit registers, such as a first power amplifier unit register for storing operational configuration information for a first power amplifier unit and a second power amplifier unit register for storing operational configuration information for a second power amplifier unit. The controller 250 uses these registers, along with the sampling point identifiers and sampling channel identifiers stored therein, to determine how to sample the reference audio data.
[0167] The first recapture point represents the recapture location at the output of the first power amplifier unit. This means that the media audio data corresponding to the recapture location represented by the first recapture point has already been fully processed by the first power amplifier unit. This ensures that the first reference audio data extracted by controller 250 from the first recapture point is highly similar to the interfering audio in the external voice data, thereby filtering out the interfering audio.
[0168] The second recapture point represents the capture location at the nonlinear optimization point of the second power amplifier unit. This means that the media audio data corresponding to the nonlinear optimization point has only been processed by the second power amplifier unit. This setup allows the second reference audio data extracted at the second recapture point to retain more of its original sound quality, ensuring the restoration of the sound quality.
[0169] It is understandable that, when audio media is output, the controller 250 may periodically read the sampling point identifiers in the first power amplifier unit register and the second power amplifier unit register to periodically obtain reference audio data.
[0170] In this way, the controller can combine the original sound quality characteristics of the media audio with the sound quality of the sound effect processing to obtain the target reference audio data, thereby making the target reference audio data closer to the interference noise, which is beneficial to the subsequent interference noise filtering process.
[0171] It should be noted that setting the first sampling point of the first power amplifier unit at the output end of the first power amplifier unit is beneficial to synchronizing the phase of the data sampled from the first power amplifier unit and the second power amplifier unit by the controller 250, thereby ensuring the data quality of the target reference audio.
[0172] It is understood that the controller 250 directly inputs the media audio to the second power amplifier unit for processing via the I2S interface, so the controller 250 and the second power amplifier unit share a single clock. Similarly, the sound processing chip also inputs the media audio to the first power amplifier unit via the I2S interface, meaning that the sound processing chip and the first power amplifier unit share another clock.
[0173] Therefore, the phases of the first reference audio data and the second reference audio data are not synchronized, making it difficult to directly synthesize the target reference audio data. If the first sampling point is set to a nonlinear sampling point, the controller 250 can only sample the first reference audio data of the digital signal from the first power amplifier unit. Furthermore, the first reference audio data and the second reference audio data are not synchronized, making it difficult for the controller 250 to accurately obtain the target reference audio data.
[0174] like Figure 9As shown, in some embodiments, the display device 200 further includes an analog-to-digital conversion unit. The analog-to-digital conversion unit is connected to the controller 250 and the first power amplifier unit, respectively. Thus, the controller 250 can control the first power amplifier unit to transmit the first reference audio data to the analog-to-digital conversion unit, which converts the first reference audio data and synchronizes the phases of the first reference audio data and the second reference data during the conversion process.
[0175] The controller 250 is configured as follows:
[0176] Controlling the first power amplifier unit to input the modulated media audio data into the analog-to-digital conversion unit to obtain third reference audio data. The second reference audio data and the third reference audio data have the same phase; the second reference audio data and the third reference audio data have different phases;
[0177] The third reference audio data output by the analog-to-digital conversion unit is received.
[0178] In some embodiments, the controller 250 can control the first power amplifier unit to transmit the modulated media audio data to the analog-to-digital conversion module. The media audio data modulated by the first power amplifier unit is actually input into the audio output device 270 to drive the audio output device 270 to output the media audio data. Therefore, this data is an analog signal. Furthermore, based on the sampling method, this data also serves as the first reference audio data.
[0179] The signals received by controller 250 from the second power amplifier unit and the microphone array are all digital signals. Therefore, the media audio data modulated by the first power amplifier unit needs to undergo analog-to-digital conversion to unify the reference audio signal. The analog-to-digital conversion unit can convert the received first reference audio data from an analog signal to a digital signal to obtain third reference audio data. In addition, because the analog-to-digital conversion unit and controller 250 can communicate directly based on the I2S interface, the analog-to-digital conversion unit and controller 250 can also be considered to share the same clock.
[0180] In this way, when the analog-to-digital conversion unit transmits the third reference audio data to the controller 250, it ensures that the third reference audio data is phase synchronized with the second reference audio data, so that the controller 250 can accurately synthesize the target reference audio data, which is beneficial to improving the accuracy and stability of subsequent interference audio filtering.
[0181] It is understandable that after obtaining the target reference audio data, the controller 250 may filter out the interference audio according to the target reference audio data.
[0182] In the process of filtering out the interfering audio, the controller 250 is specifically configured to:
[0183] The signal-to-noise ratio of the voice instruction is adjusted according to the audio characteristics of the target reference audio data and the audio data corresponding to the interference audio to obtain a target voice instruction.
[0184] In some embodiments, the signal-to-noise ratio is used to characterize the signal strength relationship between the voice command and the audio data corresponding to the interfering audio. That is, in a voice command, the stronger the signal strength of the audio signal corresponding to the voice command, the stronger the voice recognition engine of the display device 200 is able to recognize the voice command.
[0185] The target voice command is a voice command that has been filtered out of interfering audio. Therefore, the signal-to-noise ratio of the target voice command is greater than the signal-to-noise ratio of the external audio data.
[0186] In addition, since the target reference audio data is the audio data collected by the controller 250 from the power amplifier unit, and the interference audio is the background noise such as the echo formed by the media audio output by the audio output device 270, the similarity between the audio characteristics of the target reference audio data and the audio data of the interference audio is greater than the similarity between the audio characteristics of the target reference audio data and the audio data of the voice command.
[0187] Furthermore, the controller 250 can combine an echo cancellation algorithm such as AEC (Acoustic Echo Canceller) to filter out interfering audio in the external audio data. For example, the signal-to-noise ratio of the audio data corresponding to the voice command and the interfering audio in the external audio data can be adjusted based on the AEC echo cancellation algorithm. In this way, when the signal-to-noise ratio is optimized to a larger value, the voice recognition engine in the display device 200 can accurately and quickly recognize the voice command, thereby improving the recognition effect and efficiency of the voice command.
[0188] It is understood that the processing of echo cancellation algorithms such as AEC is not described in detail in the embodiments of this application. Interference audio in external audio signals can be processed using the conventional processing methods of the AEC echo cancellation algorithm. The embodiments of this application aim to improve the quality of reference audio data based on the setting of echo points, thereby providing a precise audio data comparison basis for the subsequent interference audio filtering process, thereby improving the accuracy and efficiency of voice command recognition.
[0189] In some embodiments, after filtering out interfering audio in the voice command, the controller 250 may input the target voice command into a voice recognition engine, and parse the target voice command based on the voice recognition engine.
[0190] The controller 250 is configured as follows:
[0191] The target voice command is parsed to obtain semantic information of the voice command.
[0192] The target task is determined based on the mapping relationship between semantic information and the target task.
[0193] In some embodiments, the AEC algorithm can be loaded into the controller 250 to adapt to local voice wake-up scenarios, ensuring that voice interaction scenarios such as power on and power off can operate normally even when not connected to the network.
[0194] In other embodiments, for complex voice processing tasks, the controller 250 can use the cloud server's algorithm to perform processing, such as performing echo cancellation and semantic understanding in the cloud.
[0195] like Figure 10 As shown, the mapping relationship between semantic information and target tasks is information pre-stored in the storage space of the display device 200, which can be stored in the form of a mapping table, etc. The controller 250 can query the mapping table to find the target task corresponding to the semantic information and then execute the target task.
[0196] In some embodiments, the display device 200 further includes digital silicon microphones arranged in a linear array in a preset sound collection area of the display. The digital silicon microphones are connected to a user input interface and are configured to: collect voice commands; and input the voice commands into the user input interface.
[0197] In some embodiments, the preset sound collection area can be located in the display 260 or in a base connected to the display 260. The linear array arrangement of digital silicon microphones can effectively improve the accuracy of voice command collection and ensure the stability of multi-channel far-field interaction.
[0198] In some embodiments, the number of digital silicon microphones can be four, forming a linear array arranged in the top bezel area of the display 260. This arrangement facilitates customer-facing viewing, thereby improving voice reception coverage and far-field voice capture. Furthermore, this higher-positioned arrangement effectively reduces background noise, such as desktop reflections, thereby improving the signal-to-noise ratio of voice commands.
[0199] In some embodiments, the silicon microphone array is disposed behind the micro-holes of the display screen to maintain a simple appearance.
[0200] It is understood that the arrangement of the silicon microphone array and the direct connection between the silicon microphone array and the controller 250 via the PDM interface in the present embodiment can effectively improve the efficiency of transmitting voice commands to the controller 250. By optimizing the transmission path and reducing signal delays, it ensures a rapid command response and improves the user experience.
[0201] Furthermore, this embodiment of the present application uses the controller 250 to recapture media audio data from the amplifier as target reference audio data. Based on the similarity between the media audio data corresponding to the recapture point and the media audio data ultimately output by the audio output device 270, it can effectively filter out background noise generated by the media audio data output by the audio output device 270 during voice commands. This effectively improves voice recognition accuracy while reducing hardware design costs, thereby enhancing the ubiquity of far-field voice interaction.
[0202] The present application also provides a method for canceling far-field voice echo, including:
[0203] When the audio output device outputs media audio, read the re-collection point identifier stored in the power amplifier unit register; the power amplifier unit register is used to store the operation configuration information of the power amplifier unit; the operation configuration information of the power amplifier unit includes the re-collection point identifier;
[0204] If a nonlinear re-acquisition point identifier is read, reference audio data is extracted from the power amplifier unit based on the nonlinear re-acquisition point represented by the nonlinear re-acquisition point identifier; wherein the reference audio data is the media resource audio data to be optimized that has undergone nonlinear optimization;
[0205] In response to the voice command, filtering out interference audio in external audio data based on the reference audio data to extract the voice command from the external audio data; the external audio data includes the voice command and / or interference audio corresponding to the media audio;
[0206] Execute the target task indicated by the voice instruction.
[0207] In some embodiments, when the display device 200 outputs media audio via the audio output device 270, it can read the re-sampling point identifier in the power amplifier unit register to determine the re-sampling point for reading the target reference audio data, and then extract the target reference audio data based on the re-sampling point. When the display device 200 receives a voice command, it can filter the target reference audio data containing the voice command based on the target reference audio data to extract a clearer voice command, thereby improving the recognition of the voice command.
[0208] It can be seen from the above technical content that the embodiments of the present application provide a display device and an echo cancellation method. The display device can continuously recapture the media audio data in the power amplifier unit during the process of media audio output, and use the media audio data as the target audio reference data. Furthermore, when the display device receives a voice command, it can filter out the interference audio in the voice command according to the target audio reference data to obtain a target voice command with a higher signal-to-noise ratio. This is conducive to improving the voice recognition capability of the display device, and the setting of the recapture point based on the target reference audio data can reduce the hardware design cost while ensuring the interference audio filtering capability, which is conducive to improving the universality of the far-field voice interaction function.
[0209] Similar parts between the embodiments provided in this application can be referenced to each other. The specific implementation methods provided above are only a few examples under the overall concept of this application and do not constitute a limitation on the scope of protection of this application. For those skilled in the art, any other implementation methods expanded based on the scheme of this application without expending creative work shall fall within the scope of protection of this application.
Claims
1. A display device, characterized in that: include: monitor; The audio output device is configured to: output media audio; The user input interface is configured to: receive a voice command; The user input interface is connected to the controller; the voice command includes interference audio; the interference audio is background noise corresponding to the media audio; A power amplifier unit, which is connected to the controller and the audio output device respectively; The controller is configured as: When the audio output device outputs media audio, read the recollection point identifier stored in the power amplifier unit register; the power amplifier unit register is used to store the operation configuration information of the power amplifier unit; the operation configuration information of the power amplifier unit includes the recollection point identifier; Extracting target reference audio data from a power amplifier unit based on the re-collection point represented by the re-collection point identifier; wherein the target reference audio data is at least media resource audio data that has undergone nonlinear optimization; In response to the voice instruction, filtering out interference audio in the voice instruction based on target reference audio data to obtain a target voice instruction; Execute a target task determined based on the target voice instruction.
2. The display device according to claim 1, wherein The operation configuration information further includes a sampling channel identifier; the sampling channel identifier is used to indicate a sampling data channel used by the power amplifier unit to transmit the target reference audio data; the controller extracts the target reference audio data from the power amplifier unit based on a nonlinear sampling point represented by the nonlinear sampling point identifier, and is specifically configured as follows: Read the acquisition channel identifier in the power amplifier unit register; Determine the acquisition data channel of the power amplifier unit according to the acquisition channel identifier; Extract target reference audio data from the retrieved data channel.
3. The display device according to claim 1, wherein Also includes: The sound effect processing unit is connected to the controller and the power amplifier unit respectively; When the audio output device outputs the media audio, the controller is specifically configured to: Inputting media audio data into the sound processing unit; Controlling the sound effect processing unit to input the media audio data processed by the sound effect into the power amplifier unit; The power amplifier unit is controlled to input the modulated media audio data into the audio output device.
4. The display device according to claim 3, wherein The power amplifier unit includes a first power amplifier unit and a second power amplifier unit; the sound effect processing unit is connected to the first power amplifier unit, and the controller is connected to the second power amplifier unit; The controller controls the sound effect processing unit to input the media audio data processed by the sound effect into the power amplifier unit, and is specifically configured as follows: Inputting media audio data into the sound processing unit; Controlling the audio processing unit to input the media audio data into the first power amplifier unit; The controller executes inputting the media asset audio data into the sound effect processing unit, and is further configured to: The media audio data is input to the second power amplifier unit.
5. The display device according to claim 4, wherein: The controller extracts target reference audio data from the power amplifier unit based on the nonlinear recollection point represented by the nonlinear recollection point identifier, and is specifically configured to: Extracting first reference audio data from a first power amplifier unit according to a first re-collection point represented by a first re-collection point identifier; storing the first re-collection point identifier in a first power amplifier unit register corresponding to the first power amplifier unit; extracting second reference audio data from the second power amplifier unit according to the second recollection point represented by the second recollection point identifier; The first reference audio data and the second reference audio data correspond to different sound effects; the first re-collection point and the second re-collection point represent different audio data processing stages; The first reference audio data and the second reference audio data are fused to obtain target reference audio data.
6. The display device according to claim 5, wherein: The system further includes an analog-to-digital conversion unit; the analog-to-digital conversion unit is connected to the first power amplifier unit and the controller; the controller extracts target reference audio data from the power amplifier unit based on the nonlinear recollection point identification representation, and is specifically configured as follows: Controlling the first power amplifier unit to input the modulated media audio data into the analog-to-digital conversion unit to obtain third reference audio data; the second reference audio data and the third reference audio data have the same phase; the second reference audio data and the third reference audio data have different phases; The third reference audio data output by the analog-to-digital conversion unit is received.
7. The display device according to claim 1, wherein The controller performs filtering of interference audio in the voice instruction based on target reference audio data, and is specifically configured to: The signal-to-noise ratio of the voice command is adjusted based on the audio characteristics of the target reference audio data and the audio data corresponding to the interference audio to obtain the target voice command; the signal-to-noise ratio is used to characterize the signal strength relationship between the main audio data of the voice command and the audio data of the interference audio; the signal-to-noise ratio of the target voice command is greater than the signal-to-noise ratio of the voice command.
8. The display device according to claim 7, wherein: Before executing the target task indicated by the voice instruction, the controller is further configured to: Parsing the target voice command to obtain semantic information of the voice command; The target task is determined based on the mapping relationship between semantic information and the target task.
9. The display device according to claim 1, wherein Also included are digital silicon microphones; the digital silicon microphones are arranged in a linear array in a preset sound collection area of the display; the digital silicon microphones are connected to a user input interface; The digital silicon microphone is configured as follows: Collect voice commands; The voice command is input into a user input interface.
10. A far-field speech echo cancellation method, characterized in that: include: When the audio output device outputs the media audio, read the recollection point identifier stored in the power amplifier unit register; The power amplifier unit register is used to store the operation configuration information of the power amplifier unit; the operation configuration information of the power amplifier unit includes the recovery point identifier; If the nonlinear re-collection point identifier is read, target reference audio data is extracted from the power amplifier unit based on the nonlinear re-collection point represented by the nonlinear re-collection point identifier; wherein the reference audio data is media audio data that has undergone nonlinear optimization; In response to the voice instruction, filtering out interference audio in the voice instruction based on the target reference audio data to obtain a target voice instruction; Execute a target task determined based on the target voice instruction.