Display device and method for preventing false wakeup in far-field voice
By saving and processing the original audio file in the display device, adding feature tags to distinguish between TV playback and user wake-up words, solving false wake-up problems in far-field voice interactions, ensuring that the device only responds when the user wakes up.
Patent Information
- Application Number
- CN202510279569.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-25
AI Technical Summary
There is a problem of false wakeup in far-field voice interaction scenarios, especially when the TV plays content containing wakeup words, echo cancellation fails to completely eliminate residual sound, resulting in the wakeup model being misidentified as a user wakeup command, causing self-question and self-answer.
The display device recognizes the wakeup word by saving the original audio file, performing echo cancellation processing, and adds feature tags to the original audio file to determine whether the wakeup word is the content played by the TV itself, and avoids accidental wakeup.
Ensure that only the wake-up words of external users trigger the device to wake up, avoid false wake-up caused by the content played by the TV itself, improve the accuracy of wake-up words recognition, and reduce false operations.
Smart Images

Figure CN120371249A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of display devices, and in particular, to a display device and a method for preventing false wake-up in far-field voice. Background Art
[0002] The application scenarios of voice interaction in daily life are becoming increasingly extensive. For example, it can cover multiple fields such as smart speakers, smart TVs, smart vehicles, smart homes, and smart robots. In these scenarios, the distance of human-machine voice interaction is no longer limited to the near field. For example, due to the increasing requirements for user experience in human-machine interaction, the distance of human-machine voice conversations is no longer limited to the near field. A far-field voice microphone array can extend the distance of human-machine interaction, enabling users to interact with devices more naturally through voice.
[0003] To accurately recognize wake-up words, the TV needs to preprocess these sounds first, that is, perform noise reduction. In particular, since the sound played by the TV is known, it can be distinguished from other sounds and separately undergo acoustic echo cancellation (AEC) processing to reduce its impact on wake-up word recognition. In some embodiments, the acoustic echo cancellation in the TV scenario mainly relies on traditional AEC algorithms. These algorithms use the audio received by the input microphone and the audio played by the TV (as a reference signal), and utilize specific algorithms (such as NLMS, NNAES, etc.) to eliminate the audio components of the TV playback contained in the microphone audio, thereby retaining external sounds. The far-field voice performance has a great correlation with the AEC (acoustic echo cancellation) of the native playback sound. To improve the acoustic echo cancellation effect, TV devices usually add a DSP chip to handle complex multi-channel sound effect designs. However, with the improvement of the robustness of the wake-up model and the demand for cost control, more and more multi-channel models are beginning to consider sacrificing some acoustic echo cancellation performance and ensuring the wake-up rate by increasing the compatibility of the model with echo residues.
[0004] Although this approach improves the wake-up rate to a certain extent, it may also lead to an increase in the false wake-up rate. Especially when the TV plays content containing wake-up words, if the AEC fails to completely eliminate the residual sound, these residual sounds may be misrecognized as the user's wake-up command by the wake-up model, thus triggering the phenomenon of self-answering. Therefore, there is a problem of false wake-up in the current far-field voice interaction scenario. Summary of the Invention
[0005] Some embodiments of this application provide a display device and a method for preventing false wake-up in far-field voice to solve the problem of false wake-up in the far-field voice interaction scenario.
[0006] In a first aspect, some embodiments of this application provide a display device, including:
[0007] A display configured to display a user interface;
[0008] A controller, configured to:
[0009] In response to an activation instruction of far-field voice, save an original audio file; the original audio file at least includes an audio file containing a wake word and an audio file collected by a sound collection device;
[0010] Obtain a post-echo-cancellation audio file after echo cancellation processing of the original audio file, and identify whether the post-echo-cancellation audio file contains a wake word;
[0011] When the post-echo-cancellation audio file contains a wake word, obtain the original audio file corresponding to the wake word, and identify whether the original audio file contains a feature tag; the feature tag is a signal whose frequency exceeds a preset frequency;
[0012] When the original audio file does not contain the feature tag, perform a wake-up operation according to the wake word;
[0013] When the original audio file contains the feature tag, do not respond to the wake word.
[0014] The above technical solution has the following advantages or beneficial effects: The display device can ensure that the display device is only woken up when the wake word comes from an external user by judging whether the original audio file contains a feature tag, thereby avoiding the phenomenon of false wake-up caused by the content played by the TV itself and solving the problem of false wake-up in the far-field voice interaction scenario.
[0015] In some embodiments, before the step of saving the original audio file, the controller is further configured to:
[0016] Obtain the text information of the content to be played containing a wake word;
[0017] Combine the text information into a sentence;
[0018] Input the sentence into a TTS model to output the audio corresponding to the sentence;
[0019] Generate an original audio file according to the audio.
[0020] The above technical solution has the following advantages or beneficial effects: The display device converts the text information containing a wake word into a playable audio file, and through TTS technology, can output the text content in the form of voice, so as to realize playing voice content in the display device or other voice interaction devices.
[0021] In some embodiments, the controller identifies whether the post-echo-cancellation audio file contains a wake word, and is specifically configured to:
[0022] Input the audio file after elimination frame by frame into the wake word recognition model;
[0023] In the case that the wake word recognition model recognizes the wake word, it is determined that the audio file after elimination contains the wake word;
[0024] In the case that the wake word recognition model does not recognize the wake word, it is determined that the audio file after elimination does not contain the wake word.
[0025] The above technical solution has the following advantages or beneficial effects: Through the wake word recognition model, it can accurately judge whether the audio file after elimination contains the wake word. By analyzing the audio file frame by frame, it ensures that the wake word can be detected in real time, providing a data basis for subsequent judgment of whether the audio is played by the TV itself.
[0026] In some embodiments, before the step where the controller recognizes whether the original audio file contains a feature tag, it is further configured to:
[0027] Obtain the original audio file containing the wake word;
[0028] Identify the target frame position corresponding to the wake word contained in the original audio file;
[0029] Add the feature tag at the target frame position.
[0030] The above technical solution has the following advantages or beneficial effects: After the display device recognizes the target frame position of the wake word, it can add feature tags to these frames to mark the position of the wake word, facilitating subsequent processing or retrieval.
[0031] In some embodiments, when the controller adds the feature tag at the target frame position, it is specifically configured to:
[0032] Obtain the start position and end position corresponding to the target frame position;
[0033] Insert the feature tag between the start position and the end position.
[0034] The above technical solution has the following advantages or beneficial effects: By accurately positioning the start and end positions of the wake word and inserting the feature tag between them, it can provide an accurate data basis for the processing and analysis of the audio file, providing an accurate basis for subsequent steps.
[0035] In some embodiments, when the controller recognizes whether the original audio file contains a feature tag, it is specifically configured to:
[0036] Calculate the sound energy value of the original audio file containing the wake word;
[0037] When the sound energy value is greater than or equal to the energy threshold, it is determined that the original audio file contains a feature tag;
[0038] When the sound energy value is less than the energy threshold, it is determined that the original audio file does not contain a feature tag.
[0039] The above technical solution has the following advantages or beneficial effects: By judging the source of the wake-up word, the display device can avoid unnecessary processing of false wake-up events, ensure that the device responds only when the user actively triggers the wake-up, and reduce misoperations.
[0040] In some embodiments, before the step of the controller calculating the sound energy value of the original audio file containing the wake-up word, it is further configured to:
[0041] Obtain the original audio file without echo cancellation processing;
[0042] Extract the feature tags contained in the original audio file;
[0043] Determine the original audio file corresponding to the feature tag containing the wake-up word according to the feature tag.
[0044] The above technical solution has the following advantages or beneficial effects: Through the feature tag, the display device can efficiently identify the audio file containing the wake-up word, provide a data basis for subsequent judgment of whether it is a false wake-up, and avoid unnecessary processing of irrelevant files.
[0045] In some embodiments, the controller obtains the post-echo-cancellation audio file after performing echo cancellation processing on the original audio file, and is specifically configured to:
[0046] Perform echo cancellation processing on the audio file containing the wake-up word;
[0047] Generate the post-echo-cancellation audio file according to the audio file after echo cancellation processing and the audio file collected by the sound collection device.
[0048] The above technical solution has the following advantages or beneficial effects: Through echo cancellation processing, the display device can effectively reduce the interference of the device's own played sound on wake-up word recognition, improve the accuracy of wake-up word recognition. At the same time, recognize the wake-up word in the post-echo-cancellation audio file to provide a basis for subsequent false wake-up judgment.
[0049] In some embodiments, after the step of the controller identifying whether the original audio file contains a feature tag, it is further configured to:
[0050] In the case that the feature tag is not included in the original audio file, determine that the original audio file is an audio file containing a wake-up word triggered by the user, and perform the step of performing a wake-up operation according to the wake-up word;
[0051] In the case that the feature tag is included in the original audio file, determine that the original audio file is an audio file for miswakening triggered by audio played by the display device, and perform the step of not responding to the wake-up word.
[0052] The above technical solution has the following advantages or beneficial effects: By determining whether the original audio file contains a feature tag, the display device can distinguish between a valid wake-up triggered by the user and a miswakening caused by the device's own played sound, and accordingly decide whether to perform a wake-up operation.
[0053] In a second aspect, some embodiments of the present application provide a method for preventing miswakening in far-field voice, which can be applied to the display device in the first aspect. The display device includes a display and a controller. The method includes:
[0054] In response to an enabling instruction for far-field voice, save the original audio file; the original audio file at least includes an audio file containing a wake-up word and an audio file collected by a sound collection device;
[0055] Obtain an audio file after echo cancellation processing of the original audio file, and identify whether the wake-up word is included in the audio file after cancellation;
[0056] In the case that the wake-up word is included in the audio file after cancellation, obtain the original audio file corresponding to the wake-up word, and identify whether the feature tag is included in the original audio file; the feature tag is a signal whose frequency exceeds a preset frequency;
[0057] In the case that the feature tag is not included in the original audio file, perform a wake-up operation according to the wake-up word;
[0058] In the case that the feature tag is included in the original audio file, do not respond to the wake-up word.
[0059] The above technical solution has the following advantages or beneficial effects: The method can ensure that the display device is only woken up when the wake-up word comes from an external user by determining whether the original audio file contains a feature tag, thereby avoiding the miswakening phenomenon caused by the content played by the TV itself and solving the miswakening problem in the far-field voice interaction scenario.
[0060] As can be seen from the above technical solutions, some embodiments of the present application provide a display device and a method for preventing false wake-up in far-field voice. The method includes: in response to an enabling instruction of far-field voice, the display device saves the original audio file; the original audio file at least includes an audio file containing a wake-up word and an audio file collected by a sound collection device; obtaining an audio file after echo cancellation processing of the original audio file, and identifying whether the wake-up word is included in the audio file after cancellation; in the case where the wake-up word is included in the audio file after cancellation, obtaining the original audio file corresponding to the wake-up word, and identifying whether a feature tag is included in the original audio file; the feature tag is a signal representing a frequency exceeding a preset frequency; in the case where the original audio file does not include the feature tag, performing a wake-up operation according to the wake-up word; in the case where the original audio file includes the feature tag, not responding to the wake-up word. The method can ensure that the display device is only woken up when the wake-up word comes from an external user by judging whether the original audio file includes a feature tag, thereby avoiding the false wake-up phenomenon caused by the content played by the TV itself and solving the false wake-up problem in the far-field voice interaction scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in some embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0062] Figure 1 It is a schematic diagram of an operation scenario between a display device and a control device provided by some embodiments of the present application;
[0063] Figure 2 It is a schematic diagram of the hardware configuration of a display device provided by some embodiments of the present application;
[0064] Figure 3 It is a schematic diagram of the software configuration of a display device provided by some embodiments of the present application;
[0065] Figure 4 It is a schematic diagram of an echo cancellation and noise suppression scenario shown by some embodiments of the present application;
[0066] Figure 5 It is a schematic diagram of an audio processing architecture shown by some embodiments of the present application;
[0067] Figure 6 It is a schematic diagram of the timing of a method for preventing false wake-up in far-field voice executed by a display device provided by some embodiments of the present application;
[0068] Figure 7Schematic diagram of a scenario for a display device provided in some embodiments of the present application to generate an original audio file;
[0069] Figure 8 Effect diagram of a display device provided in some embodiments of the present application to generate an original audio file;
[0070] Figure 9 Schematic diagram of a scenario for a display device provided in some embodiments of the present application to add feature tags to an original audio file;
[0071] Figure 10 Effect diagram of a display device provided in some embodiments of the present application to add feature tags to an original audio file. Detailed implementation manners
[0072] The embodiments will be described in detail below, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following embodiments do not represent all implementation manners consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application detailed in the claims.
[0073] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the following described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise stated, these terms should be understood in their ordinary and common meanings.
[0074] The terms "first", "second", "third", etc. in the specification, claims and the above accompanying drawings of the present application are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.
[0075] The terms "comprising" and "having" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device comprising a series of components does not necessarily have to be limited to all the components clearly listed, but may include other components not clearly listed or inherent to these products or devices.
[0076] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or a combination of hardware or / and software code that can perform functions related to the element.
[0077] In the embodiments of the present application, the display device 200 generally refers to a device with the capabilities of screen display and data processing. For example, the display device 200 includes but is not limited to smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc.
[0078] Figure 1 Schematic diagram of the operation scenario between the display device and the control device provided for some embodiments of the present application. As Figure 1 shown, the user can operate the display device 200 through touch operations, a mobile terminal 300, and a control device 100. Among them, the control device 100 is used to receive the operation instructions input by the user and convert the operation instructions into control instructions that the display device 200 can recognize and respond to. For example, the control device 100 can be a remote control, a stylus, a handle, etc.
[0079] The mobile terminal 300 can be used as a control device to perform human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a communication device to establish a communication connection with the display device 200 for data interaction. In some embodiments, software applications can be installed on the mobile terminal 300 and the display device 200, and the connection communication can be realized through network communication protocols to achieve the purpose of one-to-one control operation and data communication. It is also possible to transmit the audio and video content displayed on the mobile terminal 300 to the display device 200 to achieve the synchronous display function.
[0080] In some embodiments, the mobile terminal 300 or other electronic devices can also simulate the functions of the control device 100 by running an application program for controlling the display device 200.
[0081] As Figure 1 also shown, the display device 200 also communicates with the server 400 for data communication through various communication methods. The display device 200 is allowed to establish a communication connection through a local area network (LAN), a wireless local area network (WLAN), and other networks.
[0082] The display device 200 can provide a broadcast receiving TV function, and can also additionally provide an intelligent network TV function with computer support functions, including but not limited to, network TV, smart TV, Internet Protocol TV (IPTV), etc.
[0083] Figure 2 Provided for some embodiments of the present application Figure 1 Hardware configuration block diagram of the display device 200 in
[0084] In some embodiments, the display device 200 may include at least one of a tuner demodulator 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.
[0085] In some embodiments, the detector 230 is used to collect signals of the external environment or for external interactions. For example, the detector 230 includes a light receiver, a sensor for collecting the ambient light intensity; alternatively, the detector 230 includes an image collector, such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures. Or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.
[0086] In some embodiments, the display 260 includes a display functional component for presenting a picture and a driving component for driving image display. The display 260 is used to receive the image signal output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, components of a menu manipulation interface, and a user manipulation UI interface, etc.
[0087] In some embodiments, the communication device 220 is a component for communicating with an external device or server 400 according to various communication protocol types. The display device 200 can be provided with multiple communication devices 220 according to different supported communication methods. For example, when the display device 200 supports wireless network communication, the display device 200 can be provided with a communication device 220 including a WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.
[0088] The communication device 220 can enable the display device 200 to communicate with an external device or server 400 in a wireless or wired connection manner. Among them, the wired connection can connect the display device 200 with an external device through components such as a data cable and an interface. The wireless connection can connect the display device 200 with an external device through a wireless signal or a wireless network. The display device 200 can directly establish a connection relationship with an external device, or can indirectly establish a connection relationship through a gateway, a router, a connection device, etc.
[0089] In some embodiments, the controller 250 can include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and first to n interfaces for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.
[0090] In some embodiments, the controller 250 and the tuner demodulator 210 can be located in different split devices, that is, the tuner demodulator 210 can also be in an external device of the main device where the controller 250 is located, such as an external set-top box, etc.
[0091] In some embodiments, the user may input a user command through a graphical user interface (GUI) displayed on the display 260 , and the user input interface receives the user input command through the graphical user interface (GUI).
[0092] In some embodiments, the audio output device 270 may be a local speaker of the display device 200, or may be an external audio output device of the display device 200. In particular, for the external audio output device of the display device 200, the display device 200 may also be provided with an external audio output terminal, and the audio output device may be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.
[0093] In some embodiments, the user input interface 280 may be used to receive instructions from a user.
[0094] In order to perform user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program for managing and controlling hardware resources and software resources in the display device 200. The operating system can control the display device to provide a user interface, for example, the operating system can directly control the display device to provide a user interface, or can provide a user interface by running an application program. The operating system also allows the user to interact with the display device 200.
[0095] It should be noted that the operating system may be a native operating system based on a specific operating platform, or a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.
[0096] The operating system can be divided into different modules or layers according to the functions implemented, such as Figure 3 As shown, in some embodiments, the system is divided into four layers, from top to bottom, namely, the application layer (Applications) layer (referred to as "application layer"), the application framework layer (Application Framework) layer (referred to as "framework layer"), the system library layer and the kernel layer.
[0097] In some embodiments, the application layer is used to provide services and interfaces for applications so that the display device 200 can run applications and interact with users based on the applications. At least one application can be run in the application layer, and these applications can be window programs, system settings programs, clock programs, etc. that come with the operating system; they can also be applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the above examples.
[0098] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications. The application framework layer includes some predefined functions. The application framework layer is equivalent to a processing center that determines the actions of the applications in the application layer. Through the API interface, an application can access the resources in the system and obtain the services of the system during execution.
[0099] As Figure 3 shown, in the embodiments of the present application, the application framework layer includes a view system, managers, content providers, etc. Among them, the view system can design and implement the interfaces and interactions of applications. The view system includes lists, grids, text boxes, buttons, etc. The managers include at least one of the following modules: The activity manager is used to interact with all the activities running in the system; the location manager is used to provide access to the system location service for system services or applications; the package manager is used to retrieve various information related to the application packages currently installed on the device; the notification manager is used to control the display and clearing of notification messages; the window manager is used to manage the icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.
[0100] In some embodiments, the activity manager is used to manage the life cycles of various applications and the general navigation back function, such as controlling the exit, opening, and backward movement of applications. The window manager is used to manage all window programs, such as obtaining the size of the display screen, determining whether there is a status bar, locking the screen, taking screenshots, and controlling the changes of the display window. For example, shrinking the display window, jittering the display, and distorting the display, etc.
[0101] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system will run the instruction libraries included in the system runtime library layer, such as C / C++ instruction libraries, to implement the functions that the framework layer needs to achieve.
[0102] In some embodiments, the kernel layer is a functional layer between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, as Figure 3As shown, hardware drivers can be configured in the kernel layer. The drivers included in the kernel layer can be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor drivers (such as fingerprint sensors, temperature sensors, pressure sensors, etc.), and power drivers, etc.
[0103] It should be noted that the above examples are only simple divisions of the functions of the operating system, and do not limit the specific form of the operating system of the display device 200 in the embodiments of the present application. According to factors such as the functions of the display device and the type of the operating system, the number of hierarchical levels and the specific hierarchical types included in the operating system can be in other forms.
[0104] In some embodiments, the application scenarios of voice interaction in daily life are becoming increasingly extensive. For example, it can cover multiple fields such as smart speakers, smart TVs, smart vehicles, smart homes, and smart robots. In these scenarios, the distance of human-machine voice interaction is no longer limited to the near field. For example, due to the increasing requirements for user experience in human-machine interaction, the distance of human-machine voice dialogue is also less and less limited to the near field. The far-field voice microphone array can extend the distance of human-machine interaction, enabling users to interact with devices more naturally through voice.
[0105] In the far-field voice interaction technology chain in the TV scenario, the interruption wake-up rate is a crucial performance indicator. When a user wishes to interact with a TV that is playing content through voice, they need to start the interaction process by saying a preset wake-up word (such as "Hello, Xiao A"). At this time, the TV needs to respond quickly, reducing or muting its own playback volume so that its microphone can accurately capture and recognize the user's voice.
[0106] To achieve the above process, the TV needs to continuously monitor the sounds in the external environment to promptly recognize the user's wake-up word. However, during the monitoring process, the TV will receive sounds from multiple sound sources, including the user's voice, the sound played by the TV itself, environmental noises in the kitchen, and the sounds made by pets, etc. To accurately recognize the wake-up word, the TV needs to preprocess these sounds first, that is, noise reduction. In particular, since the sound played by the TV is known, it can be distinguished from other sounds and separately subjected to echo cancellation (AEC) processing to reduce its impact on wake-up word recognition.
[0107] In some embodiments, echo cancellation in a TV scenario mainly relies on traditional AEC algorithms. These algorithms use the audio received by the input microphone and the audio played by the TV (as a reference signal), and utilize specific algorithms (such as NLMS, NNAES, etc.) to eliminate the TV-played audio components contained in the microphone audio, thereby retaining external sounds. For example, commonly used algorithms include the normalized least mean square (NLMS) algorithm, which is applicable to linear echo signals; the neural network echo cancellation (NNAES) algorithm, which is applicable to processing non-linear echo signals; and methods based on power spectral density and correlation. The basic audio transmission path of these algorithms includes microphone signal acquisition, reference signal input, echo cancellation processing, and output of the processed audio signal.
[0108] Exemplarily, Figure 4 FIG. is a schematic diagram of an echo cancellation and noise suppression scenario shown in some embodiments of the present application. As Figure 4 shown, this figure shows the process of echo cancellation and noise suppression in a voice communication system, which can be applied to scenarios such as TV voice wake-up. Among them, the far-end signal s(k) is a voice signal sent remotely to the local end. The speaker is used to play the far-end signal. The voice intercom detector is used to detect whether the current is in a voice intercom state to determine whether echo cancellation and noise suppression processing are required. The residual echo canceller is used to cancel the echo of the far-end signal that is captured again by the microphone after being output by the speaker. The adaptive filter is used to estimate and cancel the echo by comparing the far-end signal and the near-end signal to generate an echo estimate x'(k). The echo estimate x'(k) can be the echo estimate signal generated by the adaptive filter. The near-end signal u(k) is the signal captured by the local microphone, including local voice and possible echoes. The microphone can be used to capture the near-end signal. The error signal e(k) is the signal obtained by subtracting the echo estimate from the near-end signal and should theoretically only contain local voice. The echo signal x(k) can be the signal output by the speaker that is captured by the microphone after being reflected by the room. The noise signal v(k) can be the background noise in the environment. The voice signal y(k) can be the signal that theoretically only contains local voice after being processed and is used for voice wake-up or other voice processing tasks. In the TV voice wake-up scenario, this flowchart describes how to extract a clear local voice signal from a complex signal containing far-end voice, echo, and noise for voice recognition or wake-up operations. Through the adaptive filter and the residual echo canceller, the system can effectively reduce the influence of echo and noise and improve the accuracy of voice recognition.
[0109] Figure 5 FIG. is a schematic diagram of an audio processing architecture shown in some embodiments of the present application. As Figure 5As shown, this figure mainly describes the acquisition, processing, and playback processes of audio signals. Among them, the microphone is used to collect sound signals and convert them into electrical signals. The main-channel PA (power amplifier) is responsible for processing the audio signals of the main channel and amplifying them to drive the speaker. The subwoofer PA (power amplifier) is responsible for processing the audio signals of the subwoofer channel and amplifying them to drive the subwoofer speaker. The main chip can include two main modules, AEC (echo cancellation) and playback service, which are responsible for the processing and playback of audio signals. AEC (echo cancellation) can be used to eliminate the echo contained in the sound signals collected by the microphone to improve the clarity of calls or recordings. The playback service can be used to be responsible for playing audio signals, such as including functions like decoding and mixing. The reference signal - left channel is the audio signal output from the main-channel PA and is used for the AEC module to perform echo cancellation processing; the reference signal - right channel is the audio signal output from the main-channel PA and is used for the AEC module to perform echo cancellation processing; the reference signal - subwoofer is the audio signal output from the subwoofer PA and is used for the AEC module to perform echo cancellation processing. Through Figure 5 the audio processing logic shown, the AEC module uses the reference signals output from the main channel and the subwoofer PA to eliminate echoes to ensure the quality of the audio output.
[0110] Combined Figure 4 with Figure 5 , it is easy to see that there is a great correlation between the far-field voice performance and the AEC (echo cancellation) of the native playback sound. To improve the echo cancellation effect, TV devices usually add a DSP chip to process complex multi-channel sound effect designs, such as 5.1.2 or 6.2.2 channel systems. However, with the improvement of the robustness of the wake-up model and the demand for cost control, more and more multi-channel models begin to consider sacrificing some echo cancellation performance to ensure the wake-up rate by increasing the compatibility of the model with echo residues. Although this approach improves the wake-up rate to a certain extent, it may also lead to an increase in the false wake-up rate. Especially when the TV plays content containing the wake-up word, if the AEC fails to completely eliminate the residual sound, these residual sounds may be misrecognized as the user's wake-up command by the wake-up model, thus causing the phenomenon of self-answering questions.
[0111] Exemplarily, in order to reduce the cost of far-field voice interaction, some devices choose not to use a DSP chip to implement echo cancellation, which will result in a decrease in the signal-to-noise ratio of the audio after echo cancellation, that is, there is a certain residual sound in the sound playback. When the TV plays content containing the wake-up word, the residual sound may be misrecognized as the user's wake-up command, thus causing the false wake-up phenomenon. For example, when the TV plays "You can shout 'Hello, Xiao A' to wake me up", due to the incomplete elimination of the echo cancellation, the residual sound is input into the wake-up model, which may trigger a false wake-up and cause the phenomenon of self-answering questions.
[0112] That is to say, in the far-field voice interaction scenario, the residual signal of echo cancellation cannot be completely avoided at present. Thus, when the display device (such as a TV) plays the audio containing the wake word, since the echo cancellation (AEC) fails to completely eliminate the residual of the played audio, and the wake-up model in the display device can only probabilistically determine whether the audio is a wake-up audio and cannot accurately distinguish the residual signal from the external user's voice, it will cause the display device to misidentify the audio played by itself as the user's wake-up instruction, thereby triggering the phenomenon of false wake-up. This false wake-up not only affects the user experience but also may cause the display device to execute instructions under unexpected circumstances, reducing the reliability of the system and user satisfaction.
[0113] To solve the problem of false wake-up in the far-field voice interaction scenario, some embodiments of the present application provide a display device 200, and the display device 200 includes a display 260 and a controller 250. Among them, the display 260 is configured to display a user interface, and the controller 250 runs an application program to enable the display device 200 to execute a method for preventing false wake-up in far-field voice. In the current mainstream solutions for preventing false wake-up in far-field voice, it completely relies on models and algorithms to prevent false wake-up, which requires high computing power and high cost. Compared with the current mainstream solutions for preventing false wake-up, the display device 200 can identify and process the audio from the source of the audio file. In the audio containing the wake word, an audio mark with a feature label is added. After identifying the wake word in the audio, by analyzing whether the feature label exists in the original audio file, it is determined whether the audio is the residual of the wake word played by the display device itself or triggered by an external user, thereby solving the problem of false wake-up in far-field voice.
[0114] To facilitate the understanding of the technical solutions in some embodiments of the present application, the following describes each step in detail with reference to some specific embodiments and drawings. Figure 6 The timing diagram of the method for preventing false wake-up in far-field voice executed by the display device provided in some embodiments of the present application is as follows Figure 6 As shown, in some embodiments, when the display device 200 executes the method for preventing false wake-up in far-field voice, it may include the following steps:
[0115] Step S1: In response to the start instruction of far-field voice, the display device 200 saves the original audio file.
[0116] In some embodiments, when the user enables the far-field voice function through a voice command or device settings, the display device 200 can first save the original audio file. The original audio file at least includes the audio file containing the wake word and the audio file collected by the sound collection device. For example, it at least includes the audio file played by the display device 200 itself, which may contain a wake word (such as "Hello, Xiao A"); it may also include the external environment audio file collected by a sound collection device such as a microphone, such as the user's voice, environmental sounds, etc. By saving the original audio file, it can provide basic data for subsequent echo cancellation and wake word recognition.
[0117] In some embodiments, before saving the original audio file, the display device 200 can perform the following steps. The display device 200 can obtain the text information of the content to be played containing the wake word, combine the text information into a sentence, input the sentence into a Text-to-Speech (TTS) model, that is, the TTS model, to output the audio corresponding to the sentence, and then generate the original audio file according to the audio.
[0118] Exemplarily, Figure 7 FIG. is a schematic diagram of a scenario for a display device provided in some embodiments of the present application to generate an original audio file. Figure 8 FIG. is a schematic diagram of the effect of a display device provided in some embodiments of the present application to generate an original audio file. Combining Figure 7 with Figure 8 , in some embodiments, it can be implemented through the TTS model. The TTS model is a technology that converts text into speech and can convert the input text into natural and fluent speech. Specifically, the display device 200 can first combine the content to be broadcast into a sentence in the form of text information, that is, combine it into a sentence, and then input the sentence into the TTS model, and the corresponding audio is output through the TTS model. For example, as Figure 8 shown, taking a TV as an example, a common wake-up statement in the TV is "You can wake me up by shouting Hello, Xiao A", and its corresponding audio signal is as Figure 8 shown. The display device 200 converts the text information containing the wake word into a playable audio file. Through the TTS technology, the text content can be output in the form of speech, so as to realize playing the speech content in the display device or other voice interaction devices. After step S1 is completed, the following step S2 can be included.
[0119] Step S2: The display device 200 obtains the audio file after echo cancellation processing of the original audio file, and determines whether the wake word is included in the audio file after cancellation.
[0120] After saving the original audio file, the display device 200 performs echo cancellation processing on it. In some embodiments, the display device 200 can obtain the audio file after the original audio file is processed by echo cancellation in the following manner: the display device 200 can first perform echo cancellation processing on the audio file containing the wake-up word, and then generate the audio file after the echo cancellation processing according to the audio file after the echo cancellation processing and the audio file collected by the sound receiving device.
[0121] Exemplarily, the purpose of echo cancellation is to remove the sound played by the device itself and only retain the sound of the external environment. Echo cancellation can use existing echo cancellation algorithms (such as NLMS, NNAES, etc.) to eliminate the TV playback sound contained in the audio collected by the microphone, thereby obtaining an audio file after elimination. Then, the display device 200 can perform wake-up word recognition on the audio file after elimination to determine whether it contains a preset wake-up word (such as "Hello Xiao A"). Through echo cancellation processing, the display device 200 can effectively reduce the interference of the device's own playback sound on the wake-up word recognition and improve the accuracy of wake-up word recognition. At the same time, the wake-up word in the audio file after elimination is identified to provide a basis for subsequent false wake-up judgments.
[0122] In some embodiments, the display device 200 may identify whether the audio file after elimination contains the wake-up word in the following manner. The display device 200 may first input the audio file after elimination into the wake-up word recognition model frame by frame, and if the wake-up word recognition model recognizes the wake-up word, it is determined that the audio file after elimination contains the wake-up word; if the wake-up word recognition model does not recognize the wake-up word, it is determined that the audio file after elimination does not contain the wake-up word.
[0123] Exemplarily, the display device 200 can decompose the audio file after elimination into frame-by-frame audio data, and then input these frames one by one into the wake-up word recognition model. The wake-up word recognition model analyzes the input audio frame to determine whether the current frame or continuous frame contains a preset wake-up word. According to the output result of the wake-up word recognition model, the display device 200 can determine whether the audio file after elimination contains the wake-up word. In this way, the wake-up word recognition model can accurately determine whether the audio file after elimination contains the wake-up word. By analyzing the audio file frame by frame, it is ensured that the wake-up word can be detected in real time, providing data basis for subsequent judgment of whether it is the audio played by the TV itself. For example, when the wake-up word is not included, the display device 200 may not perform any processing. When the wake-up word is included, it is determined whether it is the wake-up word remaining in the TV broadcast itself, providing data basis for subsequent false wake-ups. After step S2 is executed, the following step S3 may be included.
[0124] Step S3: When the wake-up word is included in the post-elimination audio file, the display device 200 acquires the original audio file corresponding to the wake-up word, and identifies whether the original audio file contains a feature tag.
[0125] In some embodiments, before identifying whether the original audio file contains a feature tag, the display device 200 may acquire the original audio file containing the wake-up word, identify the target frame position corresponding to the wake-up word included in the original audio file, and then add a feature tag at the target frame position.
[0126] Exemplarily, the display device 200 may analyze the original audio file to identify the specific frame position corresponding to the wake-up word in the audio file, facilitating subsequent processing. After the display device 200 identifies the target frame position of the wake-up word, it may add a feature tag to these frames to mark the position of the wake-up word, facilitating subsequent processing or retrieval.
[0127] In some embodiments, when adding a feature tag at the target frame position, the display device 200 may first acquire the start position and end position corresponding to the target frame position, and then insert the feature tag between the start position and the end position.
[0128] Exemplarily, Figure 9 FIG. is a schematic diagram of a scenario where a display device provided in some embodiments of the present application adds a feature tag to an original audio file. Figure 10 FIG. is a schematic diagram of the effect of a display device provided in some embodiments of the present application adding a feature tag to an original audio file. Combining Figure 9 with Figure 10 , the display device 200 can determine the start position and end position of the target frame corresponding to the wake-up word by analyzing the audio file, thereby more accurately capturing the range of the wake-up word. Then, a feature tag is inserted between the determined start position and end position to mark the range of the wake-up word, as Figure 10 shown. By accurately positioning the start and end positions of the wake-up word and inserting a feature tag therebetween, an accurate data basis can be provided for the processing and analysis of the audio file, providing an accurate basis for subsequent steps.
[0129] In some embodiments, the feature tag may be a signal with a frequency exceeding a preset frequency, and the preset frequency may be set in combination with the actual usage scenario. For example, the frequency signal of the feature tag should be inaudible to the human ear and not affect the playback and reception of audio in other frequency bands, can be played by the TV speaker, and can also be received by the microphone. Thus, the way of using a sound tag can be adopted to indicate that the wake-up word here is played by the TV.
[0130] Exemplarily, the feature tag of the preset frequency can select a high-frequency audio signal such as 7.5K Hz, 8.5K Hz, 9.5K Hz or more than ten kHz. The display device 200 can embed a high-frequency tag feature (such as a sine wave above 7.5K Hz) in the audio played locally, and form a feature tag in the form of high-frequency mixing (such as wav format) to identify that the audio contains the wake-up word.
[0131] In some embodiments, when the far-field voice switch of the display device 200 is turned on, the sound receiving device such as a microphone will always be in the sound receiving state. At this time, the microphone will receive the external sound plus the sound played by the TV itself. In the echo cancellation process, the echo cancellation algorithm will eliminate the sound played by the TV in the original audio file, and at the same time, the feature label will also be eliminated. In order to subsequently determine whether the wake-up word is the residual of the TV's own playback, the original audio file needs to be saved in advance.
[0132] Therefore, if the display device 200 recognizes the wake-up word in the eliminated audio file, it will further obtain the original audio file corresponding to the wake-up word. Then, the original audio file is subjected to feature tag recognition to determine whether it contains a preset feature tag, thereby determining whether the wake-up word is played by the display device 200 itself.
[0133] In some embodiments, the display device 200 can obtain an original audio file that has not been processed by echo cancellation, extract feature tags contained in the original audio file, and then determine the original audio file containing the wake-up word corresponding to the feature tags based on the feature tags.
[0134] Exemplarily, the display device 200 first obtains an original audio file that has not been processed by echo cancellation. The audio file can be directly collected by a recording device to provide a basis for subsequent feature extraction and wake-up word recognition. Afterwards, the display device 200 can analyze the original audio file, extract the feature tags contained therein, and then determine which original audio files contain the wake-up word based on the extracted feature tags. Through the feature tags, the display device 200 can efficiently identify audio files containing wake-up words, provide a data basis for subsequent judgment of whether it is a false wake-up, and avoid unnecessary processing of irrelevant files.
[0135] In some embodiments, the display device 200 can identify whether a feature tag is included in the original audio file in the following manner. First, the display device 200 can calculate the sound energy value of the original audio file containing the wake word. When the sound energy value is greater than or equal to the energy threshold, it is determined that the original audio file contains a feature tag; when the sound energy value is less than the energy threshold, it is determined that the original audio file does not contain a feature tag. By identifying the feature tag in the original audio file, the display device 200 can accurately determine whether the wake word is played by the device itself, thereby effectively distinguishing between user-initiated wake-up and device false wake-up, and improving the user experience.
[0136] Exemplarily, when the wake word is included in the post-echo-cancellation audio file, it is necessary to determine the specific source of the wake word to determine whether it is a user-initiated wake-up or a false wake-up caused by the sound played by the device itself. The original audio file (without echo cancellation) saved in advance can be subjected to feature analysis to extract high-frequency signals. For example, if the mixed high frequency is a 7.5KHz sine wave, through a band-pass filter, it can be determined that the sound energy in the 7.5K frequency band is significantly higher than that in other frequency bands (such as "5KHz"), then it can be determined that this audio is an audio containing a feature tag, and the wake-up event at this time is triggered by the sound played by the TV itself (false wake-up), not a user-initiated wake-up. That is, this wake-up can be ignored and no processing is performed. By determining the source of the wake word, the display device 200 can avoid unnecessary processing of false wake-up events, ensure that the device responds only when the user actively triggers wake-up, and reduce misoperations. After step S3 is completed, the following step S4 can be executed.
[0137] Step S4: When the original audio file does not contain a feature tag, the display device 200 performs a wake-up operation according to the wake word.
[0138] In some embodiments, if the display device 200 does not recognize a feature tag in the original audio file, it can be determined that the original audio file is an audio file containing a wake word triggered by the user, that is, it can be determined that the wake word is actively issued by the user. At this time, the display device can be woken up according to the wake word to enter the voice interaction mode. By determining whether a feature tag is included in the original audio file, it can be ensured that the display device 200 is only woken up when the wake word comes from an external user, thereby avoiding the false wake-up phenomenon caused by the content played by the TV itself and solving the false wake-up problem in the far-field voice interaction scenario. After step S4 is completed, the following step S5 can be included.
[0139] Step S5: When the original audio file contains a feature tag, the display device 200 does not respond to the wake word.
[0140] In some embodiments, if the display device 200 recognizes a feature tag in the original audio file, it will determine that the wake-up word is a false wake-up caused by the content played by the TV itself, that is, it is determined that the wake-up word is played by the device itself and belongs to a false wake-up. At this time, the display device 200 will ignore this wake-up event and not perform any wake-up operations to avoid the phenomenon of self-answering. By ignoring the wake-up word containing the feature tag, the display device 200 can effectively avoid the false wake-up phenomenon caused by the content played by the TV itself, and improve the robustness of the system and the user experience.
[0141] That is to say, by determining whether the original audio file contains a feature tag, the display device 200 can distinguish the valid wake-up triggered by the user from the false wake-up caused by the sound played by the device itself, and decide whether to perform a wake-up operation accordingly. Through the judgment of the feature tag, the source of the wake-up event can be accurately identified, ensuring that the device responds only when the user actively triggers the wake-up, reducing misoperations and unnecessary processing, and solving the false wake-up problem in the far-field voice interaction scenario by ignoring the false wake-up event.
[0142] As can be seen from the above technical solutions, the above embodiments provide a display device 200. In response to the start instruction of far-field voice, the display device 200 saves the original audio file; the original audio file at least includes the audio file containing the wake-up word and the audio file collected by the sound collection device; obtains the processed audio file after echo cancellation of the original audio file, and identifies whether the processed audio file after echo cancellation contains a wake-up word; in the case where the processed audio file after echo cancellation contains a wake-up word, obtains the original audio file corresponding to the wake-up word, and identifies whether the original audio file contains a feature tag; the feature tag is a signal whose frequency exceeds a preset frequency; in the case where the original audio file does not contain a feature tag, performs a wake-up operation according to the wake-up word; in the case where the original audio file contains a feature tag, does not respond to the wake-up word. The display device 200 can ensure that the display device 200 is only awakened when the wake-up word comes from an external user by determining whether the original audio file contains a feature tag, thereby avoiding the false wake-up phenomenon caused by the content played by the TV itself and solving the false wake-up problem in the far-field voice interaction scenario.
[0143] Based on the above display device 200, some embodiments of the present application further provide a method for preventing false wake-up in far-field voice. The method can be applied to the display device 200 in the above embodiments. In some embodiments, the method may include the following content:
[0144] In response to the start instruction of far-field voice, save the original audio file; the original audio file at least includes the audio file containing the wake-up word and the audio file collected by the sound collection device;
[0145] Obtain the audio file after echo cancellation of the original audio file, and identify whether the wake word is included in the audio file after cancellation;
[0146] When the wake word is included in the audio file after cancellation, obtain the original audio file corresponding to the wake word, and identify whether the original audio file includes a feature tag; the feature tag is a signal characterizing a frequency exceeding a preset frequency;
[0147] When the original audio file does not include the feature tag, perform a wake-up operation according to the wake word;
[0148] When the original audio file includes the feature tag, do not respond to the wake word.
[0149] As can be seen from the above technical solutions, the above embodiments provide a method for preventing false wake-up in far-field voice. The method can ensure that the display device 200 is only woken up when the wake word comes from an external user by judging whether the original audio file includes a feature tag, thereby avoiding the false wake-up phenomenon caused by the content played by the TV itself and solving the false wake-up problem in the far-field voice interaction scenario.
[0150] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods in various embodiments or some parts of the embodiments of the present invention.
[0151] The above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features.
[0152] For the sake of explanation, the above description has been made in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. According to the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are for better explaining the principles and actual applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.
Claims
1. A display device, characterized in that, Including: A display configured to display a user interface; A controller configured to: In response to an activation instruction of far-field voice, save an original audio file; the original audio file at least includes an audio file containing a wake word and an audio file collected by a sound collection device; Obtain an audio file after echo cancellation processing of the original audio file, and identify whether the audio file after cancellation contains a wake word; When the audio file after cancellation contains a wake word, obtain the original audio file corresponding to the wake word, and identify whether the original audio file contains a feature tag; The feature tag is a signal representing a frequency exceeding a preset frequency; When the original audio file does not contain the feature tag, perform a wake-up operation according to the wake word; When the original audio file contains the feature tag, do not respond to the wake word.
2. The display device according to claim 1, wherein Before the step of the controller saving the original audio file, it is further configured to: Obtain the text information of the content to be played containing the wake word; Combine the text information into a sentence; Input the sentence into a TTS model to output the audio corresponding to the sentence; Generate an original audio file according to the audio.
3. The display device according to claim 1, wherein The controller identifies whether the audio file after cancellation contains a wake word, and is specifically configured to: Input the audio file after cancellation into a wake word recognition model frame by frame; When the wake word recognition model recognizes the wake word, determine that the audio file after cancellation contains the wake word; When the wake word recognition model does not recognize the wake word, determine that the audio file after cancellation does not contain the wake word.
4. The display device according to claim 1, characterized in that, Before the step of the controller identifying whether the original audio file contains a feature tag, it is further configured to: Obtain the original audio file containing the wake word; Identify the target frame position corresponding to the wake word contained in the original audio file; Add the feature tag at the target frame position.
5. The display device according to claim 4, wherein The controller adds the feature tag at the target frame position, and is specifically configured to: Obtain the start position and end position corresponding to the target frame position; Insert the feature tag between the start position and the end position.
6. The display device according to claim 4, wherein The controller identifies whether the original audio file contains a feature tag, and is specifically configured to: Calculate the sound energy value of the original audio file containing the wake word; When the sound energy value is greater than or equal to an energy threshold, determine that the original audio file contains a feature tag; When the sound energy value is less than the energy threshold, determine that the original audio file does not contain a feature tag.
7. The display device according to claim 5, characterized in that, Before the step of the controller calculating the sound energy value of the original audio file containing the wake word, it is further configured to: Obtain the original audio file without echo cancellation processing; Extract the feature tag contained in the original audio file; Determine the original audio file containing the wake word corresponding to the feature tag according to the feature tag.
8. The display device according to claim 1, characterized in that, The controller obtains the audio file after echo cancellation processing of the original audio file, and is specifically configured to: Perform echo cancellation processing on the audio file containing the wake word; Generate the post-cancellation audio file based on the audio file after echo cancellation processing and the audio file collected by the sound collection device.
9. The display device according to claim 1, wherein After the step of the controller identifying whether the original audio file contains a feature label, it is further configured to: In the case that the original audio file does not contain the feature label, determine that the original audio file is an audio file containing a wake word triggered by the user, and perform the step of performing a wake-up operation according to the wake word; In the case that the original audio file contains the feature label, determine that the original audio file is an audio file for mis-wake-up triggered by the audio played by the display device, and perform the step of not responding to the wake word.
10. A method for preventing false wake-up in far-field voice, applied to the display device described in any one of claims 1-9, the display device including a display and a controller, characterized in that, The method includes: In response to an enabling instruction for far-field voice, save the original audio file; the original audio file at least includes an audio file containing a wake word and an audio file collected by the sound collection device; Obtain the post-cancellation audio file after echo cancellation processing on the original audio file, and identify whether the post-cancellation audio file contains a wake word; In the case that the post-cancellation audio file contains a wake word, obtain the original audio file corresponding to the wake word, and identify whether the original audio file contains a feature label; the feature label is a signal whose frequency exceeds a preset frequency; In the case that the original audio file does not contain the feature label, perform a wake-up operation according to the wake word; In the case that the original audio file contains the feature label, do not respond to the wake word.