Display device and voice assistant awakening method

By enhancing the wake-up word voice fragments and sending them to the server, the problem of high CPU occupancy in traditional technology is solved, and the voice assistant wake-up rate and the CPU occupancy rate are improved.

CN120201218APending Publication Date: 2025-06-24VIDAA (NETHERLANDS) INT HLDG LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510154875.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In traditional technology, the method of improving the far-field wake-up rate of voice assistants has a high CPU usage rate, resulting in a high CPU usage rate.

Method used

By enhancing the wake-up word voice fragments, the enhanced wake-up word voice fragments are obtained, and the enhanced wake-up word voice fragments and noise-reducing command voice fragments are sent to the server to reduce noise and improve the accuracy of wake-up word verification.

Benefits of technology

This method not only improves the wake-up rate of the voice assistant, but also reduces the CPU occupancy rate, solving the problem of high CPU occupancy rate in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201218A_ABST
    Figure CN120201218A_ABST
Patent Text Reader

Abstract

The invention relates to a display device and a voice assistant awakening method. Comprising the steps that an input voice data stream is acquired, a target voice data stream is determined according to the voice data stream, and the target voice data stream is the voice data stream or is obtained by performing noise reduction on the voice data stream; under the condition that the target voice data stream comprises a wake-up word voice segment, the wake-up word voice segment in the target voice data stream is enhanced, an enhanced wake-up word voice segment is obtained, the wake-up word voice segment is matched with a wake-up word, and the wake-up word is used for waking up a voice assistant; and sending the enhanced wake-up word voice segment to a server corresponding to the voice assistant, the enhanced wake-up word voice segment being used for the server to perform wake-up word verification, and if the wake-up word verification is passed, representing that the voice assistant is successfully awakened. The CPU occupancy rate can be reduced while the wake-up rate is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of display devices, and in particular to a display device and a voice assistant wake-up method. Background Art

[0002] In traditional technologies, methods for improving the far-field wake-up rate of voice assistants include: using front-end audio processing technologies such as sound source localization and speech enhancement to process far-field speech in noisy conditions to improve the wake-up rate, or obtaining acoustic features and completing modeling through the features to identify wake-up words, or weakening interference sources through one or more states of gain settings, mute settings, and enhancement settings associated with one or more microphones to improve the wake-up rate.

[0003] However, the conventional solution for improving the wake-up rate has a high CPU usage rate and consumes a lot of CPU resources. Therefore, it is necessary to propose a solution for improving the wake-up rate that can reduce the CPU resource usage rate. Summary of the invention

[0004] The present application provides a display device and a voice assistant wake-up method to solve the problem of high CPU occupancy during the process of improving the wake-up rate.

[0005] In a first aspect, some embodiments provide a display device, comprising: a display, a communication device, and a controller. The display is configured to display a user interface; the communication device is configured to establish a communication connection with a server; and the controller is configured to:

[0006] Acquire an input voice data stream, and determine a target voice data stream according to the voice data stream, wherein the target voice data stream is the voice data stream or is obtained by performing noise reduction on the voice data stream;

[0007] In the case where the target voice data stream includes a wake-up word voice segment, enhancing the wake-up word voice segment in the target voice data stream to obtain an enhanced wake-up word voice segment, wherein the wake-up word voice segment matches the wake-up word, and the wake-up word is used to wake up the voice assistant;

[0008] The enhanced wake-up word voice segment is sent to the server corresponding to the voice assistant, wherein the enhanced wake-up word voice segment is used by the server to perform wake-up word verification. If the wake-up word verification passes, it indicates that the voice assistant is successfully awakened.

[0009] In this embodiment, since the embedded system faces the dilemma of high CPU occupancy, an algorithm with low latency and low complexity is required to reduce the CPU occupancy rate. In this application, the wake-up word voice segment is enhanced to obtain an enhanced wake-up word voice segment, and the enhanced wake-up word voice segment and the noise reduction command voice segment are sent to the server. On the one hand, the enhancement process helps to reduce the noise in the enhanced wake-up word voice segment, improve the accuracy of the server's wake-up word verification, and reduce the situation of verification failure due to more noise, thereby improving the wake-up rate. On the other hand, using the enhancement process of the wake-up word voice segment to improve the wake-up rate can reduce the CPU occupancy rate compared with the traditional way of improving the wake-up rate. Therefore, this application can reduce the CPU occupancy rate while improving the wake-up rate.

[0010] In a second aspect, some embodiments further provide a method for waking up a voice assistant, which is applied to the display device provided in the first aspect. The display device includes: a display, a communication device, and a controller. The method includes:

[0011] Obtain the input voice data stream, and determine the target voice data stream according to the voice data stream, where the target voice data stream is the voice data stream or the voice data stream obtained by performing noise reduction on the voice data stream;

[0012] When the target voice data stream includes a wake-up word voice segment, perform enhancement processing on the wake-up word voice segment in the target voice data stream to obtain an enhanced wake-up word voice segment, where the wake-up word voice segment matches the wake-up word, and the wake-up word is used to wake up the voice assistant;

[0013] Send the enhanced wake-up word voice segment to the server corresponding to the voice assistant, where the enhanced wake-up word voice segment is used for the server to perform wake-up word verification. If the wake-up word verification is passed, it indicates that the voice assistant wakes up successfully.

[0014] In this embodiment, since the embedded system faces the dilemma of high CPU occupancy, an algorithm with low latency and low complexity is required to reduce the CPU occupancy rate. In this application, the wake-up word voice segment is enhanced to obtain an enhanced wake-up word voice segment, and the enhanced wake-up word voice segment and the noise reduction command voice segment are sent to the server. On the one hand, the enhancement process helps to reduce the noise in the enhanced wake-up word voice segment, improve the accuracy of the server's wake-up word verification, and reduce the situation of verification failure due to more noise, thereby improving the wake-up rate. On the other hand, on the other hand, using the enhancement process of the wake-up word voice segment to improve the wake-up rate can reduce the CPU occupancy rate compared with the traditional way of improving the wake-up rate. Therefore, this application can reduce the CPU occupancy rate while improving the wake-up rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0016] Figure 1 Schematic diagram of the operation scenario between the display device and the control device provided by some embodiments of the present application;

[0017] Figure 2 Schematic diagram of the hardware configuration of the display device provided by some embodiments of the present application;

[0018] Figure 3 Schematic diagram of the hardware configuration of the control device provided by some embodiments of the present application;

[0019] Figure 4 Schematic diagram of the software configuration of the display device provided by some embodiments of the present application;

[0020] Figure 5 Schematic diagram of the flow of the voice assistant wake-up method provided by some embodiments of the present application;

[0021] Figure 6 Schematic diagram of the interaction principle between the controller and the server during the wake-up process provided by some embodiments of the present application;

[0022] Figure 7 Schematic diagram of the principle of sending voice data provided by some embodiments of the present application;

[0023] Figure 8 Schematic diagram of the principle of sending voice data provided by other some embodiments of the present application;

[0024] Figure 9 Schematic diagram of the principle of noise reduction processing and wake-up word recognition for the voice data stream provided by some embodiments of the present application;

[0025] Figure 10 Schematic diagram of the principle of selecting the target voice data stream and performing enhancement processing provided by some embodiments of the present application;

[0026] Figure 11 Timing diagram of the voice assistant wake-up method provided by some embodiments of the present application. Detailed implementation manners

[0027] Embodiments will be described in detail below, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following examples do not represent all embodiments consistent with the present application. They are merely examples of systems and methods consistent with some aspects of the present application detailed in the claims.

[0028] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the embodiments described next, rather than intending to limit the embodiments of the present application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.

[0029] The terms "first", "second", "third", etc. in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.

[0030] The terms "comprising" and "having" and any variations thereof are intended to cover but not be exclusive of inclusion. For example, a product or device comprising a series of components does not have to be limited to all the components clearly listed, but may include other components not clearly listed or inherent to these products or devices.

[0031] The term "module" refers to any known or later-developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware or / and software code that can perform functions related to that element.

[0032] In the embodiments of the present application, the display device 200 generally refers to a device with the ability to display images and process data. For example, the display device 200 includes, but is not limited to, smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc.

[0033] Figure 1 It is a schematic diagram of the operation scenario between the display device and the control device provided for some embodiments of the present application. As Figure 1 shown, the user can operate the display device 200 through touch operations, the mobile terminal 300, and the control device 100. For example, the control device 100 can be a remote control, a stylus, a gamepad, etc.

[0034] The mobile terminal 300 can be used as a control device to perform human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a communication device to establish a communication connection with the display device 200 for data interaction. In some embodiments, software applications can be installed on the mobile terminal 300 and the display device 200, and the connection communication can be achieved through network communication protocols to achieve the purpose of one-to-one control operations and data communication. It is also possible to transmit the audio and video content displayed on the mobile terminal 300 to the display device 200 to achieve the synchronous display function.

[0035] As Figure 1 also shown in, the display device 200 also conducts data communication with the server 400 through various communication methods. The display device 200 is allowed to establish a communication connection through a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0036] The display device 200 can provide a broadcast receiving television function, and can also additionally provide an intelligent network television function with computer support functions, including but not limited to, network television, smart television, Internet Protocol Television (IPTV), etc.

[0037] Figure 2 For some embodiments of this application Figure 1 The hardware configuration block diagram of the display device 200 in.

[0038] In some embodiments, the display device 200 may include at least one of a tuner demodulator 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.

[0039] In some embodiments, the detector 230 is used to collect signals from the external environment or interact with the outside. For example, the detector 230 includes a light receiver, a sensor for collecting the intensity of environmental light; or, the detector 230 includes an image collector, such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures. Or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.

[0040] In some embodiments, the display 260 includes a display function component for presenting a picture and a driving component for driving the image display. The display 260 is used to receive the image signal output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, components of a menu manipulation interface, and a user manipulation UI interface, etc.

[0041] In some embodiments, the communication device 220 is a component for communicating with an external device or server 400 according to various communication protocol types. The display device 200 may be provided with a plurality of communication devices 220 according to different supported communication methods. For example, when the display device 200 supports wireless network communication, the display device 200 may be provided with a communication device 220 including a WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.

[0042] The communication device 220 can enable the display device 200 to communicate with an external device or server 400 through a wireless or wired connection. Among them, the wired connection can connect the display device 200 to an external device through components such as a data cable and an interface. The wireless connection can connect the display device 200 to an external device through a wireless signal or a wireless network. The display device 200 can directly establish a connection relationship with an external device, or can indirectly establish a connection relationship through a gateway, a router, a connection device, etc.

[0043] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and the first interface to the nth interface for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.

[0044] In some embodiments, the controller 250 and the tuner demodulator 210 may be located in different split devices, that is, the tuner demodulator 210 may also be in an external device of the main device where the controller 250 is located, such as an external set-top box, etc.

[0045] In some embodiments, the user can input a user command on the graphical user interface (Graphical User Interface, GUI) displayed on the display 260, and then the user input interface receives the user input command through the graphical user interface (GUI).

[0046] In some embodiments, the audio output device 270 may be a built-in speaker of the display device 200, or may be an external audio output device connected to the display device 200. Among them, for the external audio output device connected to the display device 200, the display device 200 may also be provided with an external audio output terminal, and the audio output device can be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.

[0047] In some embodiments, the user input interface 280 can be used to receive instructions from user input.

[0048] Figure 3 The hardware configuration block diagram of the control device provided by some embodiments of the present application Figure 1 is shown as follows. The control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply. Figure 3 As shown, the control device 100 is configured to control the display device 200, and can receive input operation instructions from the user, and convert the operation instructions into instructions recognizable and responsive by the display device 200, playing the role of an interaction intermediary between the user and the display device 200.

[0049] In some embodiments, the control device 100 may be an intelligent device. For example, the control device 100 can install various applications for controlling the display device 200 according to user needs.

[0050] In some embodiments, as shown, after the mobile terminal 300 or other intelligent electronic devices install the application for controlling the display device 200, they can perform similar functions to the control device 100.

[0051] In some embodiments, as Figure 1 shown, after the mobile terminal 300 or other intelligent electronic devices install the application for controlling the display device 200, they can perform similar functions to the control device 100.

[0052] The controller 110 includes a processor 112, a RAM 113, a ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, as well as the communication and cooperation between internal components and the data processing functions between the external and internal.

[0053] Under the control of the controller 110, the communication interface 130 realizes the communication of control signals and data signals with the display device 200. The communication interface 130 may include at least one of a WiFi chip 131, a Bluetooth module 132, an NFC module 133, and other near-field communication modules.

[0054] The user input / output interface 140, wherein the input interface includes at least one of a microphone 141, a touchpad 142, a sensor 143, a button 144, and other input interfaces.

[0055] In some embodiments, the control device 100 includes at least one of the communication interface 130 and the input / output interface 140. The communication interface 130 configured in the control device 100, such as WiFi, Bluetooth, NFC, etc. modules, can encode the user input instructions through the WiFi protocol, or the Bluetooth protocol, or the NFC protocol, and send them to the display device 200.

[0056] The memory 190 is used to store various operating programs, data, and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can store various control signal instructions input by the user.

[0057] A power supply 180 is used to provide operating power support for each component of the control device 100 under the control of the controller.

[0058] To perform user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program for managing and controlling the hardware resources and software resources in the display device 200. The operating system can (control the display device) provide a user interface, allowing users to interact with the display device 200 and supporting the running of various applications.

[0059] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.

[0060] The operating system can be divided into different modules or levels according to the functions implemented.

[0061] Such as Figure 4 shown Figure 4 is a schematic diagram of the software configuration in the display device provided by some embodiments of the present application. In some embodiments, the system of the display device 200 can be divided into three layers, from top to bottom are the application layer, the middleware layer, and the hardware layer. Figure 1 The application layer mainly includes common applications on the TV and an application framework (Application Framework). Among them, the common applications are mainly applications developed based on the browser Browser, such as HTML5 APPs, and native applications (Native APPs).

[0062] The application framework (Application Framework) is a complete program model with all the basic functions required by standard application software, such as file access, data exchange, etc., and the usage interfaces of these functions (toolbars, status bars, menus, dialog boxes).

[0063] Native applications (Native APPs) can support online or offline, message push or local resource access.

[0064] The middleware layer includes various middleware such as TV protocols, multimedia protocols, and system components. The middleware can use the basic services (functions) provided by the system software to connect various parts of the application system on the network or different applications, and can achieve the purpose of resource sharing and function sharing.

[0065]

[0066] ​The hardware layer mainly includes the Hardware Abstraction Layer (HAL) interface, hardware, and drivers. Among them, the HAL interface is the unified interface for all TV chips to dock, and the specific logic is implemented by each chip. The drivers mainly include: audio drivers, display drivers, Bluetooth drivers, camera drivers, WIFI drivers, USB drivers, HDMI drivers, sensor drivers (such as fingerprint sensors, temperature sensors, pressure sensors, etc.), and power drivers, etc.

[0067] It should be noted that the above examples are only simple divisions of the functions of the operating system, and do not limit the specific form of the operating system of the display device 200 in the embodiments of the present application. According to factors such as the functions of the display device and the type of the operating system, the number of levels included in the operating system and the specific level types can be in other forms.

[0068] Based on this, in some embodiments, the present application provides a method for waking up a voice assistant, which is applied to a display device, such as Figure 5 as shown, the method includes:

[0069] Step 502, obtain the input voice data stream, and determine the target voice data stream according to the voice data stream. Among them, the target voice data stream is the voice data stream or the voice data stream obtained by denoising the voice data stream.

[0070] Among them, the voice data stream can be collected through a microphone, or can be input through a microphone in a mobile device or a remote control. The microphone can be a pick-up microphone, for example, a far-field pick-up microphone.

[0071] Step 504, in the case that the target voice data stream includes a wake-up word voice segment, perform enhancement processing on the wake-up word voice segment in the target voice data stream to obtain an enhanced wake-up word voice segment. Among them, the wake-up word voice segment matches the wake-up word, and the wake-up word is used to wake up the voice assistant.

[0072] Among them, the wake-up word is preset, and can be but is not limited to at least one of the Chinese name, English name, or nickname of the voice assistant, etc. The voice assistant can be but is not limited to Amazon's Alexa intelligent voice assistant. The length of the wake-up word voice segment can be limited, for example, the upper limit can be 1.3 seconds or 2 seconds, etc. The enhancement processing includes but is not limited to at least one of echo processing, noise cancellation, or gain improvement, etc.

[0073] In some embodiments, the target voice data stream is stored in the first cache, and the controller can obtain the wake-up word position corresponding to the target voice data stream, and obtain the voice segment at the wake-up word position from the target voice data stream to obtain the wake-up word voice segment. The wake-up word position represents the storage position of the voice segment matching the wake-up word in the first cache in the target voice data stream.

[0074] Step 506: Send the enhanced wake-up word voice segment to the server corresponding to the voice assistant. The enhanced wake-up word voice segment is used by the server for wake-up word verification. If the wake-up word verification is passed, it indicates that the voice assistant is successfully awakened.

[0075] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The cloud server is used to provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0076] In some embodiments, when the wake-up word verification is passed, it indicates that the wake-up word is valid. Then the server can perform intent recognition on the subsequent voice data sent by the display device and feedback the valid intent obtained from the intent recognition to the display device.

[0077] In some embodiments, when the wake-up word verification fails, it indicates that the wake-up word is invalid. Then the server can feedback to the display device that the wake-up word is invalid. When the controller determines that the wake-up word is invalid, it can continue to perform wake-up word recognition on the input voice data stream and send a new enhanced wake-up word voice segment to the server, and the server performs wake-up word verification again. The invalid wake-up word also belongs to a valid intent.

[0078] In some embodiments, the controller can call the data sending interface corresponding to the voice assistant to send the enhanced wake-up word voice segment to the server.

[0079] In some embodiments, such as Figure 6As shown in the figure, the controller may include a recording module, an audio processing module, and an enhancement processing module, and may also include a preprocessing module. The recording module can obtain the voice data stream collected by the microphone, perform noise reduction through the audio processing module to obtain a target voice data stream, and perform wake-up word recognition on the target voice data stream. In the case of successful recognition, the enhancement processing module is used to perform enhancement processing on the wake-up word voice segment in the target voice data stream to obtain an enhanced wake-up word voice segment. The preprocessing module can write the enhanced wake-up word voice segment into the cache corresponding to the voice assistant. The controller can call the data sending interface corresponding to the voice assistant to read the enhanced wake-up word voice segment from the cache and send it to the cloud server corresponding to the voice assistant. The voice assistant has a data receiving interface, and through the data receiving interface, it can receive the data returned by the cloud server. For example, after the cloud server performs wake-up word verification, the controller can receive the verification result returned by the cloud server through the data receiving interface corresponding to the voice assistant. Among them, the recording module and the audio processing module can be implemented using the iFlytek front-end audio processing library. For example, the recording module can be implemented using the recording library in the iFlytek front-end audio processing library. Noise reduction, wake-up word recognition, and enhancement processing can be implemented using the algorithms provided in the audio algorithm processing library in the iFlytek front-end audio processing library. It should be noted that although the algorithms provided by the iFlytek front-end audio processing library can be used for implementation, it does not mean that the wake-up word recognition and enhancement processing operations in this application have been implemented in the iFlytek front-end audio processing library. It only means that the algorithms provided by it can be used for processing.

[0080] Among them, the data sending interface and the data receiving interface corresponding to the voice assistant can be provided by the SDK (Software Development Kit) of the voice assistant. For example, the data sending interface can be the Alexa SDK data sending interface, and the data receiving interface can be the Alexa SDK data receiving interface.

[0081] In this embodiment, since the embedded system faces the dilemma of high CPU occupancy, algorithms with low latency and low complexity are required to reduce the CPU occupancy rate. In this application, by performing enhancement processing on the wake-up word voice segment to obtain an enhanced wake-up word voice segment and sending the enhanced wake-up word voice segment and the noise reduction command voice segment to the server. On the one hand, the enhancement processing helps to reduce the noise in the enhanced wake-up word voice segment, improve the accuracy of the server's wake-up word verification, reduce the situation of verification failure due to more noise, and improve the verification success rate, thereby improving the wake-up rate. On the other hand, using enhancement processing on the wake-up word voice segment to improve the wake-up rate can reduce the CPU occupancy rate compared with the traditional way of improving the wake-up rate. Therefore, this application can reduce the CPU occupancy rate while improving the wake-up rate.

[0082] In some embodiments, the present application provides a display device, including a display and a controller; the controller is configured to: obtain an input voice data stream, determine a target voice data stream according to the voice data stream, where the target voice data stream is the voice data stream or the voice data stream obtained by noise reduction; in the case that the target voice data stream includes a wake-up word voice segment, perform enhancement processing on the wake-up word voice segment in the target voice data stream to obtain an enhanced wake-up word voice segment, where the wake-up word voice segment matches the wake-up word, and the wake-up word is used to wake up the voice assistant; send the enhanced wake-up word voice segment to the server corresponding to the voice assistant, where the enhanced wake-up word voice segment is used for the server to perform wake-up word verification, and if the wake-up word verification is passed, it indicates that the voice assistant is successfully woken up.

[0083] In this embodiment, since the embedded system faces the dilemma of high CPU occupancy, an algorithm with low latency and low complexity is required to reduce the CPU occupancy rate. In the present application, by performing enhancement processing on the wake-up word voice segment to obtain an enhanced wake-up word voice segment, and sending the enhanced wake-up word voice segment and the noise reduction command voice segment to the server. On the one hand, the enhancement processing helps to reduce the noise in the enhanced wake-up word voice segment, improve the accuracy of the server's wake-up word verification, and reduce the situation of verification failure due to more noise, thereby improving the wake-up rate. On the other hand, using the enhancement processing of the wake-up word voice segment to improve the wake-up rate can reduce the CPU occupancy rate compared with the traditional way of improving the wake-up rate. Therefore, the present application can reduce the CPU occupancy rate while improving the wake-up rate.

[0084] In some embodiments, the target voice data stream is stored in the first cache, and the controller is further configured to: in the first cache, insert the enhanced wake-up word voice segment into the target voice data stream to obtain an updated voice data stream, and obtain the first storage position of the enhanced wake-up word voice segment in the first cache. When the controller executes to send the enhanced wake-up word voice segment to the server corresponding to the voice assistant, it is configured to: determine the starting reading position according to the first starting storage position; read the target voice segment from the first cache and send the target voice segment to the server corresponding to the voice assistant.

[0085] Among them, the enhanced wake-up word voice segment is inserted after the wake-up word voice segment. The enhanced wake-up word voice segment is adjacent to the strongly enhanced wake-up word voice segment, or the enhanced wake-up word voice segment is separated from the wake-up word voice segment by an environmental voice segment of a preset length.

[0086] The first storage position includes a first starting storage position and a first ending storage position. The first starting storage position refers to the starting address of the enhanced wake-up word voice segment in the first cache. The first ending storage position refers to the ending address of the enhanced wake-up word voice segment in the first cache.

[0087] The starting reading position is the first starting storage position, or the starting reading position is the starting storage position of the environmental voice segment. The starting storage position of the target voice segment is the starting reading position, and the ending storage position of the target voice segment is the first ending storage position. Thus, the target voice segment is the enhanced wake-up word voice segment, or the target voice segment is a spliced ​​voice segment, the spliced ​​voice segment is the splicing of the environmental voice segment and the enhanced wake-up word voice segment, and the environmental voice segment is arranged before the enhanced wake-up word voice segment.

[0088] If the enhanced wake-up word voice segment is adjacent to the strong wake-up word voice segment, the start reading position is the first start storage position. If the enhanced wake-up word voice segment is separated from the wake-up word voice segment by an ambient voice segment of a preset length, the start reading position is the start storage position of the ambient voice segment.

[0089] In some embodiments, the controller may send the updated voice data in the voice data stream to the server corresponding to the voice assistant starting from the starting reading position. Thus, the target voice segment is sent first, and then the voice data after the target voice segment is sent. The voice data in the updated voice data stream is sent to the server corresponding to the voice assistant starting from the starting reading position, so that the wake-up word voice segment can be skipped and the voice data after the wake-up word voice segment can be sent.

[0090] In this embodiment, since the starting storage position of the target voice segment is the starting reading position and the ending storage position of the target voice segment is the first ending storage position, the target voice segment includes an enhanced wake-up word voice segment, so sending the target voice segment can achieve the purpose of sending the enhanced wake-up word voice segment. Since the enhanced wake-up word voice segment has been enhanced, it helps to improve the verification accuracy and thus improve the wake-up rate.

[0091] In some embodiments, when the controller executes inserting an enhanced wake-up word voice segment into the target voice data stream to obtain an updated voice data stream, it is configured to: splice the enhanced wake-up word voice segment with an ambient voice segment of a preset length to obtain a spliced ​​voice segment, wherein, in the spliced ​​voice segment, the ambient voice segment is arranged before the enhanced wake-up word voice segment; insert the spliced ​​voice segment into the target voice data stream to obtain an updated voice data stream, wherein, in the updated voice data stream, the spliced ​​voice segment is adjacent to the wake-up word voice segment, and the spliced ​​voice segment is located after the wake-up word voice segment.

[0092] Among them, the preset length can be set as needed, for example, it is a voice segment of 500 milliseconds. The ambient voice segment can be collected when there is no voice in the environment, or it can be generated. For example, the data in the ambient voice segment can all be 0. Or, if the length of the segment before the wake-up word voice segment in the target voice data stream is greater than or equal to the length threshold, the controller can intercept a preset length from the segment before the wake-up word voice segment in the target voice data stream to obtain the ambient voice segment, where the length threshold is greater than or equal to the preset length.

[0093] In this embodiment, a spliced voice segment is inserted after the wake-up word voice segment to obtain an updated voice data stream, and the spliced voice segment is adjacent to the wake-up word voice segment and is located after the wake-up word voice segment. Therefore, when sending, the enhanced wake-up word voice segment is sent instead of the wake-up word voice segment, which can reduce the wake-up failure caused by the inaccuracy of the wake-up word voice segment.

[0094] In some embodiments, when the controller reads the target voice segment from the first cache and sends the target voice segment to the server corresponding to the voice assistant, it is configured to: read the target voice segment from the first cache; write the target voice segment to the second cache corresponding to the voice assistant, and record the second storage location of the enhanced wake-up word voice segment in the target voice segment in the second cache; call the data sending interface of the voice assistant based on the second storage location to send the target voice segment to the server corresponding to the voice assistant through the data sending interface.

[0095] Among them, the first cache and the second cache are different caches. The second cache is a cache that can be accessed by the voice assistant SDK, so that voice data can be sent by calling the data sending interface of the voice assistant SDK.

[0096] In some embodiments, when calling the data sending interface of the voice assistant, the second storage location is set as a parameter into the data sending interface, so that the voice assistant SDK can read the target voice segment according to the second storage location.

[0097] In some embodiments, the voice assistant SDK can send a wake-up word verification request to the server through the data sending interface. The wake-up word verification request carries the target voice segment. After receiving the target voice segment, the server can convert the target voice segment into text. If there is a wake-up word in the text, it is determined that the wake-up word verification is passed; otherwise, it is determined that the wake-up word verification is not passed. After sending the target voice segment, the controller can continue to send the command voice data located after the target voice segment. When sending, it can be sent frame by frame. If the wake-up word verification is not passed, this voice data transmission is ended, that is, the sending of the command voice data is stopped. Among them, passing the wake-up word verification belongs to a valid intention.

[0098] As shown Figure 7 in the figure, in the updated voice data stream, segment S1 is the wake-up word voice segment, S2 is the ambient voice segment of a preset length, and S3 is the enhanced wake-up word voice segment. When the controller obtains the first storage location, it can trigger a wake-up event. In the figure, the start position of the wake-up word is the first start storage position in the first storage location, and the end position of the wake-up word is the first start storage position in the second storage location. After triggering the wake-up event, preprocess the wake-up word based on the first storage location to obtain the target voice segment, that is, "the ambient voice segment of preset length + the enhanced wake-up word voice segment" in the figure, and cache the target voice segment into the second cache and perform repositioning, that is, determine the second storage location. The second storage location includes a second start storage position and a second end storage position. The second start storage position refers to the start address of the enhanced wake-up word voice segment in the second cache, and the second end storage position refers to the end address of the enhanced wake-up word voice segment in the second cache. The new start position of the wake-up word refers to the second start storage position, and the new end position of the wake-up word refers to the second end storage position. Then, the data in the second cache can be sent starting from the second start storage position, so that the target voice segment can be sent first (it can be sent frame by frame), and then the command voice data can be sent (it can be sent frame by frame). Of course, the command voice data can also be sent only when the wake-up word is verified.

[0099] In this embodiment, by writing the target voice segment into the second cache corresponding to the voice assistant and recording the second storage position, the purpose of sending the target voice segment through the data sending interface of the voice assistant can be achieved.

[0100] In some embodiments, the controller is further configured to: when the wake-up word is verified, send the command voice data in the updated voice data stream that is after the target voice segment to the server; when the server recognizes a valid intent based on the command voice data, pause sending the subsequent command voice data to the server.

[0101] Among them, the command voice data after the target voice segment refers to the voice data in the second cache that is after the target voice segment. The valid intent recognized based on the command voice data can be an intent to control a display device, including but not limited to: playing a specified video, searching for a specified video, fast-forwarding a playing video, or pausing a playing video, etc. For example, the updated voice data stream includes the voice "Alexa, please search for the AABB movie and play it". Then, the enhanced sensitive word voice segment in the updated voice data stream is the voice "Alexa", and the voice "please search for the AABB movie and play it" is also included after the target voice segment in the updated voice data stream.

[0102] As Figure 8As shown, the controller carefully reports the wake word to the cloud for processing, that is, sends the target voice segment to the server for wake word verification. The server performs wake word verification and returns the verification result. If the verification result is passed, the sensitive word is valid; if the verification fails, the sensitive word is invalid. In the case of passing the verification, the command voice data can be sent frame by frame, for example, successively sending command voice data frames 2 to command voice data frame N, etc. The server performs intent recognition on the received command voice data and returns the intent recognition result corresponding to the command voice data. If the intent recognition result is valid, that is, the voice command intent is valid, then the subsequent command voice data sending is paused.

[0103] In this embodiment, only when the wake word verification is passed, the command voice data after the target voice segment is sent, so as to avoid sending data to the server when the user does not interact with the voice assistant, save computer resources, and further reduce the CPU occupancy rate.

[0104] In some embodiments, there are at least two voice data streams. Taking the existence of two voice data streams as an example, these two voice data streams are the first voice data stream and the second voice data stream respectively. When the controller executes to obtain the input voice data stream and determine the target voice data stream according to the voice data stream, it is configured as follows: obtain the input first voice data stream and determine the first candidate voice data stream according to the first voice data stream; obtain the input second voice data stream and determine the second candidate voice data stream according to the second voice data stream; perform wake word recognition on the first candidate voice data stream, and obtain the wake word confidence corresponding to the first candidate voice data stream in the case of successful recognition. Perform wake word recognition on the second candidate voice data stream, and obtain the wake word confidence corresponding to the second candidate voice data stream in the case of successful recognition; determine the one with the higher wake word confidence among the first candidate voice data stream and the second candidate voice data stream as the target voice data stream.

[0105] Among them, the time for collecting the first voice data stream is the same as the time for collecting the second voice data stream, or the interval between the time for collecting the first voice data stream and the time for collecting the second voice data stream is less than the interval threshold, and the interval threshold is set as needed. The microphones for separately collecting the first voice data stream and the second voice data stream can be pre-set in the recording module. The first voice data stream, the second voice data stream, the first candidate voice data stream and the second candidate voice data stream can be stored in the first buffer.

[0106] The first candidate voice data stream is the first voice data stream or the result of noise reduction on the first voice data stream. The second candidate voice data stream is the second voice data stream or the result of noise reduction on the second voice data stream. The first voice data stream and the second voice data stream come from different microphones.

[0107] The wake word confidence reflects the probability of the existence of a wake word. The wake word confidence corresponding to the first candidate voice data stream is used to reflect the probability of recognizing a wake word from the first candidate voice data stream. The wake word confidence corresponding to the second candidate voice data stream is used to reflect the probability of recognizing a wake word from the second candidate voice data stream.

[0108] In some embodiments, the controller includes a recording module, and a microphone is pre-set in the recording module. Among them, the recording module can store the voice data collected (picked up) by the pick-up microphone and can also store the voice data collected by the speaker. The recording module can be implemented using the iFlytek recording library, and the audio processing module can be implemented through the iFlytek audio algorithm processing library.

[0109] In some embodiments, the controller contains an audio processing module, and the audio processing module includes a wake word recognition unit. When the first candidate voice data stream is obtained, the wake word recognition unit can be used to perform wake word recognition on the first candidate voice data stream. When a wake word is recognized, the wake word confidence corresponding to the first candidate voice data stream is output. Similarly, when the second candidate voice data stream is obtained, the wake word recognition unit can be used to perform wake word recognition on the second candidate voice data stream. When a wake word is recognized, the wake word confidence corresponding to the second candidate voice data stream is output.

[0110] In some embodiments, the wake word recognition unit in the controller can search for a voice segment in the first candidate voice data stream whose wake word confidence is greater than the confidence threshold. If found, it is determined that a wake word is recognized, and the position of the found voice segment (the storage position in the first cache) is determined as the wake word position corresponding to the first candidate voice data stream, and the wake word confidence corresponding to the found voice segment is used as the wake word confidence corresponding to the first candidate voice data stream; if not found, it is determined that no wake word is recognized. Among them, the wake word confidence corresponding to the voice segment represents the degree of matching between the voice segment and the wake word. The wake word recognition unit can be the WWE (Wake Word Engine) in Amazon's Alexa intelligent voice assistant. Similarly, the wake word confidence and wake word position corresponding to the second candidate voice data stream can be obtained in the same way.

[0111] In some embodiments, the controller may intercept a reference speech segment from the first candidate speech data stream in the order from front to back. The earlier the acquisition time is for the speech data closer to the front in the first candidate speech data stream, and the length of the reference speech segment can be set according to actual needs. The controller may perform wake word recognition on the reference speech segment to obtain the wake word confidence level and the wake word position corresponding to the reference speech segment. The wake word confidence level corresponding to the reference speech segment represents the degree of matching between the speech segment at the wake word position in the reference speech segment and the wake word. When the wake word confidence level reaches the confidence threshold, it is determined that the recognition is successful. When the reference speech segment does not satisfy the condition that the wake word confidence level is greater than the confidence threshold, the next reference speech segment is intercepted until a reference speech segment that satisfies the condition that the wake word confidence level reaches the confidence threshold is obtained. The reference speech segment with the wake word confidence level reaching the confidence threshold can be referred to as the target reference speech segment. Similarly, the second candidate speech data stream can also be determined whether the recognition is successful in the same way.

[0112] In this embodiment, since the higher the wake word confidence level, the greater the degree of matching between the speech segment and the wake word, and thus the greater the probability of waking up. Therefore, determining the one with the higher wake word confidence level among the first candidate speech data stream and the second candidate speech data stream as the target speech data stream helps to improve the wake-up rate.

[0113] In some embodiments, when the controller executes to determine the first candidate speech data stream according to the first speech data stream, it is configured to: obtain a third speech data stream from the speaker, where the time interval between the playback time of the third speech data stream and the acquisition time of the first speech data stream is less than the time interval threshold; use the third speech data stream to perform noise reduction processing on the first speech data stream to obtain the first candidate speech data stream.

[0114] Similarly, the controller may use the third speech data stream to perform noise reduction processing on the second speech data stream to obtain the second candidate speech data stream.

[0115] Among them, the third speech data stream is the speech data stream played by the speaker. The time interval threshold is a value close to 0 seconds or a value less than or equal to 30 milliseconds, for example, it can be 5 milliseconds or 10 milliseconds, etc.

[0116] In some embodiments, the recording module may be preset with one or more speakers. For example, the recording module may be preset with 2, 3, or 5 speakers. The number of microphones set and the number of speakers set may be the same or different. For example, 2 microphones and 3 speakers are set. Different speakers may correspond to their respective third voice data streams. The controller may use the third voice data streams corresponding to the respective speakers to perform noise reduction processing on the first voice data stream to obtain a first candidate voice data stream. The playback times of the respective third voice data streams may be the same.

[0117] In some embodiments, the audio processing module further includes a device sound filtering unit and an echo processing unit. Among them, the device sound filtering unit is used to filter out the noise generated by the display device itself during operation. The controller may use the device sound filtering unit to perform noise reduction on the first voice data stream, the second voice data stream, and the third voice data stream respectively to obtain a first noise-reduced voice data stream, a second noise-reduced voice data stream, and a third noise-reduced voice data stream. Among them, when the device sound filtering unit performs noise reduction, at least one of beamforming or sound source localization may be used. Of course, other noise reduction means may also be used.

[0118] Then, the controller may use the third noise-reduced voice data stream to perform noise reduction on the first noise-reduced voice data stream to filter out the voice played by the speaker in the display device from the first noise-reduced voice data stream to obtain a first candidate voice data stream. For example, the controller may input the first noise-reduced voice data stream and the third noise-reduced voice data stream into the echo processing unit, and the echo processing unit uses the third noise-reduced voice data stream to perform noise reduction on the first noise-reduced voice data stream to obtain a first candidate voice data stream.

[0119] In this embodiment, since the audio played by the speaker will be collected by the microphone, and thus a part of the noise in the voice data stream collected by the microphone comes from the voice played by the speaker. By using the third voice data stream to perform noise reduction on the first voice data stream, the voice played by the speaker in the display device can be filtered out from the first voice data stream, noise reduction can be achieved, and further the noise in the enhanced wake-up word voice segment can be reduced, which helps to improve the accuracy of wake-up word verification and the wake-up rate.

[0120] Illustrate by way of example, such as Figure 9As shown, a schematic diagram from noise reduction to wake word recognition is provided. Among them, microphone B1 collects the first voice data stream, microphone B2 collects the second voice data stream, and the third voice data streams 1 to 3 are respectively collected from speakers 1 to 3. The first voice data stream and each of the third voice data streams are input into the audio processing module. The device sound filtering unit is used to filter the device sound and the echo processing unit is used to process the noise to obtain the first candidate voice data stream. Similarly, based on the second voice data stream collected by microphone B2, a second candidate voice data stream can be obtained. Then, the wake word recognition unit is used to recognize the wake word in the first candidate voice data stream, and when the recognition is successful, the wake word position and wake word confidence corresponding to the first candidate voice data stream are output. Similarly, the wake word recognition unit is used to recognize the wake word in the second candidate voice data stream, and when the recognition is successful, the wake word position corresponding to the second candidate voice data stream (i.e., the wake word position corresponding to the target reference voice segment) and the wake word confidence are output. In the figure, the gray part in the target reference voice segment is the voice segment at the wake word position.

[0121] In some embodiments, each microphone can collect a 16-bit audio data stream. Each microphone can correspond to one channel. Each speaker can correspond to one channel and play a 16-bit audio data stream. Thus, the first voice data stream and the second voice data stream can be audio data streams in channels corresponding to different microphones. Each of the third voice data streams can be an audio data stream in a different channel of the speaker.

[0122] In some embodiments, the audio processing module in the controller further includes a selection module and an enhancement processing module. After obtaining the first candidate voice data stream and the second candidate voice data stream, the selection module selects the one with the highest wake word confidence as the target voice data stream. Then, the wake word voice segment is read from the target voice data stream, and the enhancement processing module is used to enhance the wake word voice segment to obtain an enhanced wake word voice segment. Then, the enhanced wake word voice segment is inserted, that is, embedded, into the target voice data stream.

[0123] Such as Figure 10As shown, a working principle diagram of a selection module and an enhancement processing module is provided. Among them, the selection module can select the target reference speech segment in the target speech data stream from the target reference speech segments in the first candidate speech data stream and the target reference speech segments in the second candidate speech data stream according to the wake word confidence. The gray part in the target reference speech segment in the target speech data stream is the wake word speech segment. The enhancement processing module can obtain the wake word speech segment from the target reference speech segment in the target speech data stream and perform enhancement processing on the wake word speech segment to obtain an enhanced wake word speech segment. The audio processing module can also insert, that is, embed, the enhanced wake word speech segment into the target speech data stream.

[0124] In some embodiments, the controller is further configured to: perform wake word recognition on the target speech data stream, and when the recognition is successful, obtain the wake word position corresponding to the target speech data stream; obtain the speech segment at the wake word position from the target speech data stream to obtain the wake word speech segment.

[0125] Specifically, the controller can input the target speech data stream into the wake word recognition unit, and the wake word recognition unit performs wake word recognition on the target speech data stream. When the recognition is successful, the wake word position is output.

[0126] In this embodiment, by directly performing wake word recognition on the target speech data stream to determine the wake word speech segment, the effect of obtaining the wake word speech segment can be achieved.

[0127] In some embodiments, when the controller performs wake word recognition on the target speech data stream and, when the recognition is successful, obtains the wake word position corresponding to the target speech data stream, it is further configured to: intercept the reference speech segment from the target speech data stream in the order from front to back; perform wake word recognition on the reference speech segment to obtain the wake word confidence and wake word position corresponding to the reference speech segment. The wake word confidence characterizes the degree of matching between the speech segment at the wake word position in the reference speech segment and the wake word; when the wake word confidence reaches the confidence threshold, the wake word position corresponding to the reference speech segment is used as the wake word position corresponding to the target speech data stream.

[0128] Among them, the confidence threshold can be set as needed, can be a percentage, can be a value close to 100%, or can be a value in the range of 60% - 100%. The target speech data stream can be stored in the first cache. The wake word position corresponding to the reference speech segment can include the start address and end address in the first cache of the speech segment in the reference speech segment that matches the wake word.

[0129] In some embodiments, when the reference speech segment does not meet the condition that the wake-word confidence is greater than the confidence threshold, the next reference speech segment is intercepted until a reference speech segment that meets the condition that the wake-word confidence reaches the confidence threshold is obtained.

[0130] In this embodiment, when the wake-word confidence reaches the confidence threshold, the wake-word position corresponding to the reference speech segment is used as the wake-word position corresponding to the target speech data stream. Thus, through a reasonable confidence threshold, it can be ensured that the obtained wake-word speech segment truly matches the wake word.

[0131] In some embodiments, as Figure 11 shown, a timing diagram corresponding to a voice assistant wake-up method is provided. Specifically,

[0132] 1. The controller obtains the input first speech data stream, determines the first candidate speech data stream according to the first speech data stream, obtains the input second speech data stream, and determines the second candidate speech data stream according to the second speech data stream.

[0133] 2. The controller performs wake-word recognition on the first candidate speech data stream. When the recognition is successful, the wake-word confidence corresponding to the first candidate speech data stream is obtained. The controller performs wake-word recognition on the second candidate speech data stream. When the recognition is successful, the wake-word confidence corresponding to the second candidate speech data stream is obtained.

[0134] 3. The controller determines the one with the higher wake-word confidence among the first candidate speech data stream and the second candidate speech data stream as the target speech data stream.

[0135] 4. The controller inserts an enhanced wake-word speech segment into the target speech data stream in the first cache to obtain an updated speech data stream, obtains the first storage position of the enhanced wake-word speech segment in the first cache, and determines the starting reading position according to the first starting storage position in the first storage position.

[0136] 5. The controller reads the target speech segment from the first cache according to the starting reading position and the first ending storage position in the first storage position, writes the target speech segment into the second cache corresponding to the voice assistant, and records the second storage position of the enhanced wake-word speech segment in the target speech segment in the second cache.

[0137] 6. The controller calls the data sending interface of the voice assistant based on the second storage position to send a wake-word verification request to the server corresponding to the voice assistant through the data sending interface. The wake-word verification request carries the target speech segment.

[0138] 7. The communication device receives the wake-word verification request and notifies the processor to process the wake-word verification request.

[0139] 8. The processor performs wake-up word verification and returns the verification result to the controller via the communication device.

[0140] 9. If the controller determines that the verification fails based on the verification result, it regenerates a wake-up word verification request and sends it. If the controller determines that the verification passes based on the verification result, it sends the clear command voice data.

[0141] 10. The communication device receives the command voice data and notifies the processor to process it. The processor performs intent recognition on the command voice data to obtain an intent recognition result and returns the intent recognition result to the controller via the communication device.

[0142] Among them, the controller can perform intent recognition on the received command voice data segment within a specified duration. If no valid intent is recognized after exceeding the specified duration, it returns an intent recognition result indicating that no valid intent is recognized. If a valid intent is recognized within the specified duration, it returns an intent recognition result indicating that a valid intent is recognized, and the intent recognition result includes the recognized target intent such as playing a specified video, etc.

[0143] 11. If there is a valid intent in the intent recognition result, the controller controls the display to display according to the valid intent in the intent recognition result.

[0144] Among them, if there is no valid intent in the intent recognition result, a wake-up word verification request is regenerated.

[0145] The voice assistant wake-up method provided by this application can be applied to any voice assistant, including but not limited to the far-field Alexa voice assistant. It can not only improve the wake-up rate but also reduce the CPU occupancy rate during the wake-up process. Through experiments, it is found that compared with the traditional wake-up method, the wake-up rate can be increased by 3%, thus improving the user experience. In addition, the voice assistant wake-up method provided by this application meets the requirements of the embedded system for low latency and low complexity and can be better applied in the embedded system.

[0146] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in rotation with at least a part of other steps or steps or stages in other steps.

[0147] Based on the same inventive concept, an embodiment of the present application also provides a voice assistant wake-up device for implementing the above-mentioned voice assistant wake-up method. The implementation solution provided by this device to solve problems is similar to the implementation solution described in the above method. For specific limitations, reference can be made to the limitations on the voice assistant wake-up method in the above text, and details will not be repeated here.

[0148] In some embodiments, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned voice assistant wake-up method are implemented.

[0149] In some embodiments, the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the above-mentioned voice assistant wake-up method are implemented.

[0150] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, Resistive Random Access Memory (ReRAM), Magnetoresistive Random Access Memory (MRAM), Ferroelectric Random Access Memory (FRAM), Phase Change Memory (PCM), graphene memory, etc. Volatile memory can include Random Access Memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, Artificial Intelligence (AI) processors, etc., without limitation.

[0151] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0152] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A display device, characterized in that: include: Display and controller; The controller is configured to: Acquire an input voice data stream, and determine a target voice data stream according to the voice data stream, wherein the target voice data stream is the voice data stream or is obtained by performing noise reduction on the voice data stream; In the case where the target voice data stream includes a wake-up word voice segment, enhancing the wake-up word voice segment in the target voice data stream to obtain an enhanced wake-up word voice segment, wherein the wake-up word voice segment matches the wake-up word, and the wake-up word is used to wake up the voice assistant; The enhanced wake-up word voice segment is sent to the server corresponding to the voice assistant, wherein the enhanced wake-up word voice segment is used by the server to perform wake-up word verification. If the wake-up word verification passes, it indicates that the voice assistant is successfully awakened.

2. The display device according to claim 1, characterized in that The target voice data stream is stored in a first buffer, and the controller is further configured to: In the first cache, inserting the enhanced wake-up word voice segment into the target voice data stream to obtain an updated voice data stream, and obtaining a first storage position of the enhanced wake-up word voice segment in the first cache; Wherein, the first storage location includes a first starting storage location and a first ending storage location, the enhanced wake-up word voice segment is inserted after the wake-up word voice segment, the enhanced wake-up word voice segment is adjacent to the wake-up word voice segment, or the enhanced wake-up word voice segment is separated from the wake-up word voice segment by an environmental voice segment of a preset length; When the controller executes the sending of the enhanced wake-up word voice segment to the server corresponding to the voice assistant, the controller is configured to: Determining a start reading position according to the first start storage position, wherein the start reading position is the first start storage position, or the start reading position is the start storage position of the environmental voice segment; Read the target voice segment from the first cache, and send the target voice segment to the server corresponding to the voice assistant, wherein the starting storage position of the target voice segment is the starting reading position, and the ending storage position of the target voice segment is the first ending storage position.

3. The display device according to claim 2, characterized in that When the controller inserts the enhanced wake-up word voice segment into the target voice data stream to obtain an updated voice data stream, the controller is configured as follows: Splicing the enhanced wake-up word voice segment with an environmental voice segment of a preset length to obtain a spliced ​​voice segment, wherein in the spliced ​​voice segment, the environmental voice segment is arranged before the enhanced wake-up word voice segment; The spliced ​​voice segment is inserted into the target voice data stream to obtain an updated voice data stream, wherein in the updated voice data stream, the spliced ​​voice segment is adjacent to the wake-up word voice segment, and the spliced ​​voice segment is located after the wake-up word voice segment.

4. The display device according to claim 2, characterized in that When the controller executes the step of reading the target voice segment from the first cache and sending the target voice segment to the server corresponding to the voice assistant, the controller is configured as follows: Reading a target speech segment from the first buffer; Writing the target voice segment into a second cache corresponding to the voice assistant, and recording the enhanced wake-up word voice segment in the target voice segment at a second storage location in the second cache; The data sending interface of the voice assistant is called based on the second storage location to send the target voice segment to the server corresponding to the voice assistant through the data sending interface.

5. The display device according to claim 2, characterized in that The controller is also configured to: In the case of passing the wake-up word verification, sending the command voice data located after the target voice segment in the update voice data stream to the server; When the server recognizes a valid intention based on the command voice data, the sending of subsequent command voice data to the server is suspended.

6. The display device according to any one of claims 1 to 5, characterized in that: When the controller executes the step of acquiring the input voice data stream and determining the target voice data stream according to the voice data stream, the controller is configured as follows: Acquire an input first voice data stream, and determine a first candidate voice data stream according to the first voice data stream, where the first candidate voice data stream is the first voice data stream or is obtained by performing noise reduction on the first voice data stream; Acquire an input second voice data stream, and determine a second candidate voice data stream according to the second voice data stream, where the second candidate voice data stream is the second voice data stream or is obtained by performing noise reduction on the second voice data stream, and the first voice data stream and the second voice data stream come from different microphones; Performing wake-up word recognition on the first candidate voice data stream, and if the recognition is successful, obtaining the wake-up word confidence corresponding to the first candidate voice data stream, wherein the wake-up word confidence reflects the probability of the existence of the wake-up word; Performing wake-up word recognition on the second candidate voice data stream, and obtaining a wake-up word confidence level corresponding to the second candidate voice data stream if the recognition is successful; The one with a higher confidence level of the wake-up word between the first candidate voice data stream and the second candidate voice data stream is determined as the target voice data stream.

7. The display device according to claim 6, characterized in that When the controller performs the step of determining the first candidate voice data stream according to the first voice data stream, the controller is configured to: Acquire a third voice data stream from a speaker, wherein a time interval between a playback time of the third voice data stream and a collection time of the first voice data stream is less than a time interval threshold; The first voice data stream is subjected to noise reduction processing by using the third voice data stream to obtain a first candidate voice data stream.

8. The display device according to any one of claims 1 to 5, characterized in that: The controller is also configured to: Performing wake-up word recognition on the target voice data stream, and obtaining the wake-up word position corresponding to the target voice data stream if the recognition is successful; A voice segment at the position of the wake-up word is obtained from the target voice data stream to obtain the wake-up word voice segment.

9. The display device according to claim 8, characterized in that The controller performs the wake-up word recognition on the target voice data stream, and when the recognition is successful, obtains the wake-up word position corresponding to the target voice data stream, and is further configured to: Extracting reference voice segments from the target voice data stream in a sequence from front to back; Performing wake-up word recognition on the reference voice segment to obtain a wake-up word confidence and a wake-up word position corresponding to the reference voice segment, wherein the wake-up word confidence represents the degree of match between the voice segment at the wake-up word position in the reference voice segment and the wake-up word; When the wake-up word confidence reaches a confidence threshold, the wake-up word position corresponding to the reference voice segment is used as the wake-up word position corresponding to the target voice data stream.

10. A voice assistant wake-up method, characterized in that: Applied to a display device, the method comprises: Acquire an input voice data stream, and determine a target voice data stream according to the voice data stream, wherein the target voice data stream is the voice data stream or is obtained by performing noise reduction on the voice data stream; In the case where the target voice data stream includes a wake-up word voice segment, enhancing the wake-up word voice segment in the target voice data stream to obtain an enhanced wake-up word voice segment, wherein the wake-up word voice segment matches the wake-up word, and the wake-up word is used to wake up the voice assistant; The enhanced wake-up word voice segment is sent to the server corresponding to the voice assistant, wherein the enhanced wake-up word voice segment is used by the server to perform wake-up word verification. If the wake-up word verification passes, it indicates that the voice assistant is successfully awakened.

Citation Information

Cited By

  • Display apparatus and method for waking voice assistant

    WO2026171655A1