Display device, server and false wake-up diagnosis method

By using ring buffers and linear buffers in the display device to store audio data, combined with the results of the speech recognition model, the cause of false wake-up of the display device is accurately diagnosed, and the user experience problem caused by false wake-up of the display device is solved, and the accuracy of voice recognition and the reliability of wake-up are improved.

CN120602700APending Publication Date: 2025-09-05HISENSE VISUAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510724565.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The display device may experience false wake-up in a hibernation state, resulting in a decline in user experience. The existing diagnostic methods are relatively low in accuracy and cannot effectively identify the cause of false wake-up.

Method used

By setting a ring buffer and a linear buffer in the display device, audio data is collected and stored, and combined with the recognition results of the speech recognition model, the cause of false wake-up is determined, including the recognition accuracy of the speech recognition model and the integrity analysis of the wake-up word.

Benefits of technology

It achieves accurate diagnosis of false wake-up of display devices, identifies the specific cause of false wake-up, improves user experience, and ensures the accuracy of the voice recognition model and the reliability of wake-up.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602700A_ABST
    Figure CN120602700A_ABST
Patent Text Reader

Abstract

The invention provides a display device, a server and a false wake-up diagnosis method, and the display device receives a starting instruction of a far-field voice wake-up function, controls a sound collector to collect audio, and calls a voice recognition model to carry out voice recognition on the audio; storing a first audio collected by a sound collector before a target moment to the annular buffer area, wherein the target moment is a moment when the audio is recognized to comprise an initial character field of a wake-up word; storing a second audio collected by the sound collector after the target moment to a linear buffer area, wherein the maximum storage capacity of the linear buffer area is greater than or equal to the number of bytes of the wake-up word; determining a first target audio based on the first audio, the second audio and the identification result of the second audio; and determining a cause of false wake-up based on the first target audio. The first target audio can reflect the identification condition of the display equipment on the audio, so that the reason of false wake-up of the display equipment can be determined, and the requirement of diagnosing false wake-up of the display equipment is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of display devices, and in particular to a display device, a server, and a false wakeup diagnosis method. Background Art

[0002] A display device is a terminal device that can output a specific display image, typically including smart TVs, mobile terminals, smart advertising screens, and projectors. To provide a convenient operating experience, some display devices support far-field voice wake-up.

[0003] Through the far-field voice wake-up function, the display device can collect audio in the sleep state and identify whether the audio includes the preset wake-up word. If it does, the display device will switch from the sleep state to the wake-up state, thereby waking up the display device.

[0004] However, in some cases, display devices may experience false wakeups, meaning that the user doesn't issue an audio signal containing the wake-up word, but the display device still wakes up from sleep. False wakeups can degrade the user experience, so a method for diagnosing false wakeups is urgently needed to identify the cause. Summary of the Invention

[0005] In order to solve the above problems, embodiments of the present application provide a display device, a server, and a false wakeup diagnosis method to determine the cause of the false wakeup of the display device and implement diagnosis of the false wakeup of the display device.

[0006] A first aspect of an embodiment of the present application provides a display device, including a display;

[0007] Sound collector;

[0008] a controller coupled to the display and configured to:

[0009] In response to receiving a start instruction of the far-field voice wake-up function, controlling the sound collector to collect audio, and calling the voice recognition model to perform voice recognition on the audio;

[0010] Storing the first audio collected by the sound collector before a target time in a ring buffer, where the target time is the time when the speech recognition model recognizes that the audio includes an initial character segment of the wake-up word;

[0011] Storing the second audio collected by the sound collector after the target time in a linear buffer, where the maximum storage capacity of the linear buffer is greater than or equal to the number of bytes included in the wake-up word;

[0012] Determining a first target audio based on the first audio, the second audio, and a recognition result of the second audio by the speech recognition model;

[0013] A reason why the display device is falsely awakened is determined based on the first target audio.

[0014] Through the above solution, the display device can obtain the first target audio, which can reflect the display device's recognition of the audio. Based on this, the reason for the display device's false wake-up can be determined, meeting the need for diagnosing the display device's false wake-up.

[0015] In another aspect of the embodiment of the present application, the controller determines a first target audio based on the first audio, the second audio, and a recognition result of the second audio by the speech recognition model, including:

[0016] When the speech recognition model recognizes that the second audio includes the remaining character segments of the wake-up word, determining the number of first bytes of the second audio stored in the linear buffer;

[0017] Calculating a difference between a target number of bytes and the first number of bytes, where the target number of bytes is not less than the number of bytes included in the wake-up word;

[0018] intercepting a first audio subset stored most recently from the first audio stored in the ring buffer, where the number of bytes included in the first audio subset is the difference;

[0019] It is determined that the first target audio includes the first audio subset and the second audio.

[0020] Through the above scheme, when the speech recognition model recognizes that the second audio includes the remaining character segments in the wake-up word, the first target audio including the first audio subset and the second audio can be determined, and the first audio subset is recognized by the speech recognition model as including the initial character segment of the wake-up word, and the second audio is recognized by the speech recognition model as including the remaining character segments of the wake-up word. Therefore, it can be considered that the first target audio is recognized by the speech recognition model as including the complete wake-up word, and the diagnosis of false wake-up can be achieved based on the first target audio.

[0021] In another aspect of the embodiment of the present application, before the controller calculates the difference between the target number of bytes and the first number of bytes, the controller is further configured to:

[0022] Determining the number of sampling points for the wake-up word based on the sampling rate of the sound collector and the duration of speaking the wake-up word;

[0023] Determining the number of bytes contained in the wake-up word based on the number of sampling points of the wake-up word and the number of bytes of each sampling point;

[0024] The target number of bytes is obtained by multiplying the number of bytes contained in the wake-up word by a preset coefficient, where the preset coefficient is a positive number not less than 1.

[0025] Through the above scheme, the controller can determine the target number of bytes, and the determined target number of bytes is not less than the number of bytes included in the wake-up word, so that when the audio including the wake-up word is collected, the first target audio determined according to the target number of bytes can be guaranteed to include the complete wake-up word.

[0026] In another aspect of the embodiment of the present application, the controller determines the first target audio based on the first audio, the second audio, and the recognition result of the second audio by the speech recognition model, and is specifically configured to:

[0027] determining whether the second audio occupies all storage space of the linear buffer;

[0028] When the second audio occupies all storage space of the linear buffer and the speech recognition model does not recognize that the second audio includes the remaining character segments of the wake-up word, part or all of the first audio is determined to be the first target audio.

[0029] Since the maximum storage capacity of the linear buffer is greater than or equal to the number of bytes included in the wake-up word, and the circular buffer has already stored the initial character segment of the wake-up word, if the second audio has occupied all the storage space of the linear buffer, but the speech recognition model still has not recognized that the second audio includes the remaining character segments of the wake-up word, it means that the speech recognition model has not recognized the complete wake-up word. Based on this, it can be determined that part or all of the first audio is the first target audio, and the reason for the false wake-up of the display device is determined based on the first target audio.

[0030] In another aspect of the embodiment of the present application, the controller determines a reason why the display device is falsely awakened based on the first target audio, and is specifically configured to:

[0031] Determining a similarity between the first target audio and the audio of the wake-up word;

[0032] When the similarity is less than the first threshold and the display device is falsely awakened, it is determined that the reason for the false awakening of the display device is that the speech recognition model of the display device needs to be optimized.

[0033] If the similarity between the first target audio and the audio of the spoken wake-up word is less than the first threshold, and the display device is falsely awakened in this case, it indicates that the speech recognition model of the display device used for far-field voice wake-up still recognizes the first target audio including the complete wake-up word when the similarity between the first target audio and the audio of the spoken wake-up word is low. Therefore, it can be determined that the recognition accuracy of the speech recognition model is low, and the reason for the false awakening of the display device is that the speech recognition model of the display device needs to be optimized.

[0034] In another aspect of the embodiment of the present application, the controller determines a reason why the display device is falsely awakened based on the first target audio, and is specifically configured to:

[0035] Determining the number of times the first target audio is acquired within a preset time period;

[0036] When the number of times is greater than a second threshold and the display device is falsely awakened, it is determined that a reason for the false awakening of the display device is that a speech recognition model of the display device needs to be optimized.

[0037] If the number of times the first target audio is obtained within the preset time length is greater than the second threshold, it indicates that the speech recognition model has repeatedly recognized the initial character segment containing only the wake-up word in the audio, but has not recognized other character segments of the wake-up word. This may be because other audio is recognized as the initial character segment including the wake-up word, or the user has issued the wake-up word many times, but the speech recognition model has errors in recognition and cannot recognize other character segments of the wake-up word. This indicates that the recognition accuracy of the speech recognition model is low. Therefore, it can be determined that the reason for the false wake-up of the display device is that the speech recognition model of the display device needs to be optimized.

[0038] In another aspect of the embodiment of the present application, after the controller stores the audio collected by the sound collector after the target time in the linear buffer, the controller is further configured to:

[0039] When the speech recognition model recognizes that the second audio includes the remaining character segments of the wake-up word, obtaining a third audio by adding a first tag to the first audio, and obtaining a fourth audio by adding a second tag to the second audio;

[0040] controlling the display device to transmit the third audio and the fourth audio to the server, so that the server determines a cause of the false awakening of the display device based on the third audio and the fourth audio;

[0041] or,

[0042] When the second audio occupies all storage space of the linear buffer and the speech recognition model does not recognize that the second audio includes the remaining character segments of the wake-up word, obtaining a fifth audio by adding the first tag to the first audio;

[0043] The display device is controlled to transmit the fifth audio to the server, so that the server determines a reason why the display device is falsely awakened based on the fifth audio.

[0044] Through this solution, the server can receive the audio transmitted by the display device, and determine the cause of the display device's false wake-up based on the audio, thereby realizing the diagnosis of the false wake-up.

[0045] A first aspect of another embodiment of the present application provides a server, including:

[0046] The controller is configured as:

[0047] Receive a second target audio transmitted by a display device, where the second target audio includes a third audio and a fourth audio, or the second target audio is a fifth audio, the third audio is an audio obtained by adding a first tag to the first audio in the first target audio, the fourth audio is an audio obtained by adding a second tag to the second audio in the first target audio, and the fifth audio is an audio obtained by adding the first tag to the first audio, the first audio is audio collected by the display device before a target moment, the second audio is audio collected by the display device after the target moment, the target moment is the moment when the speech recognition model recognizes that the audio includes the initial character segment of the wake-up word, and the first target audio is the audio determined by the display device based on the first audio, the second audio, and the recognition result of the second audio by the speech recognition model;

[0048] A reason why the display device is falsely awakened is determined based on the second target audio and a tag in the second target audio.

[0049] Through the solution of the embodiment of the present application, the server can obtain the second target audio transmitted by the display device, and diagnose the false wake-up of the display device based on the second target audio.

[0050] Another embodiment of the present application provides a false awakening diagnosis method, which is applied to the display device as described in the above embodiment, and includes:

[0051] In response to receiving a start instruction of the far-field voice wake-up function, controlling the sound collector of the display device to collect audio, and calling the voice recognition model of the display device to perform voice recognition;

[0052] Storing the first audio collected by the sound collector before a target time in a ring buffer, where the target time is the time when the speech recognition model recognizes that the audio includes an initial character segment of the wake-up word;

[0053] Storing the second audio collected by the sound collector after the target time in a linear buffer, where the maximum storage capacity of the linear buffer is greater than or equal to the number of bytes included in the wake-up word;

[0054] Determining a first target audio based on the first audio, the second audio, and a recognition result of the second audio by the speech recognition model;

[0055] A reason why the display device is falsely awakened is determined based on the first target audio.

[0056] Another embodiment of the present application provides a false wakeup diagnosis method, which is applied to the server described in the above embodiment. The method includes:

[0057] Receive a second target audio transmitted by a display device, where the second target audio includes a third audio and a fourth audio, or the second target audio is a fifth audio, the third audio is an audio obtained by adding a first tag to the first audio in the first target audio, the fourth audio is an audio obtained by adding a second tag to the second audio in the first target audio, and the fifth audio is an audio obtained by adding the first tag to the first audio, the first audio is audio collected by the display device before a target moment, the second audio is audio collected by the display device after the target moment, the target moment is the moment when the speech recognition model recognizes that the audio includes the initial character segment of the wake-up word, and the first target audio is the audio determined by the display device based on the first audio, the second audio, and the recognition result of the second audio by the speech recognition model;

[0058] A reason why the display device is falsely awakened is determined based on the second target audio and a tag in the second target audio.

[0059] The display device provided in the embodiment of the present application can control the sound collector to start collecting audio in response to receiving a startup instruction for the far-field voice wake-up function, and control the voice recognition model to perform voice recognition on the audio collected by the sound collector. That is, after receiving the startup instruction for the far-field voice wake-up function, the display device starts collecting audio from the surrounding environment, that is, the audio is collected before the display device wakes up. In addition, the display device uses the moment when the initial character segment of the wake-up word is recognized in the audio as the target moment, marks the first audio collected before the target moment and stores it in a circular buffer, and marks the audio collected after the target moment and stores it in a linear buffer.

[0060] Since the maximum storage capacity of the linear buffer is greater than or equal to the number of bytes included in the wake-up word, if the speech recognition model recognizes that the audio includes a complete wake-up word, the audio including the complete wake-up word will be stored in the circular buffer and the linear buffer, or if the storage space of the linear buffer is full and the other character segments of the wake-up word are still not recognized, it indicates that only the initial character segment of the wake-up word is recognized. In addition, the present application also determines the first target audio based on the first audio, the second audio and the recognition result of the second audio.

[0061] That is to say, through the solution of the embodiment of the present application, the audio of the display device before and during the wake-up process can be collected, and the obtained first target audio can reflect the display device's recognition of the audio, based on which the reason for the false wake-up of the display device can be determined, meeting the need for diagnosing the false wake-up of the display device. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 is a schematic diagram illustrating an operation scenario between a display device and a control device according to some exemplary embodiments;

[0063] Figure 2 is a block diagram of a hardware configuration of a display device 200 according to some exemplary embodiments;

[0064] Figure 3 is a schematic diagram of software configuration of a display device 200 according to some exemplary embodiments;

[0065] Figure 4 A schematic diagram of a false wakeup diagnosis solution provided in the prior art;

[0066] Figure 5(a) is a schematic diagram of a linear buffer;

[0067] Figure 5(b) is a schematic diagram of a ring buffer;

[0068] Figure 6 A schematic diagram illustrating a method for diagnosing a false wakeup performed by a display device according to some exemplary embodiments;

[0069] FIG7( a ) is a schematic diagram of an interface of a remote controller of a display device according to some exemplary embodiments;

[0070] FIG7( b ) is a schematic diagram of another interface of a remote controller of a display device according to some exemplary embodiments;

[0071] Figure 8 A schematic diagram illustrating another method for diagnosing a false wakeup performed by a display device according to some exemplary embodiments;

[0072] Figure 9 A schematic diagram illustrating another method for diagnosing a false wakeup performed by a display device according to some exemplary embodiments;

[0073] Figure 10 A schematic diagram illustrating a method for diagnosing a false wakeup performed by a server according to some exemplary embodiments;

[0074] Figure 11 A schematic diagram illustrating another method for diagnosing a false wakeup performed by a server according to some exemplary embodiments;

[0075] Figure 12 A schematic diagram illustrating another method for diagnosing a false wakeup performed by a server according to some exemplary embodiments;

[0076] Figure 13 A schematic diagram illustrating another method for diagnosing a false wakeup performed by a server according to some exemplary embodiments;

[0077] Figure 14 A schematic diagram illustrating another method for diagnosing a false wakeup performed by a server according to some exemplary embodiments;

[0078] Figure 15 is a schematic diagram of modules in a display device according to some exemplary embodiments;

[0079] Figure 16 FIG. 1 is a timing diagram of a display device according to some exemplary embodiments. DETAILED DESCRIPTION

[0080] The following embodiments are described in detail, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numbers in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following embodiments are not intended to represent all possible implementations consistent with the present application. They are merely examples of systems and methods consistent with certain aspects of the present application, as detailed in the claims.

[0081] It should be noted that the brief descriptions of terms in this application are only for the purpose of facilitating the understanding of the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.

[0082] In the specification and claims of this application and the accompanying drawings, the terms "first," "second," "third," etc. are used to distinguish similar or similar objects or entities, and are not necessarily intended to limit a particular order or sequence, unless otherwise noted. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances.

[0083] The terms "comprise," "include," and "have," and any variations thereof, are intended to cover but not exclude inclusion; for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.

[0084] In the embodiments of the present application, a display device generally refers to a device capable of displaying images and processing data. For example, a display device includes but is not limited to smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc.

[0085] In the embodiments of the present application, the display device generally refers to a device with image display and data processing capabilities. The display device can have various implementation forms, for example, it can be a television, a smart TV, a laser projection device, a monitor, an electronic bulletinboard, an electronic table, etc. Figure 1 and Figure 2 This is a specific implementation of the display device of the present application.

[0086] Figure 1 Schematic diagram of an operation scenario between a display device and a control device according to an embodiment. Figure 1 As shown, the user can operate the display device 200 through the smart device 300 or the control apparatus 100 .

[0087] In some embodiments, the control device 100 may be a remote controller, and communication between the remote controller and the display device 200 may include infrared protocol communication, Bluetooth protocol communication, or other short-range communication methods, to control the display device 200 wirelessly or wiredly. The user may control the display device 200 by inputting user commands through buttons on the remote controller, voice input, control panel input, and the like.

[0088] In some embodiments, a smart device 300 (such as a mobile terminal, tablet computer, computer, laptop computer, etc.) can also be used to control the display device 200. For example, the display device 200 is controlled using an application running on the smart device 300.

[0089] In some embodiments, the display device 200 may not use the aforementioned smart device or control device to receive instructions, but may receive user control through touch or gestures.

[0090] In some embodiments, the display device 200 can also be controlled in a manner other than the control device 100 and the smart device 300. For example, the user's voice command control can be directly received through a module for obtaining voice commands configured inside the display device 200, or the user's voice command control can be received through a voice control device set outside the display device 200.

[0091] In some embodiments, the display device 200 also communicates data with the server 400. In this embodiment, the display device 200 can be connected to the server 400 via a local area network (LAN), a wireless local area network (WLAN), or other networks. The server 400 can provide various content and interactions to the display device 200. The server 400 can be a single cluster or multiple clusters, each of which can include one or more types of servers.

[0092] like Figure 2 The display device 200 includes at least one of a tuner and demodulator 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface.

[0093] In some embodiments, the controller includes a processor, a video processor, an audio processor, a graphics processor, RAM, ROM, and first to nth interfaces for input / output.

[0094] The display 260 includes a display screen component for presenting images, and a driving component for driving image display, a component for receiving image signals output from a controller, and a component for displaying video content, image content, and a menu control interface and a user control UI interface.

[0095] The display 260 may be a liquid crystal display, an OLED display, or a projection display, and may also be a projection device and a projection screen.

[0096] Communicator 220 is a component used to communicate with external devices or servers using various communication protocols. For example, the communicator may include at least one of a Wi-Fi module, a Bluetooth module, a wired Ethernet module, or other network communication protocol chip or a near-field communication protocol chip, as well as an infrared receiver. Display device 200 can use communicator 220 to send and receive control signals and data signals with external control device 100 or server 400.

[0097] The user interface can be used to receive control signals from the control device 100 (such as an infrared remote controller, etc.).

[0098] Detector 230 is used to collect signals from the external environment or external interactions. For example, detector 230 includes a light receiver, a sensor for collecting ambient light intensity; or detector 230 includes an image collector, such as a camera, for collecting external environmental scenes, user attributes, or user interaction gestures; or detector 230 includes a sound collector, such as a microphone, for receiving external sounds.

[0099] The external device interface 240 may include, but is not limited to, any one or more of the following: a high-definition multimedia interface (HDMI), an analog or digital high-definition component input interface (component), a composite video input interface (CVBS), a USB input interface (USB), an RGB port, etc. It may also be a composite input / output interface formed by multiple of the above interfaces.

[0100] The tuner-demodulator 210 receives broadcast television signals via wired or wireless reception, and demodulates audio and video signals, such as EPG data signals, from a plurality of wireless or wired broadcast television signals.

[0101] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0102] Controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory. Controller 250 controls the overall operation of display device 200. For example, in response to receiving a user command to select a UI object for display on display 260, controller 250 may perform operations related to the object selected by the user command.

[0103] In some embodiments, the controller includes at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), RAM (Random Access Memory, RAM), ROM (Read-Only Memory, ROM), a first interface to an nth interface for input / output, a communication bus (Bus), etc.

[0104] The user may input a user command through a graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input command through the graphical user interface (GUI). Alternatively, the user may input a user command through a specific voice or gesture, and the user input interface may recognize the voice or gesture through a sensor to receive the user input command.

[0105] A user interface is the medium for interaction and information exchange between an application or operating system and the user. It converts information between its internal form and a user-friendly format. A common user interface is the graphical user interface (GUI), which refers to a graphical user interface related to computer operations. It can be an icon, window, control, or other interface element displayed on an electronic device's display. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.

[0106] See also Figure 3 In some embodiments, the system is divided into four layers, from top to bottom: the application layer (referred to as the "application layer"), the application framework layer (referred to as the "framework layer"), the Android runtime and system library layer (referred to as the "system runtime library layer"), and the kernel layer.

[0107] In some embodiments, at least one application runs in the application layer. These applications can be window programs, system settings programs, clock programs, etc. that come with the operating system, or applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the above examples.

[0108] The framework layer provides applications with an application programming interface (API) and programming framework. The application framework layer includes predefined functions. The application framework layer acts as a processing center, determining the actions taken by applications in the application layer. Through the API, applications can access system resources and services during execution.

[0109] like Figure 3As shown, in the embodiment of the present application, the application framework layer includes managers, content providers, etc., wherein the manager includes at least one of the following modules: an activity manager (ActivityManager) is used to interact with all activities running in the system; a location manager (Location Manager) is used to provide system services or applications with access to system location services; a package manager (Package Manager) is used to retrieve various information related to the application packages currently installed on the device; a notification manager (NotificationManager) is used to control the display and clearing of notification messages; a window manager (Window Manager) is used to manage icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0110] In some embodiments, the activity manager is used to manage the lifecycle of each application and common navigation back functions, such as controlling the exit, opening, and backing of an application. The window manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, taking screenshots, and controlling display window changes (such as shrinking the display window, shaking the display, distorting the display, etc.).

[0111] In some embodiments, the system runtime layer provides support for the upper layer, namely the framework layer. When the framework layer is used, the Android operating system will run the C / C++ library contained in the system runtime layer to implement the functions to be implemented by the framework layer.

[0112] In some embodiments, the kernel layer is a layer between hardware and software. Figure 3 As shown, the kernel layer includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0113] Typically, users primarily use display devices for video and audio playback. When the user no longer needs the display device for video and audio playback, the display device can enter a dormant state to save energy. Alternatively, users can use the display device's far-field voice wake-up feature to wake the dormant display device.

[0114] Among them, the far-field voice wake-up function means that the display device collects audio from the surrounding environment in the sleep state, and then recognizes the audio by calling the voice recognition model to determine whether the audio includes the wake-up word. If it does, the display device enters the wake-up state from the sleep state to realize the wake-up of the display device.

[0115] However, in some cases, the user does not send audio that includes the wake-up word, but the display device still believes that the collected audio includes the wake-up word, thereby entering the wake-up state from the sleep state, that is, a false wake-up occurs, resulting in a degraded user experience.

[0116] To reduce false wakeups, some display devices currently adjust the wakeup threshold of their speech recognition models. However, this approach reduces wakeup sensitivity, increases difficulty, and has limited versatility. Therefore, a method for diagnosing false wakeups is urgently needed to identify the cause and subsequently implement targeted measures based on the cause.

[0117] A common method for diagnosing false awakening is Figure 4 See Figure 4 In this method, the same audio is played in two different environments, environment A and environment B. Then, the false awakening rate A in environment A and the false awakening rate B in environment B are obtained respectively. If the difference between the two false awakening rates is too large, for example, the false awakening rate A is much higher than the false awakening rate B, it is considered that the false awakening of the display device is caused by noise in the external environment.

[0118] However, in actual applications, there may be many reasons for display device false awakenings. For example, in addition to external environmental interference, the voice recognition model used by the display device to identify audio may have low recognition accuracy. Therefore, the accuracy of the above false awakening diagnosis method is low. Simply assuming that false awakenings in this scenario are caused by external environmental interference will bring hidden dangers to product quality.

[0119] Therefore, in response to the above situation, embodiments of the present application provide a display device and a false awakening diagnosis method to implement diagnosis of false awakening.

[0120] To clearly illustrate the embodiments of the present application, some explanations of related terms are given below.

[0121] A ring buffer (or loop buffer): Also known as a circular buffer or circular queue, it is a data structure used to represent a fixed-size, end-to-end connected buffer. A ring buffer is a circular buffer in which data is stored sequentially. When the end of the array is reached, the data automatically returns to the beginning of the array to continue storing data, forming a closed loop. Figure 5(a) shows an example of a ring buffer. The single arrow at the bottom of the figure indicates that the storage spaces are "connected end-to-end," forming a circular address space. This circular address space is the ring buffer.

[0122] A linear buffer is a contiguous data structure designed to efficiently handle sequential read and write operations. Its memory space is pre-allocated and fixed in size, supporting unidirectional movement between the front-end (write side) and the back-end (read side). Data is stored sequentially in the order it was written. Figure 5(b) shows an example of a linear buffer.

[0123] To enable false wakeup diagnosis, an embodiment of the present application provides a display device. The display device includes a display and a controller, wherein the display is configured to display a media resource interface, and the controller is coupled to the display. Furthermore, the display device may also include a sound collector, and the controller may also be coupled to the sound collector to receive audio captured by the sound collector.

[0124] See also Figure 6 , the controller can be configured to perform the following operations:

[0125] Step S100: In response to receiving a start instruction of the far-field voice wake-up function, controlling the sound collector to collect audio, and calling the voice recognition model to perform voice recognition on the audio.

[0126] In a feasible design of the present application, if the user wants the display device to activate the far-field voice wake-up function, the user will perform a touch operation on the menu option corresponding to the far-field voice wake-up function in the remote control of the display device. See the example shown in Figure 7(a), the touch operation is a touch operation on the menu option "Handsfree Function".

[0127] After receiving the touch operation, the user must also perform a long press on the remote control button corresponding to the activation function. For example, in Figure 7(b), this long press is for the case "する." In this case, the activation command is generated by the remote control after receiving the touch operation and the long press operation.

[0128] Of course, the startup instruction can also be an instruction generated by the display device or the remote control of the display device after receiving other forms of operations, and this application does not limit this.

[0129] After generating the start-up instruction, the remote controller transmits the start-up instruction to the display device so that the controller of the display device receives the start-up instruction.

[0130] After receiving the start instruction, the controller controls the sound collector to start collecting audio. In a feasible design, the sound collector may include a recording control panel set in the display device, or include Figure 2 The sound collector in the detector 230 shown. Of course, the sound collector can also be other forms of audio collection devices, which is not limited in this application.

[0131] In addition, the controller can also call the speech recognition model for speech recognition, so that the speech recognition model recognizes the audio collected by the sound collector and obtains the corresponding speech recognition results. Among them, the speech recognition model is a model used by the display device to identify whether the audio includes a wake-up word when applying the far-field wake-up function. In a feasible implementation method, the speech recognition model may include an artificial intelligence (AI) speech recognition model. Of course, the speech recognition model may also include other types of models, which is not limited in this application.

[0132] Step S200: Store the first audio collected by the sound collector before the target time into a ring buffer. The target time is the time when the speech recognition model recognizes that the audio collected by the sound collector includes the initial character segment of the wake-up word.

[0133] That is to say, from the time the sound collector starts collecting audio to the time the speech recognition model recognizes the initial character segment of the wake-up word contained in the audio collected by the sound collector, the audio collected by the sound collector is stored in a ring buffer.

[0134] The initial character segment typically includes the first r characters of the wake-up word, where r can be a preset positive integer. For example, if the wake-up word is "hello, turn on the TV," the initial character segment can be "hello." The target moment is the moment when the speech recognition model recognizes that the audio contains "hello." The audio collected before this target moment is stored in the ring buffer, that is, the audio stored in the ring buffer includes the initial character segment "hello."

[0135] A ring buffer is a circular buffer in which data is stored sequentially. When the end of the array is reached, data automatically returns to the beginning of the array to continue storage, forming a closed loop. Therefore, the audio collected by the sound collector can be stored in the ring buffer until the target time is reached. During the storage process, if the entire storage space of the ring buffer is occupied before the target time is reached, the newly collected audio can overwrite the previously stored audio, ensuring that the ring buffer can store audio including the initial character segment of the wake-up word.

[0136] In addition, in the embodiment of the present application, the maximum storage capacity of the ring buffer is generally not less than the number of bytes included in the initial character segment of the wake-up word. In one example, the maximum storage capacity of the ring buffer can be 1.5 times the total number of bytes included in the wake-up word.

[0137] Step S300: Store the second audio collected by the sound collector after the target time into a linear buffer, where the maximum storage capacity of the linear buffer is greater than or equal to the number of bytes included in the wake-up word.

[0138] That is to say, in the embodiment of the present application, during the process of the sound collector collecting audio, the speech recognition model recognizes the collected audio, and takes the moment when the speech recognition model recognizes that the audio collected by the sound collector includes the initial character segment of the wake-up word as the target moment. The target moment is the dividing point, and the first audio obtained before the target moment is stored in the circular buffer, and the second audio obtained after the target moment is stored in the linear buffer, that is, the start collection moment of the first audio is the moment when the sound collector is started, and the end collection moment of the first audio is the target moment.

[0139] In addition, the audio collected by the sound collector after the target time is the second audio, and the second audio is stored in the linear buffer, that is, the start time of collecting the second audio is the target time.

[0140] Since the maximum storage capacity of the linear buffer is greater than or equal to the number of bytes included in the wake-up word, and before the target moment, the audio collected by the sound collector already includes the initial character segment of the wake-up word, therefore, if the user around the display device utters a complete wake-up word, the audio stored in the linear storage area and the circular buffer includes the complete wake-up word.

[0141] For example, if the wake-up word is "hello, turn on the TV", the initial character segment may be "hello", and the remaining character segments of the wake-up word may be "turn on the TV". If the user sends the complete wake-up word after the display device receives the startup instruction of the far-field voice wake-up function, the audio stored in the circular buffer (i.e., the first audio) includes the initial character segment "hello", and the audio stored in the linear storage area (i.e., the second audio) includes the character segment "turn on the TV".

[0142] In addition, in some cases, the speech recognition model only recognizes the initial character segment of the wake-up word in the audio, but not other character segments of the wake-up word. In this case, the audio collected by the sound collector after the target time will be stored in the linear buffer until all the storage space of the linear buffer is occupied.

[0143] The embodiment of the present application can divide the wake-up process into two stages according to whether other character segments including the wake-up word are recognized in the audio after the target moment. The stage in which other character segments including the wake-up word are not recognized in the audio after the target moment can be called the pre-wake-up stage. In addition, after the pre-wake-up stage, if other character segments including the wake-up word are recognized in the audio, the display device will think that the complete wake-up word is received, and wake-up occurs. Therefore, the stage after recognizing other character segments including the ring word in the audio can be called the wake-up stage, but the wake-up may be a normal wake-up or a false wake-up caused by an error in the recognition of the speech recognition model.

[0144] That is, the pre-wake-up phase is entered at the target time, and the wake-up phase is entered after the target time when other character segments including the wake-up word are recognized in the audio.

[0145] Step S400: Determine a first target audio based on the first audio, the second audio, and a recognition result of the second audio by a speech recognition model.

[0146] The speech recognition model may recognize the second audio as including the remaining character segments of the wake-up word, or may not recognize the second audio as including the remaining character segments of the wake-up word. Accordingly, the controller may determine the first target audio for each of these two situations.

[0147] Step S500: Determine a reason why the display device is falsely awakened based on the first target audio.

[0148] The display device provided in the embodiment of the present application can control the sound collector to start collecting audio in response to receiving the startup instruction of the far-field voice wake-up function, and control the voice recognition model to perform voice recognition on the audio collected by the sound collector, that is, after receiving the startup instruction of the far-field voice wake-up function, it starts collecting audio from the surrounding environment, that is, the audio is collected before the display device wakes up. In addition, the display device uses the moment when the initial character segment of the wake-up word is recognized in the audio as the target moment, marks the audio collected before the target moment and stores it in the circular buffer, and marks the audio collected after the target moment and stores it in the linear buffer.

[0149] Since the maximum storage capacity of the linear buffer is greater than or equal to the number of bytes included in the wake-up word, if the speech recognition model recognizes that the audio includes a complete wake-up word, the audio including the complete wake-up word will be stored in the circular buffer and the linear buffer, or if the storage space of the linear buffer is full and the other character segments of the wake-up word are still not recognized, it indicates that only the initial character segment of the wake-up word is recognized. In addition, the present application also determines the first target audio based on the first audio, the second audio and the recognition result of the second audio.

[0150] That is to say, through the solution of the embodiment of the present application, the audio of the display device before and during the wake-up process can be collected, and the obtained first target audio can reflect the display device's recognition of the audio, based on which the reason for the false wake-up of the display device can be determined, meeting the need for diagnosing the false wake-up of the display device.

[0151] In actual application scenarios of display devices, the user may send a complete wake-up word, or may send only an initial character segment and then not send the remaining character segments of the wake-up word. For these two scenarios, the present application can perform step S400 through different embodiments.

[0152] In one embodiment provided in this application, see Figure 8 The operation of determining the first target audio based on the first audio, the second audio, and the recognition result of the second audio by the speech recognition model provided in step S400 may include the following steps:

[0153] Step S410: When the speech recognition model recognizes that the second audio includes the remaining character segments of the wake-up word, determine the first byte number of the second audio stored in the linear buffer.

[0154] The remaining character segments are the other character segments in the wake-up word except the initial character segment. For example, if the wake-up word is "hello, turn on the TV", and "hello" is the initial character segment of the wake-up word, the remaining character segments may be "turn on the TV".

[0155] Step S420: Calculate the difference between the target number of bytes and the first number of bytes, where the target number of bytes is not less than the number of bytes included in the wake-up word.

[0156] In one example, the target number of bytes may be 1.5 times the number of bytes included in the wake word.

[0157] Step S430: extract the first audio subset stored most recently from the first audio stored in the ring buffer.

[0158] The number of bytes included in the first audio subset is the difference between the target number of bytes and the first number of bytes. In other words, the sum of the number of bytes included in the first audio subset and the second audio is the target number of bytes.

[0159] Since the first audio subset is the latest stored part of the first audio, the first audio subset includes the initial character segment of the wake-up word.

[0160] Step S440: Determine whether the first target audio includes the first audio subset and the second audio.

[0161] In this step, the first audio subset includes the initial character segment of the wake-up word. When the first audio subset is collected, the display device is not actually awakened. Therefore, this first audio subset can also be called pre-wake-up word audio. In addition, the second audio includes the remaining character segment of the wake-up word. After receiving this second audio, the speech recognition model will believe that the complete wake-up word has been recognized, and the display device will wake up. Therefore, this second audio can also be called wake-up word audio.

[0162] In this embodiment, the speech recognition model recognizes that the second audio includes the remaining character segments of the wake-up word, then the display device will usually believe that the audio collected by the sound collector includes the complete wake-up word, and thus wake-up occurs, but the wake-up may be a false wake-up. Accordingly, since the first audio collected before the target moment includes the initial character segment of the wake-up word, and the second audio collected after the target moment includes the remaining character segments of the wake-up word, the audio obtained by merging the first audio and the second audio is recognized by the speech recognition model as including the complete wake-up word. In this case, the first target audio can be determined based on the first audio subset and the second audio intercepted from the first audio, and the cause of the false wake-up can be diagnosed through the first target audio.

[0163] Through this embodiment, when the speech recognition model recognizes that the second audio includes the remaining character segments in the wake-up word, the first target audio including the first audio subset and the second audio can be determined, and the first audio subset is recognized by the speech recognition model as including the initial character segment of the wake-up word, and the second audio is recognized by the speech recognition model as including the remaining character segments of the wake-up word. Therefore, it can be considered that the first target audio is recognized by the speech recognition model as including the complete wake-up word, and the diagnosis of false wake-up can be achieved based on the first target audio.

[0164] In another embodiment of the present application, the controller may further perform the following steps before performing step S420:

[0165] The first step is to determine the number of sampling points for the wake-up word based on the sampling rate of the sound collector and the duration of speaking the wake-up word.

[0166] The number of sampling points of the wake-up word is usually the ratio of the duration of speaking the wake-up word to the sampling rate of the sound collector.

[0167] The second step is to determine the number of bytes contained in the wake-up word based on the number of sampling points of the wake-up word and the number of bytes corresponding to each sampling point.

[0168] The third step is to calculate the product of the number of bytes contained in the wake-up word and the preset coefficient to obtain the target number of bytes. The preset coefficient is a positive number not less than 1.

[0169] Exemplarily, the preset coefficient is 1.5. In this case, the target number of bytes can be calculated using the following formula:

[0170] n=1.5T / S×N formula (1).

[0171] In the above formula, T represents the duration of speaking the wake-up word, S represents the sampling rate of the sound collector, N represents the number of bytes corresponding to each sampling point, and n represents the target number of bytes.

[0172] Through this embodiment, the controller can determine the target number of bytes, and the determined target number of bytes is not less than the number of bytes included in the wake-up word, so that when the audio including the wake-up word is collected, it can be ensured that the first target audio determined according to the target number of bytes can include the complete wake-up word.

[0173] In a feasible design, the maximum storage capacity of the linear buffer can also be n bytes calculated by formula (1). In addition, the maximum storage capacity of the ring buffer can also be n bytes calculated by formula (1).

[0174] In another embodiment, see Figure 9 The operation of determining the first target audio based on the first audio, the second audio, and the recognition result of the second audio by the speech recognition model provided in S400 may include the following steps:

[0175] Step S450: Determine whether the second audio occupies all storage space of the linear buffer.

[0176] Step S460: When the second audio occupies all the storage space of the linear buffer and the speech recognition model does not recognize that the second audio includes the remaining character segments of the wake-up word, determine that part or all of the first audio is the first target audio.

[0177] Since the maximum storage capacity of the linear buffer is greater than or equal to the number of bytes included in the wake-up word, and the circular buffer has already stored the initial character segment of the wake-up word, if the second audio has occupied all the storage space of the linear buffer, but the speech recognition model has not yet recognized that the second audio includes the remaining character segments of the wake-up word, it indicates that the speech recognition model has not recognized the complete wake-up word, and accordingly, the display device will not wake up.

[0178] However, if the speech recognition model fails to recognize the complete wake-up word multiple times, but only recognizes the initial character segment of the wake-up word, it may be that other audio is recognized as including the initial character segment of the wake-up word, or the user has issued the wake-up word multiple times, but the speech recognition model has errors and cannot recognize the other character segments of the wake-up word. This indicates that the recognition accuracy of the speech recognition model is low, which will increase the probability of false wake-up of the display device. In this case, part or all of the fields in the first audio can be used as the first target audio, so that the cause of the false wake-up can be determined later based on this type of first target audio.

[0179] In one embodiment of the present application, if the first target audio includes the first audio subset and the second audio, that is, the first target audio is determined through steps S410 to S440, then see Figure 10 Determining a reason why the display device is falsely awakened based on the first target audio may include the following steps:

[0180] Step S510: Determine the similarity between the first target audio and the audio of the wake-up word.

[0181] Step S520: When the similarity is less than the first threshold and the display device is falsely awakened, determine that the reason for the false awakening of the display device is that the speech recognition model of the display device needs to be optimized.

[0182] Exemplarily, the first threshold value may be 90%. Of course, the first threshold value may also be other values, which is not limited in this application.

[0183] If the first target audio includes the first audio subset and the second audio, it indicates that the display device's speech recognition model has recognized the complete wake-up word. In this case, if the similarity between the first target audio and the audio of the wake-up word is less than the first threshold, and the display device is falsely awakened in this case, it indicates that the display device's speech recognition model for far-field voice wake-up still recognizes that the first target audio includes the complete wake-up word even though the similarity between the first target audio and the audio of the wake-up word is low. Therefore, it can be determined that the recognition accuracy of the speech recognition model is low, and the reason for the false awakening of the display device is that the display device's speech recognition model needs to be optimized.

[0184] In one embodiment of the present application, if the first target audio is part or all of the fields in the first audio, that is, the first target audio is determined by steps S450 to S460, then see Figure 11 Determining a reason why the display device is falsely awakened based on the first target audio may include the following steps:

[0185] Step S530: Determine the number of times the first target audio is acquired within a preset time period.

[0186] Step S540: When the number of times is greater than the second threshold and the display device is falsely awakened, determine that the reason for the false awakening of the display device is that the speech recognition model of the display device needs to be optimized.

[0187] If the number of times the first target audio is obtained within the preset time length is greater than the second threshold, it indicates that the speech recognition model has repeatedly recognized the initial character segment containing only the wake-up word in the audio, but has not recognized other character segments of the wake-up word. This may be because other audio is recognized as the initial character segment including the wake-up word, or the user has issued the wake-up word many times, but the speech recognition model has errors in recognition and cannot recognize other character segments of the wake-up word. This indicates that the recognition accuracy of the speech recognition model is low. Therefore, it can be determined that the reason for the false wake-up of the display device is that the speech recognition model of the display device needs to be optimized.

[0188] Through the above embodiment, the display device can determine the cause of the display device's false awakening based on the first target audio, thereby diagnosing the false awakening. In addition, the server can also perform false awakening diagnosis. In this case, after the controller executes the operation of storing the audio collected by the sound collector after the target time into the linear buffer, in another embodiment of the present application, the controller can also perform the following operations:

[0189] When the speech recognition model recognizes that the second audio includes other character segments of the wake-up word, the third audio is obtained by adding the first tag to the first audio, and the fourth audio is obtained by adding the second tag to the second audio.

[0190] The display device is controlled to transmit the third audio and the fourth audio to the server, so that the server determines a cause of false awakening of the display device based on the third audio and the fourth audio.

[0191] The adding of the first mark to the first audio may be adding the first mark to all the first audio, or may be adding the first mark to a first audio subset in the first audio.

[0192] In this embodiment, the first mark and the second mark are different. For example, the first mark may be the character segment "PREWU" and the second mark may be the character segment "WU". Of course, the first mark and the second mark may also be other character segments, which is not limited in this application.

[0193] Through this embodiment, the server can receive the third audio and the fourth audio transmitted by the display device, and determine the cause of the false awakening of the display device based on the received audio, thereby achieving the diagnosis of the false awakening.

[0194] In addition, when the display device transmits the third audio and the fourth audio to the server, it can transmit the third audio and the fourth audio separately in two channels, or it can combine the third audio and the fourth audio into one channel for transmission, which is not limited in this application.

[0195] Alternatively, in another embodiment, after the controller stores the audio collected by the sound collector after the target time into the linear buffer, the controller may further perform the following operations:

[0196] When the second audio occupies all storage space of the linear buffer and the speech recognition model does not recognize that the second audio includes the remaining character segments of the wake-up word, a fifth audio is obtained by adding the first tag to the first audio;

[0197] The display device is controlled to transmit a fifth audio to the server, so that the server determines a cause of false awakening of the display device based on the fifth audio.

[0198] In this embodiment, the first mark is added to the first audio. The first mark can be added to the entire first audio or only to part of the first audio. This embodiment of the present application does not limit this.

[0199] Through this embodiment, the server can receive the fifth audio transmitted by the display device, and determine the cause of the false awakening of the display device based on the fifth audio, thereby realizing the diagnosis of the false awakening.

[0200] Corresponding to the above-mentioned display device, the following embodiment of the present application provides a server. The server can obtain audio transmitted by the display device through data interaction with the display device, and determine the reason why the display device is falsely awakened based on the audio.

[0201] In the embodiment of the present application, the server includes a controller, see Figure 12 , the controller can be configured to perform the following operations:

[0202] Step S600: Receive a second target audio signal transmitted by a display device.

[0203] Among them, the second target audio includes the third audio and the fourth audio, or the second target audio is the fifth audio, the third audio is the audio obtained by adding the first mark to the first audio in the first target audio, the fourth audio is the audio obtained by adding the second mark to the second audio in the first target audio, and the fifth audio is the audio obtained by adding the first mark to the first audio. The first audio is the audio collected by the display device before the target moment, and the second audio is the audio collected by the display device after the target moment. The target moment is the moment when the speech recognition model recognizes that the audio includes the initial character segment of the wake-up word, and the first target audio is the audio determined by the display device based on the recognition result of the second audio by the first audio, the second audio and the speech recognition model.

[0204] In addition, the third audio may be audio obtained by adding the first tag to a subset of the first audio, or may be audio obtained by adding the first tag to all of the first audio. Furthermore, the fifth audio may be audio obtained by adding the first tag to part or all of the first audio.

[0205] Step S700: Determine a reason why the display device is falsely awakened based on the second target audio and a mark in the second target audio.

[0206] Through the embodiments of the present application, the server can obtain the second target audio transmitted by the display device, and determine the reason why the display device is falsely awakened based on the second target audio and the mark included in the second target audio.

[0207] In another embodiment of the present application, see Figure 13 The server can implement the operation disclosed in step S700 through the following embodiments:

[0208] Step S710: identifying a marker in the second target audio;

[0209] Step S720: When the marker in the second target audio includes the first marker and the second marker, determine the similarity between the second target audio and the audio of the wake-up word;

[0210] If the mark in the second target audio includes the first mark and the second mark, it indicates that the second target audio includes the third audio and the fourth audio. If the display device obtains the third audio by adding the first mark to the first audio subset, this step can directly compare the similarity between the second target audio and the audio of the wake-up word. In addition, if the display device obtains the third audio by adding the first mark to all the first audios, in this step, the server can intercept the latest stored audio subset from the third audio, and the intercepted audio subset includes the initial character segment of the wake-up word, and then merge the intercepted audio subset with the fourth audio, and then compare the similarity between the merged audio and the audio of the wake-up word.

[0211] Step S730: When the similarity is less than the first threshold and the display device is falsely awakened, determine that the reason for the false awakening of the display device is that the speech recognition model of the display device needs to be optimized.

[0212] In this embodiment, if the mark in the second target audio includes the first mark and the second mark, it indicates that the second target audio includes the third audio and the fourth audio, and accordingly, it indicates that the speech recognition model recognizes the complete wake-up word. In this case, if the similarity between the second target audio and the audio of the wake-up word is less than the first threshold, and the display device is falsely awakened in this case, it indicates that the speech recognition model used for far-field voice wake-up of the display device still recognizes the audio collected by the sound collector as including the complete wake-up word when the similarity between the second target audio and the audio of the wake-up word is low, thereby determining that the recognition accuracy of the speech recognition model is low, and the reason for the false awakening of the display device is that the speech recognition model of the display device needs to be optimized.

[0213] In addition, in another embodiment of the present application, see Figure 14 The server can implement the operation disclosed in step S700 through the following embodiments:

[0214] Step S740: identifying a marker in the second target audio;

[0215] Step S750: When the mark in the second target audio is the first mark, determine the number of times the second target audio is obtained within a preset time length.

[0216] Step S760: When the number of times is greater than the second threshold and the display device is falsely awakened, determine that the reason for the false awakening of the display device is that the speech recognition model of the display device needs to be optimized.

[0217] If within the preset time period, the number of times the server obtains the second target audio is greater than the second threshold, it indicates that the speech recognition model of the display device has repeatedly recognized the initial character segment containing only the wake-up word in the audio, but has not recognized other character segments of the wake-up word. The recognition accuracy of the speech recognition model of the display device is low. In this case, the server can determine that the reason for the false wake-up of the display device is that the speech recognition model of the display device needs to be optimized.

[0218] Furthermore, after the server identifies that the speech recognition model of the display device needs to be optimized, it can also determine the optimized speech recognition model and send the optimized speech recognition model to the display device so that the display device can wake up according to the optimized speech recognition model, reducing the chance of false wake-up.

[0219] Figure 15A schematic diagram of an example of a module included in the display device provided in an embodiment of the present application is shown in FIG. Figure 15 The controller of the display device may include the following modules: a speech recognition model (if the speech recognition model is an AI speech recognition model, it may also be called AIVoiceModel), a wake-up callback module (ie WakeUP_CB), a wake-up word collection callback module (ie Echo_CB) and a middleware service module (ie MidWare).

[0220] Among them, the wake-up callback module and the wake-up word collection callback module can be registered with the speech recognition model.

[0221] The speech recognition model can obtain audio collected by the sound collector, recognize the audio, and determine the wake-up state based on the recognition results. If the speech recognition model recognizes that the audio includes the initial character segment of the wake-up word, the wake-up state is determined to have entered the pre-wake-up phase; if the speech recognition model recognizes that the audio also includes the remaining character segments of the wake-up word, the wake-up state is determined to have entered the wake-up phase. In addition, the speech recognition model can also trigger the wake-up callback module accordingly based on the recognized wake-up state.

[0222] After being triggered by the speech recognition model, the wake-up callback module can notify the wake-up word collection callback module of the corresponding wake-up status.

[0223] In addition, the speech recognition model can also obtain the audio collected by the sound collector, so that the wake-up word collection callback module can obtain the corresponding audio through the speech recognition model.

[0224] The wake-up word collection callback module can add corresponding tags to the audio according to the recognition results of the speech recognition model. For example, after recognizing that the audio includes the initial character segment and other character segments of the wake-up word, a first tag is added to the first audio to obtain the third audio, and a second tag is added to the second audio to obtain the fourth audio. Or, when all the storage space of the linear buffer is occupied, after recognizing that the audio only includes the initial character segment of the wake-up word, a first tag is added to part or all of the first audio to obtain the fifth audio. Moreover, after adding the tag, the wake-up word collection callback module can transmit the tagged audio to the middleware service module.

[0225] After receiving the tagged audio from the wake-up word collection callback module, the middleware service module can transmit the audio to the server. In one feasible design, if a third audio and a fourth audio need to be transmitted to the server, the middleware service module can transmit them in two separate channels, or combine the third and fourth audio channels into one channel for transmission. Figure 15 In the transmission mode of the middleware service module, the third audio and the fourth audio are transmitted in two channels.

[0226] After receiving the audio transmitted by the display device, the server determines the cause of the false awakening of the display device according to the audio and the mark included in the audio, thereby realizing the diagnosis of the false awakening.

[0227] Another embodiment of the present application provides a process for each module in the display device to perform operations in the process of determining the cause of the display device's false awakening through the server. In this embodiment, labels corresponding to different awakening states can be set, including: FIRST (indicating that the pre-awakening stage has not been entered), SECOND (indicating that the pre-awakening stage has been entered), and OVER (indicating that the awakening stage has been entered). Figure 16 As shown in the timing diagram, the process includes the following operations:

[0228] Step S151: After the sound collector collects the audio, the speech recognition model of the display device recognizes the audio and transmits first trigger information to the wake-up callback module.

[0229] Step S152: After receiving the first trigger information, the wakeup callback module transmits a first notification message to the wakeup word collection callback module. The first notification message is used to indicate the start of the false wakeup diagnosis so that the wakeup word collection callback module can start to perform subsequent operations.

[0230] Step S153: The speech recognition model transmits the first audio to the wake-up word collection callback module.

[0231] The first audio is the audio collected before the target moment, and the target moment is the moment when the speech recognition model recognizes that the collected audio includes the initial character segment of the wake-up word.

[0232] Step S154: The wake-up word collection callback module adds a first tag to the first audio to obtain a third audio.

[0233] Step S155: The wake-up word collection callback module stores the third audio in a ring buffer.

[0234] The operations from step S153 to step S155 may be referred to as operations in the FIRST phase. During this process, the display device is in a state where it has not yet entered the pre-wake-up phase.

[0235] Step S156: The speech recognition model continues to recognize the audio collected by the sound collector. After recognizing that the audio includes the initial character segment of the wake-up word, the audio collected by the sound collector is called the second audio. The speech recognition model transmits the second audio to the wake-up word collection callback module.

[0236] Step S157: The wake-up word collection callback module adds a second tag to the second audio to obtain a fourth audio.

[0237] Step S158: After receiving the fourth audio, the wake-up word collection callback module determines whether the linear buffer is full. If so, the linear buffer is adjusted to no longer receive the second audio. If not, step S159 is executed.

[0238] In addition, if the remaining character segments of the wake-up word are not identified in the collected audio, and the wake-up word collection callback module determines that the storage space of the linear buffer is full, the wake-up word collection callback module performs the operation of step S163.

[0239] Step S159: If the wake-up word collection callback module determines that the storage space of the linear buffer is not full, the wake-up word collection callback module stores the fourth audio into the linear buffer.

[0240] Among them, the operations from step S156 to step S159 can be called operations in the SECOND stage. In this process, since the speech recognition model recognizes the initial character segment including the wake-up word in the first audio, the display device enters the pre-wake-up stage.

[0241] Step S160: After the speech recognition model recognizes that the second audio includes the remaining character segments in the wake-up word, or when the storage space of the linear buffer is fully occupied, the speech recognition model still does not recognize that the second audio includes the remaining character segments in the wake-up word, and transmits the second trigger information to the wake-up callback module.

[0242] Step S161: After receiving the second trigger information, the wake-up callback module transmits the second notification information to the wake-up word collection callback module.

[0243] Step S162: After receiving the second notification information, the wake-up word collection callback module determines the second target audio.

[0244] The second target audio may include the third audio and the fourth audio, or the second target audio may be the fifth audio.

[0245] Step S163: The wake-up word collection callback module transmits the second target audio to the middleware service module.

[0246] Step S164: The wake-up word collection callback module adjusts its own state to a state of terminating reception of the third audio.

[0247] Step S165: The wake-up word collection callback module enters a state of stopping the current diagnosis and ends information processing.

[0248] Among them, the operations from step S156 to step S159 can be called the operations of the OVER stage. During this process, if the speech recognition model recognizes the remaining character segments of the wake-up word, it can be considered that the display device enters the wake-up stage.

[0249] Step S166: If the second target audio includes the third audio and the fourth audio, the middleware service module transmits the third audio and the fourth audio to the server. The middleware service module may transmit the third audio and the fourth audio separately in two channels, or may combine the third audio and the fourth audio into one channel and transmit the audio to the server.

[0250] Step S167: If the second target audio is the fifth audio, the middleware service module transmits the fifth audio to the server.

[0251] Through the operations of step S166 and step S167, the server can obtain the second target audio and determine the reason why the display device is falsely awakened based on the second target audio.

[0252] Step S168: The middleware service module transmits end information to the wake-up word collection callback module.

[0253] Step S169: The wake-up word collection callback module adjusts its own state to a state of terminating reception of the second audio and ends information processing.

[0254] Through the operation of the embodiments of the present application, the modules in the display device can cooperate with each other to determine the second target audio for false wake-up diagnosis, and transmit the second target audio to the server, so that the server can determine the cause of the false wake-up of the display device based on the second target audio and realize the diagnosis of false wake-up.

[0255] Referring to the above embodiments, other embodiments of the present application further provide a false awakening diagnosis method, which is applied to the display device in the above embodiments of the present application. The false awakening diagnosis method includes the following steps:

[0256] In the first step, in response to receiving a start instruction of the far-field voice wake-up function, the sound collector of the display device is controlled to collect audio, and the voice recognition model of the display device is called to perform voice recognition;

[0257] The second step is to store the first audio collected by the sound collector before the target time into the ring buffer. The target time is the time when the speech recognition model recognizes that the audio includes the initial character segment of the wake-up word;

[0258] Step 3: storing the second audio collected by the sound collector after the target time into a linear buffer, where the maximum storage capacity of the linear buffer is greater than or equal to the number of bytes included in the wake-up word;

[0259] Step 4: determining a first target audio based on the first audio, the second audio, and a recognition result of the second audio by the speech recognition model;

[0260] In the fifth step, a reason why the display device is falsely awakened is determined based on the first target audio.

[0261] Through this method, the display device can determine the cause of the false awakening and realize the diagnosis of the false awakening.

[0262] Referring to the above embodiments, some other embodiments of the present application further provide a false awakening diagnosis method, which is applied to the server in the above embodiments of the present application. The false awakening diagnosis method includes the following steps:

[0263] The first step is to receive the second target audio transmitted by the display device;

[0264] Among them, the second target audio includes the third audio and the fourth audio, or the second target audio is the fifth audio, the third audio is the audio obtained by adding the first mark to the first audio in the first target audio, the fourth audio is the audio obtained by adding the second mark to the second audio in the first target audio, and the fifth audio is the audio obtained by adding the first mark to the first audio. The first audio is the audio collected by the display device before the target time, and the second audio is the audio collected by the display device after the target time. The target time is the moment when the speech recognition model recognizes that the audio includes the initial character segment of the wake-up word, and the first target audio is the audio determined by the display device based on the recognition result of the first audio, the second audio and the speech recognition model for the second audio;

[0265] In the second step, based on the second target audio and the mark in the second target audio, the reason why the display device is falsely awakened is determined.

[0266] Through this method, the server can determine the cause of the false wakeup and diagnose the false wakeup.

[0267] For ease of explanation, the above description has been made in conjunction with specific embodiments. However, the above discussion of some embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations can be obtained. The above embodiments are selected and described to better explain the content of this disclosure, thereby enabling those skilled in the art to better use the embodiments.

Claims

1. A display device, characterized in that: include: monitor; Sound collector; a controller coupled to the display and configured to: In response to receiving a start instruction of the far-field voice wake-up function, controlling the sound collector to collect audio, and calling the voice recognition model to perform voice recognition on the audio; Storing the first audio collected by the sound collector before a target time in a ring buffer, where the target time is the time when the speech recognition model recognizes that the audio includes an initial character segment of the wake-up word; Storing the second audio collected by the sound collector after the target time in a linear buffer, where the maximum storage capacity of the linear buffer is greater than or equal to the number of bytes included in the wake-up word; Determining a first target audio based on the first audio, the second audio, and a recognition result of the second audio by the speech recognition model; A reason why the display device is falsely awakened is determined based on the first target audio.

2. The display device according to claim 1, wherein The controller determines a first target audio based on the first audio, the second audio, and a recognition result of the second audio by the speech recognition model, and is specifically configured to: When the speech recognition model recognizes that the second audio includes the remaining character segments of the wake-up word, determining the number of first bytes of the second audio stored in the linear buffer; Calculating a difference between a target number of bytes and the first number of bytes, where the target number of bytes is not less than the number of bytes included in the wake-up word; intercepting a first audio subset stored most recently from the first audio stored in the ring buffer, where the number of bytes included in the first audio subset is the difference; It is determined that the first target audio includes the first audio subset and the second audio.

3. The display device according to claim 2, wherein Before the controller calculates the difference between the target number of bytes and the first number of bytes, the controller is further configured to: Determining the number of sampling points for the wake-up word based on the sampling rate of the sound collector and the duration of speaking the wake-up word; Determining the number of bytes contained in the wake-up word based on the number of sampling points of the wake-up word and the number of bytes of each sampling point; The target number of bytes is obtained by multiplying the number of bytes contained in the wake-up word by a preset coefficient, where the preset coefficient is a positive number not less than 1.

4. The display device according to claim 1, wherein The controller determines a first target audio based on the first audio, the second audio, and a recognition result of the second audio by the speech recognition model, and is specifically configured to: determining whether the second audio occupies all storage space of the linear buffer; When the second audio occupies all storage space of the linear buffer and the speech recognition model does not recognize that the second audio includes the remaining character segments of the wake-up word, part or all of the first audio is determined to be the first target audio.

5. The display device according to claim 2 or 3, characterized in that The controller determines a reason why the display device is falsely awakened based on the first target audio, and is specifically configured to: Determining a similarity between the first target audio and the audio of the wake-up word; When the similarity is less than the first threshold and the display device is falsely awakened, it is determined that the reason for the false awakening of the display device is that the speech recognition model of the display device needs to be optimized.

6. The display device according to claim 4, wherein: The controller determines a reason why the display device is falsely awakened based on the first target audio, and is specifically configured to: Determining the number of times the first target audio is acquired within a preset time period; When the number of times is greater than a second threshold and the display device is falsely awakened, it is determined that a reason for the false awakening of the display device is that a speech recognition model of the display device needs to be optimized.

7. The display device according to claim 1, wherein After the controller executes the step of storing the audio collected by the sound collector after the target time in the linear buffer, the controller is further configured to: When the speech recognition model recognizes that the second audio includes the remaining character segments of the wake-up word, obtaining a third audio by adding a first tag to the first audio, and obtaining a fourth audio by adding a second tag to the second audio; controlling the display device to transmit the third audio and the fourth audio to the server, so that the server determines a cause of the false awakening of the display device based on the third audio and the fourth audio; or, When the second audio occupies all storage space of the linear buffer and the speech recognition model does not recognize that the second audio includes the remaining character segments of the wake-up word, obtaining a fifth audio by adding the first tag to the first audio; The display device is controlled to transmit the fifth audio to the server, so that the server determines a reason why the display device is falsely awakened based on the fifth audio.

8. A server, characterized in that: include: The controller is configured as: Receive a second target audio transmitted by a display device, where the second target audio includes a third audio and a fourth audio, or the second target audio is a fifth audio, the third audio is an audio obtained by adding a first tag to the first audio in the first target audio, the fourth audio is an audio obtained by adding a second tag to the second audio in the first target audio, and the fifth audio is an audio obtained by adding the first tag to the first audio, the first audio is audio collected by the display device before a target moment, the second audio is audio collected by the display device after the target moment, the target moment is the moment when the speech recognition model recognizes that the audio includes the initial character segment of the wake-up word, and the first target audio is the audio determined by the display device based on the first audio, the second audio, and the recognition result of the second audio by the speech recognition model; A reason why the display device is falsely awakened is determined based on the second target audio and a tag in the second target audio.

9. A false awakening diagnosis method, characterized in that: Applied to the display device according to any one of claims 1 to 7, the method comprises: In response to receiving a start instruction of the far-field voice wake-up function, controlling the sound collector of the display device to collect audio, and calling the voice recognition model of the display device to perform voice recognition; Storing the first audio collected by the sound collector before a target time in a ring buffer, where the target time is the time when the speech recognition model recognizes that the audio includes an initial character segment of the wake-up word; Storing the second audio collected by the sound collector after the target time in a linear buffer, where the maximum storage capacity of the linear buffer is greater than or equal to the number of bytes included in the wake-up word; Determining a first target audio based on the first audio, the second audio, and a recognition result of the second audio by the speech recognition model; A reason why the display device is falsely awakened is determined based on the first target audio.

10. A false awakening diagnosis method, characterized in that: Applied to the server according to claim 8, the method comprises: Receive a second target audio transmitted by a display device, where the second target audio includes a third audio and a fourth audio, or the second target audio is a fifth audio, the third audio is an audio obtained by adding a first tag to the first audio in the first target audio, the fourth audio is an audio obtained by adding a second tag to the second audio in the first target audio, and the fifth audio is an audio obtained by adding the first tag to the first audio, the first audio is audio collected by the display device before a target moment, the second audio is audio collected by the display device after the target moment, the target moment is the moment when the speech recognition model recognizes that the audio includes the initial character segment of the wake-up word, and the first target audio is the audio determined by the display device based on the first audio, the second audio, and the recognition result of the second audio by the speech recognition model; A reason why the display device is falsely awakened is determined based on the second target audio and a tag in the second target audio.