A display apparatus and a method of guiding a user using a voice command
By acquiring and filtering user operation sequences, the device generates dynamic guidance messages, which solves the problem of mismatch between guidance messages and user operations in existing technologies, and improves the user interaction and device interaction experience.
Patent Information
- Application Number
- CN202411884424.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-12-19
AI Technical Summary
Existing display devices use fixed scenarios when guiding users to use voice commands, which cannot adapt to the user's real-world usage scenarios. This results in a mismatch between the guidance and the user's operation, affecting the user's interactive experience.
The display device acquires the user's operation sequence, filters frequently occurring and highly relevant operation events, generates dynamic prompts, and displays them immediately after the last operation is performed, thereby enhancing the relevance to the user's operation.
It achieves precise matching between the guidance text and user operations, improves the user's sense of interaction, eliminates the need for users to search for guidance text, and enhances the interactive experience of the display device.
Smart Images

Figure CN119854549B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of display devices, and in particular to a display device and a method for guiding a user to use a voice instruction. BACKGROUND
[0002] A display device refers to a terminal device capable of outputting a specific display screen, and can be a smart television, a mobile terminal, a smart advertising screen, a projector, or the like. Taking a smart television as an example, a smart television is a television product based on Internet application technology, having an open operating system and a chip, and possessing an open application platform, and capable of realizing bidirectional man-machine interaction, and integrating functions such as audio, video, and entertainment, and data, and the like, and is used to meet the diversified and personalized needs of users.
[0003] The display device can be configured with a far-field voice function, for example, a “voice assistant”. Based on the far-field voice function, the display device can respond to a voice instruction of a user to execute a corresponding task, so that the user can interact with the display device in a more convenient manner through a voice. The standard degree of the voice instruction input by the user is one of important factors affecting the accuracy of triggering the display device to execute a task and determining the task to be executed by the display device. Therefore, in order to improve the standard degree of the voice instruction input by the user, the display device can be configured with a guiding function. Based on the guiding function, the display device can show a guiding sentence to the user to guide the user to learn a standard voice instruction through the guiding sentence.
[0004] The display device usually provides a guiding sentence related to the far-field voice function to the user in an operation guiding process of the display device after the first boot. Alternatively, the display device displays a guiding sentence related to the far-field voice function on a home page. Alternatively, the display device displays a guiding sentence related to a voice instruction corresponding to a specified control when the user triggers the specified control. It can be seen that the scenario of triggering the display device to guide the user to use the voice instruction is relatively fixed, and the guiding sentence displayed is also relatively fixed, and cannot well adapt to the scenario in which the user actually uses the display device. SUMMARY
[0005] The present application provides a display device and a method for guiding a user to use a voice instruction, which can dynamically display a matching guiding sentence based on a user operation to guide the user to use a voice instruction corresponding to the user operation.
[0006] In a first aspect, the present application provides a display device, comprising:
[0007] a display configured to display a user interface;
[0008] The controller is configured to: acquire at least one operation sequence when a trigger condition for generating a guide language is met, each operation sequence including at least one operation event, the operation event corresponding to a non-speech instruction of a user; filter, from a plurality of sub-sequences included in the at least one operation sequence, at least one target sub-sequence whose frequency of occurrence in the at least one operation sequence is greater than or equal to a frequency threshold, wherein the sub-sequence includes at least one operation event; filter at least one target operation event according to a correlation between the operation events in the at least one target sub-sequence; generate a guide language corresponding to the at least one target operation event according to a semantic instruction corresponding to each target operation event, the guide language being used to guide the user to use a speech instruction corresponding to the at least one target operation event; and control the display to display a target page after a final operation event is executed, and display the guide language on the target page, the final operation event being a last operation event in a last operation sequence in the at least one operation sequence.
[0009] The above technical solution has the following advantages or benefits:
[0010] The display device can dynamically generate corresponding guide languages according to operations performed by the user through non-speech instructions, the guide languages being used to guide the user to use speech instructions corresponding to the operations, and the guide languages can be accurately matched with real operations of the user. Moreover, the display device can display the guide languages on a page displayed after a last operation of the user is performed, without the user having to search for the guide languages, and the display device can enhance the relevance between the displayed guide languages and corresponding operations of the user, and effectively improve the interaction of the user.
[0011] In some embodiments of the present application, the controller acquires at least one operation sequence when a trigger condition for generating a guide language is met, and is configured to: start a task of generating a guide language when the trigger condition for generating the guide language is met; record an operation event corresponding to the non-speech instruction input by the user after the non-speech instruction is received; determine a specified type of the operation event as the final operation event when the specified type of the operation event is detected; and acquire the at least one operation sequence from the recorded operation events, wherein each operation sequence corresponds to a preset time unit, the at least one operation event included in each operation sequence is arranged in a corresponding recording order, and a preset offset distance is arranged between adjacent two operation sequences.
[0012] The above technical solution has the following advantages or benefits:
[0013] The display device can obtain operation sequences meeting the requirements of subsequent generated guiding language in a more flexible manner by configuring time units and offset distances. The display device can limit the time length of the operation sequences by configuring the time units and limit the coverage between adjacent operation sequences by configuring the offset distances.
[0014] In some embodiments of the present application, the controller is configured to: obtain a plurality of sub-sequences included in the at least one operation sequence; calculate the frequency of each of the sub-sequences in the at least one operation sequence; and filter at least one target sub-sequence whose frequency in the at least one operation sequence is greater than or equal to a frequency threshold.
[0015] The above technical solutions have the following advantages or beneficial effects:
[0016] The display device can accurately filter the target sub-sequences whose frequency meets the frequency threshold from the plurality of sub-sequences according to the frequency of each sub-sequence in the operation sequences, so that the guiding language generated based on the target sub-sequences can accurately guide the user to use the voice instructions corresponding to the series of operations performed.
[0017] In some embodiments of the present application, the controller is configured to: obtain semantic features corresponding to each of the operation events in the at least one target sub-sequence; generate semantic instructions corresponding to each of the operation events according to the semantic features corresponding to each of the operation events in the at least one target sub-sequence; calculate semantic correlations between the semantic instructions corresponding to each of the operation events; and filter at least one target operation event according to the semantic correlations between the semantic instructions.
[0018] The above technical solutions have the following advantages or beneficial effects:
[0019] The display device can accurately calculate the semantic correlations between the semantic instructions corresponding to each of the operation events according to the semantic features corresponding to each of the operation events in the target sub-sequence, and filter the operation events corresponding to the semantic instructions whose semantic correlations meet the correlation standard, i.e., the target operation events. These target operation events can reflect the real intention of the series of operations performed by the user, and the guiding language generated based on the target operation events can accurately guide the user to use the voice instructions corresponding to the series of operations performed.
[0020] In some embodiments of the present application, the controller generates the semantic instruction corresponding to each operation event in the at least one target sub-sequence according to the semantic feature corresponding to each operation event, and is specifically configured to: acquire a mapping table, the mapping table including a mapping relationship between a preset control and a guide word; determine a target guide word corresponding to each operation event in the mapping table according to the control name in the semantic feature corresponding to each operation event; determine a user intention according to the semantic feature; and process the target guide word according to the user intention to obtain the semantic instruction corresponding to each operation event.
[0021] The above technical solution has the following advantages or beneficial effects:
[0022] The display device can improve the standardization of the generated semantic instruction and effectively improve the efficiency of generating the semantic instruction by processing the guide word based on the pre-stored guide word.
[0023] In some embodiments of the present application, the controller generates the guide speech corresponding to the at least one target operation event according to the semantic instruction corresponding to each target operation event, and is configured to: generate at least one candidate path according to the at least one target operation event, each candidate path including at least one target operation event; determine a target path from the at least one candidate path, the target path including the least number of target operation events; acquire the semantic instruction corresponding to each target operation event in the target path; and generate the guide speech according to the semantic instruction corresponding to each target operation event in the target path.
[0024] The above technical solution has the following advantages or beneficial effects:
[0025] The display device can effectively reduce the number of semantic instructions used to generate the guide speech by screening out the target path including the least number of target operation events and generating the guide speech according to the semantic instruction corresponding to each target operation event in the target path, thereby effectively reducing the generated guide speech and effectively improving the generation efficiency of the guide speech.
[0026] In some embodiments of the present application, the controller controls the display to display a target page after executing the final operation event and display the guide speech on the target page, and is configured to: detect the state of a far-field voice function; if the state of the far-field voice function is an open state, display the guide speech on the target page; and if the state of the far-field voice function is a closed state, do not display the guide speech on the target page.
[0027] The above technical solution has the following advantages or beneficial effects:
[0028] The display device can determine whether to display the guide speech according to the state of the far-field voice function. In this way, the display device can display the guide speech when the state of the far-field voice function is an open state, i.e., the far-field voice function is available, so that the display device can give a response based on the far-field voice function when the user learns the corresponding voice instruction following the guide speech. When the state of the far-field voice function is a closed state, i.e., the far-field voice function is not available, the guide speech is not displayed, so as to avoid the problem that the display device cannot give a response based on the far-field voice function when the user learns the corresponding voice instruction following the guide speech.
[0029] In some embodiments of the present application, if the state of the far-field voice function is an open state, the controller is configured to display the guide speech on the target page, specifically configured to: if the state of the far-field voice function is an open state, obtain a first time length from a time corresponding to a preset node to a current time; if the first time length is less than or equal to a time length threshold, display the guide speech on the target page; if the first time length is greater than the time length threshold, do not display the guide speech on the target page.
[0030] The above technical solution has the following advantages or beneficial effects:
[0031] The display device can no longer display the dynamically generated guide speech to the user after a certain period of time, so as to avoid affecting the normal use of the display device by the user due to the display of the dynamically generated guide speech to the user for a long time.
[0032] In some embodiments of the present application, the controller generates the guide speech according to the semantic instructions corresponding to each target operation event in the target path, and is further configured to: determine a target control corresponding to the final operation event; create a mapping relationship between the target control and the guide speech; and update the mapping relationship to the mapping table.
[0033] The above technical solution has the following advantages or beneficial effects:
[0034] The display device can expand the mapping table by creating a mapping relationship between the guide speech and the target control, so that the guide speech can be directly displayed when the target control is triggered again by the user. Moreover, a more abundant guide speech basis can be provided when generating guide speeches of other operation events.
[0035] In a second aspect, the present application also provides a method for guiding a user to use a voice instruction, the method comprising: obtaining at least one operation sequence when a trigger condition for generating a guide word is met, each operation sequence comprising at least one operation event, the operation event corresponding to a non-voice instruction of the user; screening at least one target sub-sequence from a plurality of sub-sequences included in the at least one operation sequence, the frequency of occurrence of the at least one target sub-sequence in the at least one operation sequence being greater than or equal to a frequency threshold, wherein the sub-sequence comprises at least one operation event; screening at least one target operation event according to the correlation between the operation events in the at least one target sub-sequence; generating a guide word corresponding to the at least one target operation event according to a semantic instruction corresponding to each target operation event, the guide word being used to guide the user to use a voice instruction corresponding to the at least one target operation event; displaying a target page after a final operation event is executed, the final operation event being the last operation event in the last operation sequence in the at least one operation sequence, and displaying the guide word on the target page.
[0036] The above technical solution has the following advantages or beneficial effects:
[0037] The display device can dynamically generate corresponding guide words according to the operations performed by the user through non-voice instructions, the guide words being used to guide the user to use voice instructions corresponding to the operations, and the guide words can be accurately matched with the real operations of the user. Moreover, the display device can display the guide words on a page displayed after the last operation of the user is executed, without the user having to search for the guide words, and the display device can enhance the relevance between the displayed guide words and the corresponding user operations, and effectively improve the interaction of the user.
[0038] In the technical solution provided in the embodiments of the present application, the display device obtains at least one operation sequence when a trigger condition for generating a guide word is met, each operation sequence comprising at least one operation event, the operation event corresponding to a non-voice instruction of the user. The display device screens at least one target sub-sequence from a plurality of sub-sequences included in the at least one operation sequence, the frequency of occurrence of the at least one target sub-sequence in the at least one operation sequence being greater than or equal to a frequency threshold, and screens at least one target operation event according to the correlation between the operation events in the at least one target sub-sequence. The display device generates a guide word corresponding to the at least one target operation event according to a semantic instruction corresponding to each target operation event, the guide word being accurately matched with the real operations of the user. The display device displays a target page after a final operation event is executed, and displays the guide word on the target page, without the user having to search for the guide words, and the guide words displayed in real time have strong relevance with the corresponding user operations, and can effectively improve the interaction of the user. BRIEF DESCRIPTION OF DRAWINGS
[0039] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without any creative effort.
[0040] Figure 1 A schematic diagram of an operation scenario between the display device 200 and the control device 100 in the embodiment of the present application;
[0041] Figure 2 A hardware configuration block diagram of the control device 100 in the embodiment of the present application;
[0042] Figure 3 A hardware configuration block diagram of the display device 200 in the embodiment of the present application;
[0043] Figure 4 An operating system configuration diagram of the display device 200 in the embodiment of the present application;
[0044] Figure 5 A flowchart of guiding the user to use the voice instruction by the display device 200 provided in the embodiment of the present application;
[0045] Figure 6 A flowchart of obtaining the operation sequence by the display device 200 provided in the embodiment of the present application;
[0046] Figure 7 A timing diagram of guiding the user to use the voice instruction by the display device 200 provided in the embodiment of the present application;
[0047] Figure 8 A flowchart of screening the target operation event by the display device 200 provided in the embodiment of the present application;
[0048] Figure 9 A flowchart of generating the guide language by the display device 200 provided in the embodiment of the present application;
[0049] Figure 10 A schematic diagram of the candidate path provided in the embodiment of the present application;
[0050] Figure 11 A flowchart of displaying the guide language by the display device 200 provided in the embodiment of the present application;
[0051] Figure 12 Another flowchart of displaying the guide language by the display device 200 provided in the embodiment of the present application;
[0052] Figure 13 A user operation flow under the scenario 1 provided in the embodiment of the present application;
[0053] Figure 14 An interface schematic diagram provided by an embodiment of the present application for displaying a guide sentence on a target page;
[0054] Figure 15 A user operation flow provided by an embodiment of the present application under scenario 2;
[0055] Figure 16 A user operation flow provided by an embodiment of the present application under scenario 3;
[0056] Figure 17 A user operation flow provided by an embodiment of the present application under scenario 4;
[0057] Figure 18 A user operation flow provided by an embodiment of the present application under scenario 5. DETAILED DESCRIPTION
[0058] The embodiments will be described in detail with reference to the drawings, wherein the same reference numerals refer to the same or similar elements throughout. The following description is not meant to represent all embodiments in accordance with the present application. It is merely an example of a system and method in accordance with some aspects of the present application, as detailed in the claims.
[0059] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the following described embodiments, and is not intended to limit the embodiments of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and general meanings.
[0060] The terms "first", "second", "third", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar or like objects or entities, and do not necessarily mean a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchanged under appropriate circumstances.
[0061] The terms "include" and "have" and any variations thereof are intended to cover but not exclusive inclusion, for example, a product or device including a series of components does not necessarily limit to all components clearly listed, but can include other components not clearly listed or inherent to these products or devices.
[0062] In the embodiments of the present application, the display device generally refers to a device with picture display and data processing capabilities. For example, the display device includes but is not limited to smart television, mobile terminal, computer, monitor, advertising screen, wearable device, virtual reality device, augmented reality device, etc.
[0063] Figure 1The schematic diagram of the operation scene between the display device and the control device is provided for some embodiments of the present application. As shown in Figure 1 The user can operate the display device 200 through touch operation, voice, mobile terminal 300 and control device 100, as shown in the
[0064] As shown in Figure 1 It is also shown in the that the display device 200 also communicates data with the server 400 through various communication methods. The display device 200 can be allowed to be connected through LAN, WLAN and other networks.
[0065] The display device 200 can provide a broadcast receiving television function, and can additionally provide an intelligent network television function supporting computer functions, including but not limited to network television, smart television, IPTV, etc.
[0066] Figure 2 The schematic diagram of the operation scene between the display device and the control device is provided for some embodiments of the present application. As shown in Figure 1 The hardware configuration block diagram of the control device 100 is shown in the
[0067] As shown in Figure 2 The control device 100 can include a controller 110, a communication interface 120, a user input / output interface 130, a memory 140, and a power supply 150.
[0068] The controller 110 can include a processor 111, a RAM 112, a ROM 113, and a communication bus.
[0069] The communication interface 120 can include at least one of a WiFi chip 121, a Bluetooth module 122, an NFC module 123, and other near field communication modules.
[0070] The user input / output interface 130, wherein the input interface can include at least one of a microphone 131, a touchpad 132, a sensor 133, a key 134, and other input interfaces.
[0071] In some embodiments, the control device 100 can include at least one of the communication interface 120 and the input / output interface 130.
[0072] The memory 140 is used to store various running programs, data and applications for driving and controlling the control device 100 under the control of the controller 110. The memory 140 can store various control signal instructions input by the user.
[0073] The power supply 150 is configured to provide operating power support for each element of the control device 100 under the control of the controller 110.
[0074] Figure 3 The hardware configuration block diagram of the display device 200 is shown in some embodiments of the present application. Figure 1 The hardware configuration block diagram of the display device 200 is shown in some embodiments of the present application.
[0075] In some embodiments, the display device 200 can include at least one of a tuner and demodulator 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a user input interface 280, a memory, a power supply.
[0076] In some embodiments, the communication device 220 is a component for communicating with an external device or a server 400 according to various communication protocol types. The display device 200 can be provided with multiple communication devices 220 according to different supported communication manners. The communication device 220 can make the display device 200 communicatively connected with the external device or the server 400 through wireless or wired connection.
[0077] In some embodiments, the detector 230 is configured to collect signals of external environment or external interaction. For example, the detector 230 includes a light receiver configured to collect ambient light intensity; or the detector 230 includes an image collector such as a camera, which can be configured to collect external environment scene, user attribute or user interaction gesture; or the detector 230 includes a sound collector such as a microphone, which is configured to receive external sound.
[0078] In some embodiments, the device interface 240 is configured to access external devices.
[0079] In some embodiments, the controller 250 is configured to control the overall operation of the display device 200. The controller 250 can include at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), a power supply processor, a first interface to an n-th interface for input / output, and the controller 250 controls the operation of the display device 200 and responds to the user's operation through various software control programs stored in the memory.
[0080] In some embodiments, the controller 250 and the tuner and demodulator 210 can be located in different split devices, i.e., the tuner and demodulator 210 can also be located in an external device of the main body device where the controller 250 is located, such as an external set-top box.
[0081] In some embodiments, the display 260 is used to receive and display image signals output from the controller 250. The display 260 may include display function components for presenting images and driving components for driving image display.
[0082] In some embodiments, a user can input user commands on a graphical user interface (GUI) displayed on a display 260, and a user input interface 280 can receive user commands through the GUI.
[0083] In some embodiments, the audio output device 270 may be a built-in speaker of the display device 200 or an external audio output device connected to the display device 200.
[0084] In some embodiments, the user input interface 280 can be used to receive instructions from user input.
[0085] In some embodiments, the display device 200 may run an operating system to perform user interaction. An operating system is a computer program used to manage and control the hardware and software resources of the display device 200.
[0086] The operating system can be a native operating system based on a specific operating platform, a third-party operating system that is deeply customized based on a specific operating platform, or an independent operating system specifically developed for display devices 200.
[0087] An operating system can be divided into different modules or levels based on the functions it implements, for example... Figure 4 As shown, in some embodiments, the system can be divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the System Runtime Library layer, and the Kernel Layer.
[0088] In some embodiments, the application layer provides services and interfaces for applications, enabling the display device 200 to run applications and interact with users based on the applications. For example, the application layer may include a voice assistant, which provides voice interaction functionality, allowing users to interact with the display device via voice.
[0089] The framework layer can provide an application programming interface (API) and a programming framework for applications. The application framework layer includes some pre-defined functions. The application framework layer is equivalent to a processing center that decides which application in the application layer to act. The application can access the resources in the system and obtain the services of the system through the API interface during execution.
[0090] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library included in the system runtime library layer, such as the C / C++ instruction library, to implement the functions implemented by the framework layer.
[0091] In some embodiments, the kernel layer is a functional layer between the hardware and the software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, memory management, etc.
[0092] It should be noted that the above examples are only a simple division of the functions of the operating system, and do not constitute a limitation on the specific operating system form of the display device 200 in the embodiments of the present application. According to the function of the display device, the type of the operating system, and other factors, the number and specific type of the layers included in the operating system can be in other forms.
[0093] The display device 200 can be configured with a far-field voice function, such as a "voice assistant". Based on the far-field voice function, the display device 200 can respond to the voice instructions of the user to perform corresponding tasks, so that the user can interact with the display device 200 more conveniently through the voice. The standard degree of the voice instruction input by the user is one of the important factors affecting the accuracy of triggering the display device 200 to perform tasks and determining the tasks to be performed by the display device 200. Therefore, in order to improve the standard degree of the voice instruction input by the user, the display device 200 can be configured with a guidance function. Based on the guidance function, the display device 200 can show a guidance language to the user to guide the user to learn the standard voice instruction through the guidance language.
[0094] The display device 200 usually provides the user with a guide sentence related to the far-field voice function during the operation guide process of the display device 200 after the first boot. Alternatively, the display device 200 displays the guide sentence related to the far-field voice function on the home page. Alternatively, the display device 200 displays the guide sentence related to the voice instruction corresponding to the specified control when the user triggers the specified control, which is a page control of the display device 200 that is pre-specified before the display device 200 is shipped. It can be seen that the scenario of triggering the display device 200 to guide the user to use the voice instruction is relatively fixed, and the guide sentence displayed is also relatively fixed, which cannot well adapt to the scenario in which the user actually uses the display device 200.
[0095] To solve the above problems, the display device 200 provided in the present application is configured to dynamically generate a matching guide sentence according to the real operation performed by the user, so as to improve the matching degree of the guide sentence and the scenario in which the user actually uses the display device 200. Moreover, the guide sentence can be immediately displayed to the user without the user having to find the guide sentence by himself / herself, and the guide sentence displayed immediately has a strong correlation with the corresponding operation of the user, which can effectively improve the interaction between the user and the display device 200.
[0096] The display device 200 provided in the present application is schematically described in combination with the following embodiments:
[0097] Figure 5 The flowchart of guiding the user to use the voice instruction by the display device provided in the embodiments of the present application is specifically as follows:
[0098] In step S501, at least one operation sequence is obtained when a trigger condition for generating a guide sentence is met.
[0099] The display device 200 can be configured with a guide sentence generation function. If the state of the guide sentence generation function is an open state, the display device 200 starts a task of generating a guide sentence; if the state of the guide sentence generation function is a closed state, the display device 200 does not start the task of generating a guide sentence. In some embodiments, the display device 200 can be configured to set the guide sentence generation function to the open state by default.
[0100] In the case where the state of the guide sentence generation function is the open state, the display device 200 listens to whether the trigger condition for generating a guide sentence is met. For example, the trigger condition can include one or a combination of listening to a boot instruction, the execution of the previous task of generating a guide sentence being ended, detecting that the current is a specific page (for example, the home page of the display device 200), listening to a specific instruction (for example, an instruction to return to the home page) in the boot state. Of course, the trigger condition is not limited to the specific examples disclosed above.
[0101] If the display device 200 does not listen to the trigger condition meeting the generation of the guide language, the display device 200 does not start the task of generating the guide language. If the display device 200 listens to the trigger condition meeting the generation of the guide language, the display device 200 starts the task of generating the guide language.
[0102] In the case that the display device 200 listens to the trigger condition meeting the generation of the guide language and starts the task of generating the guide language, the display device 200 acquires at least one operation sequence. Each operation sequence includes at least one operation event (denoted as Do in the embodiments of the present application), and the operation event corresponds to the non-speech instruction of the user, that is, the operation event corresponds to the operation performed by the user in a non-speech manner. For example, the operation performed by the user can include the operation of controlling the display device 200 through the control device 100, such as controlling the keys (such as direction keys, confirmation keys, back keys, etc.) on the control device 100 (such as moving the focus on the user interface, selecting the page control in the user interface, adjusting the system parameters, etc.) to control the display device 200.
[0103] In some examples, the operation sequence D can be represented as D={Do1, Do2, … Do n}, the operation sequence includes n operation events, where n is an integer greater than or equal to 1.
[0104] In some examples, the at least one operation sequence acquired by the display device 200, also referred to as the operation sequence set DT, can be represented as DT={D1, D2, … D m}, the operation sequence set includes m operation sequences, where m is an integer greater than or equal to 1.
[0105] Figure 6 The flowchart of the display device provided by the embodiments of the present application for acquiring the operation sequence is as follows:
[0106] Step S601: Start the task of generating the guide language when the trigger condition meeting the generation of the guide language is met.
[0107] Figure 7 The timing diagram of the display device provided by the embodiments of the present application for guiding the user to use the speech instruction.
[0108] In combination with Figure 7 , the display device 200 can be configured with a behavior pattern mining module, and the behavior pattern mining module is controlled to start the task of generating the guide language when the trigger condition meeting the generation of the guide language is listened to.
[0109] Step S602: Record the operation event corresponding to the non-speech instruction after receiving the non-speech instruction input by the user.
[0110] In combination Figure 7 In the case of starting the task of generating the guide speech, the display device 200 can record operation events corresponding to non-speech instructions input by the user, such as non-speech instructions input by the user through the control device 100 or the mobile terminal 300, through the behavior pattern mining module.
[0111] An operation event can include elements related to a corresponding operation. In some embodiments, an operation event can include three elements, such as an event (denoted as o in the embodiments of the present application), a device state before and after the display device 200 responds to the corresponding operation (denoted as s in the embodiments of the present application).
[0112] In one example, any operation event Do i includes three elements (s i , o i , s i+1 ), where Do i represents that the operation event is the i-th operation event in the operation sequence, o i represents the event element included in the i-th operation event, s i represents the device state before the display device 200 performs the i-th operation event, and s i+1 represents the device state after the display device 200 performs the i-th operation event.
[0113] The event can include an event type and an attribute (attr) of a control that interacts with the user.
[0114] The event type corresponds to an operation performed by the user, such as clicking, exiting, etc.
[0115] The attribute of the control can include an identification (ID) of the control, a text corresponding to the control, a content description, a hint, and an image file name (src) corresponding to the control. The identification of the control can represent a function corresponding to the control; the text corresponding to the control is the name of the control displayed on the corresponding page; the content description is a note for the control, which is used to explain the function of the control; the hint is information displayed in the page when the prompt function of the corresponding control is triggered; and the image file name corresponds to a page control that displays image content, which is the file name of the image content to be displayed by the page control.
[0116] The device state can include an activity component name of a page and a UI interface hierarchy result. The activity component name can represent the main function provided by the corresponding page. The UI interface hierarchy result represents the interface hierarchy to which the activity component belongs.
[0117] Step S603: When detecting the operation event of the specified type, determine the operation event of the specified type as the final operation event.
[0118] A user can achieve an operation target by performing a series of operations, for example, the user controls the display device 200 to play the media asset A by performing a series of operations, and playing the media asset A is the operation target of the user. The display device 200 needs to generate a corresponding guide sentence for each operation target corresponding to a series of operations, so that the guide sentence can accurately match the operations performed by the user when achieving different operation targets, thereby accurately guiding the user to use the corresponding voice instruction based on the operation purpose of the user. Based on this, the display device 200 needs to accurately obtain the operation event corresponding to the same operation target.
[0119] In some embodiments, at least one type of operation event can be pre-set, which has a high probability of corresponding to an operation target. The at least one type of operation event can also be referred to as an operation event of a specified type, wherein the operation event of the specified type can include switching (such as opening, returning, etc.) to a specified page, setting system parameters, etc.
[0120] In combination Figure 7 When the behavior pattern mining module obtains the operation event corresponding to the same operation target, it can determine whether the type of the obtained operation event is the specified type according to the matching result of the type of the obtained operation event and the specified type. For example, the behavior pattern mining module can obtain the device state after performing the operation event, such as the active component name after performing the operation event, match the active component name with the pre-set active component name, and if the matching result is matched, it can be determined that the operation event of the specified type is detected; if the matching result is not matched, it can be determined that the operation event of the specified type is not detected.
[0121] When the behavior pattern mining module detects the operation event of the specified type, it can be determined that after performing the operation event of the specified type, the operation target of the user can be achieved, that is, the operation event of the specified type is the last operation event for achieving the operation target, which can also be referred to as the final operation event corresponding to the operation target.
[0122] In some embodiments, the behavior pattern mining module can mark the detected final operation event, so that in subsequent processing, the operation events corresponding to different operation targets can be accurately distinguished based on these marks.
[0123] Step S604: obtaining at least one operation sequence from the recorded operation events, wherein each operation sequence corresponds to a preset time unit, and each operation sequence includes at least one operation event in the recorded operation events, which are arranged in a corresponding recording order, and an offset distance is preset between two adjacent operation sequences.
[0124] In step S604, the recorded operation events refer to operation events corresponding to the same operation target. The first operation event in the recorded operation events can be the first operation event recorded after the task of generating the guiding sentence is started, or if there are operation events corresponding to a previous operation target, the first operation event in the recorded operation events can be the first operation event recorded after the last operation event corresponding to the previous operation target.
[0125] The time window mechanism can be used to obtain at least one operation sequence from the recorded operation events. In step S602, the behavior pattern mining module obtains the recording order of each operation event when recording the operation events. Each operation sequence includes at least one recorded operation event, which are continuous recorded operation events and are sorted in the operation sequence according to the corresponding recording order, so as to ensure the correctness of the execution order of the corresponding operation. The last operation event in the last operation sequence is the last operation event corresponding to the current operation target.
[0126] Each operation sequence corresponds to a preset time unit, and the length of the preset time unit can be adjusted by the user, so as to limit the time length of the operation sequence and control the number of operation events included in the operation sequence.
[0127] An offset distance is preset between two adjacent operation sequences. If the offset distance is small, the coverage between the two adjacent operation sequences is high, that is, the two adjacent operation sequences include more same operation events. If the offset distance is large, the coverage between the two adjacent operation sequences is low, that is, the two adjacent operation sequences include fewer same operation events, or even no same operation events. Therefore, the behavior pattern mining module can limit the coverage between adjacent operation sequences by flexibly configuring the offset distance.
[0128] In some embodiments, the offset distance can be a time distance, that is, a length. That is, in the two adjacent operation sequences, the start time of the time unit corresponding to the previous operation sequence and the start time of the time unit corresponding to the next operation sequence are separated by a preset time distance.
[0129] In some embodiments, the offset distance can be a quantity distance, i.e., a quantity. That is, in two adjacent operation sequences, a first operation event included in a previous operation sequence and a first operation event included in a next operation sequence are separated by a preset quantity distance.
[0130] Step S502: From the plurality of sub-sequences included in the at least one operation sequence, at least one target sub-sequence whose frequency of occurrence in the at least one operation sequence is greater than or equal to a frequency threshold is screened out.
[0131] Each operation sequence includes at least one sub-sequence a, which can also be referred to as a behavior sequence, and each sub-sequence includes at least one operation event belonging to one operation sequence. The at least one operation event included in the sub-sequence can be continuously recorded operation events or discontinuously recorded operation events.
[0132] In an example, the sub-sequence can be represented as a = {Do i , Do i+1 , …, Do k}, where i ≥ 1 and k ≤ n.
[0133] The frequency of occurrence of the sub-sequence in all operation sequences can represent the frequency of the user performing the operations corresponding to the operation events in the sub-sequence. Therefore, the higher the frequency of the sub-sequence, the higher the frequency of the user performing the operations corresponding to the operation events in the sub-sequence; the lower the frequency of the sub-sequence, the lower the frequency of the user performing the operations corresponding to the operation events in the sub-sequence.
[0134] In some embodiments, the frequency of occurrence can be represented by the support of the sub-sequence, and the support of the sub-sequence refers to the total number of the sub-sequence included in the set of operation sequences.
[0135] In an example, the support of the sub-sequence can be represented as: where |DT| represents the number of operation sequences in DT.
[0136] The frequency of the user performing the operations can reflect the operation characteristics, preferences, and other features of the user, which can also be referred to as the behavior pattern of the user. It can be understood that the higher the frequency of the user performing the operations, the more consistent with the behavior pattern of the user; the lower the frequency of the user performing the operations, the more deviated from the behavior pattern of the user. The sub-sequences that are more consistent with the behavior pattern of the user can be screened out by setting a frequency threshold. The sub-sequences whose frequency of occurrence is greater than or equal to the frequency threshold in each sub-sequence can be determined as target sub-sequences, which can also be referred to as frequent sub-sequences. The set of target sub-sequences can be referred to as a frequent sub-sequence set, and the operations corresponding to the operation events in the target sub-sequences are consistent with the behavior pattern of the user.
[0137] In one example, the frequent subsequence set FS can be represented as: FS = {a | support(a) ≥ min sup}, where min sup represents the frequency threshold.
[0138] The frequent subsequence set can be processed to ensure that there is no inclusion relationship between the target subsequence. The processed frequent subsequence set can be referred to as a behavior pattern set.
[0139] In one example, the behavior pattern set PS can be represented as:
[0140] The guide language generated based on the target subsequence in the behavior pattern set can effectively improve the matching degree of the guide language and the behavior pattern of the user.
[0141] Step S503: According to the correlation between each operation event in at least one target subsequence, at least one target operation event is screened out.
[0142] Among the multiple operations performed by the user, there can be operations with low correlation with the operation target, which will affect the accuracy of the guide language generated subsequently. In order to improve the accuracy of the generated guide language, the display device 200 can screen out target operation events with strong correlation with the operation target (i.e. the final operation event) from each target subsequence, so as to generate guide language based on these target operation events subsequently.
[0143] Figure 8 The flowchart of the display device provided by the embodiment of the present application for screening target operation events is as follows:
[0144] Step S801: Obtain the semantic features corresponding to each operation event in at least one target subsequence.
[0145] In combination with Figure 7 The display device 200 can be configured with a semantic event matching module. Through the semantic event matching module, target operation events with strong correlation with the operation target (i.e. the final operation event) can be screened out from each target subsequence.
[0146] The behavior pattern mining module can pass the screened target subsequence to the semantic event matching module, and the semantic event matching module can obtain the semantic features corresponding to each operation event in the target subsequence.
[0147] The semantic features corresponding to the operation event can include an activity component name included in the operation event and an attribute of a control in an event element included in the operation event, that is, the semantic features corresponding to the operation event are associated with a page displayed by the display device 200 before and after responding to the user operation, a control interacted by the user when performing the operation, and the like.
[0148] Step S802: generating a semantic instruction corresponding to each operation event according to the semantic features corresponding to each operation event in the at least one target subsequence.
[0149] In combination Figure 7 The semantic event matching module can determine a user intention corresponding to the operation event according to the semantic features, and then determine the semantic instruction according to the user intention. In some examples, the semantic event matching module can determine the user intention according to the semantic features based on a natural language processing (NLP) algorithm.
[0150] In some embodiments, the memory 140 of the display device 200 can store a mapping table (Map) including a mapping relationship between the names of the page controls and the guide words. The mapping relationship includes a first mapping relationship between the names of the page controls pre-configured by the display device 200 and the corresponding guide words. The mapping relationship can also include a second mapping relationship between the guide words obtained by performing the task of generating the guide language and the corresponding page controls.
[0151] The semantic event matching module can determine the corresponding target guide word from the mapping table according to the name of the control in the semantic features. The semantic event matching module can process the target guide word according to the user intention to obtain a semantic instruction that can meet the user intention. For example, some words in the target guide word can be replaced with words that meet the user intention, or words that meet the user intention can be added to the target guide word.
[0152] In this way, based on the pre-stored guide word, the generated semantic instruction can not only improve the standardization of the generated semantic instruction, but also effectively improve the efficiency of generating the semantic instruction.
[0153] Step S803: calculating semantic correlations between the semantic instructions corresponding to the respective operation events.
[0154] Each operation event other than the final operation event can also be referred to as a non-final operation event, and the non-final operation event corresponds to an operation process for achieving the operation target by the user.
[0155] The semantic event matching module can determine the correlation between each operation and the operation target in the operation process for achieving the operation target by the user by calculating the semantic correlations between the semantic instructions corresponding to each non-final operation event and the semantic instruction corresponding to the final operation event.
[0156] The semantic event matching module can determine the correlation between operations in the operation process of the user achieving the operation target by calculating the semantic correlation between the semantic instructions corresponding to each non-final operation event.
[0157] In combination Figure 7 The semantic correlation between the semantic instructions can reflect the correlation between the corresponding operation events. In some examples, the semantic event matching module can calculate the semantic distance between the semantic instructions corresponding to the corresponding operation events according to the semantic features, and the semantic distance is used to represent the semantic correlation between the semantic instructions. Wherein, the closer the semantic distance, the higher the semantic correlation between the semantic instructions; the farther the semantic distance, the lower the semantic correlation between the semantic instructions.
[0158] Step S804: filtering out at least one target operation event according to the semantic correlation between the semantic instructions.
[0159] A correlation threshold is set in advance to filter the target operation event through the correlation threshold. By adjusting the correlation threshold, the number and quality of the filtered target operation events can be controlled. Wherein, the higher the correlation threshold, the relatively less the number of the filtered target operation events, and the higher the quality; the lower the correlation threshold, the relatively more the number of the filtered target operation events, and the relatively lower the quality.
[0160] In combination Figure 7 The semantic event matching module can filter out the target operation event according to the correlation threshold and the semantic correlation. Wherein, the operation event corresponding to the semantic instruction with the semantic correlation greater than or equal to the correlation threshold is determined as the target operation event.
[0161] Step S504: generating at least one guiding sentence corresponding to each target operation event according to the semantic instruction corresponding to each target operation event.
[0162] The guiding sentence is used to guide the user to use the voice instruction corresponding to the at least one target operation event.
[0163] In some embodiments, the display device 200 can process the at least one target operation event before generating the guiding sentence, so as to simplify the number of target operation events and simplify the corresponding guiding sentence.
[0164] In one example, the display device 200 can determine meaningless operation events in the at least one target operation event, and eliminate the meaningless operation events to reduce the number of target operation events. In some embodiments, the display device 200 can splice semantic instructions corresponding to the at least one target operation event in the order of the record corresponding to the at least one target operation event to obtain the guidance language.
[0165] In some embodiments, the display device 200 can process the at least one target operation event, and splice semantic instructions corresponding to the processed target operation event in the order of the record corresponding to the processed target operation event to obtain the guidance language. In this way, the guidance language can be effectively simplified, and the readability of the guidance language can be improved.
[0166] Figure 9 The flowchart of generating the guidance language by the display device provided in the embodiments of the present application is as follows:
[0167] Step S901: generating at least one candidate path according to at least one target operation event, each candidate path including at least one target operation event.
[0168] The target operation events included in each candidate path are not completely the same, and thus the guidance language corresponding to each candidate path is different. Each candidate path includes at least one target operation event arranged in the order of record.
[0169] In combination with Figure 7 The display device 200 can be configured with a guidance language generation module, which can generate candidate paths according to the target operation events filtered by the semantic event matching module.
[0170] Executing the operation events included in each candidate path can make the display device 200 switch from the same initial device state to the same final device state. The initial device state is the device state of the display device 200 before the first operation event in the task of generating the guidance language, and the final device state is the device state of the display device 200 after the final operation event.
[0171] In some embodiments, the guidance language generation module can generate candidate paths through corresponding path models, wherein the initial device state and the final device state can be set for the path model, and each target operation event can be used as input data to output the candidate path through the path model.
[0172] Figure 10An example of the schematic diagram of the candidate paths provided by the embodiments of the present application is shown, taking the target operation events as O0-O7, the initial device state as S0, and the final device state as S6 as an example. According to the correlation between the target operation events, three candidate paths can be generated, which are O0→O1→O2→O3→O4, O0→O5→O4, and O6→O7.
[0173] Step S902: determining a target path from the at least one candidate path, the target path containing the least number of target operation events.
[0174] The number of target operation events contained in the candidate path can represent the degree of simplification of the guidance language corresponding to the candidate path. The fewer the number of target operation events contained in the candidate path, the higher the degree of simplification of the corresponding guidance language. In order to improve the readability of the guidance language, the guidance language with the highest degree of simplification is selected for the user to display, that is, the candidate path containing the least number of target operation events is determined as the target path, and the guidance language is generated based on the target path.
[0175] In combination with Figure 10 , the candidate path O6→O7 contains the least number of target operation events, and can be determined as the target path.
[0176] Step S903: obtaining semantic instructions corresponding to each target operation event in the target path.
[0177] The semantic instructions corresponding to each target operation event in the target path are used to generate the guidance language.
[0178] Step S904: generating the guidance language according to the semantic instructions corresponding to each target operation event in the target path.
[0179] In some embodiments, the display device 200 can splice the corresponding semantic instructions in the order recorded in the target path to obtain the guidance language.
[0180] Step S505: displaying the target page after executing the final operation event, and displaying the guidance language on the target page.
[0181] In combination with Figure 7 , the display device 200 can be configured with a voice terminal display module, and the guidance language generated by the guidance language generation module is displayed through the voice terminal display module.
[0182] In some embodiments, the guiding sentence generation module writes the guiding sentence into the guiding word database after generating the guiding sentence, that is, updates the guiding word database. The voice terminal display module can monitor the data state of the guiding word database. If the voice terminal display module monitors that the data state of the guiding word database is updated data, the voice terminal display module reads the updated data from the guiding word database, that is, the guiding sentence. The voice terminal display module determines that the updated data is the guiding sentence to be displayed, that is, determines to display the guiding sentence.
[0183] Figure 11 The flowchart for the display device to display the guiding sentence provided by the embodiments of the present application is as follows:
[0184] Step S1101: detecting the state of the far-field voice function.
[0185] In combination with Figure 7 Before displaying the guiding sentence, the display device 200 detects the state of the far-field voice function through the voice terminal display module. The far-field voice function includes an open state and a closed state. In the open state, the far-field voice function can be normally used, that is, the user can perform voice interaction with the display device 200 based on the far-field voice function. In the closed state, the far-field voice function cannot be used, that is, the user cannot perform voice interaction with the display device 200 based on the far-field voice function.
[0186] Step S1102: if the state of the far-field voice function is the open state, displaying the guiding sentence on the target page.
[0187] When the voice terminal display module detects that the state of the far-field voice function is the open state, the voice terminal display module adds the guiding sentence to be displayed on the target page.
[0188] In this way, when the user learns the corresponding voice instruction following the guiding sentence, the display device 200 can give a response based on the far-field voice function in the open state and interact with the user.
[0189] Step S1103: if the state of the far-field voice function is the closed state, not displaying the guiding sentence on the target page.
[0190] When the voice terminal display module detects that the state of the far-field voice function is the closed state, the voice terminal display module does not display the guiding sentence on the target page.
[0191] In this way, when the user learns the corresponding voice instruction following the guiding sentence, the problem that the display device 200 cannot give a response due to the unavailability of the far-field voice function can be avoided.
[0192] Figure 12 Another flowchart for the display device to display the guiding sentence provided by the embodiments of the present application is as follows:
[0193] Step S1201: If the state of the far-field voice function is the open state, a first duration from a time corresponding to the preset node to a current time is obtained.
[0194] In some embodiments, the preset node can be a task of starting up or triggering generation of the guide speech. The current time can be a time of generating the guide speech, or in other words, a time when the task of generating the guide speech is completed.
[0195] After the preset node is detected, a timer can be started, and timing can be performed by using the timer. After the guide speech is detected or the task of generating the guide speech is detected to be completed, a duration currently counted by the timer is read, and the duration is the first duration.
[0196] The display device 200 can be provided with a duration threshold, to control a task of displaying the guide speech by using the duration threshold. In this way, after a certain time is exceeded, the dynamically generated guide speech is no longer displayed to the user, so as to avoid the dynamically generated guide speech being displayed to the user at any time for a long time, and affecting normal use of the display device 200 by the user.
[0197] Step S1202: If the first duration is less than or equal to the duration threshold, the guide speech is displayed on the target page.
[0198] In combination Figure 7 If the voice terminal display module detects that the first duration is less than or equal to the duration threshold, it is indicated that the duration of the dynamically generated guide speech displayed to the user is within a reasonable duration range, and the guide speech can be displayed to the user after the guide speech is dynamically generated. Therefore, the voice terminal display module confirms the task of displaying the guide speech, and after the task of displaying the guide speech is triggered, the guide speech is displayed on the target page.
[0199] Step S1203: If the first duration is greater than the duration threshold, the guide speech is not displayed on the target page.
[0200] In combination Figure 7 If the voice terminal display module detects that the first duration is greater than the duration threshold, it is indicated that the duration of the dynamically generated guide speech displayed to the user has exceeded the reasonable duration range, and at this time, even if the guide speech has been dynamically generated, the guide speech cannot be displayed to the user, so as to avoid affecting normal use of the display device 200 by the user. Therefore, the voice terminal display module confirms that the task of displaying the guide speech is not triggered, that is, the guide speech is not displayed on the target page.
[0201] Taking the control device 100 as a remote controller and the display device 200 as a television as an example, an example of guiding a user to use a voice instruction by the display device 200 in different scenarios is given.
[0202] In Scenario 1, the prompts generated by the TV can guide the user to control the TV to perform multiple tasks at once using multiple voice commands.
[0203] Combination Figure 13 The user operation events can include: when the TV is playing media asset A, in response to the user's command input via the back button on the remote control, exiting the playback of media asset A and displaying the TV home page; the TV responding to the user's command input via the confirmation button on the remote control based on the controls for media asset B on the TV home page, playing media asset B; the TV responding to the user's command input via the settings button on the remote control, displaying the settings menu; and the TV responding to the user's command input via the confirmation button on the remote control based on the controls for eye protection mode in the settings menu, activating eye protection mode.
[0204] The TV can generate corresponding prompts based on the above operation events. This process is described above and will not be repeated here. The generated prompts could be: "Remote control operation is too slow. You can say to me: '××, close media asset A; then play media asset B; turn on eye protection mode.'" The final operation event in the above process is turning on eye protection mode. After turning on eye protection mode, the TV can display something like: Figure 14 The target page is shown. When the television detects that the conditions for displaying the introductory text are met, it can, as shown... Figure 14 As shown, a guide is displayed on the target page. For example, a guide box 1401 is displayed at the top of the target page, and the guide is displayed in the guide box 1401. In this way, users can learn the voice commands for corresponding user operations by following the guide.
[0205] In scenario 2, the prompts generated by the TV can guide users to use voice commands to quickly start playing media resources.
[0206] Combination Figure 15 The user operation events may include: when the TV is displaying the home screen, in response to a user's command input via the OK button on the remote control based on the TV series control, displaying a TV series list; in response to a user's command input via the OK button on the remote control based on the control for TV series A in the TV series list, displaying the details page for TV series A; in response to a user's command input via the OK button on the remote control based on the control for episode 6 on the details page of TV series A, playing episode 6 of TV series A; in response to a user's command input via the right button on the remote control, fast-forwarding episode 6 of TV series A to the 3-minute mark.
[0207] The television can generate corresponding guide speech according to the operation events, and the process can refer to the foregoing, and details are not described herein. The generated guide speech can be: directly starting to play the video content, try telling me: "X, the 3rd minute of the 6th episode of TV series A". The final operation event in the operation events is fast forwarding the 6th episode of the TV series A to the 3rd minute, and the television displays the target page of the 3rd minute of the 6th episode of the TV series A after performing the fast forwarding of the 6th episode of the TV series A to the 3rd minute. When the television detects that the condition of displaying the guide speech is met, the television can display the guide speech on the target page, and the display form can refer to Figure 14 , and details are not described herein.
[0208] In scene 3, the guide speech generated by the television can guide the user to use a voice instruction to control the television to play music or sing through the television.
[0209] In combination with Figure 16 , the operation events corresponding to the user operation can include: the television displays a music search page in response to an instruction input by the user through a confirmation key of a remote controller based on a music control when displaying a television home page; the television displays a search result page of a song A according to a search word input by the user in response to an instruction input by the user through the confirmation key of the remote controller based on a search control in the music search page; the display device 200 plays the song A in response to an instruction input by the user through a determination key of the remote controller based on a play option in the search result page; and the display device 200 starts a karaoke function in response to an instruction input by the user through a down key of the remote controller.
[0210] The television can generate corresponding guide speech according to the operation events, and the process can refer to the foregoing, and details are not described herein. The generated guide speech can be: experiencing the karaoke function, and can tell me: "X, I want to sing the song A". The final operation event in the operation events is starting the karaoke function, and the television displays a song play page in a karaoke state, that is, a target page, after performing the starting of the karaoke function. When the television detects that the condition of displaying the guide speech is met, the television can display the guide speech on the target page, and the display form can refer to Figure 14 , and details are not described herein.
[0211] In scene 4, the guide speech generated by the television can guide the user to use a voice instruction to control a live program to be played.
[0212] In combination with Figure 17The corresponding operation events of the user operation can include: the television displays a set-top box live page in response to an instruction input by the user through a confirmation key of a remote controller based on a High Definition Multimedia Interface (HDMI) control while displaying a signal source page; the television displays a channel list in the set-top box live page in response to an instruction input by the user through a menu key of the remote controller; and the display device 200 switches to playing channel A in response to an instruction input by the user through a determination key of the remote controller based on a control of channel A in the channel list.
[0213] The television can generate the corresponding guide language according to the operation events, and the process can refer to the foregoing, and details are not described herein again. The generated guide language can be: "watching a live program, next time, say to me: '××, watch channel A'". The final operation event in the operation events is switching to playing channel A, and the television displays a target page of playing channel A after performing the switching to playing channel A. The television can display the guide language on the target page when a condition of displaying the guide language is met, and the display form can refer to the foregoing, and details are not described herein again. Figure 14
[0214] In scenario 5, the guide language generated by the television can guide the user to use a voice instruction to control content search in a third-party application.
[0215] In combination with Figure 18 The corresponding operation events of the user operation can include: the television displays a homepage of application A in response to an instruction input by the user through a confirmation key of a remote controller based on a control of application A while displaying a "my application" page; the television displays a detail page of a TV series B in response to an instruction of searching for the TV series B by the user in the homepage of application A through the remote controller; and the display device 200 plays the 12th episode of the TV series B in response to an instruction input by the user through a determination key of the remote controller based on a control of the 12th episode in the detail page of the TV series B.
[0216] The television can generate the corresponding guide language according to the operation events, and the process can refer to the foregoing, and details are not described herein again. The generated guide language can be: "playing a video of application A, can say to me: '××, application A plays the 12th episode of the TV series B'". The final operation event in the operation events is playing the 12th episode of the TV series B, and the television displays a target page of the 12th episode of the TV series B after performing the playing of the 12th episode of the TV series B. The television can display the guide language on the target page when a condition of displaying the guide language is met, and the display form can refer to the foregoing, and details are not described herein again. Figure 14
[0217] In some embodiments, a method for guiding a user to use a voice instruction is also provided, which can be applied to the display device 200 as shown above. The method comprises:
[0218] When the trigger condition for generating the guide language is met, at least one operation sequence is acquired, each of the operation sequences including at least one operation event, the operation event corresponding to a non-voice instruction of the user; from a plurality of sub-sequences included in the at least one operation sequence, at least one target sub-sequence is screened out, the frequency of occurrence of the target sub-sequence in the at least one operation sequence being greater than or equal to a frequency threshold, the sub-sequence including at least one operation event; at least one target operation event is screened out according to the correlation between the operation events in the at least one target sub-sequence; a guide language corresponding to each of the target operation events is generated according to a semantic instruction corresponding to each of the target operation events, the guide language being used to guide the user to use a voice instruction corresponding to the at least one target operation event; a target page after execution of a final operation event is displayed, the final operation event being a last operation event in a last operation sequence in the at least one operation sequence, and the guide language is displayed on the target page.
[0219] The display device can dynamically generate corresponding guide language according to operations performed by the user through non-voice instructions, the guide language being used to guide the user to use voice instructions corresponding to the operations, and the guide language can be accurately matched with real operations of the user. Moreover, the display device can display the guide language on a page displayed after execution of a last operation of the user, without the user having to search for the guide language, and the display device can enhance the relevance between the displayed guide language and corresponding operations of the user by displaying the guide language in time, and can effectively improve the interaction of the user.
[0220] The above has been described in conjunction with specific embodiments for the convenience of explanation. However, the above description in some embodiments is not intended to be exhaustive or to limit the embodiments to the specific form disclosed. Various modifications and variations can be derived from the above teachings. The selection and description of the above embodiments are for better explanation of the content of the disclosure, so that those skilled in the art can better use the embodiments.
Claims
1. A display device, characterized by comprising: The application comprises: a display configured to display a user interface; a controller configured to: obtain at least one operation sequence when a trigger condition for generating a guide language is met, each of the operation sequences comprising at least one operation event, the operation event corresponding to a non-speech instruction of a user; screen at least one target sub-sequence from a plurality of sub-sequences included in the at least one operation sequence, the target sub-sequence having a frequency of occurrence in the at least one operation sequence greater than or equal to a frequency threshold, wherein the sub-sequence comprises at least one of the operation events; screen at least one target operation event according to a correlation between the operation events in the at least one target sub-sequence; generate a guide language corresponding to the at least one target operation event according to a semantic instruction corresponding to each of the target operation events, the guide language being used to guide the user to use a speech instruction corresponding to the at least one target operation event; control the display to display a target page after a final operation event is executed, and display the guide language on the target page, the final operation event being a last operation event in a last operation sequence in the at least one operation sequence.
2. The display device of claim 1, wherein, The controller configured to obtain at least one operation sequence when a trigger condition for generating a guide language is met, specifically configured to: start a task of generating a guide language when a trigger condition for generating a guide language is met; record the operation event corresponding to the non-speech instruction after receiving the non-speech instruction input by the user; determine the specified type of operation event as the final operation event when the specified type of operation event is detected; obtain the at least one operation sequence from the recorded operation events, wherein each of the operation sequences corresponds to a preset time unit, and the at least one operation event included in each of the operation sequences is arranged in a corresponding recording order, and a preset offset distance is arranged between adjacent two of the operation sequences.
3. The display device of claim 1, wherein, The controller configured to screen at least one target sub-sequence from a plurality of sub-sequences included in the at least one operation sequence, the target sub-sequence having a frequency of occurrence in the at least one operation sequence greater than or equal to a frequency threshold, specifically configured to: obtain the plurality of sub-sequences included in the at least one operation sequence; calculate the frequency of occurrence of each of the sub-sequences in the at least one operation sequence; screen the target sub-sequence having a frequency of occurrence in the at least one operation sequence greater than or equal to the frequency threshold.
4. The display device of claim 1, wherein, The controller configured to screen at least one target operation event according to a correlation between the operation events in the at least one target sub-sequence, specifically configured to: obtain a semantic feature corresponding to each of the operation events in the at least one target sub-sequence; generate a semantic instruction corresponding to each of the operation events according to the semantic feature corresponding to each of the operation events in the at least one target sub-sequence; calculate a semantic correlation between the semantic instructions corresponding to each of the operation events; screen the at least one target operation event according to the semantic correlation between the semantic instructions.
5. The display device of claim 4, wherein, The controller is configured to generate semantic instructions corresponding to each of the operation events according to semantic features corresponding to the operation events, and specifically configured to: obtain a mapping table, the mapping table comprising a mapping relationship between preset controls and guide words; determine a target guide word in the mapping table according to a control name in the semantic features corresponding to each of the operation events; determine a user intention according to the semantic features; process the target guide word according to the user intention to obtain semantic instructions corresponding to each of the operation events.
6. The display device of claim 1, wherein, The controller is configured to generate guide language corresponding to the at least one target operation event according to the semantic instructions corresponding to each of the target operation events, and specifically configured to: generate at least one candidate path according to the at least one target operation event, each of the candidate paths comprising at least one of the target operation events; determine a target path from the at least one candidate path, the target path comprising the least number of the target operation events; obtain semantic instructions corresponding to each of the target operation events in the target path; generate the guide language according to the semantic instructions corresponding to each of the target operation events in the target path.
7. The display device according to any one of claims 1 to 5, wherein The controller is configured to control the display to display a target page after the final operation event is executed, and display the guide language on the target page, and specifically configured to: detect a state of a far-field voice function; if the state of the far-field voice function is an open state, display the guide language on the target page; if the state of the far-field voice function is a closed state, do not display the guide language on the target page.
8. The display device of claim 7, wherein, If the state of the far-field voice function is an open state, the controller is configured to display the guide language on the target page, and specifically configured to: if the state of the far-field voice function is an open state, obtain a first time length from a time corresponding to a preset node to a current time; if the first time length is less than or equal to a time length threshold, display the guide language on the target page; if the first time length is greater than the time length threshold, do not display the guide language on the target page.
9. The display device of claim 5, wherein, The controller is configured to generate the guide language according to the semantic instructions corresponding to each of the target operation events in the target path, and specifically configured to: determine a target control corresponding to the final operation event; create a mapping relationship between the target control and the guide language; update the mapping relationship to the mapping table.
10. A method of guiding a user using voice instructions, characterized by, The method comprises: when a trigger condition for generating guide language is met, obtaining at least one operation sequence, each of the operation sequences comprising at least one operation event, the operation event corresponding to a non-voice instruction of a user; from a plurality of sub-sequences included in the at least one operation sequence, screening at least one target sub-sequence whose frequency of occurrence in the at least one operation sequence is greater than or equal to a frequency threshold, wherein the sub-sequence comprises at least one of the operation events; screening at least one target operation event according to a correlation between each of the operation events in the at least one target sub-sequence; According to the semantic instruction corresponding to each target operation event, a guide sentence corresponding to the at least one target operation event is generated, and the guide sentence is used to guide the user to use the voice instruction corresponding to the at least one target operation event; A target page after executing a final operation event is displayed, and the guide sentence is displayed on the target page, the final operation event being the last operation event in the last operation sequence in the at least one operation sequence.
Citation Information
Patent Citations
Control method and device based on voice assistant
CN116798418A
Display device and user scene creating method
CN118733917A